Skip to main content

Quick Start

Short task-oriented paths to first success. Not sure where to start? Pick your role and goal in the cookbook. For the full reference, see the CLI index.

Index a collection in 30 seconds​

pip install metawarc
metawarc index 'archives/**/*.warc*' --dbfile collection.db --resume
metawarc catalog --dbfile collection.db
metawarc stats --dbfile collection.db --mode mimes

The default sidecar directory for collection.db is collection.data. Relocate it with --data-dir. Catalog paths are stored relative to that workspace whenever possible, so moving the database and its data directory together remains supported.

Query and export payloads​

metawarc list-files --dbfile collection.db --mimes application/pdf
metawarc dump --dbfile collection.db --exts pdf --limit 100 --output exported

Exports write below the requested directory, sanitize names, refuse silent overwrite, and record SHA-256 checksums in a JSONL manifest.

Extract metadata, search, and analyze​

metawarc index-content --dbfile collection.db --type links --type pdfs
metawarc index-content --dbfile collection.db --text # populate texts sidecar
metawarc analyze summary --dbfile collection.db --output summary.json
metawarc analyze metadata --dbfile collection.db --type all --top 20
metawarc search "Welcome to the museum of modern art" --dbfile collection.db

index-content --text runs the text-extractor chain (TextExtractor for HTML, PdfTextExtractor for PDF, OoxmlTextExtractor for OOXML) and writes a texts Parquet sidecar that feeds metawarc search, GET /records/search, and the search_records MCP tool.

Replay archived websites locally​

pip install 'metawarc[replay]' # or metawarc[api]
metawarc serve --dbfile collection.db
# Home page: http://127.0.0.1:8000/
# Replay URL: http://127.0.0.1:8000/replay/<YYYYMMDDHHMMSS>mp_/https://example.com/

Long-running exports as batch jobs​

# submit a job that materialises the filtered result set to CSV
metawarc jobs submit --format csv --limit 50000 --mimes application/pdf
# -> job-7f3a...

# wait for it
metawarc jobs wait job-7f3a...

The MVP ships the export-records job kind, which produces JSON, CSV, or Parquet under <data_dir>/jobs/<job_id>/. Job state persists across server restarts; concurrency is bounded by METAWARC_JOB_MAX_CONCURRENT and per-job execution by METAWARC_JOB_TIMEOUT. The same workflow is reachable through POST /jobs on metawarc serve.

Incremental operation and recovery​

metawarc ingest 'archives/**/*.warc*' --dbfile collection.db --dry-run
metawarc ingest 'archives/**/*.warc*' --dbfile collection.db --resume
metawarc doctor --dbfile collection.db

Next steps​