Skip to main content

Indexing collections

Build a resumable index of a WARC collection without copying payloads. Source archives stay immutable; the DuckDB catalog and Parquet sidecars are the operational index.

First index​

metawarc index 'archives/**/*.warc*' --dbfile collection.db --resume
metawarc catalog --dbfile collection.db
metawarc doctor --dbfile collection.db

add, update, rescan, and force are explicit index --mode values. Default mode is update. Changed archives retain their stable catalog ID.

Incremental ingest​

Preview classification before writing:

metawarc ingest 'archives/**/*.warc*' --dbfile collection.db --dry-run
metawarc ingest 'archives/**/*.warc*' --dbfile collection.db --resume

A moved source is reported as a candidate and requires metawarc rebind ARCHIVE_ID NEW_PATH. Metawarc never silently guesses identity.

Optional payload hashes at index time​

metawarc index archives --dbfile collection.db --resume --hash-payloads

Hashes are cataloged sidecars and can also be computed later with analyze hashes.

metawarc index-content --dbfile collection.db --text

Runs the text-extractor chain (TextExtractor for HTML, PdfTextExtractor for PDF, OoxmlTextExtractor for OOXML) and writes the texts Parquet sidecar that feeds metawarc search, GET /records/search, and the search_records MCP tool. Default off so structured extraction keeps its current cost.

See index, ingest, index-content, and workspace layout. For the end-to-end index → search chain, see Index and search.