Skip to main content

index-content

Build typed derived metadata indexes from catalog records. Optional source arguments restrict extraction to matching archives; omit them to process the whole catalog.

metawarc index-content --dbfile collection.db --type links --type pdfs
metawarc index-content --dbfile collection.db --type images --type videos --type audio --type fonts
metawarc index-content --dbfile collection.db --type ooxmldocs --rescan
# Populate the 'texts' sidecar for phrase search
metawarc index-content --dbfile collection.db --text

--type (repeatable, default links): links, pdfs, images, ooxmldocs, oledocs, videos, audio, fonts.

--text (flag): also runs the text-extractor chain (TextExtractor for HTML, PdfTextExtractor for PDF, OoxmlTextExtractor for OOXML) and writes a texts Parquet sidecar with the (archive_id, warc_id, source, url, language, text) schema. The sidecar feeds search, /records/search, and the search_records MCP tool. Default off so structured extraction keeps its current cost; pass --text to opt in.

Options:

  • --rescan — rebuild even when a current sidecar exists
  • --batch-size — default 1,000
  • --silent / -s
  • --progress / --no-progress

Output is JSON: one result object per requested type (processed, skipped, failed), plus a {kind: "texts", ...} object when --text is set. The command fails if any type or text run reports failures.

Extraction uses MIME, extension, and a short signature probe. ZIP/XML, payload size, and time limits apply before a derived sidecar is published.

See metadata and analysis and Querying and export for the matching workflows.