Skip to main content

Metadata and analysis

Extract document and media metadata from indexed records, optionally populate the texts sidecar for phrase search, and run revision-scoped collection reports.

Extract derived metadata​

metawarc index-content --dbfile collection.db --type links --type pdfs
metawarc index-content --dbfile collection.db --type images --type videos --type audio --type fonts
metawarc index-content --dbfile collection.db --type ooxmldocs --rescan

--type may be repeated. Valid types: links, pdfs, images, ooxmldocs, oledocs, videos, audio, fonts. Extraction uses MIME, extension, and bounded signature signals. Results include a versioned envelope, normalized metadata, raw parser output, warnings, stable error codes, inspected bytes, and duration.

Supported families include Office Open XML documents, templates, macro-enabled files, binary workbooks, and presentations (including PPSX); PDF; GIF, SVG, WebP, icons, bitmap, camera, and editing images; common MP4/QuickTime, AVI, WebM/Matroska, Ogg, MPEG, ASF/WMV, and FLV video; MP3, WAV, AIFF, FLAC, Ogg/Opus, M4A, WMA, MIDI, and RealAudio; and TTF/OTF, font collections, WOFF, WOFF2, and EOT fonts.

Add --text to opt in to the text-extractor chain (TextExtractor for HTML, PdfTextExtractor for PDF, OoxmlTextExtractor for OOXML). The run appends a {kind: "texts", ...} entry to the JSON output, writes a (archive_id, warc_id, source, url, language, text) Parquet sidecar, and feeds metawarc search, /records/search, and the search_records MCP tool:

metawarc index-content --dbfile collection.db --text
metawarc search "Welcome to the museum of modern art" --dbfile collection.db

Collection reports​

metawarc analyze summary --dbfile collection.db --output summary.json
metawarc analyze metadata --dbfile collection.db --type all --top 20
metawarc analyze hashes --dbfile collection.db --resume
metawarc analyze duplicates --dbfile collection.db --output duplicates.csv --output-format csv
metawarc analyze links --dbfile collection.db --output links.parquet --output-format parquet
metawarc analyze integrity --dbfile collection.db --deep --max-records 1000

analyze metadata reads stored extraction envelopes. --type all covers PDF, image, OOXML, OLE, video, audio, and font sidecars; link metadata uses analyze links.

analyze summary and analyze metadata are also exposed to MCP clients as collection_stats and metadata_summary; see Agents and MCP.

See index-content, analyze, and search.