Skip to main content

Cookbook

metawarc covers indexing, query, extraction, analysis, search, batch export, and replay. This page is a task-oriented index: find the row that sounds like you, then follow the linked reference. If you are completely new, do the quick start first.

You are a…You want to…Start with
Web archivistBuild a resumable index of a WARC collection without copying payloadsindex, ingest, catalog, doctor
Search operatorRun the end-to-end index → search chain over a WARC collectionindex, index-content --text, search
Researcher / journalistFind PDFs, hosts, or date ranges and export selected payloadsstats, list-files, dump, get
Preservation engineerExtract document metadata, hashes, duplicates, and integrity evidenceindex-content, analyze
InvestigatorPhrase-search the extracted text of every HTML, PDF, and OOXML recordindex-content --text, search
Data engineerMaterialise a 50 000-record CSV that does not fit a single HTTP requestjobs submit, jobs wait
Replay operatorBrowse archived sites locally or feed pywbserve, replay, export-cdxj
Application developerExpose a read-only typed API over an indexserve, REST /records/list, /records/search, /jobs
AI / automation builderGive agents controlled metadata toolsmcp

Detailed walkthroughs​