Cookbook
metawarc covers indexing, query, extraction, analysis, search, batch export, and replay. This page is a task-oriented index: find the row that sounds like you, then follow the linked reference. If you are completely new, do the quick start first.
| You are a… | You want to… | Start with |
|---|---|---|
| Web archivist | Build a resumable index of a WARC collection without copying payloads | index, ingest, catalog, doctor |
| Search operator | Run the end-to-end index → search chain over a WARC collection | index, index-content --text, search |
| Researcher / journalist | Find PDFs, hosts, or date ranges and export selected payloads | stats, list-files, dump, get |
| Preservation engineer | Extract document metadata, hashes, duplicates, and integrity evidence | index-content, analyze |
| Investigator | Phrase-search the extracted text of every HTML, PDF, and OOXML record | index-content --text, search |
| Data engineer | Materialise a 50 000-record CSV that does not fit a single HTTP request | jobs submit, jobs wait |
| Replay operator | Browse archived sites locally or feed pywb | serve, replay, export-cdxj |
| Application developer | Expose a read-only typed API over an index | serve, REST /records/list, /records/search, /jobs |
| AI / automation builder | Give agents controlled metadata tools | mcp |