Cookbook
metawarc covers indexing, query, extraction, analysis, and replay. This page is a task-oriented index: find the row that sounds like you, then follow the linked reference. If you are completely new, do the quick start first.
| You are a… | You want to… | Start with |
|---|---|---|
| Web archivist | Build a resumable index of a WARC collection without copying payloads | index, ingest, catalog, doctor |
| Researcher / journalist | Find PDFs, hosts, or date ranges and export selected payloads | stats, list-files, dump, get |
| Preservation engineer | Extract document metadata, hashes, duplicates, and integrity evidence | index-content, analyze |
| Replay operator | Browse archived sites locally or feed pywb | serve, replay, export-cdxj |
| Application developer | Expose a read-only typed API over an index | serve, REST /records/list |
| AI / automation builder | Give agents controlled metadata tools | mcp |