Performance and large collections
metawarc is designed so indexing memory stays bounded as record count grows.
Source WARC files are never loaded whole into memory. Writers create bounded
Parquet parts under staging/, validate them, then publish atomically.
Recommended flags
# Resume interrupted indexing; skip unchanged sources in update mode
metawarc index 'archives/**/*.warc*' --dbfile collection.db --resume
# Smaller record batches if memory is tight (default 10_000)
metawarc index archives --dbfile collection.db --batch-size 2000 --resume
# Preview incremental work before writing
metawarc ingest 'archives/**/*.warc*' --dbfile collection.db --dry-run
# Bound payload export volume
metawarc dump --dbfile collection.db --exts pdf --limit 1000 --max-bytes 104857600 --output exported
# Submit a long-running export as a batch job
metawarc jobs submit --format csv --limit 50000 --mimes application/pdf
Notes
- Indexing uses
--batch-size(default 10,000 records) for Parquet parts. - Content extraction uses a smaller default batch (1,000) because parsers are heavier.
- Interrupted writes stay in
staging/and are not registered as complete sidecars. doctorreports orphans;cleanuppreviews retired/staging files (requires--applyto remove).- Progress is automatic on interactive stderr;
--no-progressor--silentdisables it. - Website replay streams payloads with a configured byte cap; it is not a bulk export path.
- Phrase search reads the active
textsParquet files via DuckDB'sread_parquet. There is no separate index to build;--textonindex-contentwrites the projection.
Performance tips
- Index once, query many times: the catalog is the expensive step;
list-filesandstatsare cheap. - Resume: pass
--resumeso a crashed run continues from checkpoints. - Incremental ingest: use
ingest --dry-runtheningest --resumeinstead ofindex --mode force. - Filter early: restrict
--mimes,--exts,--archive-ids, and--limitbeforedump. - Extract only needed types:
index-content --type pdfsis cheaper than extracting every family. - Keep workspace local: DuckDB and Parquet sidecars should live on fast local disk next to the catalog.
- Do not pre-scan for progress: unknown totals are shown as counters; Metawarc will not read a WARC twice just to compute a percentage.
- Push large exports into batch jobs: when a single HTTP request
would exceed
METAWARC_REQUEST_TIMEOUT_SECONDS(default 30 s), submitmetawarc jobs submitinstead of loopinglist-files.
See also: Indexing collections, Batch jobs, Workspace.