metawarc search
Phrase-search across the indexed texts sidecar. Requires the workspace
to have been indexed with text extraction
(metawarc index-content ... --text) so the texts Parquet sidecar is
populated.
Synopsis
metawarc search [--dbfile <path>] [--data-dir <path>] <phrase> [--limit <n>]
Options
<phrase>(positional): the phrase to search for. Must not be empty.--dbfile,--data-dir: workspace options (see workspace).--limit(default50): maximum number of matches; capped at the configuredMETAWARC_MAX_PAGE.
Behaviour
The search runs a bounded columnar scan over the active texts Parquet
files via DuckDB's read_parquet. Phrases are matched with a
case-insensitive ILIKE '%phrase%' predicate, so matches can span HTML
markup, PDF/OOXML artefacts, and arbitrary slice boundaries in the
projected text. There is no second FTS index to build or keep in sync.
DuckDB's FTS extension regressed in 1.5.x (its create_fts_index pragma
fails to materialise the virtual index). The columnar scan is the
dependable backend until a future DuckDB release fixes the pragma; the
_ensure_texts_index helper is a no-op placeholder that retains the
documented contract.
If the workspace has no texts sidecar (because the user did not run
index-content --text yet), the command prints (no matches) and exits
0 — it does not raise.
Output
Each match prints one tab-separated line:
<warc_id>\t<url>\t<snippet>
The snippet is the first 256 characters of the matched text.
Examples
# populate the texts sidecar
metawarc index-content --text sample.warc
# search for a phrase
metawarc search "Welcome to the museum of modern art"
# <urn:uuid:...>\thttps://example.test/welcome.html\tWelcome to the museum of modern art ...
# limit results
metawarc search "API" --limit 10
See also: index-content --text,
/records/search,
and the search_records MCP tool in
Agents and MCP.