Phrase search
Search the extracted plain-text projection of every indexed HTML, PDF, and
OOXML record. The pipeline writes a texts Parquet sidecar; the
backend reads it with a bounded DuckDB columnar scan.
Pipeline
-
Index the WARC collection as for any other workflow:
metawarc index 'archives/**/*.warc*' --dbfile collection.db --resume -
Populate the
textssidecar once withindex-content --text:metawarc index-content --dbfile collection.db --textThe text-extractor chain selects
TextExtractorfor HTML,PdfTextExtractorfor PDF, andOoxmlTextExtractorfor OOXML. Each successful write becomes one row in thetextsParquet with columns(archive_id, warc_id, source, url, language, text).--textis opt-in so the default structured-extraction cost does not change. -
Search from any surface that exposes the workspace.
CLI
metawarc search "Welcome to the museum of modern art" --dbfile collection.db
metawarc search "API" --dbfile collection.db --limit 10
Each match prints one tab-separated line:
<warc_id>\t<url>\t<snippet>. The snippet is the first 256 characters
of the matched text.
REST
curl -H "Authorization: Bearer $METAWARC_API_TOKEN" \
'http://127.0.0.1:8000/records/search?phrase=Welcome&limit=10'
GET /records/search accepts phrase (1–512 characters) and limit
(1 – METAWARC_MAX_PAGE, default 50). The response is a
SearchResponse carrying phrase, limit, total, and a hits list
with archive_id, warc_id, source, url, and snippet.
MCP
search_records(phrase="Welcome to the museum of modern art", limit=50)
Returns the same hits structure as the REST endpoint.
Backend
DuckDB's FTS extension regressed in 1.5.x — its create_fts_index
pragma fails to materialise the virtual index. The dependable backend
is therefore a bounded read_parquet scan with a case-insensitive
ILIKE '%phrase%' predicate. For workspaces with hundreds of
thousands of records this is well within Parquet's strengths and avoids
a second index to keep in sync.
The _ensure_texts_index helper inside Workspace is a no-op
placeholder that retains the documented contract (return True when
at least one texts sidecar is present, False otherwise) so the
caller can surface a useful "no indexable sidecar" message.
Boundaries
- Empty phrases and
limithigher than the configuredMETAWARC_MAX_PAGEare rejected with HTTP 400 (REST) orclick.UsageError(CLI). - Phrase matching is
ILIKE '%phrase%': matches can span HTML markup, PDF/OOXML artefacts, and arbitrary slice boundaries in the projected text. - A workspace with no
textssidecar returns an emptyhitslist and the CLI prints(no matches)and exits 0.
Refresh after new WARCs
Re-running metawarc index adds them, but the texts sidecar is not
automatically rebuilt. Refresh it explicitly with --rescan:
metawarc index 'archives/**/*.warc*' --dbfile collection.db --resume
metawarc index-content --dbfile collection.db --text --rescan
See search, index-content --text,
/records/search, and the search_records
tool in Agents and MCP. For the chained
index → search workflow as a single scenario, see
Index and search.