Files
gbrain/docs/guides/multi-language-fts.md
Sina Matian f1fbdfba19 v0.50.1.0 fix: community fix wave — 21 PR adoptions + 35 verified issue fixes + composite review, rebased on 0.50.0.0 (#5097)
The community fix wave, rebased onto 0.50.0.0: every open community PR triaged and every
open issue verified against master. 21 contributor pull requests adopted or reworked with
credit, 35 verified open issues fixed directly, and a hostile review pass over the composed
branch (composite review of the wave, /ship review army with 8 specialists + Claude
adversarial, five Codex structured-review rounds until 0 findings). Every adopted fix
carries a regression test proven red before the fix; the composed collector ran the full
unit suite, `verify`, and the e2e lane, and CI is green on the merged head.

Behavior changes (38, all listed under "Behavior changes" in the CHANGELOG): sync --json
emits one JSON document with cost_gate nested; sweep-only sync reports synced/deleted;
dream --dry-run skips the take/grade/calibration phases; context_pack/delta honor
budget_tokens on the rendered text and delta never budget-drops threads; codex rollouts
import one page per thread; config set registers exactly the search.* keys the search path
reads; remote search/query report degraded: [safe_index_pending] before reindex; syncEnabled:
false sources are no longer auto-synced; new sync.include_hidden key; migrateFactsToCanonical
renumbers past the canonical page's max row_num; stall-aborted embed drains fail the phase.

No new schema migration. Follow-ups filed in TODOS.md under "Community fix wave follow-ups".

Co-Authored-By: IvanPham03 <IvanPham03@users.noreply.github.com>
Co-Authored-By: Jey2311 <Jey2311@users.noreply.github.com>
Co-Authored-By: LongPV <LongPV@users.noreply.github.com>
Co-Authored-By: Masashi-Ono0611 <Masashi-Ono0611@users.noreply.github.com>
Co-Authored-By: chris-conte <chris-conte@users.noreply.github.com>
Co-Authored-By: dov-kela <dov-kela@users.noreply.github.com>
Co-Authored-By: ethanbeard <ethanbeard@users.noreply.github.com>
Co-Authored-By: garrytan-agents <garrytan-agents@users.noreply.github.com>
Co-Authored-By: howardpark <howardpark@users.noreply.github.com>
Co-Authored-By: janusch <janusch@users.noreply.github.com>
Co-Authored-By: javieraldape <javieraldape@users.noreply.github.com>
Co-Authored-By: jeanpierre121 <jeanpierre121@users.noreply.github.com>
Co-Authored-By: markkasdorf <markkasdorf@users.noreply.github.com>
Co-Authored-By: morven-ai <morven-ai@users.noreply.github.com>
Co-Authored-By: ofroiland <ofroiland@users.noreply.github.com>
Co-Authored-By: ozp <ozp@users.noreply.github.com>
Co-Authored-By: pavelpp-topia <pavelpp-topia@users.noreply.github.com>
Co-Authored-By: proxynico <proxynico@users.noreply.github.com>
Co-Authored-By: sheelcheyne <sheelcheyne@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 12:56:40 -04:00

5.4 KiB

Multi-language full-text search

GBrain's keyword search arm uses Postgres full-text search (tsvector/tsquery). The tokenizer language is configurable via the GBRAIN_FTS_LANGUAGE environment variable. Default: english.

How it works

Postgres text-search configurations control stemming and stop-word removal. GBRAIN_FTS_LANGUAGE is read by src/core/fts-language.ts and applied on both sides of the search:

  • Query side — websearch_to_tsquery('<lang>', $query) in both engines (Postgres and PGLite).
  • Write side — the update_page_search_vector and update_chunk_search_vector trigger functions that populate pages.search_vector and content_chunks.search_vector.

The value is validated against /^[a-z][a-z0-9_]*$/ before it is ever interpolated into SQL (tsvector functions don't accept parameterized config names). Invalid values fall back to english with a warning.

Built-in languages

Set the env var to any configuration your Postgres instance ships:

export GBRAIN_FTS_LANGUAGE=portuguese
export GBRAIN_FTS_LANGUAGE=spanish
export GBRAIN_FTS_LANGUAGE=german

List what's available:

SELECT cfgname FROM pg_ts_config;

PGLite (the embedded default engine) ships the same built-in snowball configurations as stock Postgres.

First install vs. changing language later

On first install (or upgrade), the configurable_fts_language schema migration reads GBRAIN_FTS_LANGUAGE and stamps the trigger functions with that language. After the migration has run, changing the env var alone does NOT retokenize existing rows — the migration shows as applied and is skipped. Use the explicit command:

export GBRAIN_FTS_LANGUAGE=portuguese
gbrain reindex-search-vector --dry-run    # preview: language + row counts
gbrain reindex-search-vector --yes        # recreate triggers + backfill

The stamp survives later schema work: initSchema() — including the replay behind gbrain init --migrate-only on every upgrade — applies the schema template under the configured language, so it re-creates the trigger functions as they already are instead of reverting them to english.

The command recreates both trigger functions under the new language and backfills every existing pages and content_chunks row in batches, streaming progress to stderr. It is idempotent: re-running with the same language produces identical vectors. --json prints a machine-readable result envelope but still requires --yes (or an interactive confirm).

No cache purge is needed. The resolved language is part of the query-cache key, so rows written under the previous language are unreachable after the switch — searches read the retokenized index immediately instead of being served pre-switch results for up to search.cache.ttl_seconds. Switching back reaches the original rows rather than rebuilding them.

Recipe: accent-insensitive Portuguese (pt_br)

Brazilian Portuguese content often mixes accented and unaccented spellings ("São Paulo" vs "Sao Paulo"). Build a custom config that folds accents via the unaccent extension, then stems with the portuguese snowball dictionary:

CREATE EXTENSION IF NOT EXISTS unaccent;

CREATE TEXT SEARCH CONFIGURATION pt_br (COPY = portuguese);

ALTER TEXT SEARCH CONFIGURATION pt_br
  ALTER MAPPING FOR hword, hword_part, word
  WITH unaccent, portuguese_stem;

Then point GBrain at it:

export GBRAIN_FTS_LANGUAGE=pt_br
gbrain reindex-search-vector --yes

Note: custom configurations require a real Postgres instance (e.g. the Supabase engine). The config must exist BEFORE the migration or the reindex command runs, or Postgres will reject the trigger recreation with text search configuration "pt_br" does not exist.

Caveats

  • One language per brain: the setting is global to the database, not per-source. Mixed-language brains should pick the dominant language (the vector-search arm is language-agnostic and covers the rest).
  • Keep GBRAIN_FTS_LANGUAGE set consistently in every environment that writes to the brain (CLI shells, MCP server, cron jobs) — a writer without the env var tokenizes new rows in english until the next reindex.
  • Interrupted reindex: the trigger flip commits before the backfill, so a reindex-search-vector run killed mid-way (crash, full disk, SIGKILL) leaves new writes in the new language and un-backfilled rows in the old one — keyword search then matches only part of the corpus. The command records an in-progress marker plus a per-batch checkpoint, gbrain doctor fails with fts_reindex_incomplete until the run completes, and re-running with the same GBRAIN_FTS_LANGUAGE resumes from the checkpoint rather than starting over. Budget minutes, not seconds, on brains with 100K+ chunks.
  • CJK (Chinese / Japanese / Korean): none of the built-in snowball configurations can tokenize CJK text, so the FTS arm would return nothing for CJK queries. Both engines detect CJK queries and route them to a term-by-term ILIKE fallback with term-frequency ranking instead (shared SQL in src/core/search/cjk-keyword-sql.ts). The fallback is correct but not index-accelerated: it scans content_chunks.chunk_text, so latency grows with corpus size. Note the routing is query-driven: any query containing CJK characters takes the fallback today, even on a Postgres instance with a CJK-aware extension (pgroonga / zhparser) installed. Wiring a CJK-capable GBRAIN_FTS_LANGUAGE config past the fallback is a filed follow-up.