The community fix wave, rebased onto 0.50.0.0: every open community PR triaged and every open issue verified against master. 21 contributor pull requests adopted or reworked with credit, 35 verified open issues fixed directly, and a hostile review pass over the composed branch (composite review of the wave, /ship review army with 8 specialists + Claude adversarial, five Codex structured-review rounds until 0 findings). Every adopted fix carries a regression test proven red before the fix; the composed collector ran the full unit suite, `verify`, and the e2e lane, and CI is green on the merged head. Behavior changes (38, all listed under "Behavior changes" in the CHANGELOG): sync --json emits one JSON document with cost_gate nested; sweep-only sync reports synced/deleted; dream --dry-run skips the take/grade/calibration phases; context_pack/delta honor budget_tokens on the rendered text and delta never budget-drops threads; codex rollouts import one page per thread; config set registers exactly the search.* keys the search path reads; remote search/query report degraded: [safe_index_pending] before reindex; syncEnabled: false sources are no longer auto-synced; new sync.include_hidden key; migrateFactsToCanonical renumbers past the canonical page's max row_num; stall-aborted embed drains fail the phase. No new schema migration. Follow-ups filed in TODOS.md under "Community fix wave follow-ups". Co-Authored-By: IvanPham03 <IvanPham03@users.noreply.github.com> Co-Authored-By: Jey2311 <Jey2311@users.noreply.github.com> Co-Authored-By: LongPV <LongPV@users.noreply.github.com> Co-Authored-By: Masashi-Ono0611 <Masashi-Ono0611@users.noreply.github.com> Co-Authored-By: chris-conte <chris-conte@users.noreply.github.com> Co-Authored-By: dov-kela <dov-kela@users.noreply.github.com> Co-Authored-By: ethanbeard <ethanbeard@users.noreply.github.com> Co-Authored-By: garrytan-agents <garrytan-agents@users.noreply.github.com> Co-Authored-By: howardpark <howardpark@users.noreply.github.com> Co-Authored-By: janusch <janusch@users.noreply.github.com> Co-Authored-By: javieraldape <javieraldape@users.noreply.github.com> Co-Authored-By: jeanpierre121 <jeanpierre121@users.noreply.github.com> Co-Authored-By: markkasdorf <markkasdorf@users.noreply.github.com> Co-Authored-By: morven-ai <morven-ai@users.noreply.github.com> Co-Authored-By: ofroiland <ofroiland@users.noreply.github.com> Co-Authored-By: ozp <ozp@users.noreply.github.com> Co-Authored-By: pavelpp-topia <pavelpp-topia@users.noreply.github.com> Co-Authored-By: proxynico <proxynico@users.noreply.github.com> Co-Authored-By: sheelcheyne <sheelcheyne@users.noreply.github.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
5.4 KiB
Multi-language full-text search
GBrain's keyword search arm uses Postgres full-text search (tsvector/tsquery).
The tokenizer language is configurable via the GBRAIN_FTS_LANGUAGE
environment variable. Default: english.
How it works
Postgres text-search configurations control stemming and stop-word removal.
GBRAIN_FTS_LANGUAGE is read by src/core/fts-language.ts and applied on
both sides of the search:
- Query side —
websearch_to_tsquery('<lang>', $query)in both engines (Postgres and PGLite). - Write side — the
update_page_search_vectorandupdate_chunk_search_vectortrigger functions that populatepages.search_vectorandcontent_chunks.search_vector.
The value is validated against /^[a-z][a-z0-9_]*$/ before it is ever
interpolated into SQL (tsvector functions don't accept parameterized config
names). Invalid values fall back to english with a warning.
Built-in languages
Set the env var to any configuration your Postgres instance ships:
export GBRAIN_FTS_LANGUAGE=portuguese
export GBRAIN_FTS_LANGUAGE=spanish
export GBRAIN_FTS_LANGUAGE=german
List what's available:
SELECT cfgname FROM pg_ts_config;
PGLite (the embedded default engine) ships the same built-in snowball configurations as stock Postgres.
First install vs. changing language later
On first install (or upgrade), the configurable_fts_language schema
migration reads GBRAIN_FTS_LANGUAGE and stamps the trigger functions with
that language. After the migration has run, changing the env var alone does
NOT retokenize existing rows — the migration shows as applied and is skipped.
Use the explicit command:
export GBRAIN_FTS_LANGUAGE=portuguese
gbrain reindex-search-vector --dry-run # preview: language + row counts
gbrain reindex-search-vector --yes # recreate triggers + backfill
The stamp survives later schema work: initSchema() — including the replay
behind gbrain init --migrate-only on every upgrade — applies the schema
template under the configured language, so it re-creates the trigger
functions as they already are instead of reverting them to english.
The command recreates both trigger functions under the new language and
backfills every existing pages and content_chunks row in batches,
streaming progress to stderr. It is idempotent: re-running with the same
language produces identical vectors. --json prints a machine-readable
result envelope but still requires --yes (or an interactive confirm).
No cache purge is needed. The resolved language is part of the query-cache
key, so rows written under the previous language are unreachable after the
switch — searches read the retokenized index immediately instead of being
served pre-switch results for up to search.cache.ttl_seconds. Switching
back reaches the original rows rather than rebuilding them.
Recipe: accent-insensitive Portuguese (pt_br)
Brazilian Portuguese content often mixes accented and unaccented spellings
("São Paulo" vs "Sao Paulo"). Build a custom config that folds accents via
the unaccent extension, then stems with the portuguese snowball dictionary:
CREATE EXTENSION IF NOT EXISTS unaccent;
CREATE TEXT SEARCH CONFIGURATION pt_br (COPY = portuguese);
ALTER TEXT SEARCH CONFIGURATION pt_br
ALTER MAPPING FOR hword, hword_part, word
WITH unaccent, portuguese_stem;
Then point GBrain at it:
export GBRAIN_FTS_LANGUAGE=pt_br
gbrain reindex-search-vector --yes
Note: custom configurations require a real Postgres instance (e.g. the
Supabase engine). The config must exist BEFORE the migration or the reindex
command runs, or Postgres will reject the trigger recreation with
text search configuration "pt_br" does not exist.
Caveats
- One language per brain: the setting is global to the database, not per-source. Mixed-language brains should pick the dominant language (the vector-search arm is language-agnostic and covers the rest).
- Keep
GBRAIN_FTS_LANGUAGEset consistently in every environment that writes to the brain (CLI shells, MCP server, cron jobs) — a writer without the env var tokenizes new rows inenglishuntil the next reindex. - Interrupted reindex: the trigger flip commits before the backfill, so a
reindex-search-vectorrun killed mid-way (crash, full disk, SIGKILL) leaves new writes in the new language and un-backfilled rows in the old one — keyword search then matches only part of the corpus. The command records an in-progress marker plus a per-batch checkpoint,gbrain doctorfails withfts_reindex_incompleteuntil the run completes, and re-running with the sameGBRAIN_FTS_LANGUAGEresumes from the checkpoint rather than starting over. Budget minutes, not seconds, on brains with 100K+ chunks. - CJK (Chinese / Japanese / Korean): none of the built-in snowball
configurations can tokenize CJK text, so the FTS arm would return nothing
for CJK queries. Both engines detect CJK queries and route them to a
term-by-term
ILIKEfallback with term-frequency ranking instead (shared SQL insrc/core/search/cjk-keyword-sql.ts). The fallback is correct but not index-accelerated: it scanscontent_chunks.chunk_text, so latency grows with corpus size. Note the routing is query-driven: any query containing CJK characters takes the fallback today, even on a Postgres instance with a CJK-aware extension (pgroonga/zhparser) installed. Wiring a CJK-capableGBRAIN_FTS_LANGUAGEconfig past the fallback is a filed follow-up.