* schema: v0.36.0.0 Hindsight calibration tables (migrations v67-v71) Foundation commit for the Hindsight-inspired calibration wave. Adds four new tables + one perf index, all source-scoped from day 1 per v0.34.1 discipline: - calibration_profiles (v67): per-holder LLM-narrative aggregation of TakesScorecard data. published BOOL gates E8 cross-brain mount sharing (default false). grade_completion REAL surfaces partial-grade state to the dashboard. active_bias_tags TEXT[] with GIN index feeds E3 (calibration- aware contradictions) and E7 (real-time nudge matching). - take_proposals (v68): propose_takes phase queue. Idempotency cache via (source_id, page_slug, content_hash, prompt_version) unique index mirrors the v0.23 dream_verdicts pattern. proposal_run_id supports --rollback by run. dedup_against_fence_rows JSONB audit column records what canonical takes the LLM was told to dedupe against at proposal time. - take_grade_cache (v69): grade_takes verdict cache. Composite PK on (take_id, prompt_version, judge_model_id, evidence_signature) — prompt edits OR evidence changes cleanly invalidate prior verdicts. applied=false default + auto-resolve-off-by-default (D17) means every fresh install needs operator opt-in before grade verdicts mutate the takes table. - take_nudge_log (v70): E7 nudge cooldown state. Polymorphic FK — a nudge fires on either a canonical take OR a pending proposal (CDX-5 fix). CHECK constraint enforces exactly-one-set. channel column lets future routing (webhook, admin SPA toast) reuse the same cooldown semantics. - takes_resolved_at_idx (v71): partial index for the Brier-trend aggregation queries. Engine-aware handler — Postgres uses CONCURRENTLY to avoid the ShareLock; PGLite uses plain CREATE. Every table carries wave_version TEXT NOT NULL DEFAULT 'v0.36.0.0' so the v0.36.0.0 calibration --undo-wave command (lands later in the wave) can reverse just this wave's writes. Plan: ~/.claude/plans/system-instruction-you-are-working-rippling-knuth.md covers the design rationale (D17/D18/D21 + CDX findings). Schema parity: - src/schema.sql for fresh Postgres installs - src/core/pglite-schema.ts for fresh PGLite installs - src/core/schema-embedded.ts auto-regenerated from schema.sql - src/core/migrate.ts for upgrade-in-place from older brains VERSION bumped to 0.36.0.0 for the wave. CHANGELOG entry lands at /ship. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * core: BaseCyclePhase abstract class enforces source-scope + budget contracts D21 from the eng review. Three new v0.36.0.0 cycle phases (propose_takes, grade_takes, calibration_profile) share enough structure that the duplication-vs-abstraction trade tips toward a shared base. Without this scaffold, source-isolation discipline would drift exactly the way it drifted in v0.34.1 — except this time across three new surfaces at once. What this enforces: 1. Phase signature is uniform: run(ctx, opts) → PhaseResult. 2. ctx.sourceId / ctx.auth.allowedSources MUST be threaded through every engine call. The base class surfaces a scope() helper that wraps sourceScopeOpts(ctx) and is the only sanctioned way to read source- scoped data. Forgetting to thread source scope becomes a TypeScript compile error, not a runtime leak. Closes the v0.34.1 leak class structurally for every new phase. 3. Budget meter wraps run() automatically. Subclass declares budgetUsdKey + budgetUsdDefault; base reads the resolved cap from config and creates the BudgetMeter. Subclass calls this.checkBudget() before each LLM submit; budget-exhausted phase still returns status='ok' (clean abort) so the cycle report shows partial completion, not failure. 4. Error envelope is uniform. Thrown errors get caught and converted to status='fail' with a phase-specific error.code via the subclass's mapErrorCode() hook. 5. Progress reporter integration. Base accepts the reporter via opts; subclasses call this.tick() instead of touching the reporter directly, so the phase name in the progress stream is always correct. Tests: 13 cases in test/core/base-phase.test.ts cover source-scope threading (5 cases including the empty-allowedSources-MUST-NOT-widen-scope regression), PhaseResult shape including the error envelope path (3 cases), dry-run propagation (2 cases), and budget meter construction (3 cases including config-key override). Synthesize.ts / patterns.ts (existing pre-v0.36 phases) deliberately do NOT retrofit to this base in v0.36.0.0 — too much churn for a refactor that doesn't pay off until v0.37+. Future phases use this by default. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * cycle: propose_takes phase + take_proposals queue write path (T3) LLM-based take extraction from markdown prose. Walks pages updated since last cycle, sends each page's body to a tuned extractor, writes the extracted gradeable claims to the take_proposals queue. User accepts / rejects via `gbrain takes propose --review` (lands in Lane C). Cycle wiring: lint → backlinks → sync → synthesize → extract → extract_facts → resolve_symbol_edges → patterns → recompute_emotional_weight → consolidate → propose_takes (NEW) → grade_takes (NEW; T4) → calibration_profile (NEW; T6) → embed → orphans → purge CyclePhase enum extended with 3 new entries; ALL_PHASES + NEEDS_LOCK_PHASES updated. All three new phases acquire the cycle lock (writes to take_proposals / take_grade_cache / calibration_profiles). Idempotency contract: The (source_id, page_slug, content_hash, prompt_version) composite unique index on take_proposals means an unchanged page never re-spends LLM tokens. Bumping PROPOSE_TAKES_PROMPT_VERSION cleanly invalidates the cache so a tuned prompt re-runs proposals on every page. Mirrors the v0.23 dream_verdicts pattern. F2 fence dedup: The phase reads the page's existing `<!-- gbrain:takes:begin -->` fence (when present) and passes the canonical take rows to the extractor as "things you have already captured." Prevents duplicate proposals when prose is appended to a page that already has takes. Records the fence rows the LLM was told to dedupe against on the take_proposals row for audit (dedup_against_fence_rows JSONB). Auto-resolve posture: propose_takes only WRITES proposals to the queue. Nothing in this phase mutates the canonical takes table. Operator opt-in via the queue review CLI (Lane C) is the only path from queue to canonical fence (D17). Prompt tuning status (v0.36.0.0 ship state): The default extractor prompt is annotated `v0.36.0.0-stub`. The real tuned prompt arrives via T19 synthetic corpus build (50 anonymized pages, 3-model parallel extraction, user reviews disagreement set, F1 ≥ 0.85 on training corpus + F1 ≥ 0.8 on ground-truth holdout). Until T19 lands, propose_takes runs but produces best-effort candidates the user reviews manually. Architecture: ProposeTakesPhase extends BaseCyclePhase (T2). Inherits source-scope threading via scope(), budget metering via this.checkBudget(), error envelope wrapping. budgetUsdKey: cycle.propose_takes.budget_usd (default $5/cycle). Budget exhaustion mid-page returns status='warn' with details.budget_exhausted=true — clean partial-completion semantics. Test seam: opts.extractor injection so the phase can run hermetically without touching the gateway. defaultExtractor (production path) calls gateway.chat with the EXTRACT_TAKES_PROMPT and parses the JSON array output via parseExtractorOutput. parseExtractorOutput defends against common LLM output sins: markdown code fence wrapping, leading prose, single-object instead of array, unknown kind values, weight out of [0,1], rows missing claim_text or exceeding 500 chars. Tests: 25 cases in test/propose-takes.test.ts cover the 4 pure helpers (parseExtractorOutput, contentHash, hasCompleteFence, extractExistingTakesForDedup) + 7 phase integration scenarios (happy path, cache hit, fence dedup, extractor failure, empty pages, skipPagesWithFence, proposal_run_id stability). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * cycle: grade_takes phase + take_grade_cache verdict pipeline (T4) Walks unresolved takes that are old enough to have outcome data, retrieves evidence from the brain, asks a judge model to verdict each one. Writes verdicts to take_grade_cache. Optionally — only when operator has flipped the opt-in config flag — auto-applies high-confidence verdicts to the canonical takes table via engine.resolveTake. Auto-resolve posture (D17 — DISABLED by default): On a fresh install, grade_takes runs and writes verdicts to the cache, but applied=false on every row. Operator reviews the queue, then flips `cycle.grade_takes.auto_resolve.enabled: true` once trust is earned. Mirrors the propose_takes review-queue posture: queue exists, mutation requires explicit opt-in. Conservative threshold (D12): When auto_resolve.enabled is true, a verdict auto-applies only when confidence >= 0.95 (single-judge path). T5 ensemble path lands next, tightening this further with 3/3 unanimous requirement. 'unresolvable' verdict NEVER auto-applies even at confidence=1.0 — there's no canonical column for "we tried and there's no evidence yet." Evidence retrieval status (v0.36.0.0 ship state): The default evidence retriever returns an "evidence-retrieval not yet wired" placeholder. Most verdicts produced by the stub-judge against the stub-evidence will be 'unresolvable'. Real retrieval (hybrid search over pages newer than the take's since_date, optionally augmented by a gateway web-search recipe in v0.37+) lands as a follow-up. Documented limitation per CDX-8 + D17 — the phase ships now so the wiring is real and the cache table accumulates verdicts even if early ones are conservative. Cache key: Composite primary key on take_grade_cache is (take_id, prompt_version, judge_model_id, evidence_signature). Prompt edits OR evidence changes OR judge swap cleanly invalidate prior verdicts. Mirrors the v0.32.6 eval_contradictions_cache pattern. evidence_signature = SHA-256 of (judge_model_id + '|' + evidence_text) so identical evidence under a different judge does NOT collide. Architecture: GradeTakesPhase extends BaseCyclePhase. Inherits source-scope threading, budget metering (cycle.grade_takes.budget_usd, default $3/cycle), error envelope. Test seam: opts.judge + opts.evidenceRetriever injection so the phase runs hermetically. parseJudgeOutput defends against fence-wrapping, leading prose, out-of-range confidence (clamps to [0,1]), invalid verdict labels, oversized reasoning (truncated at 400 chars). Returns null on unrecoverable parse — caller treats null as "judge_output_parse_failed / unresolvable at confidence 0.0" so the row still lands in cache with the parse failure surfaced via warnings. takeIsOldEnough gates on since_date (default 6 months). Tolerates YYYY-MM-DD and YYYY-MM formats. Returns false on null/unparseable since_date so takes without dates never get graded (we'd be hallucinating temporal context). Tests: 23 cases covering parseJudgeOutput (7 cases), evidenceSignature (3), takeIsOldEnough (5), and 8 phase integration scenarios — happy path, D17 auto-resolve-off default, D12 above-threshold auto-apply, below- threshold cache-only, unresolvable-NEVER-applies, cache hit, too-recent gate, judge-throw warning. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * cycle: grade_takes ensemble tiebreaker for borderline verdicts (T5 / E2) Multi-judge ensemble tiebreaker, additive on top of T4's single-judge foundation. Reuses gateway.chat as the per-model judge interface; runs three judges in parallel via Promise.allSettled. Pure aggregation logic in aggregateEnsemble() — no SQL, no LLM, hermetically testable. When ensemble fires (T5 trigger band): Only when ALL of: - opts.useEnsemble === true (default false) - opts.ensembleJudges array is non-empty - single-model confidence in [0.6, 0.95) (configurable via opts.ensembleTriggerBand) - single-model verdict !== 'unresolvable' Above 0.95 the single judge is already sufficient (T4 path). Below 0.6 the verdict is clearly review-only — ensemble wouldn't change the posture. 'unresolvable' from single-judge means no evidence yet; calling three more judges on the same evidence won't manufacture some. Conservative auto-apply (D12): Ensemble verdict auto-applies via engine.resolveTake only when ALL of: - autoResolve === true (operator opt-in per D17) - ensemble.agreement === 3 (3/3 unanimous) - ensemble.minConfidence >= ensembleThreshold (default 0.85) - winning verdict !== 'unresolvable' Schema-level monotonic-tightening guard for ensembleThreshold lives in the takes resolution layer. Cache identity: When ensemble fires, the cache row's judge_model_id becomes 'ensemble:<modelA>+<modelB>+<modelC>' — a future re-run with different ensemble membership doesn't collide with prior verdicts. evidence_signature is recomputed because it includes the judge_model_id. aggregateEnsemble (pure): - 3/3 unanimous → agreement=3, minConfidence=min across the three - 2/3 majority → agreement=2, minConfidence across the agreeing two - 1/1/1 disagreement → tie-break: prefer non-'unresolvable', then alphabetical for determinism - 'unresolvable' from one model NEVER tips a 2-vote majority toward 'unresolvable' — by-label tally only counts a model toward its own label - All three judges failing (allSettled rejected) → verdict='unresolvable' with agreement=0; auto-apply path blocked - Single judge survives + two fail → agreement=1; the lone verdict wins but auto-apply gated by the 3/3 requirement Tests: 16 cases. aggregateEnsemble (6): 3/3, 2/3, 1/1/1, unresolvable-tipping-resistance, all-failed, partial-failed-but-survives. Phase trigger conditions (5): useEnsemble=false default, useEnsemble=true in borderline band, single >= 0.95 skip, single < 0.6 skip, single = 'unresolvable' skip. Phase auto-apply rules (5): 3/3+threshold+autoResolve, 2/3 majority no apply, 3/3 below threshold no apply, one ensemble judge throws still aggregates from allSettled, empty ensembleJudges falls through to single. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * cycle: calibration_profile phase + shared voice gate across surfaces (T6) The calibration narrative layer. Reads TakesScorecard, asks an LLM to write 2-4 conversational pattern statements ("right on tactics, late on macro by 18 months"), passes them through the voice gate, derives active bias tags, writes the row to calibration_profiles. This is the read-side that E1 (think anti-bias rewrite), E3 (contradictions join), E6 (dashboard), and E7 (real-time nudges) all consume. Voice gate (D24 — single function, multiple surfaces): ALL five calibration UX surfaces import the same gateVoice() function from src/core/calibration/voice-gate.ts. Mode parameter ('pattern_statement' | 'nudge' | 'forecast_blurb' | 'dashboard_caption' | 'morning_pulse') drives surface-specific tuning via the rubric the gate ships to its Haiku judge. NO forked implementations — voice rubric drift would defeat the gate. Each mode's rubric explicitly forbids preachy / clinical / corporate voice; a structural test pins this. Anchors the cross-cutting voice rule from /plan-ceo-review D2-D8. Fallback policy (D11): Up to 2 generation attempts (configurable). On both rejects → fall back to a hand-written template from src/core/calibration/templates.ts. Templates are intentionally short and a little "robotic" — they're the safety net, not the destination. voice_gate_passed=false + voice_gate_attempts get persisted on the calibration_profiles row so the operator can review the failing examples and tune the rubric over time. Suppressing the surface silently is NEVER an option — that's how voice quality silently degrades. parseJudgeOutput defaults to 'academic' on parse failure (NEVER passes pass-through) so a Haiku output garble falls through to the template rather than letting unverified text reach the user. calibration_profile phase: Extends BaseCyclePhase. Cold-brain skip: <5 resolved takes → no row written, no LLM call. Otherwise: scorecard via engine.getScorecard() → patterns via voice-gated generator → bias tags via separate generator (best-effort; failure logs warning, phase continues). The DB INSERT lands in the v67 calibration_profiles row with source_id, holder, the patterns, voice gate audit fields, active bias tags, and grade_completion (F1 fix — partial-grade state surfaces to the dashboard "60% graded" badge). Budget gate at $0.50/cycle default (mostly Haiku). Below-budget before-LLM-call check returns status='warn' without writing the row. Per-domain scorecards are a placeholder for v0.36.0.0 ship state — the F12 batchGetTakesScorecards() engine method that powers per-domain rendering lands in Lane C alongside the CLI/MCP surface. Architecture: parsePatternStatementsOutput is tolerant of LLM emitting numbered lists / bulleted lines despite the prompt asking for plain lines. Caps at 4 patterns + drops excessively long lines (>200 chars). parseBiasTagsOutput lowercases input + drops non-kebab-case tokens (defends against the LLM emitting "Over-Confident Geography" with spaces or capitals). Caps at 4 tags. Tests: 43 cases across two new test files. voice-gate.test.ts (24): parseJudgeOutput (7), gateVoice happy path (3), fallback path (5), mode parity (2), templates (7). calibration-profile.test.ts (19): parsers (10), pickFallbackSlots (3), phase integration (6 — cold-brain skip, happy path, voice gate fallback, grade_completion plumbed through, bias-tags failure non-fatal, source_id scope reaches INSERT). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * cli: gbrain calibration + get_calibration_profile MCP op (T7) Public-facing read surface for the v0.36.0.0 calibration wave. CLI prints the active calibration profile; MCP op exposes the same data path for agents. Mirror of the v0.29 salience/anomalies shape (pure data fn + JSON formatter + human formatter + thin CLI dispatch). CLI: `gbrain calibration` Flags: --holder <id> specific holder (default 'garry') --json machine output for piping --regenerate run calibration_profile phase now --undo-wave <ver> [placeholder — wires in Lane D / T17] ab-report [placeholder — wires in Lane D / T18] Human output: Calibration profile — holder: garry, source: default Generated: <local timestamp> [Note: built on 60% graded — partial completion this cycle.] (when grade_completion < 0.9) [Note: voice gate fell back to template (2 attempts).] (when voice_gate_passed=false) Resolved: 12 takes Brier: 0.210 (lower is better) Accuracy: 60.0% Partial: 10.0% Pattern statements: • You called early-stage tactics well — 8 of 10 held up. Active bias tags: over-confident-geography Cold-brain fallback message names the exact dream command to run. MCP: `get_calibration_profile` (scope: read) Param: holder?: string (defaults to 'garry') Returns: latest CalibrationProfileRow | null Source-scoping via sourceScopeOpts(ctx): scalar source-bound clients see only their source; federated_read scopes see the union of allowed sources; no source filter when neither is set (CLI default path). Throws GBrainError('INVALID_HOLDER') on empty/non-string holder so remote callers get a structured error instead of a SQL-shape failure. Architecture: getLatestProfile is the pure data fn — engine + opts → CalibrationProfileRow | null. Reused by both the CLI and the MCP op. Source-scoped via the standard v0.34.1 spread pattern (scalar sourceId vs sourceIds array). formatProfileText is pure — null → cold-brain message, populated → full printout. Annotates partial-grade rows and voice-gate-fallback rows so the operator sees data-quality status inline. parseArgs is exported via __testing for unit coverage. Sub-command ('ab-report') vs flag distinction is intentional — keeps the surface parallel with `gbrain eval cross-modal` etc. Tests: 21 cases. parseArgs (6 cases): empty, --holder, --json, --regenerate, --undo-wave, ab-report. getLatestProfile (5 cases): happy, null, scalar source scope, federated array scope, no-source-filter default. formatProfileText (5 cases): cold-brain, happy, partial-grade note, voice-fallback note, published-to-mounts note. getCalibrationProfileOp (5 cases): default holder, scalar source scope, federated scope union, returns-null-on-unknown-holder, throws on empty holder. Lane D follow-ups: --undo-wave (T17) and ab-report (T18) print a clear "lands in Lane D" stderr line + exit 2; the surfaces exist for early testers, the implementations land next. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * think: --with-calibration + anti-bias prompt rewrite (T8 / E1, D22) Optional anti-bias rewrite mode for `gbrain think`. When set, the active calibration profile gets injected per the D22 placement spec (AFTER retrieval evidence, BEFORE the user's question). The bias filter applies to QUESTION FRAMING, not evidence interpretation — matches LLM-as-judge best practice (bias prompts near end of context perform better). Default behavior unchanged (R1 regression guard): omitting --with-calibration produces the v0.28-vintage user-message shape with the question first, then retrieval. Existing think users see no change. Two user-message shapes in buildThinkUserMessage: Default (no calibration): Question: X <pages>...</pages> <takes>...</takes> <graph>...</graph> Respond with a single JSON object... With calibration (D22): <pages>...</pages> <takes>...</takes> <graph>...</graph> <calibration holder="garry"> Track record: Brier 0.210 (lower is better). Active patterns: - You called early-stage tactics well — 8 of 10 held up. Active bias tags: over-confident-geography </calibration> Question: X Respond... Calibration block is built by buildCalibrationBlock (exported for the E3 contradictions probe to render the same shape). System prompt extension (withCalibration:true): - Names BOTH the user's PRIOR (default reasoning) AND the COUNTER-PRIOR from their hedged-domain self. - References active bias tags by name when relevant ("this fits the over-confident-geography pattern"). - Does NOT silently substitute the debiased answer. ALWAYS surfaces both priors transparently. - Adds a "Calibration" section between Conflicts and Gaps in the answer body. RunThinkOpts extension: - withCalibration?: boolean — opt-in - calibrationHolder?: string — defaults to 'garry' When withCalibration=true and no profile exists, runThink falls back to baseline behavior + pushes NO_CALIBRATION_PROFILE to warnings (visible to the operator). When the calibration fetch fails, CALIBRATION_FETCH_FAILED warning surfaces with the underlying error. Either path keeps think working; the calibration loop is enhancement, not requirement. CLI: `gbrain think "<q>" --with-calibration [--calibration-holder <id>]` Tests: 11 cases. buildThinkSystemPrompt (4 cases): R1 regression — default/false/omitted → no anti-bias rules; with calibration → adds PRIOR + COUNTER-PRIOR + bias-tag reference; preserves existing hard rules. buildCalibrationBlock (3 cases): happy path, null brier omitted (not "Brier null"), empty patterns + tags still well-formed. buildThinkUserMessage (4 cases): R1 regression — without calibration: question first; D22 placement — retrieval → calibration → question → instruction; graph + calibration ordering; empty retrieval blocks render placeholders without breaking shape. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * contradictions: calibration-profile join (T9 / E3) Cross-references each contradiction finding against the active calibration profile. When a contradiction's domain matches an active bias tag (e.g. "over-confident-geography" or "late-on-macro-tech"), the output gains a one-line bias context explaining which pattern this fits. Pure functions only — no DB writes, no LLM calls. The probe runner imports tagFindingWithCalibration() and applies it to each finding before emitting. When no profile exists or no tags match, the helper returns null and the runner emits the unchanged finding (regression R2 — contradictions output is byte-identical to v0.32.6 when no calibration profile is present). Match heuristic (v0.36.0.0 ship-state): Bias tags are kebab-case axis-then-domain slugs ('over-confident-geography'). computeDomainHint() extracts a domain hint from the finding's slugs + holder + verdict text: - wiki/companies/... → hiring | market-timing - wiki/people/... → founder-behavior - macro / geography / tactics / ai segments in slug → matching tag First-match-wins for ordering determinism. Match is intentionally fuzzy — the v0.32.6 contradictions probe doesn't yet carry structured domain metadata. v0.37+ structured-domain-on-takes (Hindsight-style enum) tightens this. Output: Returns { bias_tag: string, context: string } | null. Context format: "This contradiction fits your active bias pattern \"<tag>\" (Brier 0.31). Verdict: contradiction; severity: medium. Consider reviewing both sides through the lens of that pattern." Tests: 13 cases. R2 regression (2): null profile → null tag; empty active_bias_tags → null tag. computeDomainHint (5): companies / people / macro / geography / unknown paths produce expected hints. Match path (4): macro→late-on-macro-tech, geography→over-confident-geography, mismatch returns null, first-match-wins with multiple candidate tags. buildBiasContextString (2): emits tag+verdict+severity+Brier; omits Brier when null (no "Brier null" leak). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: Brier-trend forecast at write time (T10 / E5) Pure math layer over existing TakesScorecard data. Zero new LLM cost, zero new schema. Surfaces the user's historical Brier for the take's (holder, domain) bucket at write time so they see "your historical Brier in macro takes is 0.31" before committing the take. Voice-gate-rendered output: The user-facing string goes through gateVoice mode='forecast_blurb' via templates.ts (already in T6). This module is the pure data layer; the template renders the math into the conversational voice. v0.36.0.0 ship state: Bucket dimension is the DOMAIN (slug-prefix). The conviction-weight bucket dimension would need a new engine method (engine.batchGetTakeBucketStats per F11) — deferred to v0.37+. Until then, forecast = historical Brier in this holder's domain. resolveDomainPrefix() keeps slug-prefix-looking domain hints ('companies/', 'wiki/macro') and falls back to overall for free-form hints ('macro tech', 'geography'). Hindsight-style structured domain on takes (CDX-11 mitigation TODO) tightens this in v0.37+. MIN_BUCKET_N = 5: Below this sample size, the forecast returns predicted_brier=null with insufficient_data=true. Template renders "Forecast unavailable: only N resolved takes at this conviction yet" instead of a noisy estimate. Architecture: computeForecast(input) — pure function, takes scorecards already fetched; ideal for tests + reuse across batched paths. forecastForTake(engine, input) — convenience wrapper, 1-2 engine round-trips (no domain → 1; with domain → 2). batchForecast(engine, inputs[]) — memoizes per (holder, domainPrefix); N inputs collapse to ≤2*unique_holders unique engine calls. Used by the propose-queue review flow (50 candidates → 1-2 scorecard fetches). Tests: 14 cases. computeForecast (4): insufficient_data branch, stable forecast, overall fallback, MIN_BUCKET_N export. resolveDomainPrefix (5): undefined/empty/whitespace → undefined; slug-prefix → kept; free-form → undefined. forecastForTake (3): 1-call overall, 2-call domain, free-form fallback. batchForecast (2): cache collapse for repeat queries; different holders do not collapse. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: gstack-learnings coupling on incorrect resolutions (T11 / E4) When the grade_takes phase auto-resolves a take as 'incorrect' or 'partial', optionally write a learning entry to gstack's per-project learnings.jsonl so other gstack skills (plan-ceo-review, ship, investigate, ...) can pull it as context when relevant. The brain teaches every other tool about the user's track record. Config gate (D5 / CDX-17 mitigation): `cycle.grade_takes.write_gstack_learnings` defaults FALSE. External users may not have gstack installed; the gstack-learnings binary API isn't stable yet. Garry's brain flips it true to opt in. Quality gate: Only 'incorrect' and 'partial' verdicts trigger the write. 'correct' resolutions are noise (we expected the take to hold up — no learning). 'unresolvable' has no canonical column. Defense-in-depth runtime guard in writeIncorrectResolution() rejects ineligible qualities with reason='quality_not_eligible' so a caller misuse never surfaces a malformed learning entry. Auto-apply only: Coupling fires only when grade_takes both auto-applies AND the verdict is incorrect/partial AND the config flag is enabled. Manual resolutions via `gbrain takes resolve` intentionally DO NOT propagate to gstack — manual writes already carry operator intent; the calibration loop is the noise-prone path that earns coupling. Namespace: Every entry's key starts with 'gbrain:calibration:v0.36.0.0:'. Lane D `gbrain calibration --undo-wave v0.36.0.0` (T17) filters on this prefix for the optional gstack-scrub step. First active bias tag suffixes the key (e.g. 'take-42:over-confident-geography') so future analysis can group learnings by bias pattern. Architecture: buildLearningEntry — pure. Truncates claim at 200 chars + ellipsis; emits Pattern: line when activeBiasTags present; defaults confidence to 0.8 when caller omits it. writeIncorrectResolution — async wrapper. Honors config gate; honors quality gate; calls the injected writer (or defaultGstackWriter in production). Failures are non-fatal: returns { written: false, reason: 'write_failed' | 'binary_missing', error }. The grade_takes phase logs to result.warnings and continues — gstack coupling failure NEVER aborts a cycle. defaultGstackWriter — shells out to gstack-learnings-log binary via execFileSync. Throws GBrainError('GSTACK_BINARY_NOT_FOUND') when the binary isn't on PATH; writeIncorrectResolution classifies that error to reason='binary_missing' so the operator sees the install hint instead of a generic write_failed. Wired into grade-takes.ts after engine.resolveTake() inside the auto-apply block. Only fires when shouldApply=true. Tests: 14 cases. buildLearningEntry (7): canonical shape, partial vs incorrect wording, bias-tag suffix, no-tag fallback, claim truncation, default confidence, no-reasoning omission. writeIncorrectResolution (7): config gate, quality gate, happy path, writer-throw graceful degrade, binary-missing classification, async writer awaited, partial quality writes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * doctor: 4 calibration checks — abandoned/freshness/drift/voice (T12) Adds the four calibration doctor checks per the eng-review spec. abandoned_threads: Counts active high-conviction takes (weight >= 0.7) older than 12 months that have never been superseded. Signal, not error — always status='ok' with a count. The hint sends users to `gbrain calibration` for details. calibration_freshness: Warns when the active profile is older than 7 days (configurable via the same env-var pattern other freshness checks use). Cold-brain branch (no profile yet) returns ok without scolding. Hint points at `gbrain calibration --regenerate`. grade_confidence_drift (CDX-11 mitigation): Surfaces the count of auto-applied grade verdicts. Below 30: returns "need 30+ for drift detection". At/above 30: returns "drift math arrives in v0.37+". The surface is wired; the actual confidence-vs-accuracy correlation math is a v0.37+ follow-up once we have 30+ auto-applied verdicts to measure against. Closes the CDX-11 hole structurally — the operator sees the surface even before the math is meaningful. voice_gate_health: Tracks voice gate failure rate over the last 7 days. <30% fail rate → ok (template fallback is fine in isolation). >=30% → warn with hint to review src/core/calibration/voice-gate.ts rubric. Anchors the cross-cutting voice rule observability story. All four checks return status='warn' with a diagnostic message on engine errors — non-blocking, never throws. Matches the existing doctor check pattern (see checkSyncFreshness for prior art). Wired into runDoctor after checkRerankerHealth (the v0.35 cluster), in the canonical block 10 slot. Tests: 15 cases. 4 per check (happy path, alt-status, engine-throw diagnostic, plus boundary tests for the freshness staleness gate at exactly 7 days and the grade drift gate at 30 applied verdicts). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: E7 nudge + 14-day cooldown (T13 / D16 F3) Real-time pattern surfacing when a newly-committed high-conviction take matches an active bias pattern. Conversational nudge text via the templates module; 14-day cooldown per (take_id, nudge_pattern) via take_nudge_log to prevent the feedback loop where each cycle re-fires the same nudge on the same take. Threshold gates (D16 F3): - holder match (profile.holder === take.holder) - conviction-weight > 0.7 (strict greater than) - take's slug-derived domain hint matches an active bias tag (takeDomainHint — same heuristic as eval-contradictions/calibration-join.ts for cross-surface consistency) Cooldown gate: Before firing, probe take_nudge_log for (take_id, nudge_pattern) rows with fired_at >= now() - 14 days. Any hit → silently skip. After firing, insert a new row with channel='stderr' so the next 14 days are gated. Feedback-loop prevention: User hedges a take in response to a nudge (e.g. weight 0.85 → 0.65). Even though the take's `weight` field changed, the cooldown row for the over-confident-geography pattern is still there from the original fire — so the next cycle's evaluateAndFireNudge() silently skips. The user reset path (gbrain takes nudge --reset N) clears the cooldown to re-arm. Output channel (v0.36.0.0 ship state): STDERR only. Schema's `channel` column already supports multi-channel (webhook, admin SPA toast); routing those is a v0.37+ follow-up. Architecture: evaluateNudgeRule(take, profile) — pure rule check. Returns { matched, reason, matchedTag }. No engine call. checkCooldown(engine, takeId, pattern) — engine probe, returns boolean. recordNudgeFire(engine, opts) — INSERT into take_nudge_log. evaluateAndFireNudge(opts) — full pipeline. Returns NudgeDecision. resetNudgeCooldown(engine, takeId) — DELETE...RETURNING for the CLI. buildNudgeText delegates to templates.ts nudgeTemplate (D24 mode='nudge' voice). v0.36.0.0 ship state uses the template directly; LLM-generated nudge text via the voice gate lands in v0.37+ when we have production examples to tune from. Tests: 22 cases. takeDomainHint (5): companies/people/macro/geography/unrecognized. evaluateNudgeRule (6): no_profile, wrong_holder, conviction-at-threshold- is-NOT-eligible (strict >), no matching tag, happy match, first-match-wins for multiple candidate tags. checkCooldown (3): true on row hit, false on no row, cutoff date param verifies the 14-day boundary. evaluateAndFireNudge (4): happy fire (text contains hush command + matched tag), cooldown silent skip (no INSERT, no stderr), no_profile short-circuit, below-conviction short-circuit (no cooldown query fired). buildNudgeText (2): hush command shape, conviction value embedded. resetNudgeCooldown (2): returns count, idempotent on zero rows. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: E8 team-brain sharing + D18 cross-brain query semantics (T14) Cross-brain calibration profile resolution per the D18 4-rule contract. Pins all four cross-brain leak surfaces in dedicated unit tests so future mount features can't silently regress this security model. D18 semantics (committed): Rule 1 — LOCAL-FIRST ORDERING. Query the local brain first. If a profile exists, return it. Do NOT also query mounts (avoids stale-mount-overrides-fresh-local). Verified: mountResolver is NOT called when local has a hit. Rule 2 — MOUNT FALLBACK. Only when local has no profile AND canReadMounts=true, walk the mounts in priority order. First match wins. Each mount-side row must have published=true to be visible (D15 asymmetric opt-in). Rule 3 — CROSS-BRAIN ATTRIBUTION. Every returned profile carries source_brain_id + from_mount flag. Consumers (E1 think rewrite, E3 contradictions, E7 nudge, E6 dashboard) MUST surface this via attributionSuffix() so the user sees which brain answered. Rule 4 — SUBAGENT PROHIBITION. canReadMountsForCtx() classifier returns FALSE for subagent loops without trusted-workspace allowedSlugPrefixes. Closes the OAuth-token-to-cross-brain-leak surface — subagents see ONLY their local-brain results regardless of which holder they query. Exception: trusted cycle phases (synthesize/patterns) pass allowedSlugPrefixes set and ARE allowed to read mounts. Pinned in the classifier test. Architecture: queryAcrossBrains(localEngine, opts) — pure orchestrator. Composes getLatestProfile() from src/commands/calibration.ts. Mount engine access is via opts.mountResolver — production wires this to the v0.19+ gbrain mounts subsystem; tests inject a stub returning an ordered list of mocked engines. Decouples cross-brain LOGIC from multi-engine PLUMBING. canReadMountsForCtx(ctx) — pure classifier table. Drives the rule-4 gate. Production callers compose it from OperationContext. attributionSuffix(result) — pure formatter. Emits the "(from mounted brain: <id>)" suffix when from_mount=true; empty string when local. Mandatory for user-visible cross-brain consumers. Tests: 15 cases pinned to the 4 D18 rules + 4 supplementary structural checks. D18-1: published=false profile on mount stays hidden. D18-2/3: subagent context cannot fall back to mounts (2 cases — null on local-empty + canReadMounts=false, local hit still returned). D18-4: attribution surfaces source_brain_id (3 cases — mount answer flag, local answer flag, attributionSuffix formatter). Rule 1 local-first ordering (2 cases — mountResolver NOT called on local hit, IS called on local empty). Mount priority order (3 cases — first published=true wins, all published=false returns null, no mounts configured returns null without throwing). canReadMountsForCtx classifier (4 cases — local CLI true, MCP non-subagent true, subagent without trusted-workspace false, subagent WITH trusted-workspace true). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * admin: E6 Calibration tab + D23 server-rendered SVG + TD2 contrast bump (T15) Adds the v0.36.0.0 admin SPA Calibration tab. Per the design review, the approved variant-B (Linear calm clarity) layout: single-column flow, generous whitespace, ONE big sparkline as hero, then patterns, then domain bars, then abandoned threads. D23 server-rendered SVG architecture: src/core/calibration/svg-renderer.ts — pure functions. data → SVG string. No DOM, no React, no chart library dep. Inlines the admin design tokens (#0a0a0f bg, #3b82f6 accent, etc.) so the SVG is visually consistent with the rest of the admin SPA. Four chart renderers: - renderBrierTrend({ series }) — sparkline w/ baseline reference at 0.25 (always-50% baseline) - renderDomainBars({ bars }) — horizontal accuracy bars per domain - renderAbandonedThreadsCard(threads) — D30/TD4 'revisit now' link per row, points at /admin/calibration/revisit/<takeId> - renderPatternStatementsCard(statements) — D29/TD3 clickable drill-down links per row, point at /admin/calibration/pattern/<i> XSS posture: all caller-controlled strings pass through escapeXml(). Numeric inputs are .toFixed()-coerced. Admin SPA renders via dangerouslySetInnerHTML inside a TrustedSVG wrapper component; endpoint is gated by requireAdmin middleware. /admin/api/calibration/profile — returns the active profile row as JSON. /admin/api/calibration/charts/:type — returns image/svg+xml markup for type ∈ {brier-trend, domain-bars, pattern-statements, abandoned-threads}. Cache-Control: private, max-age=60. brier-trend currently renders a single-point series from the active profile (the time-series view across calibration_profiles.generated_at history is a v0.37 follow-up once we have multiple snapshots). abandoned-threads pulls the top 5 abandoned rows via the same SQL the doctor check uses. CalibrationPage React component (admin/src/pages/Calibration.tsx): Fetches profile + 4 charts. Loading / error / cold-brain states all handled. Layout includes the audit annotations (partial-grade badge, voice-gate-fell-back-to-template badge) per the approved mockup. TrustedSVG wrapper isolates the dangerouslySetInnerHTML to the SVG surface only. App.tsx nav: added 'calibration' page route + sidebar nav item, hash routing extended to support #calibration. TD2 contrast bump: admin/src/index.css --text-muted: #555 → #777. Old value was contrast 4.0 on the #0a0a0f bg — below WCAG AA 4.5 for body text. New value is ~5.5, passes AA. Improvement is global across Dashboard, Agents, RequestLog, and the new Calibration tab — single-line CSS change with ~10x the impact. admin/dist/ rebuilt via `bun run build` (vite). 36 modules transformed. Tests: 19 cases in test/svg-renderer.test.ts. escapeXml (1): canonical entities. renderBrierTrend (6): empty state, polyline for 2+ points, clamp beyond yMax, design tokens inlined, XSS safety on date strings, text-anchor end on right label. renderDomainBars (4): empty state, label/accuracy/n rendering, out-of-range accuracy clamp, XSS safety on labels. renderAbandonedThreadsCard (4): empty state, row rendering with revisit link, claim truncation at 70 chars, custom revisitHref override. renderPatternStatementsCard (4): empty state, anchor count matches statement count, XSS safety, custom drillHref override. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * recall: calibration footer formatter for morning pulse (T16) Pure formatter that turns a CalibrationProfileRow + optional abandoned- threads list into the conversational block the morning pulse will surface: Calibration this quarter: Brier 0.18 (solid). Right on early-stage tactics, late on macro by 18 months. Over-confident on team execution; under-calibrated on regulatory risk. Threads you opened and never came back to: · AI search platform differentiation (17 months silent) · International expansion playbook (12 months silent) Cold-brain branch: returns empty string when no profile or < 5 resolved takes. Caller decides whether to render the block; cold-brain absence is the cleanest non-event. Brier trend note maps the absolute value to conversational copy: <= 0.10 → "(strong calibration)" <= 0.20 → "(solid)" <= 0.25 → "(near baseline)" > 0.25 → "(worse than always-50% baseline — review your high-conviction calls)" v0.36.0.0 ship state has only the current profile snapshot. The "was 0.22 90d ago — improving" comparison shape arrives when we accumulate generated_at history across multiple cycles. R3 regression posture: This module is the FORMATTER only. Wiring into `gbrain recall`'s text output is intentionally NOT in this commit — runRecall's surface stays unchanged. v0.37 wires it under --show-calibration (opt-in initially, default-on later). For now the formatter is callable from the admin tab + custom CLI scripts that want it. Architecture: buildRecallCalibrationFooter(opts) — pure. opts.profile required, opts.abandonedThreads optional, opts.threadColumnWidth defaults to 50. Caps at 4 patterns + 5 abandoned threads to keep the footer scannable. Truncates long abandoned-thread claim text to fit the column width with a trailing ellipsis. Tests: 14 cases. Cold-brain branch (3): null profile, < 5 resolved, zero resolved. Happy path (7): header + Brier + patterns, trend note ranges (4 brackets), null brier omits the Brier line but keeps header, caps at 4 patterns. Abandoned threads (4): omit section when none, emit when present, cap at 5, truncate long claim with column-width override. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: --undo-wave reversal command (T17 / D18 CDX-3) Implements the undo-wave reversal flow. Every new row written by the v0.36.0.0 calibration wave carries wave_version='v0.36.0.0' so a precise revert is possible without touching pre-wave data. CLI surface (replaces the v0.36.0.0 ship-state placeholder): gbrain calibration --undo-wave v0.36.0.0 [--dry-run] [--scrub-gstack] [--json] Reversal scope (4 steps): Step 1 — UNSET takes.resolved_* columns for takes auto-applied by this wave. Identifies wave-applied takes via take_grade_cache.applied=true + wave_version match. Cross-checks resolved_by='gbrain:grade_takes' to ensure we're not un-resolving a take a manual `gbrain takes resolve` override has since claimed. Manual resolutions persist; only auto-grade resolutions revert. Step 1b — Mark take_grade_cache rows applied=false post-undo so the audit trail shows they WERE applied but this wave was reverted. The CDX-11 confidence-drift check filters on applied=true and gets a cleaner sample post-undo. Step 2 — DELETE FROM calibration_profiles WHERE wave_version = ?. Step 3 — DELETE FROM take_nudge_log WHERE wave_version = ?. Step 4 — Optional gstack-learnings-prune via the binary, scoped to the GSTACK_LEARNING_NAMESPACE prefix. Opt-in via --scrub-gstack. Best-effort: binary-missing or failure logs a warning + suggests the manual command; the rest of the undo still succeeded. Dry-run posture: --dry-run computes the counts via SELECT COUNT(*) shapes without emitting any UPDATE or DELETE. Same UndoWaveResult shape returned so operator sees exactly what would be reverted before committing. --dry-run intentionally skips the gstack scrub (filesystem write) too; ship-state safety call. Idempotency: Re-running --undo-wave on a brain that's already reverted is a no-op. Each query filters on wave_version; no matching rows → zero counts. Architecture: undoWave(engine, opts) — async, returns UndoWaveResult. Pure data layer; no stderr writes, no process exits. CLI dispatch in src/commands/calibration.ts handles printing. v0.36.0.0 ship state runs steps 1-3 sequentially (no transaction). Partial reversal is recoverable via re-run since each step is idempotent on wave_version match. A future enhancement (v0.37+) can wrap in engine.transaction once that surface lands in BrainEngine. Tests: 8 cases in test/undo-wave.test.ts. Dry-run posture (1): counts emitted, NO UPDATE/DELETE SQL fired. Happy path (3): all 4 steps execute, resolved_by filter scopes UPDATE to wave-applied resolutions, custom resolvedByLabel honored. Empty wave (2): zero counts when no matching rows, idempotent re-run. Wave-version parameter threading (2): supplied version threads through all queries, different wave versions don't collide. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: A/B harness for think + ab-report (T18 / D19 CDX-18) Structural answer to CDX-18 (anti-bias rewrite may make advice worse). We don't have to guess whether calibration helps — we measure. Architecture: runAbTrial(input) — calls thinkRunner TWICE on the same question (baseline + --with-calibration), surfaces both answers to a preferenceResolver, persists the trial to think_ab_results. buildAbReport(engine, { days }) — aggregates the table over the last N days (default 30). Computes win counts, ties, neither, and a with_calibration_win_rate over DECISIVE trials only (excludes neither/tie). Flags calibration_net_negative when n >= 20 AND win rate < 45%. formatAbReport(report, days) — pretty-prints for stdout; emits the calibration_net_negative warning block when triggered. CLI: gbrain calibration ab-report [--days N] [--json] Reads the table, prints the breakdown. Replaces the v0.36.0.0 ship-state placeholder in src/commands/calibration.ts. gbrain think --ab "<question>" Wires into runAbTrial via the dispatch in src/commands/think.ts — follow-up commit. This commit lands the harness layer + schema + report surface; the --ab flag itself flips on in a one-line wiring commit when the runRecall path is ready. Schema (migration v72 / think_ab_results): source_id, wave_version, ran_at, question, baseline_answer, with_calibration_answer, preferred (CHECK in {baseline, with_calibration, neither, tie}), model_id, notes. CHECK constraint enforces preferred enum. Default wave_version 'v0.36.0.0' stamped so --undo-wave can scrub these too. Index on (source_id, ran_at DESC) supports the report's "last N days" query. schema.sql + pglite-schema.ts both updated for fresh-install parity. schema-embedded.ts regenerated via build:schema. calibration_net_negative threshold (D19): Triggers when: - decisive_trials (baseline + with_calibration) >= 20 - with_calibration_win_rate < 0.45 (NOT <= — exact 45% is OK) Small-sample guard (n < 20) prevents the warning from firing on early data with sampling noise. Confidence-flat threshold (no Wilson CI yet) keeps the math simple; v0.37+ adds CI bounds. Tests: 12 cases in test/think-ab.test.ts. runAbTrial (4): both runner calls fire, preferenceResolver receives both answers, INSERT row params shape, throws when thinkRunner missing. buildAbReport (5): zero trials, aggregation, net_negative trigger at n>=20 + win<45%, no trigger at n<20 (small-sample guard), no trigger at exact 45% boundary. formatAbReport (3): zero-state message, decisive-trials breakdown, net_negative warning block. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: pattern drill-down route + revisit-now CLI (TD3 / D29 + TD4 / D30) TD3 (D29) — clickable pattern drill-down endpoint: GET /admin/api/calibration/pattern/:id (requireAdmin) Returns the pattern statement at index `id` plus the top 25 resolved takes for the holder, sorted by weight desc. v0.36.0.0 ship-state approximation: surfaces broad provenance evidence (top resolved takes). v0.37+ stores per-pattern source_take_ids[] on a calibration_profile_patterns join table so the drill-down shows the EXACT takes that drove the pattern. Surfaces a `provenance_note` field in the response so the operator sees the v0.36.0.0-vs-v0.37 fidelity boundary inline. The admin SPA's renderPatternStatementsCard SVG already emits anchor tags pointing at /admin/calibration/pattern/<i> (T15 ship state). This route makes those anchors clickable — closes the trust loop that was the rationale for D29 ("pattern statements without their evidence are dressed-up LLM hallucinations"). TD4 (D30) — `gbrain takes revisit <slug>` editor-open action: Adds the `revisit` subcommand to gbrain takes. Opens $EDITOR (falling back to vi) on the source markdown file for the slug. Appends a `<!-- gbrain:revisit -->` cursor marker at the bottom of the page on first invocation so the editor opens with intent visible. Reads sync.repo_path from config to locate the brain repo. Refuses to proceed with a clear error when the repo isn't configured or the page doesn't exist. spawnSync with stdio:'inherit' so the editor takes the terminal. Exit status surfaced on failure. The SVG renderer's revisit-now anchor for each abandoned thread row emits /admin/calibration/revisit/<takeId>. A small route handler that resolves take_id → page_slug then dispatches `gbrain takes revisit` via spawn is a v0.37 follow-up — the CLI command exists now so developers can wire it directly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: DESIGN.md — formalize de facto design tokens (TD1) Promotes the admin SPA's de facto design tokens (landed v0.26.0) to a canonical DESIGN.md at the repo root. This is the calibration target for /plan-design-review and /design-review going forward — when a question is "does this UI fit the system?", the answer is here. Captures the system as it stands today: Voice (5 surfaces, all routed through gateVoice() with mode-specific rubrics): pattern_statement, nudge, forecast_blurb, dashboard_caption, morning_pulse. Friend-not-doctor; concrete data over abstract metrics; no preachy / clinical / corporate language. Color tokens: 10 CSS variables from admin/src/index.css inlined into the SVG renderer (src/core/calibration/svg-renderer.ts). Dark theme is the only theme — admin is an operator tool. WCAG contrast documented per token; TD2's #555 → #777 bump on --text-muted noted. Typography: Inter for UI, JetBrains Mono for numbers/slugs/data. Type scale (18 / 14 / 13 / 12 / 11) documented as de facto, not yet formalized. Spacing scale: 4 / 8 / 16 / 24 / 32px. Linear-app density. Layout: sidebar 200px, max content 720px (text) / 960px (tables). No 3-column feature grids, no icons in colored circles, no decorative blobs. Charts: server-rendered SVG via pure functions in src/core/calibration/svg-renderer.ts. XSS posture documented: server-side escapeXml on caller-controlled strings, numeric inputs .toFixed()-coerced, admin SPA renders via <TrustedSVG> wrapper. Interaction patterns: keyboard nav required (J/K/space/u/q on the propose-queue), loading/empty/error states ARE features. v0.37+ roadmap: type scale formalization, animation tokens, component library extraction. Light mode explicitly NOT planned. The doc is a living target, not a frozen spec. Major changes route through /plan-design-review per the existing review chain. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: synthetic corpus scaffold + privacy CI guard (T19 + T20) T19 — synthetic corpus scaffold for extract-takes prompt tuning. test/fixtures/calibration/extract-takes-corpus/ — 5 representative pages across 4 genres (essay, people, companies, meetings, decisions). v0.36.0.0 ships a SMALL representative corpus as proof of structure; the full 50-page training set + 10-page holdout gets generated by the operator via `gbrain calibration build-corpus` (v0.37 follow-up subcommand) or by hand with the privacy guard catching violations either way. Privacy contract per D13': every page is SYNTHETIC. None of the names/companies/funds/deals/events refer to anything real. Placeholder names per CLAUDE.md: alice-example, charlie-example, acme-example, widget-co, fund-a/b/c, acme-seed, widget-series-a, meetings/2026-04-03. test/fixtures/calibration/README.md spells out the privacy contract, generation flow, and what the corpus is (stable regression set for the extract-takes prompt) vs is not (real anything). T20 — privacy CI guard (CDX-14 mitigation). scripts/check-synthetic-corpus-privacy.sh greps the corpus for: 1. Explicit dollar amounts ($50M, $1.2B etc) — would suggest the page memorized a real round size. 2. Out-of-range year references (informational only for v0.36.0.0; deferred to a manual review checklist). 3. Pages that reference ZERO placeholder names — suggests the page might be referring to real entities. Essay-genre fixtures exempt (they're anonymized PG-style writing by design). Wired into `bun run verify` (CI gate) so contributors can't accidentally land a synthetic fixture that leaks real-world specificity. The intent is fail-fast on accidental leakage; the operator can update the allowlist if a generic dollar amount is intentional. Closes CDX-14: 'CC reads real brain pages locally, writes nothing still risks privacy if any generated synthetic fixture memorizes structure-specific facts. Placeholder names are not enough.' The corpus shipped here is intentionally small but covers the four core gbrain page genres (essay, people, companies, meetings/decisions). The v0.37 corpus-build subcommand will fan out to 50 with the operator spot-checking + the CI guard enforcing the privacy contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test: R1-R5 IRON RULE regression inventory (T21) Per /plan-eng-review D26 IRON RULE: regressions get added to the test suite as critical requirements, no AskUserQuestion needed. Pins five regressions identified during the v0.36.0.0 wave's coverage diagram: R1: think baseline UNCHANGED when --with-calibration absent. Covered structurally by test/think-with-calibration.test.ts plus assertion-pinned in this file (default user message: question first, then retrieval; system prompt: no anti-bias section). R2: contradictions probe output UNCHANGED when no calibration profile. Covered structurally by test/eval-contradictions-calibration-join.test.ts plus pinned here (null profile → null tag, byte-identical to v0.32.6). R3: takes resolution flow works when grade_takes phase disabled. Pinned import-surface coupling: takes-resolution.ts has zero dependency on grade_takes module. If a future refactor accidentally couples them, this test fails to compile. R4: search/list_pages/get_page work identically through new source_id paths. Marker test referencing existing v0.34.1 source-isolation suite at test/source-isolation-pglite.test.ts. v0.36.0.0 does NOT modify those code paths; the existing tests catch any accidental coupling. R5: existing search modes (conservative/balanced/tokenmax) unaffected. Marker test referencing existing test/search-mode.test.ts. The calibration code DOES NOT IMPORT from src/core/search/mode.ts. Plus an inventory test that confirms all 5 regressions have an 'addressed' status — fail-loud if a future contributor removes a guard without updating the inventory. 7 tests total. Pure functions, no engine, hermetic. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: v0.36.0.0 CHANGELOG + CLAUDE.md anchors + calibration convention skill CHANGELOG entry: the user-facing release notes. Leads with the headline ("the brain learns how you tend to be wrong, then argues against your blind spots on every advice call"), 5 'what you can now do' bullets in GStack voice, itemized changes by lane, and the 'To take advantage of v0.36.0.0' upgrade checklist per the CLAUDE.md required-block contract. CLAUDE.md anchors: new 'v0.36.0.0 Hindsight calibration wave (key files cluster)' block inserted before the v0.31.1 thin-client section. 23 new files / extensions annotated with one-paragraph descriptions each, linking back to the convention skill at skills/conventions/calibration.md for the agent-facing rules. skills/conventions/calibration.md: the agent-facing convention skill. Tells future contributors which calibration touchpoint applies to their task — voice gate? BaseCyclePhase? source-scope thread? doctor warning? cross-brain query rules? auto-resolve threshold posture? Test seam patterns. Bug class to avoid (the v0.34.1 source-isolation leak shape). Version trio (per CLAUDE.md mandatory audit): VERSION: 0.36.0.0 package.json: 0.36.0.0 CHANGELOG: ## [0.36.0.0] - 2026-05-17 llms.txt + llms-full.txt regenerated via `bun run build:llms` after the CLAUDE.md edit (per the explicit CLAUDE.md mandate "Any CLAUDE.md edit MUST be followed by `bun run build:llms`"). The `test/build-llms.test.ts` guard runs in CI shard 1; the committed bundles are checked against fresh generator output. bun run verify is clean. typecheck clean. Privacy CI guard passes (0 violations across 6 corpus pages). All ready for /ship. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * cycle: wire propose_takes / grade_takes / calibration_profile into runCycle (T-fix) The three new v0.36.0.0 phases were declared in CyclePhase / ALL_PHASES / NEEDS_LOCK_PHASES but the runCycle orchestrator never dispatched them. ALL_PHASES advertised them, gbrain dream --phase propose_takes accepted them, but `gbrain dream` (default) silently skipped all three. Adds a single dispatch block between consolidate and embed that: - builds an OperationContext on the fly (trusted-workspace caller, remote: false, sourceId resolved via the same helper sync uses) - dispatches the three phases in the order ALL_PHASES declares - records the same skipped-phase shape (no_database) when engine is null Pinned by test/core/cycle.serial.test.ts "default: all 6 phases run in order" which was already failing against ALL_PHASES (the test name lags the actual phase count; left as-is since renaming churns history). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: expand synthetic corpus + add hand-labeled ground-truth (T19) Adds 8 new synthetic pages modeled on the genre mix observed in the real brain (concepts-with-timeline, meeting-notes, daily-journal, people-pages, essays). Companion .gradeable-claims.json files carry hand-labeled answer keys — what a tuned propose_takes prompt SHOULD extract per page. Closes the F1 gate gap from the plan's T19/D19: Training corpus (test/fixtures/calibration/extract-takes-corpus/): + concept-startup-market-dynamics.md (10 claims) + meeting-2026-04-10-fundraise-fund-a.md (6 claims) + daily-2026-04-15.md (5 claims) Blind holdout (test/fixtures/calibration/holdout/): + concept-founder-execution.md (6 claims, F1 >= 0.80) + daily-2026-04-18.md (4 claims, F1 >= 0.80) + meeting-2026-04-17-hiring-charlie.md (5 claims, F1 >= 0.80) + essay-on-conviction.md (7 claims, F1 >= 0.80) + people-bob-example.md (5 claims, F1 >= 0.80) Privacy: - No real-brain content read into any committed artifact. Pages written from scratch using the canonical placeholder set (alice-example, charlie-example, bob-example, acme-example, widget-co, fund-a/b/c). Real-name grep confirms zero leakage: wintermute, garrytan, paul-graham, sam-altman, etc. → 0 hits. - scripts/check-synthetic-corpus-privacy.sh passes: 0 violations across 14 pages (was 6). Genre fidelity: - concept-with-timeline pages mirror the dated-assertion structure real brain uses (verb framing varies: "argues / predicts / I think / I bet / strong conviction / moderate conviction"). - meeting-notes pages carry both prose claims (extracted via hedging language) and explicit ## Takes sections. - daily-journal pages test probabilistic framing ("75/25 in favor", "call it ~0.5") and self-tagged conviction values. - essay-on-conviction is the meta-page that names the author's own bias patterns — primary signal for calibration_profile. - people pages test claim-about-third-party extraction. Each JSON ground-truth lists per-claim: - claim_text + kind (prediction|judgment|bet) + domain - conviction (0..1) - since_date - rationale (why this claim is gradeable + how a tuned prompt should infer conviction from the prose) This is the corpus that gates the T19 prompt-tune iteration: - F1 >= 0.85 on training (10+6+5 = 21 claims across 3 pages plus the existing 5 fixtures already shipped) - F1 >= 0.80 on holdout (27 claims across 5 pages) Plan reference: ~/.claude/plans/system-instruction-you-are-working-rippling-knuth.md Privacy gate: scripts/check-synthetic-corpus-privacy.sh (wired into bun run verify). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * calibration: tune propose_takes prompt against synthetic corpus (cat15 F1 0.92+) The v0.36.1.0 ship state shipped propose_takes with a stub prompt that the docs flagged as "tune via T19 corpus build before relying on propose_takes in production." T19's corpus was built in commit 69a71c9d (14 synthetic pages + 48 hand-labeled claims). The matching gbrain-evals cat15 runner validates extraction quality against that corpus. This commit back-ports the tuned prompt validated by cat15's first live run: training avg F1: 0.952 (target 0.85, +10 points) holdout avg F1: 0.922 (target 0.80, +12 points) train-holdout gap: 0.03 (well below 0.10 overfitting threshold) 8/8 probes pass their individual F1 targets Per-genre F1 floor: 0.80 (people-pages, the hardest genre). Concept- with-timeline and meeting-notes genres scored at 1.00 on holdout pages. The tuned prompt design changes vs the stub: - Worked example list seeds the "gradeable claim" notion so the model doesn't drift into pure-fact extraction. - NOT-gradeable list catches the most common over-extraction modes (pure facts, direct quotes, restatements). - Conviction inference rules anchored to specific hedging language so the model produces consistent weight values. - kind enum narrowed to 'prediction' | 'judgment' | 'bet' — the v1 stub's 4-tag enum bled into noise classification on the corpus. PROPOSE_TAKES_PROMPT_VERSION bumped 'v0.36.1.0-stub' → 'v0.36.1.0-tuned-cat15'. The bump invalidates the take_proposals idempotency cache so existing proposal rows stay as audit history but the next cycle re-extracts against the new prompt — exactly the design contract this version field is for. Re-tuning protocol: run cat15 in gbrain-evals against the fixtures BEFORE bumping the version string. The train-holdout gap should stay < 0.10. If a future tune drops below the cat15 gate, revert. Source of evidence: - cat15 runner: ~/git/gbrain-evals/eval/runner/cat15-propose-takes.ts - Fixture corpus: test/fixtures/calibration/ (this repo, commit 69a71c9d) - Live run dumps: ~/git/gbrain-evals/eval/reports/cat15-propose-takes/*.json Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: link cat14/cat15 benchmark report from CHANGELOG + README Adds the "Validated by published benchmarks" subsection to the v0.36.1.0 CHANGELOG entry and a "Calibration loop" section to the README's "Receipts on the evals" surface. Both link to the new benchmark report at gbrain-evals/docs/benchmarks/2026-05-18-brainbench-cat14-cat15-calibration.md. CHANGELOG: also updates the propose_takes bullet to reflect that the v0.36.1.0 ship state now includes the tuned 'v0.36.1.0-tuned-cat15' prompt (back-ported in 04dbab44), not the v1 stub the original entry described. README: adds a Calibration loop entry to the receipts table sitting between source-aware ranking and prompt compression. Frames the cat14 + cat15 numbers as "first published benchmark for AI memory systems that reason about user track records" — honest SOTA framing since Hindsight introduced the concept without quantified evaluation. llms.txt + llms-full.txt regenerated. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: fix benchmark-report links — gbrain-evals uses main not master 7 links to gbrain-evals/blob/master/docs/benchmarks/ were broken — the gbrain-evals repo uses 'main' as its default branch, not 'master'. Surfaced when I checked that the new cat14/cat15 link resolved post-PR-9 merge. Turned out 4 pre-existing links to longmemeval, brainbench-v0.20, brainbench-cat13b-source-swamp, and comparison-systems were all broken for the same reason — I just added a fifth by following the same wrong pattern. Sweep: gbrain-evals/blob/master/ → gbrain-evals/blob/main/ across both README.md (5 links) and CHANGELOG.md (2 links). llms.txt + llms-full.txt regenerated. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
289 KiB
CLAUDE.md
GBrain is a personal knowledge brain and GStack mod for agent platforms. Pluggable engines: PGLite (embedded Postgres via WASM, zero-config default) or Postgres + pgvector
- hybrid search in a managed Supabase instance.
gbrain initdefaults to PGLite; suggests Supabase for 1000+ files. GStack teaches agents how to code. GBrain teaches agents everything else: brain ops, signal detection, content ingestion, enrichment, cron scheduling, reports, identity, and access control.
Two organizational axes (read this first)
GBrain knowledge is organized along two orthogonal axes. Users AND agents must understand both, or queries misroute silently.
- Brain — WHICH DATABASE. Your personal brain is
host. You can mount additional brains (team-published, each with their own DB and access policy) viagbrain mounts add(v0.19+). Routing:--brain,GBRAIN_BRAIN_ID,.gbrain-mountdotfile. - Source — WHICH REPO INSIDE THE DATABASE. A brain can hold many sources
(wiki, gstack, openclaw, essays). Slugs scope per source. Routing:
--source,GBRAIN_SOURCE,.gbrain-sourcedotfile.
Both axes follow the same 6-tier resolution pattern. Read
docs/architecture/brains-and-sources.md for topology diagrams (personal, team
mount, CEO-class with multiple team brains) and
skills/conventions/brain-routing.md for the agent-facing decision table.
Architecture
Contract-first: src/core/operations.ts defines ~47 shared operations (v0.29 adds get_recent_salience, find_anomalies, get_recent_transcripts). CLI and MCP
server are both generated from this single source. Engine factory (src/core/engine-factory.ts)
dynamically imports the configured engine ('pglite' or 'postgres'). Skills are fat
markdown files (tool-agnostic, work with both CLI and plugin contexts).
Trust boundary: OperationContext.remote distinguishes trusted local CLI callers
(remote: false set by src/cli.ts) from untrusted agent-facing callers
(remote: true set by src/mcp/server.ts). Security-sensitive operations like
file_upload tighten filesystem confinement when remote=true and default to
strict behavior when unset.
Key files
src/core/operations.ts— Contract-first operation definitions (the foundation). Also exports upload validators:validateUploadPath,validatePageSlug,validateFilename, plusmatchesSlugAllowList(slug, prefixes)(v0.23 glob matcher:<prefix>/*matches recursive children; bare<prefix>matches exact only).OperationContext.remoteflags untrusted callers;OperationContext.allowedSlugPrefixes(v0.23) is the trusted-workspace allow-list set by the dream cycle.put_pageenforces: whenviaSubagentandallowedSlugPrefixesis set, slug must match the allow-list; else the legacywiki/agents/<id>/...namespace check applies. Auto-link enabled for trusted-workspace writes (skipped only whenremote=true && !trustedWorkspace). As of v0.26.0, everyOperationalso carriesscope?: 'read' | 'write' | 'admin'+localOnly?: boolean. All ops are annotated;sync_brain,file_upload,file_list, andfile_urlareadmin + localOnly(rejected over HTTP).OperationContext.auth?: AuthInfois threaded through HTTP dispatch for scope enforcement inserve-http.tsbefore the op runs. v0.26.9 (D12 + F7b):OperationContext.remoteis now a REQUIRED field in the TypeScript type — the compiler is the first defense against transports that forget to set it. Four trust-boundary call sites (put_pageallowlist, file_upload trust-narrowing, submit_job protected-name guard, auto-link skip) flipped from falsy-default (!ctx.remote) to fail-closed semantics (ctx.remote === falsefor "trusted-only" sites andctx.remote !== falsefor "untrust unless explicit-false"). Anything that isn't strictlyfalseis now treated as remote. Closed an HTTP MCP shell-job RCE: aread+write-scoped OAuth token could submitshelljobs because the HTTP request handler's literal context skippedremote: trueandsubmit_job's protected-name guard saw a falsy undefined. Stdio MCP set the field correctly via dispatch.ts; HTTP inlined a parallel context-builder for several releases and lost it. v0.34.1.0 (#861 + #876): new helpersourceScopeOpts(ctx)encodes the precedence ladder for source-scoped reads — federated array (ctx.auth.allowedSources) wins over scalar (ctx.sourceId/ctx.auth.sourceId) over nothing. Every read-side op handler routes through it so future ops can't silently drift from the canonical v0.31.8 thread. Closes the source-isolation leak on the read path: aread+write-scoped OAuth client bound to--source dept-xno longer sees rows from neighboring sources viasearch/query/list_pages/get_page/find_experts/query's image path.src/core/engine.ts— Pluggable engine interface (BrainEngine).clampSearchLimit(limit, default, cap)takes an explicit cap so per-operation caps can be tighter thanMAX_SEARCH_LIMIT. ExportsLinkBatchInput/TimelineBatchInputfor the v0.12.1 bulk-insert API (addLinksBatch/addTimelineEntriesBatch). As of v0.13.1,BrainEnginehas areadonly kind: 'postgres' | 'pglite'discriminator so migrations (src/core/migrate.ts) and other consumers can branch on engine withoutinstanceof+ dynamic imports. v0.29: four new methods —batchLoadEmotionalInputs(slugs?)(CTE-shaped read with per-table aggregates so a page × N tags × M takes never produces N×M rows),setEmotionalWeightBatch(rows)(UPDATE FROM unnest($1::text[], $2::text[], $3::real[])composite-keyed on(slug, source_id)for multi-source safety),getRecentSalience(opts),findAnomalies(opts).PageFiltersextended withsort?: 'updated_desc' | 'updated_asc' | 'created_desc' | 'slug'+PAGE_SORT_SQLwhitelist consumed by both engines (was hardcodedORDER BY updated_at DESC). v0.32.8 (PR #860): newlistAllPageRefs(): Promise<Array<{slug, source_id}>>ordered by(source_id, slug). Cheap cross-source enumeration for hot loops on large brains — replaces thegetAllSlugs()→getPage(slug)N+1 pattern in extract-takes, extract, integrity, which silently defaulted tosource_id='default'for non-default-source pages. Implementation parity across postgres-engine.ts + pglite-engine.ts. Pinned bytest/e2e/multi-source-bug-class.test.ts. v0.34.1.0 (#861):SearchOpts+PageFiltersaddsourceIds?: string[]for the federated read axis; both engines applyWHERE source_id = ANY($N::text[])when the array is set and preserve the scalarsourceIdfast path when unset.traverseGraph(slug, depth, opts?)andtraversePaths(slug, opts?)acceptopts.sourceId/opts.sourceIdsso graph walks respect the caller's scope. v0.35.6.0: two new methods supporting the phantom-redirect cycle pass —refreshPageBody(slug, sourceId, compiled_truth, timeline, content_hash)narrow-UPDATEs three columns + updated_at, skipping soft-deleted rows (codex #7: content_hash refresh is required sogbrain syncsees the canonical as unchanged after fence merge);migrateFactsToCanonical(phantomSlug, canonicalSlug, sourceId)UPDATEsentity_slug+source_markdown_slugon every active fact row keyed on the phantom, preserving embedding/validUntil/kind/status/source_session/confidence — codex #3 fix for the writeFactsToFence lossy-migration trap. Both methods have engine parity tests attest/phantom-redirect-engine-parity.test.ts.src/core/engine-factory.ts— Engine factory with dynamic imports ('pglite'|'postgres')src/core/pglite-engine.ts— PGLite (embedded Postgres 17.5 via WASM) implementation, all 40 BrainEngine methods.addLinksBatch/addTimelineEntriesBatchuse multi-rowunnest()with manual$Nplaceholders. As of v0.13.1,connect()wrapsPGlite.create()in a try/catch that emits an actionable error naming the macOS 26.3 WASM bug (#223) and pointing atgbrain doctor; the lock is released on failure so the next process can retry cleanly. v0.22.0:searchKeywordandsearchKeywordChunksmultiplyts_rankby the source-factor CASE expression at the chunk-grain level;searchVectorbecomes a two-stage CTE — inner CTE keepsORDER BY cc.embedding <=> vecso HNSW stays usable, outer SELECT re-ranks byraw_score * source_factor. Inner LIMIT scales with offset to preserve pagination contract. As of v0.22.6.1,initSchema()callsapplyForwardReferenceBootstrap()BEFORE replaying SCHEMA_SQL — probes for the specific forward-referenced state the embedded schema blob needs (pages.source_id,links.link_source,links.origin_page_id,content_chunks.symbol_name,content_chunks.language,sourcesFK target table) and adds only what's missing. Closes the upgrade-wedge bug class that bit users 10+ times across 6 schema versions over 2 years (#239/#243/#266/#357/#366/#374/#375/#378/#395/#396). No-op on fresh installs and modern brains. v0.35.5.0: probe set extended in parity with postgres-engine.ts —files.source_id,files.page_id,oauth_clients.source_id,oauth_clients.federated_read,sources.archived,sources.archived_at,sources.archive_expires_at. Bootstrap also threads the DDL connection frominitSchemaso probes run inside the advisory-lock scope. Closes #1018, #974, #820.src/core/pglite-schema.ts— PGLite-specific DDL (pgvector, pg_trgm, triggers)src/core/postgres-engine.ts— Postgres + pgvector implementation (Supabase / self-hosted).addLinksBatch/addTimelineEntriesBatchuseINSERT ... SELECT FROM unnest($1::text[], ...) JOIN pages ON CONFLICT DO NOTHING RETURNING 1— 4-5 array params regardless of batch size, sidesteps the 65535-parameter cap. As of v0.12.3,searchKeyword/searchVectorscopestatement_timeoutviasql.begin+SET LOCALso the GUC dies with the transaction instead of leaking across the pooled postgres.js connection (contributed by @garagon).getEmbeddingsByChunkIdsusestryParseEmbeddingso one corrupt row skips+warns instead of killing the query. v0.22.0:searchKeyword,searchKeywordChunks, andsearchVectorapply source-aware ranking by inlining the source-factor CASE andNOT (col LIKE …)hard-exclude clause fromsrc/core/search/sql-ranking.ts.searchVectorswitches to a two-stage CTE (HNSW-safe inner ORDER BY, source-boost re-rank in the outer SELECT) and carriesp.source_idthrough inner→outer for v0.18 multi-source callers. v0.22.1 (#406):_savedConfigretains the connect config;reconnect()tears down + recreates the pool from saved config (called by supervisor watchdog after 3 consecutive health-check failures).executeRawis a single-statement passthrough — no per-call retry (D3 dropped that as unsound for non-idempotent statements; recovery is supervisor-driven). v0.22.1 (#363, contributed by @orendi84):connect()appliesresolveSessionTimeouts()fromdb.tsas connection-time startup parameters (statement_timeout,idle_in_transaction_session_timeout) so orphan pgbouncer backends can't hold locks for hours. v0.22.1 (#409, contributed by @atrevino47):countStaleChunks()+listStaleChunks()server-side-filter onembedding IS NULLforembed --stale, eliminating ~76 MB/call client-side pull on a fully-embedded brain;upsertChunks()resets bothembeddingANDembedded_atto NULL when chunk_text changes without a new embedding (consistency). As of v0.22.6.1,initSchema()callsapplyForwardReferenceBootstrap()BEFORE replaying SCHEMA_SQL on the same forward-reference probe set as the PGLite engine, so old Postgres brains pinned at v0.13/v0.18/v0.19 walk forward cleanly instead of wedging oncolumn "..." does not exist. v0.35.5.0: probe set extended for the column-only forward-reference cases the original v0.22.6.1 sweep missed —files.source_id,files.page_id(pre-v0.18 brains whereidx_files_source_idwas the choke point),oauth_clients.source_id,oauth_clients.federated_read(pre-v0.34 brains where v60+v61+v65 chain failed), andsources.archived+sources.archived_at+sources.archive_expires_at(pre-v0.26.5 brains whereCREATE TABLE IF NOT EXISTS sourceswas a no-op on existing tables so the archive lifecycle columns never landed). Also (Codex P1 from pre-landing review): the entire probe path now runs on the DDL connection threaded down frominitSchema— previously probes ran through the instance pool while the advisory lock sat on a different connection, opening a concurrent-bootstrap race for Supabase pooler users. Closes #1018, #974, #820. v0.28.1:disconnect()is now idempotent. New_connectionStyleinstance field tracks whether the engine owns its pool (worker engines) or shares the module-level singleton; second call on an instance-pool engine is a no-op rather than falling through todb.disconnect()and clobbering the singleton. Pinned bytest/e2e/postgres-engine-disconnect-idempotency.test.ts(2 cases). Closes the bug class where any test sharing an engine across multipleworker.start()/worker.stop()cycles silently broke its own DB connectivity.src/core/cjk.ts(v0.32.7 CJK wave) — Single source of truth for CJK detection across the codebase. ExportsCJK_RANGES_REGEX,CJK_SLUG_CHARS(character-class fragment for embedding inside other regexes),CJK_SENTENCE_DELIMITERS(。!?),CJK_CLAUSE_DELIMITERS(;:,、),CJK_DENSITY_THRESHOLD = 0.30,hasCJK(s),countCJKAwareWords(s)(30% density threshold — English docs with one Japanese term stay whitespace-tokenized; Chinese-dominant docs get char-counted), andescapeLikePattern(s)(escapes%,_,\\forILIKE ... ESCAPE '\\'). Replaces the inline hasCJK regex previously duplicated atexpansion.ts:58. BMP-only ranges (Han / Hiragana / Katakana / Hangul Syllables); widening to Unicode property escapes is a v0.33+ TODO. Consumers:expansion.ts,sync.ts:slugifySegment,operations.ts:validatePageSlug + validateFilename,chunkers/recursive.ts:countWords + DELIMITERS,pglite-engine.ts:searchKeyword + searchKeywordChunks.src/core/audit-slug-fallback.ts(v0.32.7 CJK wave) — Weekly ISO-week-rotated audit JSONL at~/.gbrain/audit/slug-fallback-YYYY-Www.jsonl.logSlugFallback(slug, sourcePath)fires whenimportFromFilefalls back to a frontmatter slug becauseslugifyPathreturned empty (emoji / Thai / Arabic / non-CJK exotic-script filenames).readRecentSlugFallbacks(days)reads the last N days forgbrain doctor'sslug_fallback_auditcheck. HonorsGBRAIN_AUDIT_DIRvia the sharedresolveAuditDir()from shell-audit.ts. Separate surface fromsync-failures.jsonlper codex outside-voice review — that file carries bookmark-gating semantics that info events shouldn't trigger.src/core/embedding-pricing.ts(v0.32.7 CJK wave) —EMBEDDING_PRICINGmap keyedprovider:modelfor the post-upgrade reindex cost estimate. Sibling toanthropic-pricing.ts. Entries: OpenAI text-embedding-3-large ($0.13/1M), 3-small ($0.02/1M), ada-002 ($0.10/1M), Voyage 3-large ($0.18/1M), 3 ($0.06/1M).lookupEmbeddingPrice(modelString)returns a tagged union (knownwith price +unknownwith provider name);estimateCostFromChars(charCount, pricePerMTok)uses 3.5 chars/token approximation. Unknown providers degrade gracefully to "estimate unavailable" instead of fabricating numbers.src/core/post-upgrade-reembed.ts(v0.32.7 CJK wave) — Pure functions backing thegbrain upgradechunker-bump cost prompt.computeReembedEstimate(engine, model)queries real SQL (COUNT(*)+COALESCE(SUM(LENGTH(compiled_truth)) + SUM(LENGTH(timeline)), 0)) onpages WHERE chunker_version < MARKDOWN_CHUNKER_VERSION.formatReembedPrompt(est, graceSeconds)is the stderr-line formatter.runPostUpgradeReembedPrompt(engine, model, opts)orchestrates the 10-second Ctrl-C window; TTY-only wait (non-TTY auto-proceeds for CI / cron);GBRAIN_NO_REEMBED=1bails out with a doctor-warning marker;GBRAIN_REEMBED_GRACE_SECONDS=0skips the wait.src/commands/reindex.ts(v0.32.7 CJK wave) —gbrain reindex --markdown [--limit N] [--dry-run] [--json] [--no-embed] [--repo PATH]. Walkspages WHERE page_kind = 'markdown' AND chunker_version < MARKDOWN_CHUNKER_VERSIONin 100-row batches, ordered by id. Rows with non-nullsource_pathre-import viaimportFromFile; rows without fall back toimportFromContentagainst the storedcompiled_truth. Both paths passforceRechunk: trueto bypassimportFromContent'scontent_hashshort-circuit — without that flag (codex post-merge F1), the chunker version bump never reaches pages whose source content hasn't changed since last sync, AND master's v0.32.2 stripFactsFence privacy strip never applies to pre-strip chunks. Idempotent — partial-completion re-runs pick up where they left off via id-ordered batches. Wired intosrc/commands/upgrade.ts:runPostUpgradeafterapply-migrations.src/commands/sync.ts:resolveSlugByPathOrSourcePath(v0.32.7 CJK wave, codex post-merge F4) — Resolves a slug bypages.source_pathfirst (returns the stored slug for frontmatter-fallback pages whose path doesn't derive a slug), then falls back toresolveSlugForPath(path). Threaded into all 4 delete/rename call sites (performSync's un-syncable cleanup at ~:531, deletes at ~:603, rename oldSlug at ~:622). Without this, emoji-only / Thai / Arabic filenames whose slug came from frontmatter would orphan on delete/rename (the delete path would compute the wrong path-derived slug). Best-effort query — pre-migration brains fall through to the legacy path.src/core/utils.ts— Shared SQL utilities extracted from postgres-engine.ts. ExportsparseEmbedding(value)(throws on unknown input, used by migration + ingest paths where data integrity matters) and as of v0.12.3tryParseEmbedding(value)(returnsnull+ warns once per process, used by search/rescore paths where availability matters more than strictness). v0.26.9 (D14): addsisUndefinedColumnError(err)predicate — pattern-matches Postgres SQLSTATE 42703 / "column ... does not exist" with engine-driver shape variation tolerated. Replaces barecatch {}blocks inoauth-provider.tsso genuine errors (lock timeout, network blip, permission denied) propagate while column-missing falls through to the legacy fallback path. Reusable from any future code that needs the same column-existence probe semantics. v0.32.8 (PR #860): addsvalidateSourceId(id)that throws on anything outside^[a-z0-9_-]+$. Used by the per-source disk-layout fix in patterns.ts/synthesize.ts before anyjoin(brainDir, '.sources', source_id, slug+'.md')call so source_id can't traverse out of brainDir.rowToPageupdated to populate the now-requiredPage.source_idfield from the SELECT projection (scripts/check-source-id-projection.shenforces that every projection feedingrowToPageincludes the column).src/core/db.ts— Connection management, schema initialization. v0.22.1 (#363, contributed by @orendi84):resolveSessionTimeouts()returnsstatement_timeout+idle_in_transaction_session_timeout(defaults: 5min each, env-overridable viaGBRAIN_STATEMENT_TIMEOUT/GBRAIN_IDLE_TX_TIMEOUT/GBRAIN_CLIENT_CHECK_INTERVAL). Bothconnect()(module singleton) andPostgresEngine.connect()(worker pool) consume the result via postgres.js'sconnectionoption, sending GUCs as startup parameters that survive PgBouncer transaction mode (unlike the priorsetSessionDefaultspost-pool SET, kept as a back-compat no-op shim).src/commands/migrate-engine.ts— Bidirectional engine migration (gbrain migrate --to supabase/pglite)src/core/import-file.ts— importFromFile + importFromContent (chunk + embed + tags)src/core/sync.ts— Pure sync functions (manifest parsing, filtering, slug conversion). v0.35.5.0: new exportedpruneDir(name: string): booleanhelper is the single source of truth for descent-time directory exclusion across walkers. Blocksnode_modules(no leading dot, so pre-v0.35.5 walkers slipped through and inflated MISSING_OPEN counts via vendor packages), dot-prefix dirs,ops/, and*.rawsidecars.isSyncablenow applies it per path segment;walkMarkdownFilesinsrc/commands/extract.tsandlistTextFilesinsrc/core/cycle/transcript-discovery.tsconsult it BEFORE recursing so the IO cost of walking thousands of vendor files is saved. Closes #923 + #202.manageGitignoreworktree fix in same wave: discriminator now matches the gitdir path segment (/modules/<name>= submodule,/worktrees/<name>= worktree, per Git's documented layout) instead of the legacy absolute-vs-relative check that misclassified absorbed submodules and worktrees both. Conductor worktrees are first-class repos and now get.gitignoremanagement for storage-tiering. Closes #889. v0.22.12 (#500, foundation by @wintermute via #501):classifyErrorCode(errorMsg)regex-based classifier with 12 codes (SLUG_MISMATCH,YAML_PARSE,YAML_DUPLICATE_KEY,MISSING_OPEN,MISSING_CLOSE,NESTED_QUOTES,EMPTY_FRONTMATTER,NULL_BYTES,INVALID_UTF8,STATEMENT_TIMEOUT,FILE_TOO_LARGE,SYMLINK_NOT_ALLOWED) plusUNKNOWNfallback.summarizeFailuresByCode(failures)returns sorted[{code, count}].code?optional field onSyncFailure; backfilled at ack time on pre-v0.22.12 entries.acknowledgeSyncFailures()returnsAcknowledgeResult { count, summary }. Three regexes (MISSING_OPEN,MISSING_CLOSE,EMPTY_FRONTMATTER) broadened to match actualmarkdown.ts:159-244validator message strings, not just the literal code-name prefix.FILE_TOO_LARGEcovers all three production size sites inimport-file.ts:199, 352, 401;SYMLINK_NOT_ALLOWEDcovers the rejection at:347. Closes the silent-skip pattern that motivated #500.src/core/storage.ts— Pluggable storage interface (S3, Supabase Storage, local)src/core/storage-config.ts(v0.22.11) — Storage tiering:loadStorageConfigreadsgbrain.yml, normalizes deprecated keys (git_tracked/supabase_only) to canonical (db_tracked/db_only) with once-per-process deprecation warning, and runsnormalizeAndValidateStorageConfig(auto-fixes missing trailing/, throwsStorageConfigErroron tier overlap). Path-segment matcher:media/x/does NOT matchmedia/xerox/foo. Replaces gray-matter (broken on delimiter-less YAML) with a dedicated parser for thegbrain.ymlshape.src/core/disk-walk.ts(v0.22.11) —walkBrainRepo(repoPath)returnsMap<slug, {size, mtimeMs}>from one recursivereaddirSync. Skips dot-dirs,node_modules, non-.mdfiles. Used bygbrain storage statusto replace per-pageexistsSync + statSync(~400K syscalls on 200K-page brains → tens).src/core/git-remote.ts(v0.35.3.0) — SSRF-hardened git invocations for remote-sourcecloneRepoandpullRepo. Exports two distinct flag constants becausegit's argv grammar treats them differently:GIT_SSRF_FLAGS(3-cconfig flags —protocol.allow=user,protocol.file.allow=never,http.allowRedirects=false) is global config, spread BEFORE the subcommand verb. NewGIT_SSRF_SUBCOMMAND_FLAGS = ['--no-recurse-submodules']is subcommand-scoped, spread AFTER the verb. Pre-v0.35.3 a single combinedGIT_SSRF_FLAGSarray spread--no-recurse-submodulesbefore the verb where real git rejects it with exit 129 ("unknown option"); the fake-git test harness exited 0 regardless of argv shape, so CI missed it for ~7 months and every remote-source clone/pull was silently broken.cloneRepoargv:git <GIT_SSRF_FLAGS> clone <GIT_SSRF_SUBCOMMAND_FLAGS> --depth=1 [--branch X] -- <url> <dir>.pullRepoargv:git <GIT_SSRF_FLAGS> -C <dir> pull <GIT_SSRF_SUBCOMMAND_FLAGS> --ff-only. Pinned bytest/git-remote.test.tsposition-anchored regression guard (argv.indexOf('--no-recurse-submodules') > argv.indexOf(verb)).src/commands/storage.ts(v0.22.11) —gbrain storage status [--repo P] [--json]. Split into pure data (getStorageStatus) + JSON formatter + human formatter (ASCII-only per D10) matching theorphans.tspattern.PageCountsByTierandDiskUsageByTierare distinct nominal types so swaps fail at compile time.gbrain.yml(brain repo root, v0.22.11) — Optional storage tiering config. Top-levelstorage:section withdb_tracked:anddb_only:array-valued keys.gbrain syncauto-manages.gitignorefordb_onlypaths on successful sync (skips on dry-run, blocked-by-failures, submodule context, orGBRAIN_NO_GITIGNORE=1).gbrain export --restore-only [--repo P] [--type T] [--slug-prefix S]repopulates missingdb_onlyfiles from the database.src/core/supabase-admin.ts— Supabase admin API (project discovery, pgvector check)src/core/file-resolver.ts— File resolution with fallback chain (local -> .redirect.yaml -> .redirect -> .supabase)src/core/chunkers/— 3-tier chunking (recursive, semantic, LLM-guided). v0.19.0 addscode.ts— tree-sitter-based semantic chunker for 29 languages with embedded-asset WASMs (src/assets/wasm/),@dqbd/tiktokencl100k_base tokenizer, small-sibling merging.CHUNKER_VERSIONconstant folded intoimportCodeFile'scontent_hashso chunker shape changes force clean re-chunks across releases.src/core/errors.ts(v0.19.0) —StructuredAgentError+buildError+serializeError. Every new v0.19.0 agent-facing surface (code-def, code-refs, usage errors) uses this envelope; matches v0.17.0CycleReport.PhaseResult.errorshape.src/assets/wasm/(v0.19.0) — 36 tree-sitter grammar WASMs + tree-sitter runtime. Committed to the repo sobun --compileembeds them deterministically viaimport path from ... with { type: 'file' }. The CI guardscripts/check-wasm-embedded.shfails the build if the compiled binary ever silently falls through to recursive chunks.src/commands/code-def.ts+src/commands/code-refs.ts(v0.19.0) — symbol definition + references lookup. Querycontent_chunks.symbol_nameor chunk_text ILIKE withpage_kind='code'filter. Auto-JSON when stdout is not a TTY (gh-CLI convention). Bypass the standardsearchKeywordDISTINCT ON (slug)collapse so multiple call-sites from the same file surface.src/core/search/— Hybrid search: vector + keyword + RRF + multi-query expansion + dedup. As of v0.22.0,searchKeyword/searchKeywordChunks/searchVectorapply source-aware ranking at the SQL layer (curated content likeoriginals/,concepts/,writing/outranks bulk content likewintermute/chat/,daily/,media/x/).searchVectoruses a two-stage CTE so source-boost re-ranking doesn't kill the HNSW index. Hard-exclude prefixes (test/,archive/,attachments/,.raw/by default) filter at retrieval, not post-rank. Both gates honordetail !== 'high'so temporal queries surface chat pages normally.src/core/search/intent.ts— Query intent classifier (entity/temporal/event/general → auto-selects detail level)src/core/search/eval.ts— Retrieval eval harness: P@k, R@k, MRR, nDCG@k metrics + runEval() orchestratorsrc/core/search/source-boost.ts(v0.22.0) — Source-type boost map keyed by slug prefix.DEFAULT_SOURCE_BOOSTS(originals/ 1.5, concepts/ 1.3, writing/ 1.4, people/companies/deals/ 1.2, daily/ 0.8, media/x/ 0.7, wintermute/chat/ 0.5) andDEFAULT_HARD_EXCLUDES(test/, archive/, attachments/, .raw/).parseSourceBoostEnv/parseHardExcludesEnvparse comma-separatedprefix:factorpairs fromGBRAIN_SOURCE_BOOST/GBRAIN_SEARCH_EXCLUDEenv vars.resolveBoostMapandresolveHardExcludesmerge defaults + env + callerSearchOpts.exclude_slug_prefixes/include_slug_prefixes.src/core/search/sql-ranking.ts(v0.22.0) — Pure SQL string builders.buildSourceFactorCase(slugColumn, boostMap, detail)emits a CASE expression with longest-prefix-match wins (returns literal'1.0'whendetail === 'high'for temporal-bypass parity with COMPILED_TRUTH_BOOST).buildHardExcludeClause(slugColumn, prefixes)emitsNOT (col LIKE 'p1%' OR col LIKE 'p2%')— OR-chain wrapped in NOT, NOTNOT LIKE ALL/ANY(those quantifiers don't express set-exclusion). LIKE meta-character escape covers all three of%,_, AND\(backslash matters because it's Postgres LIKE's default escape char). Single-quote doubling on SQL string literals so injection-style inputs are inert text.src/commands/eval.ts—gbrain evalcommand: single-run table + A/B config comparison. v0.25.0 adds sub-subcommand dispatch onargs[0]sogbrain eval export+gbrain eval prune+gbrain eval replayroute into session-capture handlers; baregbrain eval --qrels …fall-through preserves the legacy IR-metrics flow. v0.27.x addsgbrain eval cross-modalto the dispatch (the user-facing path is the cli.ts no-DB branch —src/commands/eval.ts:cross-modalonly fires when callers re-enter with an existing engine).src/commands/eval-cross-modal.ts(v0.27.x) — multi-model quality gate. Three different-provider frontier models score the OUTPUT against the TASK on a 5-dim list. Verdictpass(exit 0) /fail(exit 1) /inconclusive(exit 2; <2/3 model successes per Q3=A in plans/radiant-napping-lerdorf.md). Reusessrc/core/ai/gateway.ts:chat()so config/auth/aliasing comes from the gateway recipe registry — no parallel provider stack. Self-configures the gateway (configureGateway(loadConfig() + process.env)) since the cli.ts dispatch bypassesconnectEngine(). Default cycles 3 in TTY, 1 in non-TTY (T11=B partial cost guardrail). Receipts land atgbrainPath('eval-receipts')/<slug>-<sha8-of-output>.json. The full--budget-usdcap is a v0.27.x follow-up TODO.src/core/cross-modal-eval/json-repair.ts(v0.27.x) —parseModelJSON(raw)named export with a 4-strategy fallback chain (direct parse → fence-strip → trailing-comma + single-quote + embedded-newline repair → regex nuclear option). Adversarial input throws rather than fabricating scores — the aggregator treats a throw as "this model contributed nothing this cycle" so the gate stays correct at >=2/3 successes.src/core/cross-modal-eval/aggregate.ts(v0.27.x) — pure verdict logic. Pass criterion:(successes >= 2) AND (every dim mean >= 7) AND (every dim min across models >= 5)(Q2=A floor). Inconclusive when <2/3 models returned parseable scores (Q3=A regression guard for the v1 .mjsObject.values({}).every(...) === trueempty-array PASS bug).src/core/cross-modal-eval/runner.ts(v0.27.x) — orchestrator. Each cycle runsPromise.allSettled([gwChat(slotA), gwChat(slotB), gwChat(slotC)])(T4=A — bare allSettled, no rate-leases for the CLI path; minion-integration TODO recovers cross-process concurrency). Stops early on PASS or INCONCLUSIVE; runs up to 3 cycles. Default slots:openai:gpt-4o/anthropic:claude-opus-4-7/google:gemini-1.5-pro.estimateCost()exports a small per-model pricing table (drifts; refresh alongside model-family bumps).src/core/cross-modal-eval/receipt-name.ts(v0.27.x) — receipt filename binds (slug, SKILL.md sha-8).findReceiptForSkill(skillPath, receiptDir)returns'found' | 'stale' | 'missing'(T10=A). Skillify-check item 11 surfaces the status as informational (T7=C); the audit does NOT fail on missing/stale receipts.src/core/cross-modal-eval/receipt-write.ts(v0.27.x) — wrapsfs.writeFileSyncwithmkdirSync({recursive:true})ahead of every write (T5 correction;gbrainPath()does NOT auto-mkdir).src/commands/eval-export.ts(v0.25.0) — streamseval_candidatesrows as NDJSON to stdout withschema_version: 1prefix on every line. EPIPE-safe, progress heartbeats on stderr, stable id-desc tiebreaker so--sincewindows never dupe/miss rows.src/commands/eval-prune.ts(v0.25.0) — explicit retention cleanup. Requires--older-than DUR.--dry-runreports would-delete count.src/commands/eval-replay.ts(v0.25.0) — contributor-facing replay tool. Reads NDJSON fromgbrain eval export, re-runs each capturedquery/searchop against the current brain, computes set-Jaccard@k between captured + currentretrieved_slugs, top-1 stability rate, and latency Δ. Stable JSON shape (schema_version: 1) for CI gating; human mode prints a regression table. Pure Bun, zero new deps. The dev-loop half of BrainBench-Real that closes the gap between "data captured" and "data used to gate a PR." Seedocs/eval-bench.mdfor the workflow.src/commands/eval-trajectory.ts+src/commands/founder-scorecard.ts+src/core/trajectory.ts(v0.35.7) — temporal trajectory + founder scorecard. The wave that turns the v0.35.3.1 date-aware contradiction probe into a useful temporal substrate.gbrain eval trajectory <entity>shows the chronological typed-claim history (mrr/arr/team_size/etc) with regressions auto-flagged inline;gbrain founder scorecard <entity>rolls up claim_accuracy / consistency / growth_trajectory / red_flags into one JSON. Pure-function math lives intrajectory.ts:detectRegressions(points, threshold)walks consecutive metric-value pairs per metric (10% drop default, env overrideGBRAIN_TRAJECTORY_REGRESSION_THRESHOLD);computeDriftScore(points)returns1 - mean(cosine(emb[i], emb[i-1]))over existing embeddings (null when <3 embedded points). Backed byBrainEngine.findTrajectory(opts)— both Postgres and PGLite, single SQL query, deterministicORDER BY valid_from ASC, id ASC(R3). Source-scoped via the v0.34.1.0sourceIdscalar /sourceIdsarray dual pattern (D-CDX-6); visibility-filtered for remote callers (D-CDX-1) —recall-equivalent posture. MCP opfind_trajectory(read scope, NOT localOnly) registered afterfind_experts. Migration v67 adds four optional typed-claim columns (claim_metric,claim_value,claim_unit,claim_period) + a partial index on(entity_slug, claim_metric, valid_from) WHERE claim_metric IS NOT NULL. Fence widens from 10 to 14 cells when any row has typed data; renderer stays at 10 cells when none do (no churn diff on existing fences). Metric labels normalize to lowercase snake_case vianormalizeMetricLabel(15-entry seed map for common founder metrics). Theconsolidatecycle phase gains semantic upsert keyed on(page_id, claim, since_date)— fixes the pre-existing F4 duplicate-takes bug where re-running the full cycle afterextract_factsclearedconsolidated_atwould silently append duplicate takes viaMAX(row_num)+1. Also writes chronologicalvalid_untilon each cluster's older facts. Theextract_factscycle phase batch-embeds viagateway.embed()before insert AND threadspages.effective_dateas thepageEffectiveDatefallback forvalid_from(precedence chain: fence-row > pageEffectiveDate > now()). The contradiction probe MUST NOT writevalid_until— R1+R8 grep guard attest/eval-contradictions/no-valid-until-write.test.tspins this. Codex outside-voice round caught F1 (v66 collision → v67), F2 (Haiku lives infacts/extract.tsnotextract-facts.tscycle phase), F3 (cycle didn't embed before insert), F4 (idempotency bug), F5+F6 (missedfence-write.tscaller + no Page object there → pageEffectiveDate is OPTIONAL), F7 (privacy regression — visibility filter added), F8 (ParsedFact needed typed-field extension for markdown system-of-record), F9 (dual scalar+federated sourceId). Plan:~/.claude/plans/system-instruction-you-are-working-curious-jellyfish.md. Tests: 258 across 12 files.src/commands/eval-suspected-contradictions.ts+src/core/eval-contradictions/{judge,runner,types,date-filter,cost-tracker,cache,severity-classify,cross-source,trends,calibration,judge-errors,auto-supersession,fixture-redact}.ts(v0.32.6) —gbrain eval suspected-contradictions [run|trend|review]. Probe samples top-K retrieval pairs per query (cross-slug + intra-page chunk-vs-take), date pre-filters (3-rule layered — same-paragraph-dual-date overrides separation rule), LLM judge (query-conditioned per Codex; UTF-8-safe truncation; C1 confidence-floor double-enforcement; resolution_kind output drives M7 paste-ready commands), persistent cache keyed on(chunk_a_hash, chunk_b_hash, model_id, prompt_version, truncation_policy)(Codex outside-voice fix — prompt edits cleanly invalidate prior verdicts), Wilson 95% CI calibration on the headline percentage withsmall_sample_notewhen n<30, judge_errors as first-class typed counters (parse_fail/refusal/timeout/http_5xx/unknown — Codex fix to bias from silent skip), M5 trend writes toeval_contradictions_runs, M6 source-tier breakdown reusesDEFAULT_SOURCE_BOOSTSprefix logic, deterministic sampling (combined_score DESC + lex tiebreaker — stable cache hit-rate across re-runs). Hermetic viajudgeFn+searchFnDI in the runner; never touches the real gateway in tests. Engine surface:BrainEngine.listActiveTakesForPages(P1 batched),writeContradictionsRun+loadContradictionsTrend(M5),getContradictionCacheEntry+putContradictionCacheEntry+sweepContradictionCache(P2). Schema migrations v51 + v52. MCP opfind_contradictions(read scope, NOT localOnly, NOT in subagent allowlist — user-initiated only). M1 doctor check surfaces high-severity findings with paste-ready resolution commands. M2 synthesize phase pre-fetches latest probe's top-5-by-severity findings and threads them intobuildSynthesisPromptas an informational block. 226 hermetic unit tests + 12 real-Postgres E2E. Plan:~/.claude/plans/system-instruction-you-are-working-hashed-dewdrop.md. Architecture doc:docs/contradictions.md.src/core/think/index.ts(v0.35.5.0 — gateway adapter) —runThinkno longer instantiatesnew Anthropic()directly. The internalLLMClientinstance is now built by a small adapter that wrapsgateway.chat()fromsrc/core/ai/gateway.ts, the canonical AI seam v0.31.12 established for chat/embed/expansion. Closes #952: stdio MCP launches (Claude Desktop, Cursor) don't inherit shell env, so the Anthropic SDK's env-only key resolution lost the key any user had set viagbrain config set anthropic_api_key. The gateway reads from~/.gbrain/config.jsonAND from env, so both paths work. Test seam preserved:opts.client?: ThinkLLMClientinjection still works for the 12+ existing tests (test/think-pipeline.serial.test.ts,test/think-gateway-adapter.test.ts, etc.);opts.stubResponsecontinues to short-circuit before any LLM call. When neither key nor client is available, the graceful "no LLM available" stub still fires with the sameNO_ANTHROPIC_API_KEYwarning. v0.36.x TODO: dropThinkLLMClientindirection entirely, migrate tests to__setChatTransportForTestsseam fromsrc/core/ai/gateway.ts.src/core/operations.tsextension (v0.35.5.0 orphans fix) —findOrphanPages(both engines) now filtersp.deleted_at IS NULLon the candidate side AND addsJOIN pages src ON src.id = l.from_page_id WHERE src.deleted_at IS NULLto the EXISTS subquery on the link-source side. Pre-v0.35.5 the query filtered nothing ondeleted_at, so soft-deleted pages (v0.26.5 soft-delete shipped without updating this query) appeared as orphans AND links from soft-deleted source pages still suppressed live pages from orphan results. Closes #1021. Pinned bytest/orphans.test.ts's soft-delete cases.src/commands/eval-longmemeval.ts+src/eval/longmemeval/{harness,adapter,sanitize}.ts(v0.28.1) —gbrain eval longmemeval <dataset.jsonl>runs the public LongMemEval benchmark against gbrain's hybrid retrieval. Architecture: one in-memory PGLite per benchmark run created viacreateBenchmarkBrain+withBenchmarkBrain(NOEphemeralBrainclass). Between questions,TRUNCATEover runtime-enumeratedpg_tablesso future schema migrations don't silently leak data across questions; infrastructure tables (sources,config,gbrain_cycle_locks,subagent_rate_leases) are preserved.cli.tshas a pre-dispatch bypass soeval longmemevalskipsconnectEngine()— the user's~/.gbrainbrain is never opened.--expansiondefaults to OFF (deterministic, no per-query Haiku call); pass--expansionto opt in. Default model resolves throughresolveModel()6-tier chain withmodels.eval.longmemevalas the new config key. Sanitization parity:harness.tsre-usesINJECTION_PATTERNSfromsrc/core/think/sanitize.ts(now exported, line 22) so adding a pattern automatically covers takes AND benchmarks. Retrieved chat content is wrapped in<chat_session id="..." date="...">framing; the answer-gen system prompt declares the content UNTRUSTED. LLM injection seam:runEvalLongMemEval(args, {client?: ThinkLLMClient})lets tests stub the client so the full pipeline runs without an Anthropic API key. p50 25.9ms / p99 30.3ms warm reset+import+search on Apple Silicon (pertest/eval-longmemeval.test.tsperf gate). Hand the JSONL output to LongMemEval'sevaluate_qa.pyto score (their published evaluator, not bundled — needs OpenAI gpt-4o per their spec).docs/eval-bench.md(v0.25.0) — contributor guide for using captured data to benchmark retrieval changes before merging. Linked from CONTRIBUTING.md under "Running real-world eval benchmarks (touching retrieval code)".src/core/eval-capture.ts(v0.25.0) — op-layer capture wrapper called fromsrc/core/operations.tsquery+searchhandlers. Catches MCP + CLI + subagent tool-bridge from one site. Fire-and-forget; failures route toengine.logEvalCaptureFailuresogbrain doctorsees drops cross-process. Capture is off by default —isEvalCaptureEnabledresolution: explicitconfig.eval.capture(true/false) wins, elseprocess.env.GBRAIN_CONTRIBUTOR_MODE === '1', else off. Production users get a quiet brain; contributors setexport GBRAIN_CONTRIBUTOR_MODE=1in.zshrcto enable the dev loop. PII scrubber gate is independent and defaults to true regardless of CONTRIBUTOR_MODE.src/core/eval-capture-scrub.ts(v0.25.0) — zero-deps PII scrubber: emails, phones, SSN, Luhn-verified credit cards, JWT-shaped tokens, bearer tokens.src/core/search/hybrid.ts— Cathedral IIPromise<SearchResult[]>return shape unchanged in v0.25.0. AddsonMeta?: (m: HybridSearchMeta) => voidcallback so op-layer capture can record what hybridSearch actually did. Existing callers leave it undefined. v0.33:HybridSearchOpts.types?: PageType[](defined onSearchOpts) threads a multi-type filter through to per-enginesearchKeyword+searchVector+searchKeywordChunks, where it lands asAND p.type = ANY($N::text[]). Primary consumer isgbrain whoknows(filters to['person','company']). AND-applies alongside the existing single-valuetypefilter; either or both can be used.docs/eval-capture.md(v0.25.0) — stable NDJSON schema reference for gbrain-evals consumers.test/public-exports.test.ts(v0.25.0 / R2) — runtime contract test. Imports each of the 17 public subpaths via package name and pins a canary symbol per module. Paired withscripts/check-exports-count.sh.src/core/embedding.ts— OpenAI text-embedding-3-large, batch, retry, backoff. v0.28.7:BATCH_SIZEreverted 50→100 — the original Voyage safety guard halved OpenAI throughput on every page. Per-recipe pre-split + recursive halving + adaptive shrink-on-miss now live in the gateway, so the outer paginator goes back to its original purpose: progress-callback granularity, not batch protection.src/core/ai/dims.ts(v0.33.1.1, PR #962 + #866) — per-providerproviderOptionsresolver for embed-time dimension passthrough. The single source of truth for "which provider needs which knob to producevector(N)". ExportsdimsProviderOptions(implementation, modelId, dims)(called byembed()ingateway.ts),VOYAGE_OUTPUT_DIMENSION_MODELS(private const — the 7 hosted Voyage models that acceptoutput_dimension:voyage-4-large,voyage-4,voyage-4-lite,voyage-3-large,voyage-3.5,voyage-3.5-lite,voyage-code-3— nano deliberately excluded),VOYAGE_VALID_OUTPUT_DIMS = [256, 512, 1024, 2048] as const,supportsVoyageOutputDimension(modelId), andisValidVoyageOutputDim(dims). Voyage path uses the SDK-supporteddimensionsfield ({ openaiCompatible: { dimensions: N } }), NOT Voyage'soutput_dimensionwire-key — the existingvoyageCompatFetchshim ingateway.ts:541translatesdimensions → output_dimensionbefore the HTTP body is built. The reverse (sendingoutput_dimensionfrom here) was the v0.33.1.0 bug class: the AI SDK's openai-compatible adapter doesn't recognize the wire-key so it was silently dropped, Voyage returned its default 1024-dim, and the gateway dimension check threw on every embed call. Runtime guard: when a Voyage flexible-dim model is configured withdimsoutsideVOYAGE_VALID_OUTPUT_DIMS, throwsAIConfigErrorwith a paste-readygbrain config set embedding_dimensions <256|512|1024|2048>hint at the embed boundary — fail-loud instead of opaque Voyage HTTP 400. Most common trigger:embedding_model: voyage:voyage-4-largeconfigured withoutembedding_dimensions(falls back toDEFAULT_EMBEDDING_DIMENSIONS=1536, an OpenAI default not a Voyage one). Eva (@100yenadmin) shipped the wire-key fix in #866; Codex P3 follow-up landed the validator + valid-dims allowlist in #962.src/core/ai/types.ts— provider/recipe types. v0.28.7 (#680):EmbeddingTouchpointextended with optionalchars_per_token(default 4 chars/token, matching OpenAI tiktoken on English) andsafety_factor(default 0.8, budget-utilization ceiling). Both consulted only whenmax_batch_tokensis also set. Voyage declareschars_per_token=1+safety_factor=0.5to handle dense payloads (CJK/JSON/base64) that overshoot tiktoken. The pre-split budget ismax_batch_tokens × safety_factor / chars_per_token. v0.28.11 (#719):EmbeddingTouchpoint.multimodal_models?: string[]model-level allow-list for recipes that mix text-only + multimodal models under one touchpoint (Voyage's 12 models sharesupports_multimodal: truebut onlyvoyage-multimodal-3accepts/multimodalembeddings). When omitted, recipe-levelsupports_multimodalis sufficient.AIGatewayConfig.embedding_multimodal_model?: stringletsembedMultimodal()route to a different model thanembedding_model— brains using OpenAI for text can use Voyage for images without flipping the primary embedding pipeline.src/core/ai/gateway.ts— unified seam for every AI call. v0.35.0.0: ZeroEntropy support lands. NewzeroEntropyCompatFetchshim (sibling tovoyageCompatFetch) handles ZE's non-OpenAI-compatible wire shape — rewrites the request URL from/embeddingsto/models/embed, injectsinput_type(default'document';'query'when threaded viaproviderOptions.openaiCompatible.input_type) and explicitencoding_format: 'float', and rewrites the response from{results: [{embedding}], usage: {total_bytes, total_tokens}}to{data: [{embedding, index}], usage: {prompt_tokens, total_tokens}}so the SDK's openai-compatible Zod schema validates (Voyage's shim hit the sameprompt_tokensrequirement at:655). Layer 1 (Content-Length) + Layer 2 (per-embedding) OOM caps via a new taggedZeroEntropyResponseTooLargeErrorclass (kept separate fromVoyageResponseTooLargeErrorbecausetest/voyage-response-cap.test.tsdoes structural source-text greps pinning the Voyage name — class unification is a deferred cleanup). Wired ininstantiateEmbedding()via the samerecipe.id === 'zeroentropyai'branch pattern Voyage uses. Newgateway.rerank()native HTTP path (no AI-SDK reranking abstraction): resolves the configured reranker model viagetRerankerModel(), posts to${recipe.base_url}/models/rerankwith bearer auth, returnsRerankResult[]sorted by relevance score.RerankError.reasonclassifier:auth | rate_limit | network | timeout | payload_too_large | unknown. 5s default timeout (search hot path). Pre-flight payload guard rejects bodies overrecipe.touchpoints.reranker.max_payload_byteswithreason: 'payload_too_large'so callers can fail-open without an HTTP call._rerankTransporttest seam mirrors_embedTransport. Newgateway.embedQuery(text)companion threadsinputType: 'query'throughdimsProviderOptions()(now 4-arg).getRerankerModel()accessor +isAvailable('reranker')branch added.configureGateway+reconfigureGatewayWithEnginethreadreranker_modelthrough the same path as embedding/expansion/chat.applyResolveAuth+defaultResolveAuthwiden touchpoint param to include'reranker'. v0.34.1.0 (#875): newembedMultimodalOpenAICompat()routes recipes withimplementation: 'openai-compatible'(LiteLLM, Anyscale, vLLM, Gemini multimodal via proxy) through the standard/embeddingsendpoint with content arrays carryingimage_urlentries. The pre-existing Voyage/multimodalembeddingspath is unchanged; the gateway selects by recipeimplementationtag. Runtime dimension validation throwsAIConfigError(with model id + observed + expected) before the vector reaches storage when the provider returns a width that doesn't match the recipe'sdefault_dimsor the brain'sembedding_dimensionsconfig — no more crypticvector dimension mismatchat INSERT time. Pinned by 11 cases intest/openai-compat-multimodal.test.ts. v0.28.7 (#680): module-scoped_embedTransportdefaulting to AI SDKembedMany, with__setEmbedTransportForTests(fn)test seam so tests drive the publicembed()function with a stubbed transport instead of probing private helpers.splitByTokenBudgetandisTokenLimitErrorare now exported@internal— pure functions reused directly by the test file. Module-level_shrinkState: Map<recipeId, {factor, consecutiveSuccesses}>halves the recipe's effectivesafety_factoron token-limit miss (floor 0.05) and heals back ×1.5 toward the ceiling afterSHRINK_HEAL_AFTER=10consecutive successes.configureGateway()walks every registered recipe at construction time and emits a once-per-process stderr warning for any embedding touchpoint missingmax_batch_tokens(excluding the canonical OpenAI fast-path recipe).resetGateway()clears_shrinkState, the warned-set, and restores the real transport. ASCII flow diagram embedded in theembed()JSDoc covers the routing decision, recursion + halving, and shrinkState lifecycle. v0.28.11 (#719):embedMultimodal()readscfg.embedding_multimodal_modelfirst (falls back tocfg.embedding_modelfor single-model setups). After the existing recipe-levelsupports_multimodalfast-fail, validates the resolved model againsttouchpoint.multimodal_modelswhen declared — closes the Voyage-text-only-model-into-multimodal-endpoint footgun before any HTTP call (Codex F1 from PR review). NewgetMultimodalModel()accessor mirrorsgetEmbeddingModel/getChatModelso doctor and integration tests can read the gateway state. v0.33.1.1 (#962, Codex P3 follow-up): new exportedVoyageResponseTooLargeErrortagged class at the top of the file.voyageCompatFetch's two OOM-defense caps (Layer 1 Content-Length check at:595, Layer 2 per-embedding base64 cap at:619) now throwVoyageResponseTooLargeErrorinstead of a genericError. The inbound response-rewriter's surrounding try/catch (which intentionally swallows parse failures so misshaped Voyage responses fall through to the SDK's JSON parser) checksinstanceof VoyageResponseTooLargeErrorand rethrows. Pre-fix, the Layer 2 throw was silently swallowed and the oversized response returned to the AI SDK anyway — Layer 2 was theatrical. Source-shape regression assertion intest/voyage-response-cap.test.tspins theinstanceof ⇒ throw errline.src/core/ai/recipes/zeroentropyai.ts(v0.35.0.0) — ZeroEntropy openai-compatible recipe declaring BOTHembedding(zembed-1, 7 Matryoshka dims: 2560/1280/640/320/160/80/40) ANDreranker(zerank-2flagship +zerank-1+zerank-1-small, 5MB payload cap) touchpoints.implementation: 'openai-compatible'(NOT the misspelled'openai-compat'the original plan draft had — pinned by F1 regression intest/ai/zeroentropy-recipe.test.ts).base_url_default: 'https://api.zeroentropy.dev/v1'already ends with/v1, so thezeroEntropyCompatFetchURL rewrite/embeddings → /models/embedproduces…/v1/models/embed(NOT…/v1/v1/…— pinned by F2 regression).chars_per_token: 1+safety_factor: 0.5match Voyage's dense-content hedge.src/core/rerank-audit.ts(v0.35.0.0) — failure-only JSONL audit at~/.gbrain/audit/rerank-failures-YYYY-Www.jsonl(ISO-week rotation, mirrorssrc/core/audit-slug-fallback.ts). ExportslogRerankFailure({reason, model, query_hash, doc_count, error_summary})+readRecentRerankFailures(days). Deliberately nologRerankSuccess(CDX2-F22 in plan review): writing once per tokenmax search is hot-path I/O churn AND success events leak query volume + timing into a local audit file.gbrain doctor'sreranker_healthcheck readssearch.reranker.enabledfirst so "no events in window" is interpreted correctly (disabled → ok; enabled → ok). Query text is SHA-256-prefix-hashed (8 hex chars) for privacy.GBRAIN_AUDIT_DIRenv override honored via the sharedresolveAuditDir().src/core/search/rerank.ts(v0.35.0.0) — the call-site abstraction.applyReranker(query, results, opts)slots betweendedupResults()andenforceTokenBudget()insrc/core/search/hybrid.ts. Slicesopts.topNIn(default 30) by current RRF order, sends togateway.rerank(), reorders byrelevanceScoredesc, and appends the un-reranked tail unchanged (recall protection). Fail-open on everyRerankError.reason: any error logs vialogRerankFailureand returns the input array unchanged. Stampsrerank_scoreonto reordered items so downstream telemetry sees the new ordering signal.topNOut: nullis the explicit "don't truncate" signal — semantically distinct fromundefinedwhich means "fall through to mode bundle" (CDX2-F16). Test seam:opts.rerankerFnlets tests stubgateway.rerankwithout touching the network.src/core/ai/recipes/voyage.ts— Voyage AI openai-compatible recipe. v0.28.7 (#680): declareschars_per_token=1+safety_factor=0.5so the gateway pre-splits Voyage batches at a 60K-character budget (50% of 120K-token cap with the dense-tokenizer ratio). Closes the v0.27 backfill loop where ~26% of the corpus stayed un-embedded because tiktoken-grounded budgeting silently undercounted Voyage's actual token usage. v0.28.11 (#719): declaresmultimodal_models: ['voyage-multimodal-3']so the gateway rejects text-only Voyage models pointed at the multimodal endpoint with a clearAIConfigErrorinstead of waiting for Voyage's HTTP 400. v0.33.1.1 (#962, fixup): recipe docstring at:7-16tightened to name the seven hosted flexible-dim models that acceptoutput_dimensionexplicitly (voyage-4-large,voyage-4,voyage-4-lite,voyage-3-large,voyage-3.5,voyage-3.5-lite,voyage-code-3) and call out thatvoyage-4-nanois the open-weight variant listed separately by Voyage as fixed 1024-dim — does NOT accept the parameter. The "all v4 variants are flexible" misread is what caused the original PR to include nano inVOYAGE_OUTPUT_DIMENSION_MODELS; the negative regression assertion intest/ai/gateway.test.ts(dimsProviderOptionsreturnsundefinedforvoyage-4-nano) pins the contract.src/core/ai/recipes/anthropic.ts— Anthropic recipe (chat + expansion touchpoints). v0.31.12: chat and expansionmodels:lists drop the v0.31.6 phantomclaude-sonnet-4-6-20250929date suffix — canonical id isclaude-sonnet-4-6. The wrong-direction aliasclaude-sonnet-4-6 → claude-sonnet-4-6-20250929is removed; a reverse aliasclaude-sonnet-4-6-20250929 → claude-sonnet-4-6keeps stale user configs working (rescuesfacts.extraction_modelandmodels.dream.synthesizeset by v0.31.6 installs). Recipe-shape regression pinned bytest/anthropic-model-ids.test.ts(6 cases, verbatim cherry-pick of PR #830 plus the reverse-alias rescue case).src/core/anthropic-pricing.ts— Single source of truth for Anthropic model pricing (per-MTok input/output). v0.31.12: Opus 4.7 corrected from$15/$75to$5/$25(the old number was from Opus 4 generation, never refreshed when 4.7 shipped); Opus 4.6 also corrected. Consumed bysrc/core/budget-meter.tsandsrc/core/cross-modal-eval/runner.ts— the cross-modal estimator now readsANTHROPIC_PRICINGfor Anthropic models instead of duplicating the table, killing the v0.31.6 drift bug class.src/core/model-config.ts— Model-string resolution (the seam every internal LLM call walks through). v0.31.12: four-tier system (ModelTier = 'utility' | 'reasoning' | 'deep' | 'subagent') withTIER_DEFAULTS(utility→haiku-4-5, reasoning→sonnet-4-6, deep→opus-4-7, subagent→sonnet-4-6) andtier?: ModelTieronResolveModelOpts. Resolution chain is now 8 steps: cliFlag → deprecated key → config key →models.default→models.tier.<tier>→ env var →TIER_DEFAULTS[tier]→ caller fallback. Two new exports —isAnthropicProvider(modelString)checksprovider:modelprefix ORclaude-bare-id pattern, andenforceSubagentAnthropic()is the layer-2 runtime guard: whentier === 'subagent'resolves to a non-Anthropic provider, it emits a once-per-(source, model)stderr warn AND falls back toTIER_DEFAULTS.subagentinstead of letting the Anthropic Messages API tool-loop attempt to run on OpenAI/Gemini._resetDeprecationWarningsForTest()now also clears_subagentTierWarningsEmittedso tests re-emit.src/core/ai/model-resolver.ts— Recipe-touchpoint validator. v0.31.12:assertTouchpoint(recipe, touchpoint, modelId, extendedModels?)gains an optional 4thextendedModels: ReadonlySet<string>argument. When the modelId is in that set, the native-recipe allowlist throw is bypassed — the user explicitly opted into this model via config so we let provider rejection surface asmodel_not_foundat HTTP call time (andgbrain models doctorcatches it earlier). Default code paths with hardcoded model strings MUST NOT passextendedModels— typos in source code still fail fast. Replaces the earlier plan to soften the validator wholesale (Codex F4/F5 in plan review flagged that as too broad — it would have removed the fail-fast contract for chat + expand + embed all three).src/core/ai/gateway.tsextension (v0.31.12) — new module-scoped_extendedModels: Map<providerId, Set<modelId>>registry feedsassertTouchpoint's 4th-arg path. NewreconfigureGatewayWithEngine(engine)async function is called fromcli.tsafterengine.connect()(and before every command exceptCLI_ONLYno-DB commands) — re-resolves expansion + chat defaults throughresolveModel()somodels.tier.*andmodels.defaultoverrides apply to expansion + chat both.DEFAULT_CHAT_MODELcorrected toanthropic:claude-sonnet-4-6(was the v0.31.6 phantom-20250929). New__setChatTransportForTestsseam mirrors__setEmbedTransportForTestsso tests drivechat()with a stubbed transport.src/core/minions/queue.tsextension (v0.31.12) —MinionQueue.add()now rejectssubagentjobs whosedata.modelresolves throughisAnthropicProvider()to a non-Anthropic provider. Lazy-importsmodel-config.tsto avoid pulling engine types into queue's eager-load surface. Layer 1 of the three-layer subagent provider enforcement (Codex F1+F2 in plan review). Layers 2 + 3 live insrc/core/model-config.ts(enforceSubagentAnthropicruntime fallback) andsrc/commands/doctor.ts(subagent_providercheck). Pinned by 3 cases intest/agent-cli.test.ts.src/commands/models.ts(v0.31.12) —gbrain models [--json]read-only routing dashboard: prints tier defaults (utility/reasoning/deep/subagent), the resolved value for each (re-walking the resolution chain to attribute properly), every per-task override (11PER_TASK_KEYSentries —models.dream.synthesize,models.dream.patterns,models.drift,models.auto_think,models.think,models.subagent,facts.extraction_model,models.eval.longmemeval,models.expansion,models.chat,models.dream.synthesize_verdict), the alias map (defaults + user overrides), and a source-of-truth column showingdefault/config: <key>/env: <VAR>.gbrain models doctor [--skip=<provider>] [--json]fires a 1-tokengateway.chat()probe against each configured chat + expansion model and classifies failures into{model_not_found, auth, rate_limit, network, unknown}— the structural fix for the v0.31.6 silent-no-op bug class. Wired intocli.tsdispatch table +CLI_ONLYset. v0.33.1.1 (#962, Codex P3 follow-up): doctor gains a zero-tokenembedding_configprobe that runs FIRST, before any chat/expansion probes spend money.probeEmbeddingConfig()readsgetEmbeddingModel()+getEmbeddingDimensions()from the gateway, parses the model id, and (for Voyage flexible-dim models) checksisValidVoyageOutputDim(dims)againstVOYAGE_VALID_OUTPUT_DIMS. NewProbeStatusvariant'config'and optionalfix?: stringfield onProbeResult— surfaced in both human output (paste-readygbrain config set ...line under the bad probe) and JSON output. New touchpoint label'embedding_config'joins'chat'and'expansion'in the probe-row taxonomy. Closes the Voyage flexible-dim bug class at config time, not first-embed.src/commands/doctor.tsextension (v0.31.12) — newsubagent_providercheck (layer 3 of 3 — Codex F13). Warns whenmodels.tier.subagentis explicitly set to a non-Anthropic provider (fail-loud since the user clearly meant it — message names the bad value and prints the paste-ready fix commandgbrain config set models.tier.subagent anthropic:claude-sonnet-4-6); also warns whenmodels.defaultwould sneaksubagentinto a non-Anthropic provider via tier inheritance. OK status when subagent tier resolves to Anthropic. Tests cover all three paths intest/doctor.test.ts.src/core/check-resolvable.ts— Resolver validation: reachability, MECE overlap, DRY checks, structured fix objects. v0.14.1:CROSS_CUTTING_PATTERNS.conventionsis an array (notability gate accepts bothconventions/quality.mdand_brain-filing-rules.md). NewextractDelegationTargets()parses> **Convention:**,> **Filing rule:**, and inline backtick references. DRY suppression is proximity-based viaDRY_PROXIMITY_LINES = 40.src/core/repo-root.ts— SharedfindRepoRoot(startDir?)(v0.16.4): walks up fromstartDir(defaultprocess.cwd()) looking forskills/RESOLVER.md. Zero-dependency module imported by bothdoctor.tsandcheck-resolvable.ts. ParameterizedstartDirmakes tests hermetic. v0.31.7: read-path / write-path split.autoDetectSkillsDir(shared, read+write-safe) gains tier-0$GBRAIN_SKILLS_DIRexplicit operator override (Docker mounts, CI, monorepo subdirs) ahead of the existing 4-tier chain. NewautoDetectSkillsDirReadOnlywraps it with a tier-5 install-path fallback that walks up fromfileURLToPath(import.meta.url)and gates onisGbrainRepoRootso unrelated repos can't false-positive. Read-path callers (doctor,check-resolvable,routing-eval) use the read-only variant; write-path callers (skillpack install,skillify scaffold,post-install-advisory) deliberately stay on the shared function sogbrain skillpack installfrom~cannot silently retarget the bundled gbrain repo'sskills/instead of the user's actual workspace. Two newSkillsDirSourcevariants:'env_explicit','install_path'. NewAUTO_DETECT_HINT_READ_ONLYdocuments the extra tier. The D6--fixsafety gate indoctor.ts+check-resolvable.tsrefuses auto-repair whendetected.source === 'install_path'sogbrain doctor --fixfrom~cannot silently rewrite the bundled install tree.src/commands/check-resolvable.ts— Standalone CLI wrapper (v0.16.4) overcheckResolvable(). ExportsparseFlags,resolveSkillsDir,DEFERRED,runCheckResolvable. Exit rule: 1 on any issue (warnings OR errors), stricter than doctor'sokflag — honors README:259. Stable JSON envelope{ok, skillsDir, report, autoFix, deferred, error, message}— same shape on success and error paths.--fixpath runsautoFixDryViolationsBEFOREcheckResolvable(same ordering as doctor).scripts/skillify-check.tssubprocess-callsgbrain check-resolvable --json(cached per process) and fails loud on binary-missing — no silent false-pass. v0.19: AGENTS.md workspaces now resolve natively (seesrc/core/resolver-filenames.ts) — gbrain inspects the 107-skill OpenClaw deployment whether the routing file isRESOLVER.mdorAGENTS.md.DEFERRED[]is empty — Checks 5 + 6 shipped as real code, not issue URLs. v0.31.7: the resolver lookup switched from first-match-wins to the multi-file merge insrc/core/check-resolvable.ts— entries collected from everyRESOLVER.md/AGENTS.mdacross the skills dir AND its parent, deduped byskillPath(first occurrence wins). Lifted reachable skills on the reference OpenClaw layout from 37/224 to 200/224 — the deployment ships a thinskills/RESOLVER.md(~40 entries from skillpack) plus a fat../AGENTS.md(200+ entries, the real dispatcher), and the previous code only saw the first one. The CLI also switched toautoDetectSkillsDirReadOnlysocd ~ && gbrain check-resolvablefinds the bundled skills via the install-path fallback.--fixcarries the same D6 safety gate asgbrain doctor --fix: refuses to write whendetected.source === 'install_path'.src/core/resolver-filenames.ts(v0.19) — central list of accepted routing filenames (RESOLVER.md,AGENTS.md). Shared byfindRepoRoot,check-resolvable, and skillpack install so every code path walks the same fallback chain.src/commands/skillify.ts+src/core/skillify/{generator,templates}.ts(v0.19) —gbrain skillify scaffold <name>creates all stubs for a new skill in one command: SKILL.md, script, tests, routing-eval.jsonl, resolver entry, filing-rules pointer.gbrain skillify check <script>runs the 10-step checklist (LLM evals, routing evals, check-resolvable gate, filing audit) against a candidate skill before it lands.src/commands/skillify-check.ts(v0.19) —gbrain skillpack-checkagent-readable health report. Exit 0/1/2 for CI pipeline gating; JSON for debugging. Wrapscheck-resolvable --json,doctor --json, and migration ledger into one payload so agents can decide whether a human action is required.src/commands/book-mirror.ts(v0.25.1) —gbrain book-mirror --chapters-dir <path> --slug <slug> [flags]. Flagship of the v0.25.1 skills wave. Submits N read-only subagent jobs (one per chapter;allowed_tools: ['get_page', 'search']), waits for all viawaitForCompletion, reads each child'sjob.result, assembles two-column markdown CLI-side, writes a single operator-trustput_pagetomedia/books/<slug>-personalized.md. Codex HIGH-1 fix applied: trust narrowing happens at the tool-allowlist layer (subagents can't call put_page) instead of allowedSlugPrefixes — untrusted EPUB content cannot prompt-inject any people page. Cost-estimate prompt before launching; refuses to spend in non-TTY without--yes. Per-chapter idempotency keys (book-mirror:<slug>:ch-<N>) for retry-friendly re-runs. Partial-failure handling: assembles with completed chapters and a## Failed chapterssection listing retries. Test surface:test/book-mirror.test.ts(9 cases — CLI registration + source invariants).src/commands/skillpack.ts+src/core/skillpack/{bundle,scaffold,reference,migrate-fence,scrub-legacy,harvest,harvest-lint,copy,apply-hunks,diff-text,installer}.ts(v0.19 → v0.36) — v0.36 contract change: managed-block install model retired.installanduninstallremoved (clean break, no alias; both exit non-zero with a hint pointing at the replacement command). New surface:scaffold(one-time additive copy via sharedcopyArtifactshelper incopy.ts; refuses to overwrite existing files; partial-state fills missing paired sources declared in SKILL.md frontmattersources:),reference(read-only diff lens with agent-readable framing line +--apply-clean-hunkstwo-way auto-apply via pure-JS unified-diff parser/applier inapply-hunks.ts+diff-text.ts),migrate-fence(one-shot strip of legacy fence; cumulative-slugs receipt → row-parsing fallback; preserves rows verbatim as user-owned routing),scrub-legacy-fence-rows(opt-in row cleanup with skill-present + non-empty-triggers gate),harvest(host→gbrain inverse with symlink-reject + canonical-path containment viavalidateUploadPath-style gate + default-on privacy linter inharvest-lint.tsagainst~/.gbrain/harvest-private-patterns.txtplus built-in\bWintermute\b+ email + Slack-channel patterns; rollback on match). Paired-source declarations moved fromopenclaw.plugin.jsonto each SKILL.md's frontmattersources:array (D2; validated byloadSkillSourcesinbundle.ts).autoDetectSkillsDir(insrc/core/repo-root.ts) gains acwd_walk_uptier ahead of~/.openclaw/workspace(D3; non-OpenClaw hosts like~/git/wintermuteauto-detect; R5 regression preserves$OPENCLAW_WORKSPACEprecedence).gbrain skillpack check --strictexits non-zero on drift (CI gate); top-levelgbrain skillpack-checkkeeps exit-1-on-issues for cron compat. Companion editorial skillskills/skillpack-harvest/SKILL.mddrives the genericization checklist before the CLI runs. Design + workflow doc:docs/guides/skillpacks-as-scaffolding.md. ~600 LOC of managed-block machinery deleted; ~400 LOC of new modules + ~1000 LOC of new test coverage acrosstest/skillpack-{copy,scaffold,reference,reference-apply,apply-hunks,migrate-fence,scrub-legacy,harvest,harvest-lint,frontmatter-sources}.test.ts+ 9-case E2E intest/e2e/skillpack-flow.test.ts.installer.ts+test/skillpack-install.test.tssurvive for now —gbrain skillpack diffstill usesdiffSkillfrom there; slated for v0.37 cleanup. Historical (v0.19-v0.35.1): managed-block model with<!-- gbrain:skillpack:begin -->/end -->fence,cumulative-slugs="..."receipt, content-hash gates, lockfile;install --allprune;uninstallwith D8 receipt gate + D11 atomic-refusal content-hash pre-scan. Replaced wholesale in v0.36.src/core/archive-crawler-config.ts(v0.25.1) — D12 + codex HIGH-4 safety gate for thearchive-crawlerskill. Refuses to run unlessarchive-crawler.scan_paths:is explicitly set in the brain repo'sgbrain.yml. Mirrors the storage-config.ts parsing pattern (sibling file; separate concern from storage tiering).loadArchiveCrawlerConfig(repoPath)throwsArchiveCrawlerConfigError(missing_section | empty_scan_paths | invalid_path | parse_error).normalizeAndValidateArchiveCrawlerConfigrejects relative paths and..traversal;~is expanded; trailing-slash normalized for unambiguous prefix matching.isPathAllowed(candidate, config)is the runtime per-file gate (scan_paths prefix-match with directory-boundary correctness; deny_paths overrides). Tests intest/archive-crawler-config.test.ts(19 cases).test/helpers/cli-pty-runner.ts(v0.25.1) — generic real-PTY harness ported from gstack and trimmed to ~470 lines. Uses pureBun.spawn({terminal:})(Bun 1.3.10+; engines.bun pin in package.json). Generic primitives only — no plan-mode orchestrators. Exports:launchPty,resolveBinary,stripAnsi,parseNumberedOptions,optionsSignature,isNumberedOptionListVisible,isTrustDialogVisible. Self-tests intest/cli-pty-runner.test.ts(24 cases).src/core/skill-manifest.ts(v0.19) — parser forskill-manifest.jsonrecords. Used by skillpack installer to detect drift between the shipped bundle and the user's local edits, so updates merge instead of overwriting.src/commands/routing-eval.ts+src/core/routing-eval.ts(v0.19) —gbrain routing-evalcatches user phrasings that route to the wrong skill. Readsskills/<name>/routing-eval.jsonlfixtures ({intent, expected_skill, ambiguous_with?}). Structural layer runs incheck-resolvableby default (zero API cost). The--llmflag is accepted as a placeholder for a future LLM tie-break layer; in v0.24.0 it emits a stderr notice and runs structural only. False positives surface before users hit them. v0.31.7: switched toautoDetectSkillsDirReadOnlyand the same multi-file resolver merge ascheck-resolvable, so on OpenClaw layouts (skills/RESOLVER.md+../AGENTS.md) all three commands see the same trigger index — previouslyrouting-evalread only the first resolver file it found. The v0.25.1 wave skills' RESOLVER.md rows were also synced to include the full frontmattertriggers:arrays (was only the first trigger), so the structural matcher actually sees the realistic phrasings; ambiguous-fixture annotations cover deliberate skill chains likeenrich → article-enrichment.src/core/filing-audit.ts+skills/_brain-filing-rules.json(v0.19) — Check 6 ofcheck-resolvable. Parses newwrites_pages:/writes_to:frontmatter on skills and audits their filing claims against the filing-rules JSON. Warning-only in v0.19, upgrades to error in v0.20.src/core/dry-fix.ts—gbrain doctor --fixengine.autoFixDryViolations(fixes, {dryRun})rewrites inlined rules to> **Convention:** see [path](path).callouts via three shape-aware expanders (bullet / blockquote / paragraph). Five guards: working-tree-dirty (getWorkingTreeStatus()returns 3-state'clean' | 'dirty' | 'not_a_repo'), no-git-backup, inside-code-fence, already-delegated (40-line proximity, consistent with detector), ambiguous-multi-match, block-is-callout.execFileSyncarray args (no shell — no injection surface). EOF newline preserved.src/core/backoff.ts— Adaptive load-aware throttling: CPU/memory checks, exponential backoff, active hours multipliersrc/core/fail-improve.ts— Deterministic-first, LLM-fallback loop with JSONL failure logging and auto-test generationsrc/core/transcription.ts— Audio transcription: Groq Whisper (default), OpenAI fallback, ffmpeg segmentation for >25MBsrc/core/enrichment-service.ts— Global enrichment service: entity slug generation, tier auto-escalation, batch throttlingsrc/core/data-research.ts— Recipe validation, field extraction (MRR/ARR regex), dedup, tracker parsing, HTML strippingsrc/commands/embed.ts—gbrain embed [--stale|--all] [--slugs ...]. v0.22.1 (#409, contributed by @atrevino47):--stalepath now starts withengine.countStaleChunks()(single SELECT count(*) WHERE embedding IS NULL, ~50 bytes wire). On a fully-embedded brain that's a 1-line short-circuit — no further reads. When stale chunks exist,engine.listStaleChunks()returns just the chunks needing embeddings (slug + chunk_index + chunk_text + metadata, novector(1536)payload). Caller groups by slug, embeds via OpenAI, re-upserts viaupsertChunks. Replaces the prior page-walk that pulled every chunk's embedding column over the wire and discarded most.src/commands/extract.ts—gbrain extract links|timeline|all [--source fs|db]: batch link/timeline extraction. fs walks markdown files, db walks pages from the engine (mutation-immune snapshot iteration; use this for live brains with no local checkout). As of v0.12.1 there is no in-memory dedup pre-load — candidates are buffered 100 at a time and flushed viaaddLinksBatch/addTimelineEntriesBatch;ON CONFLICT DO NOTHINGenforces uniqueness at the DB layer, and thecreatedcounter returns real rows inserted (truthful on re-runs). v0.22.1 (#417):ExtractOpts.slugs?: string[]enables incremental extract — when set,extractForSlugs()reads ONLY those slugs' files (single combined links+timeline pass) instead of the full directory walk. CLIgbrain extractkeeps full-walk behavior; the cycle path threads sync'spagesAffectedthrough.walkMarkdownFiles(brainDir)still runs at line 455 to buildallSlugsfor link resolution — seeTODOS.mdfor replacing it withengine.getAllSlugs().src/commands/graph-query.ts—gbrain graph-query <slug> [--type T] [--depth N] [--direction in|out|both]: typed-edge relationship traversal (renders indented tree)src/core/link-extraction.ts— shared library for the v0.12.0 graph layer. extractEntityRefs (canonical, replaces backlinks.ts duplicate) matches both[Name](people/slug)markdown links and Obsidian[[people/slug|Name]]wikilinks as of v0.12.3. extractPageLinks, inferLinkType heuristics (attended/works_at/invested_in/founded/advises/source/mentions), parseTimelineEntries, isAutoLinkEnabled config helper.DIR_PATTERNcoverspeople,companies,deals,topics,concepts,projects,entities,tech,finance,personal,openclaw. Used by extract.ts, operations.ts auto-link post-hook, and backlinks.ts.src/core/zombie-reap.ts(v0.28.1) — idempotentinstallSigchldHandler()so JS-spawned children get reaped via Bun's internalwaitpid(). Bun (like Node) only auto-reaps when a SIGCHLD listener is registered; without it, every child the worker spawns (shell jobs, embed batches, sub-agents) becomes a zombie on exit and holds connection slots. Called once at module load fromsrc/cli.ts(with Windows platform guard — SIGCHLD doesn't exist on Windows). Cross-file leak guard via_uninstallSigchldHandlerForTests()for tests. Layer 1 of the three-layer zombie defense; Layer 2 is tini-as-PID-1 wrapping the worker subtree (viasrc/core/minions/spawn-helpers.ts); Layer 3 is the container's own tini for hard Bun crashes.src/core/minions/— Minions job queue: BullMQ-inspired, Postgres-native (queue, worker, backoff, types, protected-names, quiet-hours, stagger, handlers/shell).src/core/minions/queue.ts— MinionQueue class (submit, claim, complete, fail, stall detection, parent-child, depth/child-cap, per-job timeouts, cascade-kill, attachments, idempotency keys, child_done inbox, removeOnComplete/Fail).add()takes a 4thtrustedarg (separate fromoptsto prevent spread leakage); protected names inPROTECTED_JOB_NAMESrequire{allowProtectedSubmit: true}and the check runs trim-normalized (whitespace-bypass safe). v0.14.1 #219:add()plumbsmax_stalledthrough with a[1, 100]clamp; omitted values let the schema DEFAULT (5) kick in. v0.19.0:handleWallClockTimeouts(lockDurationMs)is Layer 3 kill shot for jobs whereFOR UPDATE SKIP LOCKEDstall detection and the timeout sweep both fail to evict (wedged worker holding a row lock via a pending transaction). v0.19.1:maxWaitingcoalesce path now usespg_advisory_xact_lockkeyed on(name, queue)to serialize concurrent submits for the same key, and filters onqueuein addition tonameso cross-queue same-name jobs don't suppress each other.src/core/minions/worker.ts— MinionWorker class (handler registry, lock renewal, graceful shutdown, timeout safety net). v0.14.0 abort-path fix: aborted jobs now callfailJobwith reason (timeout/cancel/lock-lost/shutdown) instead of returning silently.shutdownAbort(instance field) fires on process SIGTERM/SIGINT and propagates toctx.shutdownSignal— shell handler listens to it; non-shell handlers don't. v0.22.1 (#403): per-job timeout firesabort.abort(new Error('timeout'))then a 30-second grace-then-evict safety net force-evicts the job frominFlightand marks it dead in DB if the handler ignores the abort signal — frees the slot even when a handler wedges (the 98-waiting-0-active prod incident driver). v0.28.1 engine-ownership invariant:start()no longer callsengine.disconnect()on shutdown — that was a leaky abstraction (the worker disconnected an engine it didn't own). The CLI handler insrc/commands/jobs.ts case 'work'now owns engine lifecycle via try/finally with loud error logging on disconnect failure. Pinned bytest/worker-shutdown-disconnect.test.tsasserting the inverse (disconnectSpy).not.toHaveBeenCalled()). v0.34.3.0: RSS watchdog metric switched to non-file-backed pages on Linux. New exportsparseRssFromProcStatus(status)(pure parser, exported for unit tests) andgetAccurateRss(readStatus?)(reads/proc/self/statusforRssAnon + RssShmem, falls back toprocess.memoryUsage().rsson macOS / restricted containers / kernel <4.5). The defaultgetRssinjected intoWorkerOptsis nowgetAccurateRssinstead ofprocess.memoryUsage().rss. Closes the prod incident where VmRSS inflated to 7GB on a 96K-page brain (file-backed git packfile mmaps) while heap stayed at ~100MB; the watchdog was firing every autopilot cycle. M1 parser fix uses field-presence regex checks soRssAnon: 0 + RssShmem: 512(shmem-only worker case) parses correctly instead of falling through to VmRSS. Pinned bytest/worker-rss.test.ts(11 cases).src/core/minions/supervisor.ts— MinionSupervisor process manager. Spawnsgbrain jobs workas a child, restarts on crash with exponential backoff, periodic health check. v0.22.1 (#406):consecutiveHealthFailurescounter; on 3 consecutive failures emitshealth_warnwithreason: 'db_connection_degraded'and callsengine.reconnect()to swap in a fresh pool, then resets the counter. Worker exit classifier emitslikely_causefield onworker_exitedevents:oom_or_external_kill(SIGKILL),graceful_shutdown(SIGTERM),runtime_error(code 1),clean_exit(code 0),unknown. v0.28.1: consumesdetectTini()+buildSpawnInvocation()fromsrc/core/minions/spawn-helpers.tsto wrap the worker subtree in tini-as-PID-1 when tini is onPATH(handles native-addon zombie reaping that the in-process SIGCHLD reaper can't reach). ExposesisTiniDetectedread-only accessor for tests. v0.34.3.0: spawn-and-respawn loop extracted into the sharedChildWorkerSupervisorcore (see entry below). MinionSupervisor now composes the inner class viarunSuperviseLoop()→new ChildWorkerSupervisor({...})and mapsChildSupervisorEventshapes back through the existingemit()SupervisorEvent channel — JSONL audit consumers see byte-compatible output across the rename. PID lock, signal handlers, health check, andprocess.exiton max-crashes stay in MinionSupervisor (standalone-daemon concerns). The pre-shipped reset-to-0-on-code=0 hunk that originally fixed the prod crash-counter incident is gone; the same fix lives in the shared core under the D1 amendment (code=0 leavescrashCountuntouched, so a worker alternating real crashes + watchdog drains still tripsmax_crashes). D2cleanRestartBudget(default 10 restarts per 60s) caps the macOS/non-Linux-fallback tight-loop by emittinghealth_warn { reason: 'clean_restart_budget_exceeded' }plus backoff after the threshold trips.shutdown()drains viachildSupervisor.killChild('SIGTERM')+awaitChildExit(35_000)instead of reaching intothis.childdirectly. Pinned bytest/supervisor.test.ts(16 cases; existing tests that previously relied on clean-exit-as-crash semantics now use exit-1 workers since clean exits no longer count) andtest/supervisor-tini.test.ts.src/core/minions/child-worker-supervisor.ts(v0.34.3.0) — shared spawn-and-respawn core extracted fromMinionSupervisorso it can be reused by bothMinionSupervisor(standalonegbrain jobs supervisordaemon) andsrc/commands/autopilot.ts(autopilot daemon). Pre-v0.34.3.0 the two consumers maintained parallel spawn loops that drifted into the same bug class — Codex caught it during plan-eng-review on PR #1003. Pure class: NO PID file, NO signal handlers, NOprocess.exit, NO health check. Lifecycle events fire via injectedonEvent: (ChildSupervisorEvent) => voidcallback so each composer routes to its own log/audit channel. D1 exit classifier:code === 0leavescrashCountUNCHANGED (preserves flap detection across mixed exit sequences — a worker that alternatesexit 1 / exit 0 / exit 1 / exit 0correctly tripsmax_crashesafter 10 real crashes regardless of intervening clean exits).code != 0follows the existingrunDuration > stableRunResetMs ? 1 : ++crashCountrule. D2 clean-restart budget: sliding window tracks code=0 exits; when count exceedscleanRestartBudget(default 10) insidecleanRestartWindowMs(default 60s), emitshealth_warn { reason: 'clean_restart_budget_exceeded' }and appliescleanRestartBudgetBackoffMs(default 1s) before the next spawn. Caps the worst-case tight-loop on macOS / restricted containers / kernel <4.5 where the worker's RSS watchdog falls back to VmRSS. Public read-only accessorschildAlive,inBackoff,crashCountfor composer health checks;killChild(signal)+awaitChildExit(timeoutMs)for shutdown paths.awaitChildExitshort-circuits whenchild.exitCode !== null || child.signalCode !== null(regression caught in pre-landing /review: pre-fix, fast-SIGTERM responders caused a 35-second shutdown hang because the lateonce('exit', ...)listener never fired). Test hooks:_backoffFloorMsskips the real backoff curve,_nowinjects a fake clock. Pinned bytest/child-worker-supervisor.test.ts(7 cases: D1 classifier with code=0 not counted, interleaved exits still trip max_crashes, stable-run + clean-exit interaction across faked 6-minute run, D2 budget triggers backoff + health_warn, budget config is per-instance, awaitChildExit short-circuit, event-shape regression). Plan that produced the design lives at~/.claude/plans/this-is-a-real-sleepy-sketch.md.src/core/minions/spawn-helpers.ts(v0.28.1) — puredetectTini()+buildSpawnInvocation()helpers consumed by bothsupervisor.tsandautopilot.ts. Resolves the DRY violation between the two spawn sites and makes the tini wrapping testable withoutmock.module()(rule R2 ofscripts/check-test-isolation.sh).detectTini()callsexecFileSync('which', ['tini'])with explicitenv: process.envso Bun sees runtime PATH mutations (the env-snapshot bug fix).buildSpawnInvocation(tiniPath, cmd, args)returns{cmd, args}with tini prepended when present, or the bare invocation otherwise. Pinned bytest/spawn-helpers.test.ts(5 cases) andtest/supervisor-tini.test.ts(4 cases).src/core/minions/types.ts—MinionJobInput+MinionJobStatus+ handler context types.MinionJobInput.max_stalled(new in v0.14.1) is optional; omitted values let the schema DEFAULT (5) kick in, provided values are clamped to[1, 100].src/core/minions/protected-names.ts— side-effect-free constant module exportingPROTECTED_JOB_NAMES+isProtectedJobName(). Kept pure so queue core can import without loading handler modules.src/core/minions/handlers/shell.ts—shelljob handler. Spawns/bin/sh -c cmd(absolute path, PATH-override-safe) orargv[0] argv[1..](no shell). Env allowlist:PATH, HOME, USER, LANG, TZ, NODE_ENV+ callerenv:overrides. UTF-8-safe stdout/stderr tail viastring_decoder.StringDecoder. Abort (eitherctx.signalorctx.shutdownSignal) fires SIGTERM → 5s grace → SIGKILL on child. RequiresGBRAIN_ALLOW_SHELL_JOBS=1on worker (gated byregisterBuiltinHandlers).src/core/minions/handlers/shell-audit.ts— per-submission JSONL audit trail at~/.gbrain/audit/shell-jobs-YYYY-Www.jsonl(ISO-week rotation; override viaGBRAIN_AUDIT_DIR). Best-effort:mkdirSync(recursive)+appendFileSync; failures logged to stderr, submission not blocked. Logs cmd (first 80 chars) or argv (JSON array). Never logs env values.src/core/minions/handlers/supervisor-audit.ts— supervisor lifecycle JSONL audit at~/.gbrain/audit/supervisor-YYYY-Www.jsonl(ISO-week rotation; sharescomputeIsoWeekName()helper withshell-audit.ts).writeSupervisorEvent(emission, supervisorPid)appends one line per supervisor event (started,worker_spawned,worker_exited,backoff,health_warn,health_error,max_crashes_exceeded,shutting_down,stopped,worker_spawn_failed).readSupervisorEvents({sinceMs})is the readback path forgbrain doctor. v0.35.5.0: new exportsisCrashExit(event),summarizeCrashes(events),CrashSummarytype, andCLEAN_EXIT_CAUSESdenylist ('clean_exit' | 'graceful_shutdown'). Single regression point — bothgbrain doctor(Lane D supervisor check atdoctor.ts:1011-1043) andgbrain jobs supervisor status(jobs.ts:803-826) import from here so the two CLI surfaces cannot drift.isCrashExitclassifies a singleworker_exitedevent against the denylist:clean_exit/graceful_shutdownare NON-crashes; everything else (runtime_error,oom_or_external_kill,unknown, AND any futurelikely_causevalue added upstream inchild-worker-supervisor.ts) is a crash. Pre-v0.34 audit lines lackinglikely_causefall back tocode !== 0.summarizeCrashesreturns{total, by_cause: {runtime_error, oom_or_external_kill, unknown, legacy}, clean_exits}so dashboards bind to named buckets — thelegacybucket catches BOTH pre-v0.34 fallback entries AND future unrecognizedlikely_causevalues, fail-loud instead of silent underreport. Denylist-over-allowlist was a codex outside-voice catch during/plan-eng-review— the bug being fixed (read sites counting everyworker_exitedas a crash, inflating to 120+/day on healthy brains after v0.34.3.0 watchdog drains) was itself an allowlist-of-event-names. Pinned bytest/supervisor-audit.test.ts(14 cases: 9-caseisCrashExitbranch matrix including denylist regression guard for unrecognized future causes + non-exit-event defensive case, 5-casesummarizeCrashesaggregator including unrecognized-cause routing to legacy + null-code edge case) and 4 source-grep wiring assertions intest/doctor.test.tsguarding both surfaces against drift.src/core/minions/backpressure-audit.ts(v0.19.1) — sibling of shell-audit.ts formaxWaitingcoalesce events. JSONL at~/.gbrain/audit/backpressure-YYYY-Www.jsonl. Fires one line per coalesce with(queue, name, waiting_count, max_waiting, returned_job_id, ts). Closes the silent-drop vector the v0.19.0 maxWaiting guard introduced.src/core/minions/handlers/subagent.ts(v0.15) — LLM-loop handler. Two-phase tool persistence (pending → complete/failed), replay reconciliation for mid-dispatch crashes, dual-signal abort (ctx.signal+ctx.shutdownSignal), Anthropic prompt caching on system + tool defs.makeSubagentHandler({engine, client?, ...})factory;MessagesClientis an injectable interface the real SDK implements structurally. ThrowsRateLeaseUnavailableError(renewable) when rate-lease capacity is full. v0.30.2: Anthropic 400prompt is too longresponses (status 400 + body matches/prompt is too long|prompt_too_long|context.*length/i) classify asUnrecoverableErrorso the job goes straight todeadon first attempt instead of stalling three times before dead-lettering. Catches both initial-prompt overflow and turn-N tool-loop accumulation that the chunker insynthesize.tscan't bound ahead of time.src/core/minions/handlers/subagent-aggregator.ts(v0.15) —subagent_aggregatorhandler. Claims AFTER all children resolve (queue changes guarantee every terminal child posts achild_doneinbox message with outcome). Reads inbox viactx.readInbox(), builds deterministic mixed-outcome markdown summary. No LLM call in v0.15.src/core/minions/handlers/subagent-audit.ts(v0.15) — JSONL audit + heartbeat writer at~/.gbrain/audit/subagent-jobs-YYYY-Www.jsonl. Events:submission(one line per submit) +heartbeat(per turn boundary:llm_call_started | llm_call_completed | tool_called | tool_result | tool_failed). Never logs prompts or tool inputs.readSubagentAuditForJob(jobId, {sinceIso})is the readback path forgbrain agent logs.src/core/minions/rate-leases.ts(v0.15) — lease-based concurrency cap for outbound providers (default keyanthropic:messages, max viaGBRAIN_ANTHROPIC_MAX_INFLIGHT). Owner-tagged rows withexpires_atauto-prune on acquire;pg_advisory_xact_lockguards check-then-insert; CASCADE on owning job deletion.renewLeaseWithBackoffretries 3x (250/500/1000ms).src/core/minions/wait-for-completion.ts(v0.15) — poll-until-terminal helper for CLI callers.TimeoutErrordoes NOT cancel the job;AbortSignalexits without throwing. DefaultpollMs: 1000 on Postgres, 250 on PGLite inline.src/core/minions/transcript.ts(v0.15) — renderssubagent_messages+subagent_tool_executionsto markdown. Tool rows splice under their owning assistanttool_usebytool_use_id. UTF-8-safe truncation; unknown block types fall through to fenced JSON.src/core/minions/plugin-loader.ts(v0.15) —GBRAIN_PLUGIN_PATHdiscovery. Absolute paths only, left-wins collision,gbrain.plugin.jsonwithplugin_version: "gbrain-plugin-v1", plugins ship DEFS only (no new tools),allowed_tools:validated at load time against the derived registry.src/core/minions/tools/brain-allowlist.ts(v0.15, extended v0.23, v0.29, v0.35.3.0) — derives subagent tool registry fromsrc/core/operations.ts. 13-name allow-list as of v0.29 (was 11). By defaultput_pageschema is namespace-wrapped per subagent (^wiki/agents/<subagentId>/.+). v0.23 trusted-workspace path: whenBuildBrainToolsOpts.allowedSlugPrefixesis set, the put_page schema instead describes the prefix list to the model and the OperationContext is threaded withallowedSlugPrefixes. Trust comes fromPROTECTED_JOB_NAMESgating subagent submission — MCP cannot reach this field. Only cycle.ts (synthesize/patterns) and direct CLI submitters set it. v0.29:get_recent_salience+find_anomaliesadded to the allow-list.get_recent_transcriptsdeliberately NOT added — all subagent calls run withctx.remote === true, and the v0.29 trust gate rejects remote callers, so adding it would always reject (footgun). The cycle synthesize phase already callsdiscoverTranscriptsdirectly. v0.35.3.0:paramsToInputSchema()now consumesparamDefToSchemafromsrc/mcp/tool-defs.tsinstead of its own inline destructure. Required-aggregation at the tool-def level stays here (out of scope for the shared helper, which is per-param). Closes the third drift site in the ParamDef→JSON Schema bug class.src/mcp/tool-defs.ts(v0.15, v0.35.3.0) — extractedbuildToolDefs(ops)helper. MCP server + subagent tool registry both call it; byte-for-byte equivalence pinned bytest/mcp-tool-defs.test.ts. v0.35.3.0: exports the new recursiveparamDefToSchema(p: ParamDef)helper — single source of truth for ParamDef→JSON Schema mapping. Three consumers now share one mapper:buildToolDefs(stdio MCP),src/commands/serve-http.ts:837(HTTP MCPtools/list), andsrc/core/minions/tools/brain-allowlist.ts:84(subagent tool registry). Pre-v0.35.3, three inline destructures had drifted across the surface — the live HTTP MCP path droppeditemson every array param after a v0.32 review caught only the stdio side. Recursive onitemsso nested array-of-arrays preserves inner shape on the wire. Key ordering (type, description, enum, default, items) is intentional — matches the pre-v0.35.3 inline mappers so JSON.stringify output stays byte-stable.test/mcp-tool-defs.test.tsadds afindArrayWithoutItemswalker that fails the suite with a property path on any futuretype: 'array'lackingitems.type.src/core/minions/attachments.ts— Attachment validation (path traversal, null byte, oversize, base64, duplicate detection)src/commands/agent.ts(v0.16) —gbrain agent run <prompt> [flags]CLI. Submitssubagent(or N children + 1 aggregator) under{allowProtectedSubmit: true}. Single-entry--fanout-manifestshort-circuits. Children geton_child_fail: 'continue'+max_stalled: 3.--followis the default on TTY; streams logs + pollswaitForCompletionin parallel. Ctrl-C detaches, does not cancel.src/commands/agent-logs.ts(v0.16) —gbrain agent logs <job> [--follow] [--since]. Merges JSONL heartbeat audit +subagent_messagesinto a chronological timeline.parseSinceaccepts ISO-8601 or relative (5m,1h,2d). Transcript tail renders only for terminal jobs.src/commands/jobs.ts—gbrain jobsCLI subcommands +gbrain jobs workdaemon. v0.28.1:case 'work'now wrapsworker.start()in try/finally and owns engine lifecycle — callsengine.disconnect()on shutdown with loud error logging on failure. Replaces the prior call insideMinionWorker.start()(which violated engine ownership: the worker disconnected an engine it didn't own, and clobbered the module-level singleton on PostgresEngine via the now-fixed idempotency bug). Pool slots now free immediately on shutdown instead of waiting for TCP keepalive (~minutes). v0.13.1 surfaces the fullMinionJobInputretry/backoff/timeout/idempotency surface as first-class CLI flags onjobs submit:--max-stalled,--backoff-type fixed|exponential,--backoff-delay,--backoff-jitter,--timeout-ms,--idempotency-key.jobs smoke --sigkill-rescueis the opt-in regression guard for #219. v0.16 wiresregisterBuiltinHandlersto always registersubagent+subagent_aggregator(no env flag —ANTHROPIC_API_KEYis the natural cost gate, trust is viaPROTECTED_JOB_NAMES) and loadsGBRAIN_PLUGIN_PATHplugins at worker startup with a loud startup-line per plugin.shellhandler still gated byGBRAIN_ALLOW_SHELL_JOBS=1(RCE surface, separate concern). v0.22.10 (#521): theautopilot-cyclehandler now forwardsjob.data.phasestorunCycle(was previously discarded — caller-supplied phase selection silently became a full cycle). Phases are validated againstALL_PHASESfromsrc/core/cycle.ts; invalid names are filtered out and an empty/missing array falls back to the default 6-phase cycle. v0.22.13 (PR #490 CODEX-1+CODEX-4):synchandler now resolvessourceIdat entry by looking upsources.local_path(mirrorscycle.ts:480's autopilot fix from PR #475) so multi-source brains read the per-sourcelast_commitanchor instead of the global config key. Concurrency routed through the sharedautoConcurrency()policy insrc/core/sync-concurrency.tsinstead of the prior hardcoded4; PGLite stays serial.noEmbeddefault istrue(embed is a separate job — submitgbrain embed --staleafter sync, or rely on the autopilot cycle's embed phase). v0.35.5.0:gbrain jobs supervisor statusatjobs.ts:803-826now consumessummarizeCrashes()fromsrc/core/minions/handlers/supervisor-audit.tsfor cross-surface parity withgbrain doctor. JSON output addscrashes_by_cause: {runtime_error, oom_or_external_kill, unknown, legacy}+clean_exits_24hfields so dashboards bind to named buckets; human output gains a per-cause line underCrashes (24h)plus aClean exits (24h)line. Pre-fix the read site atjobs.ts:805counted everyworker_exitedevent as a crash regardless oflikely_cause— the same bug class the v0.35.5.0 doctor fix closes. Pinned by the 4 source-grep wiring assertions intest/doctor.test.tsthat require the per-cause breakdown substrings (crashes_by_cause,clean_exits_24h=) to appear in BOTHdoctor.tsandjobs.ts.src/commands/features.ts—gbrain features --json --auto-fix: usage scan + feature adoption salesmansrc/commands/autopilot.ts—gbrain autopilot --install: self-maintaining brain daemon (sync+extract+embed). v0.28.1: consumesdetectTini()fromsrc/core/minions/spawn-helpers.tsand resolves it once at startup instead of per worker respawn (was paying anexecFileSynccost on every restart). v0.34.3.0: inline spawn-and-respawn loop replaced with aChildWorkerSupervisorinstance. DropscrashCount,lastWorkerStartTime,STABLE_RUN_RESET_MS,startWorker, and the inlinechild.on('exit')block — all consolidated into the shared core.--max-rss 2048andmaxCrashes: 5preserved from the legacy loop.onMaxCrashesExceedednow routes through autopilot's ownshutdown('max_crashes')so the autopilot lockfile gets cleaned up (pre-refactor the inline loop calledprocess.exit(1)directly and bypassed cleanup).shutdown()drains viachildSupervisor.killChild('SIGTERM')+awaitChildExit(35_000)instead ofworkerProc.kill(). Pinned bytest/autopilot-supervisor-wiring.test.ts(6 static-shape regression guards: composes ChildWorkerSupervisor not the legacy inline names,--max-rss 2048in argv,maxCrashes: 5literal, shutdown-via-callback wiring, no workerProc reference). Closes the parallel-supervisor bug class Codex flagged during plan-eng-review.src/mcp/server.ts— MCP stdio server (generated from operations). v0.22.7: tool-call handler delegates todispatchToolCallfromsrc/mcp/dispatch.tsso stdio + HTTP transports share one validation, context-build, and error-format path. v0.34.1.0 (#870): stdin'end'/'close'shutdown hooks are skipped whenprocess.env.MCP_STDIO === '1'. Gateway-piped stdio MCP wrappers (OpenClaw'sbundle-mcp, similar) pipe the JSON-RPC handshake then close their stdin half; pre-fix this killed the server before the first tool call landed. Signal handlers (SIGTERM / SIGINT / SIGHUP) and the parent-process watchdog still cover legitimate disconnects.src/commands/serve.tsexposesServeOptions.mcpStdio?: booleanas a test seam so the runtime guard is exercisable without process.env mutation. Pinned bytest/serve-stdio-lifecycle.test.ts.src/mcp/dispatch.ts(v0.22.7) — Shared tool-call dispatch consumed by both stdio (server.ts) and HTTP transports. ExportsdispatchToolCall(engine, name, params, opts),buildOperationContext(engine, params, opts), andvalidateParams(op, params). Single source of truth for(ctx, params)handler arg order and the 5-fieldOperationContextshape (engine + config + logger + dryRun + remote). Defaults toremote: true(untrusted); local CLI callers passremote: false. Closed F1/F2/F3 drift bugs in the original v0.22.5 HTTP transport. v0.26.9 (F8): addssummarizeMcpParams(opName, params)— privacy-preserving redactor formcp_request_logand the admin SSE feed. Returns{redacted, kind, declared_keys, unknown_key_count, approx_bytes}. Intersects submitted top-level keys against the operation's declaredparamsallow-list (declared keys preserved as a sorted array for debug visibility; unknown keys counted but never named, closing the attacker-controlled-key-name leak). Byte counts bucketed up to nearest 1KB so an attacker can't binary-search secret-content sizes via repeated probes. Operators on a personal laptop who want raw payload visibility opt back in withgbrain serve --http --log-full-params(loud stderr warning at startup). Canonical helper — new logging code paths route through it rather thanJSON.stringify(params).src/mcp/rate-limit.ts(v0.22.7) — Bounded-LRU token-bucket limiter.buildDefaultLimiters()returns the two-bucket pipeline: pre-auth IP (30/60s, fires BEFORE the DB lookup so brute-force load againstaccess_tokensis actually capped) + post-auth token-id (60/60s). TrackslastTouchedMsseparately fromlastRefillMsso an exhausted key can't be reset by hammering past the TTL. LRU cap bounds memory under attacker-controlled key growth.src/commands/serve-http.ts(v0.26.0) — Express 5 HTTP MCP server with OAuth 2.1, admin dashboard, and SSE live activity feed. Started viagbrain serve --http [--port N] [--token-ttl N] [--enable-dcr] [--public-url URL] [--log-full-params]. Supersedes the v0.22.7src/mcp/http-transport.tssimple bearer-auth path. Combines MCP SDK'smcpAuthRouter(authorize / token / register / revoke endpoints), a customclient_credentialshandler (SDK's token endpoint throwsUnsupportedGrantTypeErrorfor CC; the custom handler runs BEFORE the router and falls through forauth_code/refresh_token),requireBearerAuthmiddleware for/mcpwith scope enforcement before op dispatch,localOnlyrejection, andexpress-rate-limitat 50 req / 15 min on/token. Serves the built admin SPA fromadmin/dist/with SPA fallback./admin/eventsSSE endpoint broadcasts every MCP request to connected admin browsers.cookie-parsermiddleware wired (Express 5 has no built-in). Startup logging prints port, engine, configured issuer URL (honors--public-url), registered-client count, DCR status, and admin bootstrap token. v0.26.9 hardening pass: F7 setsremote: trueexplicitly on the/mcprequest handler's OperationContext literal (closes the HTTP shell-job RCE — without this,submit_job's protected-name guard atoperations.ts:1391saw a falsy undefined and skipped, letting aread+write-scoped OAuth token submitshelljobs). F8 wiressummarizeMcpParamsfromsrc/mcp/dispatch.tsinto bothmcp_request_logwrites and the admin SSE feed by default (raw payloads opt-in via--log-full-paramswith stderr warning). F9 sets cookieSecureflag when behind HTTPS or a public-URL proxy. F10 caps the magic-link nonce store with an LRU bound. F12 routes DCR disable through theGBrainOAuthProviderconstructor'sdcrDisabledoption instead of the prior monkey-patch on the express router. F14 wrapstransport.handleRequestin try/catch so SDK throws return a JSON-RPC 500 envelope instead of express's default HTML error page. F15 unifies OperationError + unexpected exceptions throughbuildError/serializeErrorso/mcpalways returns the same envelope shape. v0.28.1:/healthendpoint extracted into pureprobeHealth(engine)async function withHEALTH_TIMEOUT_MS = 3000exported constant — drops the timeout from 5s to 3s so Fly.io's 5s health-check deadline gets 2s of headroom for TCP, response framing, and clock skew. Racesengine.getStats()against the timeout viaPromise.race; saturated pool returns 503 withHealth check timed out (database pool may be saturated)instead of hanging.clearTimeoutin finally block prevents pending-timer pile-up under high probe rates (race-leak fix from adversarial review). v0.28.10:/healthis now liveness-only via the newprobeLiveness(sql, engineName, version, timeoutMs)helper that racessql\SELECT 1`againstHEALTH_TIMEOUT_MSand returns the sameProbeHealthResulttagged-union asprobeHealth(single timer-cleanup site, single 503 envelope). Body shape:{status, version, engine}only — engine stats are no longer spread on the public route. Full stats moved to a new admin endpoint/admin/api/full-stats(sibling to/admin/api/statsand/admin/api/health-indicators) gated by the existingrequireAdminmiddleware; that route callsprobeHealth(engine, ...)and returns the original spread-stats body.?full=truequery param removed entirely. Closes the original DoS surface wheregetStats()'s 6× count(*) on 96K-page brains through PgBouncer exceededHEALTH_TIMEOUT_MSand triggered orchestrator restart cascades (Fly.io / k8s seeing 503 → restart loop → advisory-lock pile-up on the migration lock). Outside-voice review (Codex) caught that/admin/api/health-indicatorsis NOT a full-stats endpoint (returns only{expiring_soon, error_rate}), and that an alternative loopback-IP gate would have depended onapp.set('trust proxy', 'loopback')semantics holding under proxy/XFF misconfiguration; the shipped admin-cookie design avoids both. **v0.31.3 (#681):** every OAuth/admin/audit SQL call routes throughsqlQueryForEngine(engine)fromsrc/core/sql-query.tssogbrain serve --httpworks against PGLite brains. The fourmcp_request_log.paramsINSERT sites (success path, auth_failed path, scope_denied path, server-error path) all go throughexecuteRawJsonb(engine, ...)so the JSONB column stores real objects, not JSON-encoded strings — closes the bug whereparams->>'op'returned the encoded string"search"(with quotes) instead ofsearch. Migration v46 normalizes any pre-v0.31.3 string-shaped backlog rows on first start. **v0.34.1.0 (#864):** new--bind HOSTCLI flag with default127.0.0.1. Personal-laptop installs no longer publish the brain to the LAN by accident. Self-hosted operators pass--bind 0.0.0.0(or a specific interface IP) once to accept remote connections. A stderr WARN fires when--public-urlis set without--bindso the operator sees the binding before the first request (common cause of "ngrok forwards to me but the agent can't reach the upstream" misconfigurations). The startup banner prints aBind:line. **v0.34.1.0 (#861):** drops the(authInfo as AuthInfo & {sourceId?: string}).sourceId ?? env ?? 'default'cast chain —AuthInfo.sourceIdandAuthInfo.allowedSourcesare now the typed source of truth, populated byoauth-provider.ts:verifyAccessTokenfrom theoauth_clientsrow. **v0.35.3.0:** the inline ParamDef→schema mapper at:837-849(HTTP MCPtools/listhandler) is replaced withparamDefToSchema(v)fromsrc/mcp/tool-defs.ts. Pre-fix this site silently droppeditemson every array param so strict-mode OAuth clients (Gemini Pro structured outputs, OpenAI strict tool defs) rejected the whole tool list. Single mapper now serves stdio MCP, HTTP MCPtools/list`, and the subagent registry.src/core/sql-query.ts(v0.31.3) — Engine-aware tagged-template SQL adapter for OAuth/admin/auth infrastructure.sqlQueryForEngine(engine)returns aSqlQuery((strings, ...values) => Promise<rows[]>) that walks the template, builds$Npositional SQL, asserts every value is aSqlValue(string | number | bigint | boolean | Date | null), and routes throughengine.executeRaw(sql, params)so Postgres goes via postgres.js'sunsafe(sql, params)path and PGLite via its embeddeddb.query(sql, params). Deliberately narrower than postgres.js'ssqltag: no nested fragments, nosql.json(), nosql.unsafe(), nosql.begin(), no array binding. The narrow surface is the feature — codex finding #7 from the v0.31 plan review argued the adapter should stay scalar-only or it drifts into a partial postgres.js clone. JSONB writes go through the separateexecuteRawJsonb(engine, sql, scalarParams, jsonbParams)helper that composes positional$N::jsonbcasts and passes JS objects through; the v0.12.0 double-encode bug class doesn't apply because positional binding throughunsafe()reaches the wire protocol with the correct type oid (verified bytest/sql-query.test.tson PGLite andtest/e2e/auth-permissions.test.ts:67on Postgres).scripts/check-jsonb-pattern.shdoesn't fire becauseexecuteRawJsonb(...)is a method call, not the banned literal-template-tag interpolation pattern. Consumed bysrc/commands/auth.ts,src/commands/serve-http.ts,src/core/oauth-provider.ts,src/commands/files.ts, andsrc/mcp/http-transport.tsso all five sites work against PGLite and Postgres uniformly. Closes the bug wheregbrain auth+gbrain serve --httpwere silently Postgres-only because they routed every SQL through the postgres.js singleton (community PR #681).src/commands/serve.ts(v0.31.3) —gbrain servestdio MCP entrypoint with idempotent shutdown across every parent-disconnect signal. Stdio EOF, SIGTERM, SIGINT, SIGHUP, and parent-process death (every reparent case — PID 1, launchd subreaper, systemd, tmux, or a parent shell withPR_SET_CHILD_SUBREAPER) all funnel into onecleanup(reason)path that releases the engine and the PGLite write-lock dir within 5 seconds. Pre-v0.31.3 the stdio MCP server held the lock indefinitely after Claude Desktop / Cursor / launchd-managed gateways disconnected, forcing a 5-minute stale-lock wait on the next start. Watchdog reparent check isgetParentPid() !== initialParentPid(capturing the initial ppid once at install time and firing on any change); the previous=== 1check missed the subreaper case under launchd / systemd. Bun'sprocess.ppidcache is stale across reparenting (see oven-sh/bun#30305) sogetParentPid()runsspawnSync('ps', ['-o', 'ppid=', '-p', PID])per tick to read the live kernel PPID. Startup probe verifiespsis on PATH; if not (stripped containers, busybox without procps), the watchdog skips installing AND emits a loud[gbrain serve] watchdog disabled: ps unavailable, parent-death detection unavailable — child will rely on stdin EOF / signals onlystderr line so operators see the degraded mode at boot. Pinned bytest/serve-stdio-lifecycle.test.ts(22 cases). Closes #413, #446. Credit @Aragorn2046 (origin features in #591) and @seungsu-kr (rebased submitter, Bun ppid workaround).src/core/oauth-provider.ts(v0.26.0) —GBrainOAuthProviderimplementing the MCP SDK'sOAuthServerProvider+OAuthRegisteredClientsStoreinterfaces. Backed by raw SQL (works on both PGLite and Postgres — OAuth is infrastructure, not a BrainEngine concern). Full OAuth 2.1 spec:authorize+exchangeAuthorizationCodewith PKCE (for ChatGPT),client_credentials(for Perplexity / Claude),refresh_tokenwith rotation,revokeToken,registerClient(DCR path validates redirect_uri must behttps://or loopback per RFC 6749 §3.1.2.1). All tokens + client secrets SHA-256 hashed before storage. Auth codes single-use with 10-minute TTL via atomicDELETE...RETURNING(closes RFC 6749 §10.5 TOCTOU race). Refresh rotation alsoDELETE...RETURNING(closes §10.4 stolen-token detection bypass).pgArray()escapes commas/quotes/braces in elements so a comma-bearing redirect_uri can't smuggle a second array element. Legacyaccess_tokensfallback inverifyAccessTokengrandfathers pre-v0.26 bearer tokens asread+write+admin.sweepExpiredTokens()runs on startup wrapped in try/catch. v0.26.9 RFC 6749/7009 hardening pass: F1+F2 foldclient_idatomically into theDELETE WHEREclauses for both auth-code exchange and refresh rotation — pre-fix the post-hoc client compare burned the row on wrong-client paths so the legitimate client couldn't retry. F3 enforces refresh-scope-subset against the original grant on the row (RFC 6749 §6), not the client's currently-allowed scopes — fixes the case where revoking a scope from a client wouldn't shrink the agent's existing refresh tokens. F4 bindsclient_idonrevokeTokenso a client can only revoke its own tokens (RFC 7009 §2.1). F7c validates the/tokenrequest'sredirect_uriagainst the value stored at/authorize(RFC 6749 §4.1.3) — empty-string treated as missing rather than wildcard match (adversarial-review fix). F5 swaps barecatch {}blocks inverifyAccessTokenandgetClientforisUndefinedColumnErrorfromsrc/core/utils.ts— only SQLSTATE 42703 falls through to legacy fallback; lock timeouts and network blips throw and surface. F6 makessweepExpiredTokens()actually return the count viaRETURNING 1+ array length, not a fire-and-forget zero. F12 addsdcrDisabledconstructor option soserve-http.tscan disable the/registerendpoint without monkey-patching the router. v0.26.2: module-privatecoerceTimestamp()boundary helper at the top of the file normalizes postgres-driver-as-string BIGINT columns to JS numbers at every read site (5 call sites:getClientL112+L113 for DCR/registerRFC 7591 §3.2.1 numeric timestamps,exchangeRefreshTokenL274 +verifyAccessTokenL296+L303 for the SDK'stypeof === 'number'bearerAuth check). Throws on non-finite input (NaN/Infinity) so corrupt rows fail loud at the boundary instead of riding through asexpiresAt: NaN; returns undefined for SQL NULL so callers decide NULL semantics explicitly (refresh + access token paths treat NULL as expired). Helper intentionally NOT promoted tosrc/core/utils.ts— codex review flagged repo-wide BIGINT precision-loss risk for a generic helper. v0.34.1.0 (#909):registerClienthonorstoken_endpoint_auth_method: "none"(RFC 7591 §3.2.1) — public PKCE clients (Claude Code, Cursor, every other PKCE-first MCP client) storeclient_secret_hash = NULLand the response payload omitsclient_secretentirely. Confidential clients (defaultclient_secret_postand explicitclient_secret_basic) keep their one-time-reveal shape.getClientcorrectly normalizes a NULLclient_secret_hashto JSundefinedso the SDK's clientAuth path accepts the public client at/token. v0.34.1.0 (#861 + #876):verifyAccessTokenJOINsoauth_clients.source_id(write scope, scalar) +oauth_clients.federated_read(read scope, TEXT[]) and surfaces both on the returnedAuthInfo. Pre-v60 / pre-v61 brains degrade gracefully viaisUndefinedColumnErrorfallback so the upgrade chain is non-blocking on legacy DBs.admin/(v0.26.0) — React 19 + Vite + TypeScript admin SPA embedded in the binary viaadmin/dist/served byserve-http.ts. 7 screens: Login (bootstrap token → session cookie), Dashboard (metrics + SSE feed + token health), Agents (sortable table + sparklines + Register button), Register (modal with scope checkboxes + grant type selector), Credentials reveal (full-screen modal with Copy + Download JSON + yellow one-time-only warning), Request Log (filterable paginated), Agent Detail drawer (Details / Activity / Config Export tabs + Revoke). Design tokens:#0a0a0fbg, Inter for UI, JetBrains Mono for data, 4-32px spacing scale, rounded pill badges. HTTP-only SameSite=Strict cookie auth. 65KB gzip. Build:cd admin && bun install && bun run build; output atadmin/dist/is committed for self-contained binaries.src/commands/auth.ts— Token management.gbrain auth create/list/revoke/testfor legacy bearer tokens (v0.22.7 wired as a first-class CLI subcommand) plusgbrain auth register-client(v0.26.0) andgbrain auth revoke-client <client_id>(v0.26.2) for OAuth 2.1 client lifecycle.revoke-clientruns an atomicDELETE...RETURNINGonoauth_clients; FKON DELETE CASCADEonoauth_tokens.client_idandoauth_codes.client_idpurges every active token + authorization code in a single transaction.process.exit(1)on no-such-client (idempotent — re-running on the same id produces the same exit-1 message). Legacy tokens stored as SHA-256 hashes inaccess_tokens; OAuth clients inoauth_clients. As of v0.26.0, legacy tokens grandfather toread+write+adminscopes on the OAuth HTTP server, so pre-v0.26 deployments keep working with no migration. v0.31.3 (#681): every SQL site routes throughsqlQueryForEngine(engine)fromsrc/core/sql-query.ts(andexecuteRawJsonbfor the takes-holderspermissionsJSONB column) sogbrain authworks against PGLite brains. Pre-fix, every call hit the postgres.js singleton viagetConn()and silently failed (or wrote to the wrong DB) when the active engine was PGLite. The takes-holders write goes throughexecuteRawJsonb(engine, sql, [name, hash], [{takes_holders:[...]}])which round-trips withjsonb_typeof = 'object'instead of the pre-v0.31.3 quoted-string shape. v0.34.1.0 (#876):register-clientaccepts--source <id>(write authority, scalar) and--federated-read <S1,S2,...>(read scope, array). The output prints the resolvedWrite sourceandFederated readsfor the registered client. Pre-v0.34 clients backfill tosource_id='default'via migration v60 so existing deployments keep their v0.33 effective behavior verbatim.src/commands/upgrade.ts— Self-update CLI.runPostUpgrade()enumerates migrations from the TS registry (src/commands/migrations/index.ts) and tail-callsrunApplyMigrations(['--yes', '--non-interactive'])so the mechanical side of every outstanding migration runs unconditionally.src/commands/migrations/— TS migration registry (compiled into the binary; no filesystem walk ofskills/migrations/*.mdneeded at runtime).index.tslists migrations in semver order.v0_11_0.ts= Minions adoption orchestrator (8 phases).v0_12_0.ts= Knowledge Graph auto-wire orchestrator (5 phases: schema → config check → backfill links → backfill timeline → verify).phaseASchemahas a 600s timeout (bumped from 60s in v0.12.1 for duplicate-heavy brains).v0_12_2.ts= JSONB double-encode repair orchestrator (4 phases: schema → repair-jsonb → verify → record).v0_14_0.ts= shell-jobs + autopilot cooperative (2 phases: schema ALTER minion_jobs.max_stalled SET DEFAULT 3 — superseded by v0.14.3's schema-level DEFAULT 5 + UPDATE backfill; pending-host-work ping for skills/migrations/v0.14.0.md). All orchestrators are idempotent and resumable frompartialstatus. As of v0.14.2 (Bug 3), the RUNNER owns all ledger writes — orchestrators returnOrchestratorResultandapply-migrations.tspersists a canonical{version, status, phases}shape after return. Orchestrators no longer callappendCompletedMigrationdirectly.statusForVersionpreferscompleteoverpartial(never regresses). 3 consecutive partials → wedged →--force-retry <version>writes a'retry'reset marker. v0.14.3 (fix wave) ships schema-only migrations v14 (pages_updated_at_index) + v15 (minion_jobs_max_stalled_default_5with UPDATE backfill) via theMIGRATIONSarray insrc/core/migrate.ts— no orchestrator phases needed.src/commands/repair-jsonb.ts—gbrain repair-jsonb [--dry-run] [--json]: rewritesjsonb_typeof='string'rows in place across 5 affected columns (pages.frontmatter, raw_data.data, ingest_log.pages_updated, files.metadata, page_versions.frontmatter). Fixes v0.12.0 double-encode bug on Postgres; PGLite no-ops. Idempotent.src/commands/orphans.ts—gbrain orphans [--json] [--count] [--include-pseudo]: surfaces pages with zero inbound wikilinks, grouped by domain. Auto-generated/raw/pseudo pages filtered by default. Also exposed asfind_orphansMCP operation. Shipped in v0.12.3 (contributed by @knee5).src/commands/salience.ts(v0.29) —gbrain salience [--days N] [--limit N] [--kind PREFIX] [--json]: pages ranked by emotional + activity salience over a recency window. Mirrors orphans.ts shape (pure data fn + JSON formatter + human formatter). Callsengine.getRecentSalience(opts). Score formula:(emotional_weight × 5) + ln(1 + active_take_count) + 1/(1 + days_since_update).src/commands/anomalies.ts(v0.29) —gbrain anomalies [--since YYYY-MM-DD] [--lookback-days N] [--sigma N] [--json]: cohort-level activity outliers. Callsengine.findAnomalies(opts). Two cohort kinds in v1: tag, type. Year cohort deferred to v0.30.src/commands/whoknows.ts(v0.33) —gbrain whoknows <topic> [--explain] [--limit N] [--json]: expertise + relationship-proximity routing. Mirrors v0.29 salience/anomalies shape (purerankCandidates()+findExperts()orchestrator +runWhoknows()CLI dispatch + thin-client routing). MCP op =find_experts(scope: read, localOnly: false) per ENG-D5. Ranking formula (ENG-D1 locked):score = log(1 + raw_match) × max(0.1, exp(-days/180)) × (0.5 + 0.5 × salience)whereraw_matchis hybridSearch's RRF+source-boost score. Filters at SQL via the newSearchOpts.types: ['person', 'company'](no post-filter waste). hybridSearch's internal salience+recency boosts are intentionally disabled — the locked formula applies on a clean signal. Floors prevent multiplicative-zero edge cases (cold-start people stay visible); ties break alphabetically by slug for determinism. 16 unit tests intest/whoknows.test.tspin the math.src/commands/eval-whoknows.ts(v0.33, v0.33.1.3 thin-client wiring) —gbrain eval whoknows <fixture.jsonl> [--json] [--skip-replay]: two-layer eval gate (ENG-D2). Layer 1 quality (hand-labeled fixture, top-3 hit rate ≥ 0.8). Layer 2 regression (eval_candidatesreplay set-Jaccard@3 ≥ 0.4). Sparseness fallback: < 20 replay-eligible rows → Layer 2 auto-skips with stderr warning. Stable JSON envelope withschema_version: 1. Exit 0/1/2 for pass/fail/usage so CI can gate. Mirrors v0.27.x cross-modal + v0.28.1 longmemeval dispatch shape undersrc/commands/eval.ts. v0.33.1.3:WhoknowsFncallable abstraction lets the gates be impl-agnostic.runEvalWhoknows(engine: BrainEngine | null, args)picks the impl at entry — thin-client mode (isThinClient(cfg)) routes per-query throughcallRemoteTool(cfg, 'find_experts', {topic, limit})via the v0.31.1 seam; local mode callsfindExperts(engine, ...)directly. cli.ts adds a thin-client bypass beforeconnectEngineforgbrain eval whoknows, matching the longmemeval/cross-modal no-DB pattern. Regression gate auto-skips in thin-client mode (no DB access toeval_candidates). Public exportsjaccardAtK,topKHit,readFixture,WhoknowsFn, threshold constants are pinned bytest/eval-whoknows.test.ts(25 cases, +2 for the null-engine signature contract).test/fixtures/whoknows-eval.jsonl(v0.33) — 10-row synthetic placeholder demonstrating the eval-fixture schema ({query, expected_top_3_slugs, notes?}JSONL). End users replace with their own real queries before shipping; the placeholder uses obviously-example slugs (wiki/people/example-alice) so production data isn't conflated with the test fixture. Drivestest/e2e/whoknows.test.ts(which seeds a matching synthetic brain and asserts the >=80% gate) and thewhoknows_healthdoctor check.src/commands/transcripts.ts(v0.29) —gbrain transcripts recent [--days N] [--full] [--json]: recent raw.txttranscripts from the dream-cycle corpus dirs. ImportslistRecentTranscriptsfromsrc/core/transcripts.ts(the same library the gatedget_recent_transcriptsMCP op uses). Local-only by construction — the CLI always runs withctx.remote=false.src/commands/integrity.ts—gbrain integrity check|auto|review|extract: bare-tweet detection, dead-link detection, three-bucket repair (auto-repair / review-queue / skip).scanIntegrity()is the shared library function called fromgbrain doctor(sampled at limit=500) andcmdCheck(full scan). v0.22.8: batch-load fast path on Postgres uses a single SQL query to fix the PgBouncer round-trip timeout (60s → ~6s). Gated byengine.kind === 'postgres'at the call site so PGLite never enters batch; fallbackcatchlogs atGBRAIN_DEBUG=1so real Postgres errors are diagnosable. v0.32.8 (PR #860): batch projection switched fromSELECT DISTINCT ON (slug)toSELECT ... ORDER BY source_id, slugso multi-source brains scan each(source, slug)row independently (pre-fix the DISTINCT collapsed same-slug-different-source pages into one scan, the same bug class this PR fixes). Sequential and auto-repair loops uselistAllPageRefs()to enumerate(slug, source_id)pairs and threadsourceIdtogetPage. Batch + sequential paths now report the same page count on multi-source brains.src/commands/doctor.ts—gbrain doctor [--json] [--fast] [--fix] [--dry-run] [--index-audit]: health checks. v0.12.3 addedjsonb_integrity+markdown_body_completenessreliability checks. v0.14.1:--fixdelegates inlined cross-cutting rules to> **Convention:** see [path](path).callouts (pipes DRY violations intosrc/core/dry-fix.ts);--fix --dry-runpreviews without writing. v0.14.2:schema_versioncheck fails loudly whenversion=0(migrations never ran — the #218bun install -gsignature) and routes users togbrain apply-migrations --yes; new opt-in--index-auditflag (Postgres-only) reports zero-scan indexes frompg_stat_user_indexes(informational only, no auto-drop). v0.15.2: every DB check is wrapped in a progress phase;markdown_body_completenessruns under a 1s heartbeat timer so 10+ min scans are observable on 50K-page brains. v0.19.1 addedqueue_health(Postgres-only) with two subchecks: stalled-forever active jobs (started_at > 1h) and waiting-depth-per-name > threshold (default 10, override viaGBRAIN_QUEUE_WAITING_THRESHOLD). Worker-heartbeat subcheck intentionally deferred to follow-up B7 because it needs aminion_workerstable to produce ground-truth signal. Fix hints point atgbrain repair-jsonb,gbrain sync --force,gbrain apply-migrations, andgbrain jobs get/cancel <id>. v0.22.12 (#500):sync_failurescheck shows[CODE=N, ...]breakdown for both unacked entries (warn) and acked-historical entries (ok), surfacing systemic failure modes (SLUG_MISMATCH=2685) instead of a bare count. v0.26.7 (#612):rls_event_triggercheck (post-install drift detector for migration v35's auto-RLS event trigger). Lives outside the// 5. RLSslice that the structural doctor.test.ts guards anchor on, so the existing test guards stay intact. Healthyevtenabledset is('O','A')only —Ris replica-only and would not fire in normal sessions;Dis disabled. Fix hint isgbrain apply-migrations --force-retry 35. v0.30.2:queue_healthgains a fourth subcheck — surfaces dead-lettered subagent jobs withlast_errormatching theprompt_too_longclassifier within the last 24h. Fix hint points atgbrain dream --phase synthesize --dry-run --jsonto identify the offending transcript andgbrain jobs prune --status dead --queue defaultto clean up. Postgres-only. v0.31.7:runDoctorswitches toautoDetectSkillsDirReadOnly(fromsrc/core/repo-root.ts) sobun install -g github:garrytan/gbrain && cd ~ && gbrain doctorfinds the bundledskills/via the install-path fallback instead of warning "Could not find skills directory" + docking the health score.--fixcarries a D6 safety gate: whendetected.source === 'install_path', the command refuses auto-repair with a stderr message pointing at$GBRAIN_SKILLS_DIR/$OPENCLAW_WORKSPACE/--skills-dir, becauseautoFixDryViolationswrites to SKILL.md files and would otherwise silently rewrite the install tree. Thegraph_coveragecheck now short-circuits took: 'No entity pages — graph_coverage not applicable (markdown-only brain)'whenSELECT COUNT(*) FROM pages WHERE type IN ('entity','person','company','organization')returns 0 (closes #530); the entity count is woven into the warn message and the WARN hint switches from the long-deprecatedgbrain link-extract && gbrain timeline-extract(gone since v0.16) to the canonicalgbrain extract all. Pinned by an IRON-RULE regression assertion intest/doctor.test.tsthat bans the stale verb names from the source string. v0.32.4: newsync_freshnesscheck (exportedcheckSyncFreshnessat the same file) added to bothrunDoctor(local) anddoctorReportRemote(thin-client). Pure staleness probe — queriessources.last_sync_atonly, no filesystem access. Warns at 24h, fails at 72h (or never-synced). Future-last_sync_atwarns ("clock skew or corrupted timestamp") instead of silently falling through as ok — codex outside-voice caught the negative-ageMs bug pre-merge. Env-var overridesGBRAIN_SYNC_FRESHNESS_WARN_HOURS/GBRAIN_SYNC_FRESHNESS_FAIL_HOURS; invalid values fall back to defaults with a once-per-process stderr warn (_resolveSyncFreshnessHours). Failure messages embedsource.id(notsource.name) so the printed fix commandgbrain sync --source <id>matches what the user copy-pastes. Filesystem-vs-DB page drift detection was deliberately stripped from the v0.32.4 scope —doctorReportRemoteruns in the HTTP MCP server (src/commands/serve-http.ts), and walking DB-suppliedlocal_pathfrom a remote-callable endpoint crosses a trust boundary (OAuth write scope could mutatesources.local_path). Drift detection will resurface in a separate PR routed throughmulti_source_drift's existing guard infrastructure (GBRAIN_DRIFT_LIMIT/GBRAIN_DRIFT_TIMEOUT_MS) with slug normalization tests and a meta-file allow-list. Pinned by 12 cases intest/doctor.test.ts("v0.32.4 — sync_freshness check" describe block): empty sources, never-synced fail, >72h fail, exact 72h boundary, 24h-72h warn, exact 24h boundary, <24h ok, future-timestamp warn, mixed sources (highest severity wins),executeRawthrows → outer-catch warn,GBRAIN_SYNC_FRESHNESS_FAIL_HOURS=6override fires at 7h, source.id-in-message regression. v0.35.5.0: the Lane D supervisor check atdoctor.ts:1011-1043now consumessummarizeCrashes(events)fromsrc/core/minions/handlers/supervisor-audit.tsinstead of the pre-fixevents.filter(e => e.event === 'worker_exited').length. The warn threshold drops from>3to>=1(any real crash is signal now that the counter is calibrated against clean exits). The ok message gainsclean_exits_24h=N; the warn message gainsruntime=A oom=B unknown=C legacy=Dper-cause breakdown so an operator triages OOM vs runtime-error vs unknown-future-cause at a glance without grep'ing the JSONL audit. Closes the "Supervisor crashes: 120x/24h, was 62x — nearly doubled" alarm class that bit users on healthy brains after v0.34.3.0's RSS-watchdog work added more code=0 worker drains — bothdoctorandgbrain jobs supervisor statuswere counting everyworker_exitedevent as a crash regardless of cause. Cross-surface parity is the regression guard: 4 source-grep wiring assertions intest/doctor.test.tsban the ad-hoc filter pattern, pin the>=1threshold, and require the per-cause breakdown substrings (runtime=,oom=,unknown=,legacy=,clean_exits_24h=,crashes_by_cause) to appear in BOTHdoctor.tsandjobs.ts.src/core/migrate.ts— schema-migration runner. Owns theMIGRATIONSarray (source of truth for schema DDL). v40 (v0.29):pages_emotional_weightaddspages.emotional_weight REAL NOT NULL DEFAULT 0.0. Column-only (no index). On Postgres 11+ and PGLite,ADD COLUMNwith a constant DEFAULT is metadata-only — instant on tables of any size. v0.14.2 extended theMigrationinterface withsqlFor?: { postgres?, pglite? }(engine-specific SQL overridessql) andtransaction?: boolean(set to false forCREATE INDEX CONCURRENTLY, which Postgres refuses inside a transaction; ignored on PGLite since it has no concurrent writers). Migration v14 (fix wave) uses a handler branching onengine.kindto run CONCURRENTLY on Postgres (with a pre-drop of any invalid remnant viapg_index.indisvalid) and plainCREATE INDEXon PGLite. v15 bumpsminion_jobs.max_stalleddefault 1→5 and backfills existing non-terminal rows. v0.22.6.1: migration v24 (rls_backfill_missing_tables) usessqlFor: { pglite: '' }to no-op on PGLite — PGLite has no RLS engine and is single-tenant by definition, and the v24 ALTERs target subagent tables that don't exist in pglite-schema.ts. Closes #395 (contributed by @jdcastro2). v30 (v0.23): createsdream_verdicts (file_path TEXT, content_hash TEXT, worth_processing BOOL, reasons JSONB, judged_at TIMESTAMPTZ, PK(file_path, content_hash)). RLS-enabled when running as a BYPASSRLS role. The synthesize phase reads/writes this table to avoid re-judging on backfill re-runs. v35 (v0.26.7): auto-RLS event trigger + one-time backfill.auto_rls_on_create_tablefires onddl_command_endforWHEN TAG IN ('CREATE TABLE','CREATE TABLE AS','SELECT INTO')and runsALTER TABLE … ENABLE ROW LEVEL SECURITYon every newpublic.*table — no FORCE (matches v24/v29/schema.sql posture so non-BYPASSRLS apps can still read their own tables). The same migration backfills RLS on every existingpublic.*base table whose comment doesn't match the doctor regex (^GBRAIN:RLS_EXEMPT\s+reason=\S.{3,}). Per-table failure aborts the offending CREATE TABLE (event triggers fire inside the DDL transaction); no EXCEPTION wrap — that would convert loud rollback into silent permissive default. PGLite no-op viasqlFor.pglite: ''. Breaking change: operators with intentionally-RLS-off public tables must add the GBRAIN:RLS_EXEMPT comment BEFORE upgrade or the backfill will flip them on. v46 (v0.31.3):mcp_request_log_params_jsonb_normalizerewrites pre-v0.31.3 rows wheremcp_request_log.paramswas stored as a JSON-encoded string (jsonb_typeof = 'string') up to a real JSONB object viaUPDATE ... SET params = params::text::jsonb WHERE jsonb_typeof(params) = 'string'. Single statement, idempotent — second-run finds no string-shaped rows and is a no-op. Closes the bug where/admin/api/requestsreturned a quoted string instead of the parsed object. v0.34.1.0 (#861 + #876, v60-v65): six-migration chain wires source-scoping into the OAuth client table. v60 (oauth_clients_source_id_fk) addsoauth_clients.source_id TEXTwith NULL→'default'backfill and an FK tosources(id) ON DELETE SET NULL. v61 (oauth_clients_federated_read_column) addsfederated_read TEXT[] NOT NULL DEFAULT '{}'. v62 (oauth_clients_federated_read_backfill) explicit-CASE backfills sosource_id IS NULLproduces'{}'not an array-containing-NULL. v63 (oauth_clients_federated_read_validate) is the fail-loud check that every row's source_id is in its federated_read array post-backfill. v64 (oauth_clients_source_id_fk_restrict) flips the FK toON DELETE RESTRICTnow that federated_read provides the alternative scope-loss path — source delete is refused if any client references it. v65 (oauth_clients_federated_read_gin_index) is the GIN index for the array-containment queries the read paths run. PGLite parity viasqlFor.pglitewhere needed.src/core/progress.ts— Shared bulk-action progress reporter. Writes to stderr. Modes:auto(TTY:\r-rewriting; non-TTY: plain lines),human,json(JSONL),quiet. Rate-gated byminIntervalMsandminItems.startHeartbeat(reporter, note)helper for single long queries.child()composes phase paths. Singleton SIGINT/SIGTERM coordinator emitsabortevents for every live phase. EPIPE defense on both sync throws and stream'error'events. Zero dependencies. Introduced in v0.15.2.src/core/cli-options.ts— Global CLI flag parser.parseGlobalFlags(argv)returns{cliOpts, rest}with--quiet/--progress-json/--progress-interval=<ms>stripped.getCliOptions()/setCliOptions()expose a module-level singleton so commands reach the resolved flags without parameter threading.cliOptsToProgressOptions()maps to reporter options.childGlobalFlags()returns the flag suffix to append toexecSync('gbrain ...')calls in migration orchestrators.OperationContext.cliOptsextends shared-op dispatch for MCP callers.src/core/db-lock.ts(v0.22.13) — generictryAcquireDbLock(engine, lockId, ttlMinutes)over the existinggbrain_cycle_lockstable. Parameterized lock id so different scopes can nest cleanly:gbrain-cyclefor the broad cycle (held bycycle.ts) andgbrain-sync(SYNC_LOCK_IDconstant) forperformSync's narrower writer window. Same UPSERT-with-TTL semantics as the prior cycle-only helper, just generalized. Survives PgBouncer transaction pooling (unlike session-scopedpg_try_advisory_lock); crashed holders auto-release once their TTL expires.src/core/sync-concurrency.ts(v0.22.13) — single source of truth for the parallel-sync policy. ExportsautoConcurrency(engine, fileCount, override?)(PGLite always serial; explicit override clamped to >=1; auto path returnsDEFAULT_PARALLEL_WORKERS=4whenfileCount > AUTO_CONCURRENCY_FILE_THRESHOLD=100),shouldRunParallel(workers, fileCount, explicit)(Q1: explicit--workersbypasses the >50-file floor), andparseWorkers(s)(rejects'0','-3','foo','1.5', trailing chars — replaces the prior parseInt-with-no-validation in bothsync.tsandimport.ts). Used byperformSync,performFullSync,runImport, and the Minionsynchandler so the three sites can no longer drift.src/commands/sync.ts—gbrain syncCLI + theperformSync/performFullSynclibrary entrypoints (consumed by the autopilot cycle and the Minion sync handler). v0.22.13 (PR #490):performSyncwraps its body in agbrain-syncwriter lock so two concurrent syncs (manual + autopilot, two terminals, two Conductor workspaces) cannot both writelast_commitand let the last writer win. Head-drift gate after the import phase re-checksgit rev-parse HEAD; if HEAD moved (someone rangit checkout/git pullmid-sync), the bookmark refuses to advance. Vanished files now record a failedFiles entry instead of silent-skip — the silent-skip-then-advance pathology that survived prior hardening passes is dead. Worker engines wrap in try/finally so disconnect always fires (panic-path leak fix). Both PGLite-detection sites useengine.kind === 'pglite'. CLI accepts--workers N(alias--concurrency N), validated viaparseWorkers. Explicit--workersbypasses the auto-path file-count floor; auto path defers toautoConcurrency(). Banner moved to stderr. v0.34.2.0: the inline.sort()over add/mod paths is replaced withsortNewestFirst(addsAndMods)fromsrc/core/sort-newest-first.ts, so the newest-first descending-lex policy lives in one helper shared withgbrain importinstead of drifting across two files.src/commands/import.ts—gbrain importCLI +runImportlibrary entrypoint. v0.34.2.0 replaces the prior positional-index checkpoint (processedIndex: Ninto a sorted file list) with a path-set checkpoint viasrc/core/import-checkpoint.ts. The walk still appliessortNewestFirst()for embed-cost ordering, but checkpoint correctness no longer depends on sort order. A file enterscompleted: Set<relativePath>only when itsprocessFilereturns success (including content-hash short-circuit no-ops); failed files never enter the set, so the next run retries them automatically with no manual~/.gbrain/import-checkpoint.jsondelete. Three bug classes died: parallel-import-with-slow-worker drops the slow file on crash-resume (closed — the slow file isn't incompleteduntil its ownprocessFileresolves), failed-file-bumps-counter-past-itself (closed — failures don't add tocompleted), and v0.33.x sort-flip-drops-newest-N-on-cross-version-resume (closed — order is no longer part of the checkpoint). Old positional checkpoints are detected and discarded with a stderr line on first resume; re-walking is cheap becausecontent_hashshort-circuits unchanged files. Checkpoint persists every 100 successful adds, not every 100 processed files, so a long failure tail doesn't churn the JSON. Pinned bytest/import-checkpoint.test.ts(18 unit cases over the helpers) +test/import-resume.test.ts(5 integration cases under PGLite, including the SLUG_MISMATCH retry regression codex caught during plan-eng-review).src/core/import-checkpoint.ts(v0.34.2.0) —loadCheckpoint(brainDir),saveCheckpoint(brainDir, completed),resumeFilter(files, completed, brainDir),clearCheckpoint(), plus theImportCheckpointtype. Path-set checkpoint format ({schema_version, brainDir, completed: string[]}) replaces the v0.33.x positional{processedIndex: N}format. Atomic write via.tmp+rename()so a mid-write crash never leaves a partial JSON.loadCheckpointreturnsnullon: missing file, malformed JSON, brainDir mismatch (you ran import against a different brain), and the old positional format (logged to stderr before being discarded).resumeFilterreturns{toProcess, skippedCount}— pure, no I/O, deterministic.clearCheckpointis no-op-on-missing for clean-exit cleanup. HonorsGBRAIN_HOMEviagbrainPath()so test isolation viawithEnv({GBRAIN_HOME: tmpdir})works without monkey-patching the fs layer. Best-effort persistence —saveCheckpointlogs warnings on write errors but never throws, so import keeps making progress even if disk is full.src/core/sort-newest-first.ts(v0.34.2.0) — single source of truth for the descending-lex sort thatgbrain importandgbrain syncboth apply. Mutates in place (Array.prototype.sort semantics), returns the same array reference for fluent chaining. Empty/single-element inputs short-circuit. Future ordering changes flip one line in this helper instead of touching two CLI commands. Pinned bytest/sort-newest-first.test.ts(5 hermetic cases: descending order, mixed prefixes, empty input, single-element input, in-place-mutation contract).src/core/cycle.ts— v0.17 brain maintenance cycle primitive (extended to 9 phases in v0.29).runCycle(engine: BrainEngine | null, opts: CycleOpts): Promise<CycleReport>composes phases in semantically-driven order: lint → backlinks → sync → synthesize → extract → patterns → recompute_emotional_weight → embed → orphans. v0.29 adds therecompute_emotional_weightphase between patterns and embed; it sees the union ofsyncPagesAffected+synthesizeWrittenSlugsfor incremental mode, or all pages when neither anchor is set (full backfill viagbrain dream --phase recompute_emotional_weight). v0.29 also extendsCycleReport.totalswithpages_emotional_weight_recomputed(additive, schema_version stays "1"). v0.23'ssynthesizephase runs after sync (cross-references see fresh brain) and before extract (auto-link materializes its writes);patternsruns after extract so it reads a fresh graph (codex finding #7 — subagent put_page setsctx.remote=trueand skips auto-link/timeline by default; extract is the canonical materialization). Three callers:gbrain dreamCLI,gbrain autopilotdaemon's inline path, and the Minionsautopilot-cyclehandler. Coordination viagbrain_cycle_locksDB table +~/.gbrain/cycle.lockfile lock with PID-liveness for PGLite.CycleReport.schema_version: "1"is stable; totals additively grew in v0.23 (transcripts_processed,synth_pages_written,patterns_written).yieldBetweenPhasesruns between phases. v0.23 addedyieldDuringPhasefor in-phase keepalive — synthesize/patterns call it during long waits to renew the cycle-lock TTL. Engine nullable; lock-skip on read-only phase selections. v0.22.1 (#403):CycleOpts.signal?: AbortSignalpropagates the worker's abort signal;checkAborted()fires between every phase. v0.22.1 (#417):runPhaseSyncreturnspagesAffectedviaSyncPhaseResult;runCyclecaptures it and threads torunPhaseExtractas the 4th arg. v0.22.1 (Codex F2):runPhaseSynctakeswillRunExtractPhase: booleanand setsnoExtract: phases.includes('extract')sogbrain dream --phase syncdoesn't silently lose extraction. v0.22.5 (#475):resolveSourceForDir(engine, brainDir)threadssourceIdtoperformSync()so sync reads the per-sourcesources.last_commitanchor instead of the drift-prone globalconfig.sync.last_commitkey.src/core/cycle/synthesize.ts(v0.23) — Synthesize phase: conversation-transcript-to-brain pipeline. Reads fromdream.synthesize.session_corpus_dir, runs cheap Haiku verdict (cached indream_verdicts), then fans out one Sonnet subagent per worth-processing transcript withallowed_slug_prefixes(sourced fromskills/_brain-filing-rules.jsondream_synthesize_paths.globs). Orchestrator collects slugs fromsubagent_tool_executions(NOTpages.updated_at— codex finding #2) and reverse-renders DB → markdown viaserializeMarkdown. Cooldown viadream.synthesize.last_completion_ts, written ONLY on success. Idempotency keydream:synth:<file_path>:<content_hash>. Auto-commit deferred to v1.1 (codex #5).--dry-runruns Haiku, skips Sonnet (codex #8). Subagent never gets fs-write access. v0.23.2:renderPageToMarkdown(now exported) stampsdream_generated: trueanddream_cycle_dateinto every reverse-write's frontmatter;writeSummaryPagedoes the same on the dream-cycle summary index. The marker is the explicit identity surface checked byisDreamOutputintranscript-discovery.ts— replaces the v0.23.1 content-prefix heuristic that could miss real output (serializeMarkdowndoesn't embed slugs in body) and false-positive on user transcripts citing brain pages.judgeSignificanceandJudgeClientare exported;judgeSignificanceaccepts averdictModelparameter (defaultclaude-haiku-4-5-20251001) loaded fromdream.synthesize.verdict_modelvialoadSynthConfig. v0.30.2: model-aware chunkersplitTranscriptByBudget(content, contentHash, maxChars)splits oversized transcripts at paragraph boundaries (## Topic:→---→\nladder) using a deterministic offset seeded from the first 32 bits ofcontentHashso retries chunk identically. Per-chunk char budget computed fromMODEL_CONTEXT_TOKENS[resolvedModel] × 0.9 × 3.5 chars/token; non-Anthropic ids fall back to a 180K-token safe default with a once-per-process stderr warning. Operator overrides:dream.synthesize.max_prompt_tokens(floor 100K, wins when set) anddream.synthesize.max_chunks_per_transcript(default 24). Per-chunk idempotency keysdream:synth:<filePath>:<hash16>:c<i>of<n>; single-chunk transcripts preserve the legacydream:synth:<filePath>:<hash16>key byte-for-byte (D8 lookup), so existing brains skip withalready_synthesized_legacy_single_chunkinstead of re-spending Sonnet on upgrade.collectChildPutPageSlugsraw-fetches every (job_id, slug) pair (notSELECT DISTINCT) and rewrites bare-hash6 slugs to<hash6>-c<idx>for chunked children (D6 — orchestrator-side, zero Sonnet trust). Cap-hit skips don't write todream_verdicts, so raising the cap on next run re-attempts cleanly. D7 scope: bounds INITIAL prompt size only; tool-loop turn-N accumulation is caught by the v0.30.2 terminal-error classification insubagent.ts, not bounded ahead of time.src/core/cycle/patterns.ts(v0.23) — Patterns phase: cross-session theme detection over reflections withindream.patterns.lookback_days(default 30). Names a pattern only when ≥dream.patterns.min_evidence(default 3) reflections support it. Single Sonnet subagent; same allow-list path as synthesize. Runs AFTERextractso the graph is fresh.src/core/cycle/extract-facts.ts(v0.32.2, extended v0.35.6.0) — extract_facts cycle phase. v0.32.2 contract: fence is canonical; per-page wipe (deleteFactsForPage) + reinsert fromparseFactsFence+extractFactsFromFenceText+engine.insertFacts. Empty-fence guard refuses when v0.31 legacy rows (row_num IS NULL AND entity_slug IS NOT NULL) pend the v0_32_2 backfill (status: warn, hint:gbrain apply-migrations --yes). v0.35.6.0 adds a phantom-redirect pre-pass that runs AFTER the legacy-row guard, BEFORE the main reconcile loop. Whenopts.brainDiris set,runPhantomRedirectPass(engine, brainDir, sourceId, dryRun)walks unprefixed-slug pages capped byGBRAIN_PHANTOM_REDIRECT_LIMIT(default 50). The pass returnstouched_canonicals— canonical slugs whose disk fence was merged with phantom rows;runExtractFactsUNIONs them into the main reconcile slug set so canonical's DB facts derive from the merged fence in the same cycle (round-14 scenario-B fix: phantom had only-on-disk fence, no DB facts).ExtractFactsResultgains six phantom fields:phantomsScanned,phantomsRedirected,phantomsAmbiguous,phantomsSkippedDrift,phantomsLockBusy,phantomsMorePending. Three of those bubble toCycleReport.totals(phantoms_redirected,phantoms_ambiguous,phantoms_skipped_drift).src/core/entities/resolve.ts(v0.30+, extended v0.35.6.0) — Free-form entity name → canonical slug resolution.resolveEntitySlug(engine, source_id, raw): exact slug → fuzzy (pg_trgm @ 0.4 threshold) → bare-name prefix expansion (people/<token>-%thencompanies/<token>-%using correlated-subqueryconnection_countfor tiebreaker) → deterministicslugifyfallback. v0.35.6.0 exports two new helpers for the phantom-redirect pass:resolvePhantomCanonical(engine, sourceId, phantomSlug)— variant that SKIPS the exact-slug step (codex #1: phantom slug'alice'exact-matches itself, would make the redirect handler a no-op); returns the canonical only when result is non-null AND contains/.findPrefixCandidates(engine, sourceId, token)— standalone SQL query returning ALL candidates acrossPREFIX_EXPANSION_DIRS(currently hardcoded['people', 'companies']) usingslug LIKE ANY($N::text[])over patternsdir/token+dir/token-%; cap of 10 ordered byconnection_count DESC, slug ASC. NOT a wrapper aroundtryPrefixExpansionbecause that path returns per-dir top-1 and suppresses ambiguity by design (codex #11). Pinned bytest/phantom-redirect.test.tsresolvePhantomCanonical describe (3 cases) + findPrefixCandidates describe (6 cases including multi-dir ambiguity and thepeople/aliceberg-doesn't-match-alicefalse-positive guard).src/core/cycle/phantom-redirect.ts(v0.35.6.0) — Phantom-redirect orchestrator. ExportsrunPhantomRedirectPass(engine, brainDir, sourceId, dryRun): Promise<PhantomPassResult>(the per-cycle wrapper that acquiresgbrain-syncwriter lock once for the entire pass, 30s bounded retry, walks up-to-GBRAIN_PHANTOM_REDIRECT_LIMITunprefixed phantoms) +tryRedirectPhantom(engine, page, sourceId, brainDir, dryRun): Promise<RedirectResult>(single-phantom orchestrator) +stripFenceAndFrontmatterAndLeadingH1(pure helper for the body-shape gate — strips facts fence including the preceding## Factsheading and the leading H1; zero residue = phantom). Handler order: body-shape gate →resolvePhantomCanonical(codex #1: bypasses exact-self-match) →findPrefixCandidatesambiguity check (codex #11: standalone query, not per-dir top-1) →fenceDbDriftbi-directional check (rounds 27/29/30) → dry-run early exit → materialize canonical viaserializeMarkdownif DB-only (codex #6) → append phantom fence rows to canonical's disk fence with(claim, valid_from)dedup-guard + row_num continuation →engine.refreshPageBodywith SHA-256 content_hash recomputed via the import-file shape (codex #7) →engine.migrateFactsToCanonical(codex #3/#4/#12 lossless preservation) →engine.rewriteLinks(DB FK rewrite; wiki-link text rewrite is a documented follow-up per codex #5) →engine.softDeletePage+engine.deleteFactsForPage(phantom)+fs.unlinkSync(phantomPath)(rounds 19/20).RedirectResult.canonicalpopulated on outcome'redirected'(incl. dry-run preview) so the caller can populatetouched_canonicals. Idempotent on re-run: phantom soft-deleted → predicate fails (deleted_at IS NULLfilter); migrate UPDATE matches no rows; dedup-guard prevents double-append.src/core/facts/phantom-audit.ts(v0.35.6.0) — JSONL audit at${resolveAuditDir()}/phantoms-YYYY-Www.jsonl. Pattern copy ofsrc/core/audit-slug-fallback.ts(ISO-week rotation, honorsGBRAIN_AUDIT_DIR). ExportslogPhantomEvent(record)+readRecentPhantomEvents(days)+computePhantomAuditFilename(now?). Records every outcome:redirected | ambiguous | drift | no_canonical | not_phantom_has_residue | pass_skipped_lock_busy. Best-effort writes — stderr warn on failure, never throws. Separate file fromstub-guard-audit.tsbecause the consumer + lifecycle are distinct (stub-guard logs PREVENTIVE blocks; phantom-audit logs CLEANUP decisions, will be read by a future T9 doctorphantoms_pendingcheck).src/core/cycle/emotional-weight.ts(v0.29) — Pure functioncomputeEmotionalWeight({tags, takes}, {highEmotionTags?, userHolder?}). Deterministic 0..1 score: tag-emotion boost (max 0.5, case-insensitive match againstHIGH_EMOTION_TAGSseed list), take density (0.1/take, capped at 0.3), take avg weight (0..0.1), user-holder ratio (0..0.1 over active takes; default holder = 'garry'). Total clamped to [0..1]. Anglocentric / personal-life-biased seed list intentional; users override via config keyemotional_weight.high_tags(JSON array).userHolderoverridable viaemotional_weight.user_holder.src/core/cycle/anomaly.ts(v0.29) — Pure stats helpers forfind_anomalies.meanStddevreturns sample stddev (n-1 denominator) and (0,0) for empty input.computeAnomaliesFromBuckets(baseline, today, sigma, limit)takes densified daily-count buckets + today's counts per cohort, returnsAnomalyResult[]. Zero-stddev fallback: cohort fires whencount > mean + 1, withsigma_observed = count - meanas a finite sort proxy (no NaN). Brand-new cohorts (no baseline) havemean=0, stddev=0so the fallback fires at count >= 2. Sorted bysigma_observeddesc, toplimit(default 20).page_slugscapped at 50 per cohort.src/core/cycle/recompute-emotional-weight.ts(v0.29) — Cycle phase orchestrator. Two SQL round-trips total:engine.batchLoadEmotionalInputs(slugs?)→computeEmotionalWeight(per-row pure function) →engine.setEmotionalWeightBatch(rows). Reads config keysemotional_weight.high_tags(JSON array, falls back to default seed list on parse error) andemotional_weight.user_holder. EmptyaffectedSlugsarray short-circuits with zero-work success. dry-run mode reports the would-write count without touching the DB. Engine throw bubbles intostatus: 'fail'with codeRECOMPUTE_EMOTIONAL_WEIGHT_FAILso the cycle continues.src/core/transcripts.ts(v0.29) —listRecentTranscripts(engine, opts)library reused by both thegbrain transcripts recentCLI and theget_recent_transcriptsMCP op. Readsdream.synthesize.session_corpus_dir+dream.synthesize.meeting_transcripts_dirconfig keys (same asdiscoverTranscripts); walks for.txtfiles withindays; appliesisDreamOutputguard fromtranscript-discovery.ts(skips dream-generated files); returns{path, date, mtime, length, summary}[]sorted newest-first. Summary mode (default true) returns first non-empty line + ~250 trailing chars. Full mode caps at 100KB/file. Missing/non-existent corpus dirs return[], not error. Trust gate lives in the op handler, not here: the op throwspermission_deniedforctx.remote === true; this library is a trusted library function used by both the gated op and the local CLI.src/core/operations-descriptions.ts(v0.29) — Constants module for tool descriptions. Pinned viatest/operations-descriptions.test.ts. HousesGET_RECENT_SALIENCE_DESCRIPTION,FIND_ANOMALIES_DESCRIPTION,GET_RECENT_TRANSCRIPTS_DESCRIPTIONplus the redirect-editedLIST_PAGES_DESCRIPTION,QUERY_DESCRIPTION,SEARCH_DESCRIPTION. Stable surface for the Tier-2 LLM routing eval — extracting them keeps the test from binding to whatever was inoperations.tsat test-run time.src/core/cycle/transcript-discovery.ts(v0.23) — Pure filesystem walk for synthesize.discoverTranscripts(opts)filters.txtfiles by date range, min_chars, and word-boundary regexexcludePatterns(Q-3:medicalmatches "medical advice" but NOT "comedical"; power users may pass full regex).readSingleTranscript(path)is thegbrain dream --input <file>ad-hoc path. v0.23.2 self-consumption guard:DREAM_OUTPUT_MARKER_RE(anchored at frontmatter open---\n, optional BOM + CRLF tolerance, scans first 2000 chars fordream_generated: truewith case-insensitive value and word boundary ontrue) drivesisDreamOutput(content, bypass=false). BothdiscoverTranscriptsandreadSingleTranscriptskip matching files and emit a[dream] skipped <basename>: dream_generated markerstderr log (no more silent skips).bypassGuard?: booleanonDiscoverOptsandreadSingleTranscript's opts disables the guard for the explicit--unsafe-bypass-dream-guardescape hatch only — never auto-applied for--input. Replaces v0.23.1'sDREAM_OUTPUT_SLUGScontent-prefix list.src/commands/dream.ts— v0.17gbrain dreamCLI; ~80-line thin alias overrunCycle. brainDir resolution requires explicit--dirORsync.repo_pathconfig. Flags:--dry-run,--json,--phase <name>,--pull,--dir <path>. v0.23 added--input <file>(ad-hoc transcript, implies--phase synthesize),--date YYYY-MM-DD,--from <d> --to <d>(backfill range). Conflict detection:--input+--dateexits 2. ISO date validation.--dry-runruns Haiku significance verdict but skips Sonnet synthesis (codex finding #8 — NOT zero LLM calls). Exit code 1 on status=failed. v0.23.2 added--unsafe-bypass-dream-guard(long-form intentional, plumbed throughrunCycle.synthBypassDreamGuard→SynthesizePhaseOpts.bypassDreamGuard→discoverTranscripts({bypassGuard})andreadSingleTranscript({bypassGuard})). Loud stderr warning fires at synthesize-phase entry when set. Never auto-applied for--inputso any caller can't silently re-trigger the loop bug.src/commands/friction.ts+src/core/friction.ts(v0.23) —gbrain friction {log,render,list,summary}reporter. Append-only JSONL under$GBRAIN_HOME/friction/<run-id>.jsonl. Schema is a flat extension ofStructuredAgentError(D20). Render groups by severity → phase, defaults to--redactfor md output (strips$HOME/$CWDto placeholders so reports paste safely in PRs). Run-id resolves from--run-id>$GBRAIN_FRICTION_RUN_ID>standalone.jsonl. Skills the claw-test exercises gain a_friction-protocol.mdcallout so agents know when to log friction.src/commands/claw-test.ts+src/core/claw-test/(v0.23) —gbrain claw-test [--scenario <name>] [--live --agent openclaw]. End-to-end "fresh user" friction harness. Two modes: scripted (CI gate, agent-free) and live (real openclaw subprocess, $1–2 in tokens). SetsGBRAIN_HOME=<tempdir>for hermeticity and captures gbrain's--progress-jsonevents from each child's stderr to verify expected phases ran (import.files,extract.links_fs,doctor.db_checks). Phases for scripted mode: setup → install_brain (gbrain init --pglite) → import (--no-embed) → query → extract → verify (gbrain doctor --json, assertsstatus: 'ok') → render. Live mode handsBRIEF.mdfromtest/fixtures/claw-test-scenarios/<name>/to the agent runner. v1 ships with the OpenClaw runner only (src/core/claw-test/runners/openclaw.ts, invokesopenclaw agent --local --agent <name> --message <brief>); hermes runner deferred to v1.1. Transcript capture (transcript-capture.ts) usesfs.createWriteStreamwith'drain'-event backpressure — D17 fix for the 256KB-burst child-stall scenario. v0.18 upgrade scenario seeded viaseed-pglite.tsSQL replay.skills/_friction-protocol.md(v0.23) — shared cross-cutting convention skill (like_brain-filing-rules.md). Tells agents when to callgbrain friction logand how to choose a severity. Routes to friction CLI from any skill the claw-test exercises.scripts/check-progress-to-stdout.sh— CI guard against regressing to\r-on-stdout progress. Wired intobun run testviascripts/check-progress-to-stdout.sh && bun testin package.json.docs/progress-events.md— Canonical JSON event schema reference. Stable from v0.15.2, additive only.src/core/markdown.ts— Frontmatter parsing + body splitter.splitBodyrequires an explicit timeline sentinel (<!-- timeline -->,--- timeline ---, or---immediately before## Timeline/## History). Plain---in body text is a markdown horizontal rule, not a separator.inferTypeauto-types/wiki/analysis/→ analysis,/wiki/guides/→ guide,/wiki/hardware/→ hardware,/wiki/architecture/→ architecture,/writing/→ writing (plus the existing people/companies/deals/etc heuristics).scripts/check-jsonb-pattern.sh— CI grep guard. Fails the build if anyone reintroduces (a) the${JSON.stringify(x)}::jsonbinterpolation pattern (postgres.js v3 double-encodes it), or (b)max_stalled INTEGER NOT NULL DEFAULT 1in any schema source file (v0.15.1 #219 regression guard — must be DEFAULT 5 to preserve SIGKILL-rescue). Wired intobun test.scripts/check-source-id-projection.sh(v0.32.8, PR #860) — CI grep guard for the multi-source bug class. Grepssrc/core/postgres-engine.ts+src/core/pglite-engine.tsforSELECT.*FROM pagesprojections matching therowToPagefeeder shape (id + slug + type + title) and fails ifsource_idis missing. After v0.32.8Page.source_idis required at the type level; a projection that drops the column producesPagerows withsource_id: undefinedwhile TypeScript's: stringlies about it. Codex's outside-voice review caught two pre-existing projections (getPage,putPage RETURNING) that lacked the column. Wired intobun run verify+bun run check:all.docker-compose.ci.yml+scripts/ci-local.sh(v0.23.1) — Local CI gate.bun run ci:localspins uppgvector/pgvector:pg16+oven/bun:1with named volumes (gbrain-ci-pg-data,gbrain-ci-node-modules,gbrain-ci-bun-cache), runs gitleaks on host, smoke-testsscripts/run-e2e.shargv handling, runs unit tests withDATABASE_URLunset (matches GH Actions structure), then runs all 29 E2E files sequentially.--diffswaps in the diff-aware selector;--no-pullskips upstream pulls;--cleannukes named volumes. Postgres host port defaults to 5434 (avoids 5432 manualgbrain-test-pgand 5433 sibling-project conflict); override withGBRAIN_CI_PG_PORT=NNNN. Stronger gate than current PR CI's 2-file Tier 1 set — closes the "push-and-wait" feedback loop pre-push.scripts/select-e2e.ts+scripts/e2e-test-map.ts(v0.23.1) — Diff-aware E2E test selector. Reads three git sources (committedorigin/master...HEAD, working-treeHEAD, andgit ls-files --others --exclude-standardfor untracked, NOT-gitignored files), classifies as EMPTY / DOC_ONLY / SRC. Fail-closed by design: EMPTY → all 29 files (clean branch shouldn't run nothing), DOC_ONLY (every path matches the README/CLAUDE/AGENTS/CHANGELOG/TODOS allowlist) → empty stdout, SRC → escape-hatch paths (schema, package.json, skills/) trigger all; otherwise the hand-tunedE2E_TEST_MAPglob → tests narrows; an unmapped src/ change still emits ALL files, never silently nothing. Pure-function exports (selectTests,classify,matchGlob) so it's trivial to test and fork.bun run ci:select-e2eprints the current selection on stdout, pipe-friendly.test/select-e2e.test.tscovers all 4 branches plus 3 codex regression guards (skills/, untracked files, unmapped src/) — 24 cases.scripts/run-e2e.sh(v0.23.1 update) — Sequential E2E runner. Now accepts an optional argv-driven file list (used byci:local:diffto pipe in selector output) and a--dry-run-listflag that prints the resolved file list and exits (used byci-local.sh's startup smoke-test). Falls back totest/e2e/*.test.tswhen invoked with no args.scripts/llms-config.ts+scripts/build-llms.ts— Generator forllms.txt(llmstxt.org-spec web index) +llms-full.txt(inlined single-fetch bundle). Curated config drives both. Runbun run build:llmsafter adding a new doc.LLMS_REPO_BASEenv var lets forks regenerate with their own URL base.FULL_SIZE_BUDGET(600KB) caps the inline bundle; generator WARNs if exceeded. Committed output is not analogous toschema-embedded.ts(no runtime consumer); we commit for GitHub browsing and fork-safe fetching.AGENTS.md— Local-clone entry point for non-Claude agents (Codex, Cursor, OpenClaw, Aider). MirrorsCLAUDE.mdintent via relative links. Claude Code keeps usingCLAUDE.md.docs/UPGRADING_DOWNSTREAM_AGENTS.md— Patches for downstream agent skill forks to apply when upgrading. Each release appends a new section. v0.10.3 includes diffs for brain-ops, meeting-ingestion, signal-detector, enrich.src/core/schema-embedded.ts— AUTO-GENERATED from schema.sql (runbun run build:schema)src/schema.sql— Full Postgres + pgvector DDL (source of truth, generates schema-embedded.ts)src/commands/integrations.ts— Standalone integration recipe management (no DB needed). ExportsgetRecipeDirs()(trust-tagged recipe sources), SSRF helpers (isInternalUrl,parseOctet,hostnameToOctets,isPrivateIpv4). Only package-bundled recipes areembedded=true;$GBRAIN_RECIPES_DIRand cwd./recipes/are untrusted and cannot runcommand/http/string health checks.src/core/search/expansion.ts— Multi-query expansion via Haiku. ExportssanitizeQueryForPrompt+sanitizeExpansionOutput(prompt-injection defense-in-depth). Sanitized query is only used for the LLM channel; original query still drives search.recipes/— Integration recipe files (YAML frontmatter + markdown setup instructions)docs/guides/— Individual SKILLPACK guides (broken out from monolith)docs/integrations/— "Getting Data In" guides and integration docsdocs/architecture/infra-layer.md— Shared infrastructure documentationdocs/ethos/THIN_HARNESS_FAT_SKILLS.md— Architecture philosophy essaydocs/ethos/MARKDOWN_SKILLS_AS_RECIPES.md— "Homebrew for Personal AI" essaydocs/guides/repo-architecture.md— Two-repo pattern (agent vs brain)docs/guides/sub-agent-routing.md— Model routing table for sub-agentsdocs/guides/skill-development.md— 5-step skill development cycle + MECEdocs/guides/idea-capture.md— Originality distribution, depth test, cross-linkingdocs/guides/quiet-hours.md— Notification hold + timezone-aware deliverydocs/guides/diligence-ingestion.md— Data room to brain pages pipelinedocs/designs/HOMEBREW_FOR_PERSONAL_AI.md— 10-star vision for integration systemdocs/mcp/— Per-client setup guides (Claude Desktop, Code, Cowork, Perplexity)- BrainBench (benchmark suite + corpus): lives in the separate gbrain-evals repo. Not installed alongside gbrain.
skills/_brain-filing-rules.md— Cross-cutting brain filing rules (referenced by all brain-writing skills)skills/RESOLVER.md— Skill routing table (based on the agent-fork AGENTS.md pattern)skills/conventions/— Cross-cutting rules (quality, brain-first, model-routing, test-before-bulk, cross-modal)skills/_output-rules.md— Output quality standards (deterministic links, no slop, exact phrasing)skills/signal-detector/SKILL.md— Always-on idea+entity capture on every messageskills/brain-ops/SKILL.md— Brain-first lookup, read-enrich-write loop, source attributionskills/idea-ingest/SKILL.md— Links/articles/tweets with author people page mandatoryskills/media-ingest/SKILL.md— Video/audio/PDF/book with entity extractionskills/meeting-ingestion/SKILL.md— Transcripts with attendee enrichment chainingskills/citation-fixer/SKILL.md— Citation format auditing and fixingskills/repo-architecture/SKILL.md— Filing rules by primary subjectskills/skill-creator/SKILL.md— Create conforming skills with MECE checkskills/daily-task-manager/SKILL.md— Task lifecycle with priority levelsskills/daily-task-prep/SKILL.md— Morning prep with calendar contextskills/cross-modal-review/SKILL.md— Quality gate via second modelskills/cron-scheduler/SKILL.md— Schedule staggering, quiet hours, idempotencyskills/reports/SKILL.md— Timestamped reports with keyword routingskills/testing/SKILL.md— Skill validation frameworkskills/soul-audit/SKILL.md— 6-phase interview for SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.mdskills/webhook-transforms/SKILL.md— External events to brain signalsskills/data-research/SKILL.md— Structured data research: email-to-tracker pipeline with parameterized YAML recipesskills/minion-orchestrator/SKILL.md— Unified background-work skill (v0.20.4 consolidation of the formerminion-orchestrator+gbrain-jobssplit). Two lanes: shell jobs viagbrain jobs submit shell --params '{"cmd":"..."}'(operator/CLI only; MCP throwspermission_deniedfor protected names) and LLM subagents viagbrain agent run(user-facing entrypoint). Shared Preconditions block, parent-child DAGs with depth/cap/timeouts,child_doneinbox for fan-in, PGLite--followinline path for dev. Triggers narrowed from bare"gbrain jobs"to"gbrain jobs submit"+"submit a gbrain job"sostats/prune/retryquestions fall through togbrain --help.templates/— SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.md templatesskills/migrations/— Version migration files with feature_pitch YAML frontmattersrc/commands/publish.ts— Deterministic brain page publisher (code+skill pair, zero LLM calls)src/commands/backlinks.ts— Back-link checker and fixer (enforces Iron Law)src/commands/lint.ts— Page quality linter (catches LLM artifacts, placeholder dates)src/commands/report.ts— Structured report saver (audit trail for maintenance/enrichment)src/core/destructive-guard.ts(v0.26.5) — three-layer protection against accidental data loss in gbrain.assessDestructiveImpact(engine, sourceId)counts pages/chunks/embeddings/files for a source.checkDestructiveConfirmation(impact, opts)is the fail-closed gate (--confirm-destructiverequired when data is present;--yesalone is rejected).softDeleteSource/restoreSource/listArchivedSources/purgeExpiredSourcesdrive the source-level archive lifecycle via the column shape introduced in migration v34 (sources.archived BOOLEAN,archived_at TIMESTAMPTZ,archive_expires_at TIMESTAMPTZ). v0.26.5 added the page-level analog throughBrainEngine.softDeletePage/restorePage/purgeDeletedPagespluspages.deleted_at TIMESTAMPTZand a partial purge index. The MCPdelete_pageop rewires tosoftDeletePage; new opsrestore_page(scope: write) andpurge_deleted_pages(scope: admin,localOnly: true) round out the surface. Search visibility (buildVisibilityClauseinsrc/core/search/sql-ranking.ts) hides soft-deleted pages and archived sources fromsearchKeyword/searchKeywordChunks/searchVectorin both engines. The autopilot cycle's new 9thpurgephase callspurgeExpiredSources+engine.purgeDeletedPages(72)so the 72h TTL is real, not honor-system.src/commands/pages.ts(v0.26.5) —gbrain pages purge-deleted [--older-than HOURS|Nd] [--dry-run] [--json]operator escape hatch. Mirror ofgbrain sources purgefor the page-level lifecycle. Hard-deletes pages whosedeleted_atis older than the cutoff; cascades to content_chunks/page_links/chunk_relations.openclaw.plugin.json— ClawHub bundle plugin manifest
BrainBench — in a sibling repo (v0.20+)
BrainBench — the public benchmark for personal-knowledge agent stacks — lives in github.com/garrytan/gbrain-evals. It depends on gbrain as a consumer; gbrain never pulls in the ~5MB eval corpus or the pdf-parse dev dep at install time.
gbrain's public API surface (the exports map in package.json) is what
gbrain-evals consumes: gbrain/engine, gbrain/types, gbrain/operations,
gbrain/pglite-engine, gbrain/link-extraction, gbrain/import-file,
gbrain/transcription, gbrain/embedding, gbrain/config, gbrain/markdown,
gbrain/backoff, gbrain/search/hybrid, gbrain/search/expansion,
gbrain/extract. Removing any of these is a breaking change for the
gbrain-evals consumer.
v0.36.1.0 Hindsight calibration wave (key files cluster)
The wave that taught gbrain to know how the user tends to be wrong + use
that knowledge at every advice surface. Six-migration schema (v67-v72),
three new cycle phases, eight expansions, one admin tab. Plan persisted
at ~/.claude/plans/system-instruction-you-are-working-rippling-knuth.md.
Convention skill at skills/conventions/calibration.md has the agent-
facing rules.
src/core/cycle/base-phase.ts— abstractBaseCyclePhaseclass. EnforcessourceScopeOpts(ctx)threading at the type level; closes the v0.34.1 source-isolation leak class structurally for every new phase. Inherits source-scope, budget meter, error envelope, progress reporter. propose_takes / grade_takes / calibration_profile all extend it.src/core/cycle/propose-takes.ts— LLM scans markdown prose, proposes gradeable claims totake_proposalsqueue. Idempotency cache on(source_id, page_slug, content_hash, prompt_version)composite unique index. F2 fence-dedup: existing canonical takes passed to extractor as context. v0.36.1.0 ships a stub prompt; tuned prompt arrives via the T19 synthetic corpus build.src/core/cycle/grade-takes.ts— walks unresolved takes older than 6 months, retrieves evidence, asks judge model, caches verdict. Auto-resolve DISABLED by default (D17). Conservative thresholds:=0.95 single OR >=0.85 ensemble 3/3 unanimous. T5 ensemble (
aggregateEnsemble) reuses v0.27.x cross-modal substrate; fires on borderline 0.6-0.95 band. Writes totake_grade_cache.src/core/cycle/calibration-profile.ts— aggregates resolved takes into 2-4 narrative pattern statements + active bias tags. Voice-gated viagateVoice(). Cold-brain skip when <5 resolved. Writes tocalibration_profileswith audit columns (voice_gate_passed,voice_gate_attempts,grade_completion).src/core/calibration/voice-gate.ts— singlegateVoice()function (D24), mode parameter (pattern_statement|nudge|forecast_blurb|dashboard_caption|morning_pulse). 2 regens then hand-written template fallback fromsrc/core/calibration/templates.ts. Haiku judge with mode-specific rubrics; all rubrics structurally forbid clinical/preachy voice.src/core/calibration/cross-brain.ts— D18 4-rule contract for cross-brain calibration reads. Local-first → mount-fallback (only withcanReadMountsForCtx(ctx)true) → cross-brain attribution viasource_brain_id+from_mount→ subagent prohibition closes the OAuth-token-to-cross-brain-leak surface. All 4 rules pinned intest/cross-brain-calibration.test.ts.src/core/calibration/nudge.ts— E7 real-time pattern surfacing.evaluateAndFireNudge(opts)is the full pipeline: threshold check (conviction > 0.7, holder match, slug-derived domain hint matches active bias tag), cooldown probe (14d via take_nudge_log), fire + log. STDERR-only output for v0.36.1.0; multi-channel deferred.src/core/calibration/take-forecast.ts— E5 Brier-trend at write time. Pure math over existingTakesScorecard; no LLM. Returnspredicted_brier,bucket_n,overall_brier. Insufficient-data branch atMIN_BUCKET_N = 5.batchForecastmemoizes per (holder, domain) tuple.src/core/calibration/gstack-coupling.ts— E4 outcome-driven learnings coupling.writeIncorrectResolution(opts)shells out togstack-learnings-logbinary. Config gate (cycle.grade_takes.write_gstack_learnings, default false for external users). Namespace prefixgbrain:calibration:v0.36.1.0:so--undo-wavecan scrub.src/core/calibration/svg-renderer.ts— D23 server-rendered SVG for the admin SPA Calibration tab. Pure functions: data → SVG string. Inlines design tokens; XSS-safe viaescapeXml(). Four chart renderers:renderBrierTrend,renderDomainBars,renderAbandonedThreadsCard,renderPatternStatementsCard. SPA renders via<TrustedSVG>wrapper behindrequireAdmin.src/core/calibration/undo-wave.ts— D18 CDX-3 resolution.undoWavereverses the wave's mutations: unsetstakes.resolved_*for wave-applied resolutions (cross-checks resolved_by so manual writes persist), deletes calibration_profiles, purges nudge logs, marks grade-cache rows applied=false.--dry-runshows counts without writing. Idempotent on wave_version match.src/core/calibration/think-ab.ts— D19 A/B harness.runAbTrialcalls thinkRunner twice (baseline + with-calibration), records preference tothink_ab_results.buildAbReportaggregates over 30-day window; flagscalibration_net_negativewhen n>=20 + win rate < 45% on decisive trials.src/core/calibration/recall-footer.ts— formatter for the morning pulse calibration block. Cold-brain branch when <5 resolved. v0.36 ship state: opt-in via the wiring layer; auto-on in v0.37+.src/core/eval-contradictions/calibration-join.ts— E3 cross- reference.tagFindingWithCalibration(finding, profile)returns bias-tag context for contradictions that match active patterns. Returns null when profile missing (R2 regression — output byte-identical to v0.32.6).src/core/think/prompt.tsextension — E1 anti-bias rewrite.withCalibrationoption onbuildThinkSystemPromptadds anti-bias rules. NewbuildCalibrationBlock()emits the<calibration>XML.buildThinkUserMessagehas TWO shapes: default (question first) for R1 regression, with-calibration (retrieval → calibration → question per D22) when opt-in. Wired intorunThinkviaopts.withCalibration+opts.calibrationHolder.src/commands/calibration.ts— CLI:gbrain calibration(read + print),--regenerate,--undo-wave <ver>(T17),ab-report(T18). MCP opget_calibration_profile(scope: read) backs the same data path. Source-scoped viasourceScopeOpts(ctx).src/commands/serve-http.tsextension — three new admin routes:/admin/api/calibration/profile,/admin/api/calibration/charts/:type(image/svg+xml; type in {brier-trend, domain-bars, pattern-statements, abandoned-threads}), and/admin/api/calibration/pattern/:id(TD3 drill-down).src/commands/takes.tsextension —gbrain takes revisit <slug>(TD4 / D30) opens $EDITOR on the source page with a<!-- gbrain:revisit -->cursor marker.src/commands/doctor.tsextension — 4 new checks:abandoned_threads,calibration_freshness,grade_confidence_drift(CDX-11 mitigation surface; math arrives v0.37+),voice_gate_health.admin/src/pages/Calibration.tsx— Calibration tab. Single-column Linear-calm-clarity layout matching the approved variant-B mockup.<TrustedSVG>wrapper handlesdangerouslySetInnerHTMLfor the server-rendered SVG.admin/src/index.cssextension —--text-muted: #555 → #777(TD2, WCAG AA contrast bump from 4.0 to ~5.5 on the #0a0a0f bg).test/fixtures/calibration/extract-takes-corpus/— synthetic prompt- tuning corpus. v0.36.1.0 ships 5 representative pages; full 50-page- 10-page holdout generated by
gbrain calibration build-corpus(v0.37+ subcommand). All anonymized per CLAUDE.md placeholder list.
- 10-page holdout generated by
scripts/check-synthetic-corpus-privacy.sh— CDX-14 mitigation. CI guard inbun run verify. Greps for explicit dollar amounts + verifies non-essay fixtures reference at least one placeholder name.test/regressions/v0.36.1.0-iron-rule.test.ts— R1-R5 regression inventory test file. Pins all 5 IRON-RULE regressions in one place for future bisects.DESIGN.md— repo-root design system. Formalizes the de facto admin tokens that landed v0.26.0. Calibration target for future/plan-design-reviewand/design-review.
Thin-client routing (v0.31.1, Issue #734)
gbrain init --mcp-only (v0.29.2) sets up a thin-client install: no local
brain content, just an OAuth client pointing at a remote gbrain serve --http.
v0.29.2/v0.30.0 only refused 9 obvious local-only commands; the other ~25
silently fell through to connectEngine() and opened the empty local PGLite,
returning "No results." against a populated remote brain. v0.31.1 fixes the
silent-empty-results bug class for every operation surface.
Key files:
src/cli.ts— Routing seam INSIDE the existing op-dispatch path (CDX-1: no parallelsrc/core/thin-client/module; routing is a ~80-line conditional inrunThinClientRouted). DetectsisThinClient(cfg)BEFOREconnectEngineso thin-client installs never open the empty PGLite. localOnly ops on thin-client refuse viarefuseThinClient(with pinpoint hint tableTHIN_CLIENT_REFUSE_HINTS). Banner viaprintIdentityBannerBestEffortbefore each routed call (suppressed by--quiet,GBRAIN_NO_BANNER=1, non-TTY default). Exhaustive TSneverswitch onRemoteMcpError.reasonfor canned, actionable error messages. ENG-2 renderer parity: local-engine path runsJSON.parse(JSON.stringify(result))so renderers see the same shape on both paths (kills Date/bigint/Buffer drift class).src/core/mcp-client.ts—callRemoteTool(config, toolName, args, opts). Hardened in v0.31.1 (CDX-4): all transport errors normalized toRemoteMcpErrorvia thetoRemoteMcpErrorfunnel. NewCallRemoteToolOptions {timeoutMs, signal};buildAbortControllercomposes external signal with timeout. NewRemoteMcpErrorReasonstable union,RemoteMcpErrorDetail.kind('timeout' | 'aborted' | 'unreachable') sub-tag,RemoteMcpErrorDetail.codefield carrying server-supplied error codes (e.g.missing_scope).extractToolErrorCodeparses JSON envelopes first, falls back to substring detection for legacy server messages.unpackToolResult<T>(res)unchanged (parses tool-call JSON content)._clearMcpClientTokenCache()test escape.src/core/cli-options.ts—parseGlobalFlagsadds--timeout=Ns(accepts30s,2m,500ms, plain ms). Defaultnull= per-command default (30s for most ops, 180s forthink).parseTimeout(s)exported helper.src/core/doctor-remote.ts—gbrain remote doctoradds theoauth_client_scopes_probecheck (CDX-5). Probes the read tier viaget_brain_identityand admin tier viaget_health; reports per-tier status with pinpoint remediation when admin is missing.buildScopeCheckScopeProbeResultexported for test access. Skippable viaGBRAIN_DOCTOR_SKIP_SCOPE_PROBE=1for fixtures that mock /mcp at JSON-RPC initialize level only (MCP SDK Client hangs on shape mismatch).
src/core/operations.ts—get_brain_identityop (read scope, no params, banner-only): cheap counter packet{version, engine, page_count, chunk_count, last_sync_iso}for the thin-client identity banner. Reusesengine.getStats(); banner's 60s client-side TTL bounds frequency to ≤1/60s per CLI process (well below the Fly.io health-check cadence that motivated the originalgetStatscost warning).src/commands/{salience,anomalies,graph-query,think}.ts— Per-command thin-client routing branches. These commands bypass the operation-layer dispatch in cli.ts (callengine.foo()directly), so each gets its ownif (isThinClient(cfg)) { callRemoteTool(...) }branch that maps CLI flags to op params.thinkis a special case: the server'sthinkop intentionally disables--save/--takefor remote callers (operations.ts:1103-1135 trust-boundary gate); thin-clientthinkwarns loudly when those flags are set.
Commands
Run gbrain --help or gbrain --tools-json for full command reference.
Key commands added in v0.7:
gbrain init— defaults to PGLite (no Supabase needed), scans repo size, suggests Supabase for 1000+ filesgbrain migrate --to supabase/gbrain migrate --to pglite— bidirectional engine migration
Key commands added for Minions (job queue):
gbrain jobs submit <name> [--params JSON] [--follow] [--dry-run]— submit a background job. v0.13.1 adds first-class flags for everyMinionJobInputtuning knob:--max-stalled N,--backoff-type fixed|exponential,--backoff-delay Nms,--backoff-jitter 0..1,--timeout-ms N,--idempotency-key K.gbrain jobs list [--status S] [--queue Q]— list jobs with filtersgbrain jobs get <id>— job details with attempt historygbrain jobs cancel/retry/delete <id>— manage job lifecyclegbrain jobs prune [--older-than 30d]— clean old completed/dead jobsgbrain jobs stats— job health dashboardgbrain jobs smoke [--sigkill-rescue]— health smoke test.--sigkill-rescueis the v0.13.1 regression guard for #219: simulates a killed worker and asserts the stalled job is requeued instead of dead-lettered on first stall.gbrain jobs work [--queue Q] [--concurrency N]— start worker daemon (Postgres only)
Key commands added in v0.32.7 (CJK fix wave):
gbrain reindex --markdown [--limit N] [--dry-run] [--json] [--no-embed] [--repo PATH]— operator-facing markdown re-chunk sweep. Walks pages withchunker_version < MARKDOWN_CHUNKER_VERSION(currently 2) and re-imports each withforceRechunk: trueso the new chunker shape actually applies. Run automatically bygbrain upgrade's post-upgrade hook; available manually for triage.gbrain doctorlearns a newslug_fallback_auditcheck: surfaces info-severity entries from~/.gbrain/audit/slug-fallback-YYYY-Www.jsonl(last 7 days) as anokcount when CJK / emoji / exotic-script filenames imported via the frontmatter-slug fallback path.gbrain search "<CJK substring>"on PGLite brains now uses anILIKE-based fallback with bigram-frequency-count ranking when the query contains Han / Hiragana / Katakana / Hangul Syllables. ASCII queries continue throughwebsearch_to_tsquery('english')unchanged. Postgres-side CJK FTS still requires an extension (pgroonga / zhparser) — see v0.33+ TODO.gbrain upgradepost-upgrade flow now prints a cost estimate before re-embedding:[chunker-bump] Will re-embed ~N markdown pages via <provider:model>, est. ~$X.XX, ~Ymin. Press Ctrl-C within 10s to abort.Sourced from real SQL counts + char totals; TTY-only wait (non-TTY auto-proceeds for CI / cron). Env overrides:GBRAIN_NO_REEMBED=1bails out entirely with a doctor-warning marker;GBRAIN_REEMBED_GRACE_SECONDS=0skips the wait.
Key commands added in v0.33.1.1 (Voyage 2048-dim correctness wave):
gbrain models doctorlearns a new zero-tokenembedding_configprobe that runs FIRST, before any chat/expansion probes spend money. Catches Voyage flexible-dim misconfigs at config time, not first-embed:embedding_model: voyage:voyage-4-largewithembedding_dimensionsoutside{256, 512, 1024, 2048}(most commonly:embedding_dimensionsleft unset, falling back to the OpenAI default 1536 which Voyage rejects with an opaque HTTP 400). Surfaces a paste-readygbrain config set embedding_dimensions <256|512|1024|2048>fix in both human and JSON output. New probe status'config'joins{ok, model_not_found, auth, rate_limit, network, unknown}; new touchpoint label'embedding_config'joins'chat'and'expansion'.- Voyage 2048-dim brains now actually embed at 2048 dims.
embedding_model: voyage:voyage-4-large+embedding_dimensions: 2048routes through the SDK-supporteddimensionsfield, whichvoyageCompatFetchtranslates to Voyage'soutput_dimensionon the wire. Same fix coversvoyage-3-large,voyage-3.5,voyage-3.5-lite,voyage-4,voyage-4-lite,voyage-code-3.voyage-4-nano(open-weight, fixed 1024-dim) intentionally NOT in the flexible-dim allowlist — sendingoutput_dimensionto nano's endpoint produces an error. - Runtime validator:
dimsProviderOptions()throwsAIConfigErrorat the embed boundary with a paste-ready fix hint when a Voyage flexible-dim model is configured with an invalid dim — fail-loud even if you skippedgbrain models doctor. VoyageResponseTooLargeError(new tagged class exported fromsrc/core/ai/gateway.ts): the 256 MB per-response cap insidevoyageCompatFetchwas previously throwing a genericErrorthat the surrounding parse-error try/catch silently swallowed, returning the oversized response to the AI SDK anyway. Now thrown at both cap sites (Content-Length Layer 1, per-embedding base64 Layer 2) and rethrown from the catch viainstanceofcheck — the cap is now actually effective.
Key commands added in v0.31.12 (model tier system + routing CLI):
gbrain models [--json]— read-only routing dashboard. Prints the four tier defaults (utility/reasoning/deep/subagent), the resolved value for each (after re-walkingmodels.default→models.tier.<tier>→ env →TIER_DEFAULTS), every per-task override (models.dream.synthesize,models.dream.patterns,models.drift,models.auto_think,models.think,models.subagent,facts.extraction_model,models.eval.longmemeval,models.expansion,models.chat,models.dream.synthesize_verdict), the alias map (defaults + user overrides), and a source-of-truth column (default/config: <key>/env: <VAR>).gbrain models doctor [--skip=<provider>] [--json]— 1-token reachability probe against each configured chat + expansion model. Classifies failures into{model_not_found, auth, rate_limit, network, unknown}. The structural fix for the bug class that motivated v0.31.12 (v0.31.6'sclaude-sonnet-4-6-20250929chat default 404'd silently on every install).- Power-user model routing via config keys:
gbrain config set models.default opus— route every internal call (chat, expansion, synthesis, classification) through Opus 4.7. Subagent loop still falls back toclaude-sonnet-4-6automatically (Anthropic-only by construction).gbrain config set models.tier.<tier> <model>— override one tier independently (utility/reasoning/deep/subagent).gbrain config set models.aliases.frontier anthropic:claude-opus-4-7— define an alias, thengbrain config set models.default frontier.- Per-task keys (e.g.
gbrain config set models.dream.synthesize <model>) still beat tier overrides because they are more specific.
- New
subagent_providercheck ingbrain doctorsurfaces config drift ifmodels.tier.subagentormodels.defaultwould route the Anthropic Messages API tool-loop to a non-Anthropic provider. - The skill at
skills/conventions/model-routing.mdwas rewritten to cover both the new tier system AND the existing subagent spawn routing in one canonical doc (power-user recipes, three-layer enforcement explanation, override priority chain).
Key commands added in v0.28.1 (LongMemEval in the box):
gbrain eval longmemeval <dataset.jsonl>— run the public LongMemEval benchmark against gbrain hybrid retrieval. Flags:--limit N,--model M,--retrieval-only,--keyword-only,--expansion,--top-k K,--output FILE. One in-memory PGLite per benchmark run;TRUNCATEbetween questions over runtime-enumeratedpg_tables(schema-migration-safe);~/.gbrainnever opened.--expansiondefaults OFF (deterministic, no per-query Haiku). Default model resolves throughresolveModel()6-tier chain with newmodels.eval.longmemevalconfig key.gbrain eval longmemeval --helpworks without a configured brain (hermeticity gate).- Sanitization parity with takes:
INJECTION_PATTERNSexported fromsrc/core/think/sanitize.ts. The benchmark harness re-uses the same pattern set so adding a new injection pattern automatically covers takes AND benchmarks. - Hand the resulting JSONL to LongMemEval's published
evaluate_qa.pyto score (not bundled — needs OpenAI gpt-4o per their spec). Dataset: https://huggingface.co/datasets/xiaowu0162/longmemeval.
Key commands added in v0.26.5 (destructive-guard, end-to-end):
gbrain sources archive <id>— soft-delete a source. Hides from search via the newsources.archivedcolumn + cascading visibility filter. Preserves data for 72h. (PR #595 cherry-pick.)gbrain sources restore <id> [--no-federate]— un-archive a soft-deleted source. Re-federates by default.gbrain sources archived [--json]— list soft-deleted sources with their TTL.gbrain sources purge [<id>] [--confirm-destructive]— permanent delete; with no id, purges all sources whose TTL expired.gbrain sources remove <id> [--confirm-destructive] [--dry-run]—--yesalone no longer enough on populated sources. Boxed impact preview before destruction.gbrain pages purge-deleted [--older-than HOURS|Nd] [--dry-run] [--json]— operator escape hatch for page-level soft-delete cleanup. Mirror ofgbrain sources purge. The autopilot cycle's newpurgephase calls the same library function automatically every run.- MCP
delete_pageop semantically shifts from hard-delete to soft-delete. New ops:restore_page(scope: write),purge_deleted_pages(scope: admin,localOnly: true). get_pageandlist_pagesextended withinclude_deleted: boolean(default false).- New autopilot cycle phase
purge(9th, runs afterorphans).gbrain dream --phase purgeruns only the purge sweep. - Index strategy note: the partial index
pages_deleted_at_purge_idx ON pages (deleted_at) WHERE deleted_at IS NOT NULLsupports the autopilot purge query. Search filters (WHERE deleted_at IS NULL) do NOT need their own index — soft-deleted cardinality stays low and Postgres won't use the partial index for the negative predicate. Don't add a regular(deleted_at)index without measuring. - Schema migration v34 (
destructive_guard_columns) addspages.deleted_at+ the partial purge index; promotesarchivedfromsources.configJSONB to real columns; backfills any pre-v0.26.5 JSONB shape.
Key commands added in v0.25.0:
gbrain eval export [--since DUR] [--limit N] [--tool query|search]— stream capturedeval_candidatesrows as NDJSON to stdout. Every line starts with"schema_version": 1per the stable contract indocs/eval-capture.md. EPIPE-safe, progress heartbeats on stderr, deterministic ordering. Primary consumer is the siblinggbrain-evalsrepo for BrainBench-Real replay.gbrain eval prune --older-than DUR [--dry-run]— explicit retention cleanup foreval_candidates. Requires--older-than(never deletes without a window). Duration strings: 30d, 7d, 1h, 90m, 3600s.gbrain eval replay --against FILE.ndjson [--limit N] [--top-regressions K] [--json] [--verbose]— contributor-facing dev loop. Reads a captured NDJSON snapshot, re-runs eachquery/searchop against the current brain, computes mean set-Jaccard@k between captured + currentretrieved_slugs, top-1 stability rate, and latency Δ. JSON mode (schema_version: 1) for CI gating; human mode prints a regression table sorted worst-first. Closes the gap between "data captured" and "data used to gate a PR." Seedocs/eval-bench.mdfor the workflow.gbrain eval cross-modal --task "..." --output <path> [--cycles N] [--slot-a-model ID] [--slot-b-model ID] [--slot-c-model ID] [--receipt-dir DIR] [--json](v0.27.x) — multi-model quality gate. Three different-provider frontier models score the OUTPUT against the TASK on 5 documented dimensions. Pass criterion: every dim mean >=7 AND no model scored any dim <5. Exit codes: 0 PASS, 1 FAIL, 2 INCONCLUSIVE (<2/3 models returned parseable scores). Default cycles=3 in TTY, cycles=1 in non-TTY (limits accidental scripted bulk spend). Default slots:openai:gpt-4o/anthropic:claude-opus-4-7/google:gemini-1.5-pro— refresh alongside model-family bumps. Receipts land at~/.gbrain/.gbrain/eval-receipts/<slug>-<sha8-of-output>.json(gbrainPath honors GBRAIN_HOME). BypassesconnectEngine()via the cli.ts no-DB branch — runs cleanly beforegbrain init. Reusessrc/core/ai/gateway.ts:chat()for config/auth (no parallel provider stack). Cost-estimate prints to stderr before each cycle (T11=B partial cost guardrail; full--budget-usd Nis a follow-up TODO).gbrain doctorgains aneval_capturecheck: readseval_capture_failuresfor the last 24h, groups by reason, warns when non-zero. Cross-process visibility (doctor runs in a separate process from MCP). Pre-v31 brains getSkipped (table unavailable)— non-fatal.- Config addition:
eval: { capture?: boolean, scrub_pii?: boolean }in~/.gbrain/config.json. File-plane only —gbrain config setwrites the DB plane and does NOT control capture. GBRAIN_CONTRIBUTOR_MODE=1env var is the contributor-facing toggle. Capture is off by default as of v0.25.0; production users get a quiet brain. Resolution order: expliciteval.captureconfig wins both directions, then env var, then off. Documented in README.md, CONTRIBUTING.md, anddocs/eval-bench.md.
Key commands added in v0.12.2:
gbrain repair-jsonb [--dry-run] [--json]— repair double-encoded JSONB rows left over from v0.12.0-and-earlier Postgres writes. Idempotent; PGLite no-ops. Thev0_12_2migration runs this automatically ongbrain upgrade.
Key commands added in v0.12.3:
gbrain orphans [--json] [--count] [--include-pseudo]— surface pages with zero inbound wikilinks, grouped by domain. Auto-generated/raw/pseudo pages filtered by default. Also exposed asfind_orphansMCP operation. The natural consumer of the v0.12.0 knowledge graph layer: once edges are captured, find the gaps.gbrain doctorgains two new reliability detection checks:jsonb_integrity(v0.12.0 Postgres double-encode damage) andmarkdown_body_completeness(pages truncated by the old splitBody bug). Detection only; fix hints point atgbrain repair-jsonbandgbrain sync --force.
Key commands added in v0.14.2:
gbrain sync --skip-failed— acknowledge the current set of failed-parse files recorded in~/.gbrain/sync-failures.jsonlso the sync bookmark advances past them. Doctor'ssync_failurescheck shows previously-skipped as "all acknowledged" instead of warning.gbrain sync --retry-failed— re-walk the unacknowledged failures and re-attempt parsing. If the files now succeed, they clear from the set and the bookmark advances naturally.gbrain apply-migrations --force-retry <version>— reset a wedged migration (3 consecutive partials with no completion) by appending a'retry'marker. Nextapply-migrations --yestreats the version as fresh.completestatus never regresses topartialeither before or after a retry marker.GBRAIN_POOL_SIZEenv var — honored by both the singleton pool (src/core/db.ts) and the parallel-import worker pool (src/commands/import.ts). Default is 10; lower to 2 for Supabase transaction pooler to avoid MaxClients crashes duringgbrain upgradesubprocess spawns. Read at call time viaresolvePoolSize().gbrain doctorgains two new checks:sync_failures(surfaces unacknowledged parse failures with exact paths + fix hints) andbrain_score(renders the 5-component breakdown when score < 100: embed coverage / 35, link density / 25, timeline coverage / 15, orphans / 15, dead links / 10 — sum equals total).
Key commands added in v0.26.0 (OAuth 2.1 + HTTP server + admin dashboard):
gbrain serve --http [--port 3131] [--token-ttl 3600] [--enable-dcr] [--log-full-params]— HTTP MCP server with OAuth 2.1, admin dashboard at/admin, SSE activity feed at/admin/events, health check at/health. Prints admin bootstrap token on first start. Alongside (not replacing) stdiogbrain serve. As of v0.26.9,mcp_request_log.paramsand the SSE feed default to a redacted summary ({redacted, kind, declared_keys, unknown_key_count, approx_bytes}); pass--log-full-paramsto log raw payloads on a personal laptop with a startup warning.- OAuth client registration — three paths:
- CLI:
gbrain auth register-client <name> --grant-types <types> --scopes <scopes>(wired intosrc/commands/auth.tsas a thin wrapper overGBrainOAuthProvider.registerClientManual). Default grant types:client_credentials. Default scopes:read. - Admin dashboard: Register client modal → credential reveal with Copy + Download JSON.
- SDK:
oauthProvider.registerClientManual(name, grantTypes, scopes, redirectUris)for programmatic wrappers.--enable-dcronserve --httpopens the/registerendpoint for RFC 7591 self-service registration (off by default).
- CLI:
gbrain auth create|list|revoke|test— legacy bearer tokens still work and grandfather toread+write+adminscopes on the OAuth server.authis wired as a first-classgbrainsubcommand in v0.26.0 (previously only invokable viabun run src/commands/auth.ts). No migration required to keep pre-v0.26 clients working.
Key commands added in v0.14.3 (fix wave):
gbrain doctor --index-audit— opt-in Postgres-only check reporting zero-scan indexes frompg_stat_user_indexes. Informational only; never auto-drops.gbrain doctorschema_version check fails loudly whenversion=0— catchesbun install -g github:...postinstall failures (#218) and routes users togbrain apply-migrations --yes.gbrain jobs submitgains--max-stalled,--backoff-type,--backoff-delay,--backoff-jitter,--timeout-ms,--idempotency-key— exposing existingMinionJobInputfields as first-class CLI flags.gbrain jobs smoke --sigkill-rescue— opt-in regression smoke case simulating a killed worker; asserts the v0.14.3 schema default (max_stalled=5) actually rescues on first stall.
Key commands added in v0.22.13 (PR #490):
gbrain sync --workers N(alias--concurrency N) — parallelize the import phase using per-worker Postgres engines (small pool of 2 each) with an atomic queue index. Auto-concurrency: defaults to 4 workers when the diff exceeds 100 files. Smaller diffs stay serial. Explicit--workersalways wins (even on a 30-file diff). PGLite forces serial regardless. Validation rejects0, negatives, non-integers loud (replaces the prior silent fall-through to auto-concurrency).gbrain import --workers N— sameparseWorkers()validation as sync; same try/finally worker-engine cleanup. Behavior surface unchanged.
Key commands added in v0.22.16 (claw-test friction loop):
gbrain claw-test [--scenario fresh-install|upgrade-from-v0.18] [--keep-tempdir]— scripted-mode CI gate that runs the full canonical first-day flow against a fresh tempdir. Asserts every expected--progress-jsonphase fired and doctor'sstatus === 'ok'. ~30s, no API keys.gbrain claw-test --live --agent openclaw— friction-discovery mode. Spawns real openclaw, hands itBRIEF.md, captures stdin/stdout/stderr to<run>/transcript.jsonl, lets the agent log friction via the friction CLI. Run on demand; ~5–10 min and ~$1–2 in tokens.gbrain claw-test --list-agents— reports which agent runners are registered + their detection state (binary path or unavailable reason).gbrain friction log --severity {confused|error|blocker|nit} --phase <name> --message <text> [--hint ...] [--kind {friction|delight}] [--run-id ...]— append a friction or delight entry to the active run JSONL.gbrain friction render --run-id <id> [--json] [--transcripts] [--no-redact]— markdown report grouped by severity + phase;--redactis the default for md output (strips$HOME/$CWDplaceholders so reports paste safely in PRs/issues).gbrain friction list [--json]— recent run-ids with friction/delight counts; interrupted runs marked(interrupted).gbrain friction summary --run-id <id> [--json]— two-column friction + delight summary.GBRAIN_HOMEenv override is now honored uniformly across every gbrain write site (config, audit, friction, sync-failures, import checkpoint, integrity log, integrations heartbeat, migration rollback, etc.) —gbrainPath(...)fromsrc/core/config.tsis the canonical helper. Read-side host-fingerprint detection (~/.claude/~/.openclawetc.) intentionally NOT confined in v1; that's a v1.1 follow-up.
Testing
Test command tiers (v0.26.4 — parallel fast loop)
Five tiers of test commands, each with a clear scope:
| Command | What it runs | Wallclock | When to use |
|---|---|---|---|
bun run test |
Parallel unit-test fast loop. 8-shard fan-out via scripts/run-unit-parallel.sh, then a serial pass over *.serial.test.ts. Excludes *.slow.test.ts and test/e2e/*. No pre-checks, no typecheck. |
~85s on a Mac dev box (3650+ tests) | Inner edit loop. Default. |
bun run verify |
CI's authoritative pre-test gate set: check:privacy && check:jsonb && check:progress && check:wasm && bun run typecheck. The 4 checks .github/workflows/test.yml runs on shard 1 + typecheck. Single source of truth — CI literally calls bun run verify. |
~12s (wasm-compile dominates) | Before pushing; before /ship. |
bun run test:full |
verify && bun run test && bun run test:slow && [smart e2e]. The local equivalent of "everything CI runs." Smart e2e: runs e2e only when DATABASE_URL is set; else loud skip notice to stderr. |
~3-5min depending on slow + e2e | Pre-merge sanity, before opening a PR. |
bun run test:slow |
Just the *.slow.test.ts set (intentional cold-path correctness checks). |
seconds-to-minutes | When touching slow-path code. |
bun run test:serial |
Just the *.serial.test.ts set (cross-file-contention quarantine; runs at --max-concurrency=1). |
~1s per quarantined file | Debugging a specific quarantined file. |
bun run test:e2e |
Real Postgres E2E. Requires Docker + DATABASE_URL. Sequential (template-DB parallelization is a v0.27+ TODO). |
~5-10min | Pre-ship; nightly. |
bun run check:all |
All 7 historical pre-checks (privacy + jsonb + progress + no-legacy-getconnection + trailing-newline + wasm + exports-count). Superset of verify. |
~10s | Local-only sweep. The 4 not in verify are nice-to-haves. |
CI vs local: intentionally divergent file sets
- CI matrix (
.github/workflows/test.yml) runsscripts/test-shard.sh4-way, which uses FNV-1a hash bucketing and INCLUDES*.slow.test.ts. As of v0.31.4.1, CI EXCLUDES*.serial.test.tsfrom the hash buckets and runs them on shard 1 viabun run test:serialat--max-concurrency=1. Before that, serial files were hashed in alongside parallel files, which broke themock.modulequarantine (top-level mocks in serial files leaked into the parallel files they shared a shard process with — most visibly,eval-takes-quality-runner.serial.test.tsstubbedgateway.tsand broke everygateway.embedMultimodaltest invoyage-multimodal.test.tson shard 2). CI is the ground truth for "did everything pass." - Local fast loop (
scripts/run-unit-shard.shvia the parallel wrapper) uses round-robin-by-index sharding and EXCLUDES*.slow.test.tsAND*.serial.test.ts. Local trades coverage for inner-loop speed; CI catches what local skips.
This divergence is intentional. Don't try to make them equal — the two scripts deliberately solve different problems. The regression test at test/scripts/run-unit-shard.test.ts pins what the local fast loop should and shouldn't include.
Failure-first logging
When bun run test finds any failure, the wrapper:
- Writes failure blocks (each prefixed with
--- shard N: <test name> ---) to.context/test-failures.log(workspace-local, gitignored). On systems without a writable.context/, falls back to/tmp/gbrain-test-failures.log. - Prints a loud stderr banner with the absolute log path, plus the last 30 lines of the failure log inlined. Banner survives
| head/| tail/ agent-side log truncation. - Writes a one-line-per-shard summary to
.context/test-summary.txt(shard N/M: pass=X fail=Y skip=Z rc=W). - Exits non-zero. Empty failure log + non-zero exit = infrastructure problem (wedged shard, killed child); the banner says so.
If a shard wedges (per-shard GBRAIN_TEST_SHARD_TIMEOUT cap, default 600s), the wrapper writes --- shard N: WEDGED after ${SHARD_TIMEOUT}s --- to the failure log, includes the last 50 lines of the shard log, and proceeds with other shards' results.
File taxonomy
*.test.ts→ fast loop (parallel 8-shard fan-out).*.slow.test.ts→ run viabun run test:slowonly (intentional cold-path tests; would dominate the fast loop's wallclock).*.serial.test.ts→ run viabun run test:serialafter the parallel pass completes; uses--max-concurrency=1. Quarantine for tests that share file-wide state and race when run alongside other files in the samebun testprocess. Currently:test/brain-registry.serial.test.ts,test/reconcile-links.serial.test.ts,test/core/cycle.serial.test.ts,test/embed.serial.test.ts(the latter two added in v0.26.7 — they usemock.module(...)which leaks across files in the shard process). Do not put the parallelism back on a serial file unless you've fixed the contention root cause (it just re-introduces the flake).test/e2e/*.test.ts→ real-Postgres E2E. Skipped whenDATABASE_URLis unset.
The intra-file parallelism project (turn bun test into bun test --concurrent after sweeping shared-state contention sites) is sliced across v0.26.7 (foundation), v0.26.8 (env-mutation sweep), and v0.26.9 (PGLite sweep + codemod + measurement). v0.26.4 ships file-level parallelism only.
Test-isolation lint and helpers (v0.26.7)
The cross-file flake class is enforced statically by scripts/check-test-isolation.sh, wired into bun run verify and bun run check:all. Rules (non-serial unit files only; *.serial.test.ts and test/e2e/* are skipped):
| Rule | What it bans | Fix |
|---|---|---|
| R1 | process.env.X = ..., bracket assignment, delete process.env.X, Object.assign(process.env, ...), Reflect.set(process.env, ...) |
Use withEnv() from test/helpers/with-env.ts, OR rename file to *.serial.test.ts |
| R2 | mock.module(...) anywhere in the file |
Rename file to *.serial.test.ts (no DI on production code for testability) |
| R3 | new PGLiteEngine( outside ~50 lines after a beforeAll( line |
Use the canonical block (below) inside beforeAll( |
| R4 | Files creating new PGLiteEngine( without engine.disconnect( inside an afterAll( block |
Add afterAll(() => engine.disconnect()) |
Files that violated these rules at the v0.26.7 baseline are listed in scripts/check-test-isolation.allowlist. The allow-list MUST shrink over time — never add new entries. v0.26.8 (env sweep) and v0.26.9 (PGLite sweep) remove entries as files get fixed.
Canonical PGLite block (R3 + R4 compliant)
Every test file that needs a PGLite engine should use this exact pattern:
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await resetPgliteState(engine);
});
Why this exact shape: beforeAll creates a single engine per file (PGLite WASM cold-start + initSchema is ~20s); beforeEach truncates user data via resetPgliteState ("two orders of magnitude faster" than fresh-engine-per-test); afterAll disconnects so the engine doesn't leak across file boundaries within a shard process.
withEnv pattern (R1 fix)
import { withEnv } from './helpers/with-env.ts';
test('reads OPENAI_API_KEY', async () => {
await withEnv({ OPENAI_API_KEY: 'sk-test' }, async () => {
expect(loadConfig().openai_key).toBe('sk-test');
});
});
// Delete a var (override is undefined):
await withEnv({ GBRAIN_HOME: undefined }, fn);
// Multiple keys:
await withEnv({ A: '1', B: '2', C: undefined }, fn);
withEnv saves the prior value of every key it touches and restores via try/finally — including when the callback throws. It is cross-test safe but NOT intra-file concurrent-safe. process.env is process-global; two test.concurrent() calls in the same file both touching the same key will race. Files using withEnv stay outside the future test.concurrent() codemod's eligibility filter.
When to quarantine instead of fix
Rename to *.serial.test.ts when:
- The file uses
mock.module(...)(R2 — there's no clean fix without changing production code). - The file is genuinely env-coupled (e.g.
gbrain-home-isolation.test.ts,claw-test-cli.test.ts) — module-load env readers + ESM caching defeat dynamic-import-after-env tricks. - The file's tests intentionally share state across
it()boundaries.
Quarantine count cap: 10 (informational). Beyond that, push back on the design.
Inventory (legacy)
bun test runs all tests. After the v0.12.1 release: ~75 unit test files + 8 E2E test files (1412 unit pass, 119 E2E when DATABASE_URL is set — skip gracefully otherwise). Unit tests run
without a database. E2E tests skip gracefully when DATABASE_URL is not set.
Unit tests: test/markdown.test.ts (frontmatter parsing), test/chunkers/recursive.test.ts
(chunking), test/parity.test.ts (operations contract
parity), test/cli.test.ts (CLI structure), test/config.test.ts (config redaction),
test/files.test.ts (MIME/hash), test/import-file.test.ts (import pipeline),
test/upgrade.test.ts (schema migrations),
test/file-migration.test.ts (file migration), test/file-resolver.test.ts (file resolution),
test/import-resume.test.ts (import checkpoints), test/migrate.test.ts (migration; v8/v9 helper-btree-index SQL structural assertions + 1000-row wall-clock fixtures that guard the O(n²)→O(n log n) fix + v0.13.1 assertions on v12/v13 SQL shape, sqlFor + transaction:false runner semantics, the max_stalled DEFAULT 1 regression guard, and v0.22.6.1 v24 sqlFor.pglite: '' no-op assertion),
test/bootstrap.test.ts (v0.22.6.1 — bootstrap contract: no-op on fresh install, idempotent across two initSchema() calls, no-op on modern brain that already has every probed column, full bootstrap path on simulated pre-v0.18 brain, fresh-install regression guard, pre-v0.13 links shape coverage),
test/schema-bootstrap-coverage.test.ts (v0.22.6.1 CI guard — REQUIRED_BOOTSTRAP_COVERAGE lists every forward reference in PGLITE_SCHEMA_SQL; the test fails loudly if applyForwardReferenceBootstrap skips one. When you add a column-with-index to the embedded schema blob, you extend both arrays or this guard fails. The pattern that broke gbrain ten times in two years is now structurally prevented. v0.35.5.0: test now also parses src/core/migrate.ts source text for every ALTER TABLE ... ADD COLUMN (top-level sql:, sqlFor.{postgres,pglite} overrides, AND handler-body engine.runMigration(N, \ALTER TABLE ...`)), and asserts each (table, column) pair is covered by the bootstrap OR by the schema blob's CREATE TABLE bodies. Catches the column-only forward-reference class (e.g. sources.archivedshape from v0.26.5,oauth_clients.source_idfrom v0.34.1) that the pre-existing CREATE INDEX parser couldn't see. Pre-existing parser bug fixed in same wave:parseBaseTableColumnsnow strips SQL line + block comments before identifying column names so commented-out lines no longer hide adjacent columns from coverage.),test/helpers/schema-diff.ts+test/helpers/schema-diff.test.ts+test/e2e/schema-drift.test.ts(v0.26.6 #588 — cross-engine schema parity gate. Helper exports puresnapshotSchema(query)/diffSnapshots(pg, pglite, opts)/formatDiffForFailure(diff)/isCleanDiff(diff) over a four-tuple per column (data_type, udt_name, is_nullable, column_default). E2E test spins up fresh PGLite + Postgres, runs engine.initSchema()on each (bootstrap + schema replay + migrations), snapshotsinformation_schema.columns, then diffs. 2-table allowlist (files, file_migration_ledger) — every other Postgres table must reach PGLite via PGLITE_SCHEMA_SQL or a migration's sqlFor.pglitebranch. Sentinels foroauth_clients, mcp_request_log, access_tokens, eval_candidatesgive tighter blame messages. Skip-gracefully withoutDATABASE_URL. Wired into scripts/e2e-test-map.tsso changes tosrc/schema.sql, src/core/pglite-schema.ts, or src/core/migrate.tstrigger it. The failure message names every drift with a paste-ready hint pointing atsrc/core/pglite-schema.ts.), test/setup-branching.test.ts(setup flow),test/slug-validation.test.ts(slug validation),test/storage.test.ts(storage backends),test/supabase-admin.test.ts(Supabase admin),test/yaml-lite.test.ts(YAML parsing),test/check-update.test.ts(version check + update CLI),test/pglite-engine.test.ts(PGLite engine, all 40 BrainEngine methods including 11 cases foraddLinksBatch/addTimelineEntriesBatch: empty batch, missing optionals, within-batch dedup via ON CONFLICT, missing-slug rows dropped by JOIN, half-existing batch, batch of 100 + v0.13.1 connect()error-wrap assertion (original error nested, #223 link in message, lock released)),test/engine-factory.test.ts(engine factory + dynamic imports),test/integrations.test.ts(recipe parsing, CLI routing, recipe validation),test/publish.test.ts(content stripping, encryption, password generation, HTML output),test/backlinks.test.ts(entity extraction, back-link detection, timeline entry generation),test/lint.test.ts(LLM artifact detection, code fence stripping, frontmatter validation),test/report.test.ts(report format, directory structure),test/skills-conformance.test.ts(skill frontmatter + required sections validation),test/resolver.test.ts(RESOLVER.md coverage, routing validation + v0.20.4 round-trip: every quoted RESOLVER.md trigger must match a frontmattertriggers:entry in the target skill, and everyname=""reference in any SKILL.md must resolve to a declared op insrc/core/operations.tsor a Minions handler inPROTECTED_JOB_NAMES), test/search.test.ts(RRF normalization, compiled truth boost, cosine similarity, dedup key),test/sql-ranking.test.ts(v0.22.0 source-boost helpers: 39 cases covering longest-prefix-match in SQL CASE, detail=high temporal-bypass, three-meta-char LIKE escape (%, _, \\), single-quote SQL-literal doubling, env override parsing for GBRAIN_SOURCE_BOOST + GBRAIN_SEARCH_EXCLUDE, resolveBoostMap / resolveHardExcludes merge semantics),test/dedup.test.ts(source-aware dedup, compiled truth guarantee, layer interactions),test/intent.test.ts(query intent classification: entity/temporal/event/general),test/eval.test.ts(retrieval metrics: precisionAtK, recallAtK, mrr, ndcgAtK, parseQrels),test/check-resolvable.test.ts(resolver reachability, MECE overlap, gap detection, DRY checks + v0.14.1 proximity-based DRY detection +extractDelegationTargetscoverage — 13 DRY cases),test/dry-fix.test.ts(v0.14.1 auto-fix: three shape-aware expander pure-function tests, five guards — working-tree-dirty, no-git-backup, inside-code-fence, already-delegated within 40 lines, ambiguous-multi-match, block-is-callout — 28 cases),test/doctor-fix.test.ts(v0.14.1gbrain doctor --fixCLI integration: dry-run preview, apply path, JSON output shape — 3 cases),test/backoff.test.ts(load-aware throttling, concurrency limits, active hours),test/fail-improve.test.ts(deterministic/LLM cascade, JSONL logging, test generation, rotation),test/transcription.test.ts(provider detection, format validation, API key errors),test/enrichment-service.test.ts(entity slugification, extraction, tier escalation),test/data-research.test.ts(recipe validation, MRR/ARR extraction, dedup, tracker parsing, HTML stripping),test/minions.test.ts(Minions job queue v7: CRUD, state machine, backoff, stall detection, dependencies, worker lifecycle, lock management, claim mechanics, depth/child-cap, timeouts, cascade kill, idempotency, child_done inbox, attachments, removeOnComplete/Fail + v0.13.1max_stalledclamp/default/plumbing coverage),test/extract.test.ts(link extraction, timeline extraction, frontmatter parsing, directory type inference),test/extract-db.test.ts(gbrain extract --source db: typed link inference, idempotency, --type filter, --dry-run JSON output),test/extract-fs.test.ts(gbrain extract --source fs: first-run inserts + second-run reports zero, dry-run dedups candidates across files, second-run perf regression guard — the v0.12.1 N+1 dedup bug),test/link-extraction.test.ts(canonical extractEntityRefs both formats, extractPageLinks dedup, inferLinkType heuristics, parseTimelineEntries date variants, isAutoLinkEnabled config),test/graph-query.test.ts(direction in/out/both, type filter, indented tree output),test/features.test.ts(feature scanning, brain_score calculation, CLI routing, persistence),test/file-upload-security.test.ts(symlink traversal, cwd confinement, slug + filename allowlists, remote vs local trust),test/query-sanitization.test.ts(prompt-injection stripping, output sanitization, structural boundary),test/search-limit.test.ts(clampSearchLimit default/cap behavior across list_pages and get_ingest_log),test/repair-jsonb.test.ts(v0.12.2 JSONB repair: TARGETS list, idempotency, engine-awareness),test/migrations-v0_12_2.test.ts(v0.12.2 orchestrator phases: schema → repair → verify → record),test/markdown.test.ts(splitBody sentinel precedence, horizontal-rule preservation, inferType wiki subtypes),test/orphans.test.ts(v0.12.3 orphans command: detection, pseudo filtering, text/json/count outputs, MCP op),test/postgres-engine.test.ts(v0.12.3 statement_timeout scoping:sql.begin+SET LOCALshape, source-level grep guardrail against reintroduced bareSET statement_timeout), test/sync.test.ts(sync logic + v0.12.3 regression guard asserting top-levelengine.transactionis not called),test/sync-concurrency.test.ts(v0.22.13 PR #490: 17 cases coveringautoConcurrency()thresholds + PGLite-forces-serial + explicit-override clamping,shouldRunParallel()Q1 explicit-bypasses-floor contract, andparseWorkers()validation that rejects'0'/'-3'/'foo'/'1.5'/trailing chars), test/sync-parallel.test.ts(v0.22.13 PR #490: PGLite-routed coverage of the bookmark gate under concurrency request, head-drift gate, vanished-file failure capture, PGLite-stays-serial, and thegbrain-syncwriter-lock contract — 7 cases),test/sync-failures.test.ts(v0.22.12: 28 cases pinningclassifyErrorCoderegex coverage for all 12 codes against literal production message strings frommarkdown.ts:159-244andimport-file.ts:199, 347, 352, 401; summarizeFailuresByCodesort + pre-classified-honor;recordSyncFailurescode-field persistence;acknowledgeSyncFailuresAcknowledgeResult shape + backfill on pre-v0.22.12 entries),test/doctor.test.ts(doctor command + v0.12.3 assertions thatjsonb_integrityscans the four v0.12.0 write sites andmarkdown_body_completenessis present),test/utils.test.ts(shared SQL utilities +tryParseEmbeddingnull-return and single-warn semantics),test/build-llms.test.ts(llms.txt/llms-full.txt generator: path resolution, idempotence, spec shape, regen-drift guard, content contract, AGENTS.md install-path mirror, size-budget enforcement — 7 cases),test/oauth.test.ts(v0.26.0 OAuth 2.1 provider — 27 cases: register, getClient,client_credentialsgrant exchange,authorization_codeflow with PKCE challenge / verifier, refresh token rotation,verifyAccessTokenwith both OAuth + legacyaccess_tokensfallback,revokeToken, sweepExpiredTokens, and a contract test asserting scope+localOnlyannotations are set correctly on all 30 operations; **v0.26.2** adds 5coerceTimestamp unit cases (null/undefined/string/number/throw-on-NaN), NULL-expires_at-as-expired contract tests for both refresh + access token paths, and a cascade-delete contract test asserting revoke-clientpurgesoauth_tokens+oauth_codesrows via FK CASCADE; **v0.26.9** adds 14 cases pinning the F1/F2/F3/F4/F5/F6/F7c/F12 invariants, including the F1/F4 cross-client isolation pattern (wrong-client attempt MUST reject AND rightful owner MUST still succeed atomically afterward) and the empty-stringredirect_uribypass guard surfaced during adversarial review),test/mcp-dispatch-summarize.test.ts(v0.26.9 — 7 cases pinning F8summarizeMcpParamsinvariants: declared-keys allow-list intersection, attacker-key-name leak guard (unknown keys counted not named), 1KB byte bucketing for size-probe defense, missing op falls through to fully-redacted shape, declared-keys sorted for deterministic output),test/trust-boundary-contract.test.ts(v0.26.9 — 4 cases pinning F7b fail-closed semantics under cast bypass:ctx.remote === undefinedtreated as remote/untrusted at every flipped call site,as anyandPartial<>spreads can't downgrade trust by accident),test/check-resolvable-cli.test.ts(v0.19 CLI wrapper: exit codes, JSON envelope shape, AGENTS.md fallback chain),test/regression-v0_16_4.test.ts(findRepoRoot regression guard — hermetic startDir parameterization),test/repo-root.test.ts(v0.16.4 / v0.19 / v0.31.7 — 20 cases:findRepoRootwalk semantics + default-arg parity, the 4-tierautoDetectSkillsDir fallback chain ($OPENCLAW_WORKSPACE→~/.openclaw/workspace→ repo-root →./skills), W1 RESOLVER.md/AGENTS.md filename precedence, D-CX-4 explicit-env-wins-over-repo-root, and 8 new v0.31.7 D3+D5 cases pinning tier-0 $GBRAIN_SKILLS_DIRvalid/invalid/precedence-over-OPENCLAW_WORKSPACE, the install-path walk inautoDetectSkillsDirReadOnly, no-drift on primary success, AUTO_DETECT_HINT+AUTO_DETECT_HINT_READ_ONLYcontent, and the D5 regression guard asserting the sharedautoDetectSkillsDirMUST NEVER return'install_path'source — that's how the read-path/write-path split stays safe),test/resolver-merge.test.ts(v0.31.7 — 8 cases pinning the multi-file resolver merge:findAllResolverFilesempty / RESOLVER.md-only / AGENTS.md-only / both-present (RESOLVER.md first), andcheckResolvablemerge semantics acrossskills/RESOLVER.md+../AGENTS.mdfor the OpenClaw layout where the skillpack ships a thin RESOLVER.md and the real dispatcher lives at the workspace root — dedup byskillPath(first occurrence wins), AGENTS.md-at-workspace-root works alone, and the previously-unreachable 187/224 OpenClaw skills become reachable),test/filing-audit.test.ts(v0.19 Check 6:writes_pages/writes_tofrontmatter, filing-rules JSON validation),test/routing-eval.test.ts(v0.19 Check 5: fixture parsing, structural routing, ambiguous_with, Haiku tie-break layer),test/skill-manifest.test.ts(v0.19 skill manifest parser: drift detection, managed-block markers),test/skillify-scaffold.test.ts(v0.19gbrain skillify scaffoldstubs: SKILL.md, script, tests, routing-eval fixtures),test/skillpack-install.test.ts(v0.19gbrain skillpack installmanaged-block install / update / no-clobber semantics),test/skillpack-sync-guard.test.ts(v0.19 sync-guard: bundled skills stay byte-identical toskills/source),test/http-transport.test.ts(v0.22.7 HTTP transport: 23 unit cases covering bearer auth + missing/no-Bearer/unknown/revoked +/healthbypass, F1+F2 round-trip via dispatch.ts, F3 invalid_params, application/json response shape (not SSE), CORS default-deny + allowlist, body cap on Content-Length AND chunked, two-bucket rate limit (refill, exhaust+Retry-After, LRU eviction, TTL prune, pre-auth IP fires before DB), andmcp_request_logaudit on success + auth_failed),test/restart-sweep.test.ts(v0.28.3 — 27 bun:test cases for therecipes/restart-sweep.mdinlined script: sentinel-anchored fenced-block extraction with salted tmp filenames to bypass ESM cache; constructor-time env reads (proves no module-load snapshot); idempotency layer load/save/atomic-tmp-rename/corrupt-JSON-recovery/30-day-prune;(sessionKey, lastAlertedAt)cooldown gate with 6h threshold (the C1 fix that survives synthesized restartTime); AGGRESSIVE-gate two-state tests; execFile argv shape proving shell metachars inOPENCLAW_TELEGRAM_GROUPcannot reach/bin/sh; real-\n-not-literal alert formatting; GBRAIN_HOMEstate path override),test/eval-longmemeval.test.ts(v0.28.8 LongMemEval harness — 12 hermetic cases with noDATABASE_URLand no API keys: PGLite create + reset over runtime-enumeratedpg_tables, infrastructure-table preservation across resets, JSONL question parsing, retrieval-only and answer-gen modes via stubbed ThinkLLMClient, --limitcutoff,--keyword-onlyvs hybrid, default--expansion=offbehavior, perf gate (p50 < 30ms / p99 < 50ms warm reset+import+search on Apple Silicon),--helpworks without a configured brain, fixture round-trip viatest/fixtures/longmemeval-mini.jsonl), test/longmemeval-sanitize.test.ts(v0.28.8 sanitization parity: 12 cases pinning thatINJECTION_PATTERNSfromsrc/core/think/sanitize.tsis the single source of truth — adding a pattern there must cover bothframing and<chat_session>` framing, no per-surface regex drift).
E2E tests (test/e2e/): Run against real Postgres+pgvector. Require DATABASE_URL.
bun run test:e2eruns Tier 1 (mechanical, all operations, no API keys). Includes 9 dedicated cases for the postgres-engineaddLinksBatch/addTimelineEntriesBatchbind path — postgres-js'sunnest()binding is structurally different from PGLite's and gets its own coverage.test/e2e/search-quality.test.tsruns search quality E2E against PGLite (no API keys, in-memory)test/e2e/graph-quality.test.tsruns the v0.10.3 knowledge graph pipeline (auto-link via put_page, reconciliation, traversePaths) against PGLite in-memorytest/e2e/postgres-jsonb.test.ts— v0.12.2 regression test. Round-trips all 5 JSONB write sites (pages.frontmatter, raw_data.data, ingest_log.pages_updated, files.metadata, page_versions.frontmatter) against real Postgres and assertsjsonb_typeof='object'plus->>'key'returns the expected scalar. The test that should have caught the original double-encode bug.test/e2e/integrity-batch.test.ts(v0.22.8) — parity tests forscanIntegrity's batch-load fast path vs sequential. Four cases (dedup, hits, validate, topPages) seed a fixture and assert both paths return identical results. Dedup case uses raw SQL viagetConn().unsafe()to seed a(test-source-2, people/alice)row alongside the default-source row, sinceengine.putPagedoesn't take asource_id. Pins the codex-caught multi-source overcounting regression.test/e2e/jsonb-roundtrip.test.ts— v0.12.3 companion regression against the 4 doctor-scanned JSONB sites. Assertion-level overlap withpostgres-jsonb.test.tsis intentional defense-in-depth: if doctor's scan surface ever drifts from the actual write surface, one of these tests catches it.test/e2e/sync.test.ts(v0.22.12 —--skip-failedfailure-loop test, alongside the existing 13 happy-path tests): exercises the full chain — broken file →performSyncreturnsblocked_by_failureswith grouped breakdown →performSync({skipFailed: true})advances bookmark and returnsAcknowledgeResultwith code summary → second broken file → second cycle. Saves and restores the user's real~/.gbrain/sync-failures.jsonlso the test is hermetic on a developer machine. Asserts bookmark gating, JSONL state, dedup across paths, summary aggregation, and the literal doctor-rendering string format. This is the integration test that proves the v0.22.12 chain holds together — unit tests cover the pure functions in isolation, this covers the integration.test/e2e/upgrade.test.tsruns check-update E2E against real GitHub API (network required)test/e2e/minions-shell-pglite.test.ts(v0.20.4) exercises the PGLite--followinline shell-job path (in-memory, noDATABASE_URLrequired) — the path the consolidated minion-orchestrator skill documents for dev usetest/e2e/openclaw-reference-compat.test.ts(v0.19) — exercisescheck-resolvable+skillpack installagainst a minimal AGENTS.md workspace fixture (test/fixtures/openclaw-reference-minimal/), regression guard for the 107-skill OpenClaw deployment shapetest/e2e/search-swamp.test.ts(v0.22.0) — reproduces the headline source-swamp case. Seeds a curatedoriginals/talks/article-outline-fat-codepage against twowintermute/chat/pages stuffed with the same multi-word phrase. Asserts the article wins keyword AND vector ranking, thatdetail=highlets the chat swamp re-surface (temporal-query workflow preserved), and thatsource_idpasses through the two-stage CTE intact. PGLite in-memory.test/e2e/search-exclude.test.ts(v0.22.0) — verifiestest/+archive/pages are hidden by default, thatinclude_slug_prefixesopts back in, and that caller-suppliedexclude_slug_prefixesadds to defaults. Both keyword and vector search paths covered.test/e2e/engine-parity.test.ts(v0.22.0) — Postgres ↔ PGLite top-result and result-set parity forsearchKeyword+searchVector. Codex flagged that Postgres ranks pages then picks best chunk while PGLite returns chunks directly — without parity coverage the source-boost fix could pass on PGLite and fail on Postgres. Skips gracefully whenDATABASE_URLis unset.test/e2e/postgres-bootstrap.test.ts(v0.22.6.1) — exercisesPostgresEngine.initSchema()directly against a fresh real Postgres database. Asserts the bootstrap path is no-op on fresh installs and that SCHEMA_SQL replays cleanly through the engine path (not via the standalonedb.initSchemafromsrc/core/db.ts, which would have produced false-positive coverage). Codex caught the E2E-shape gap during plan review.test/e2e/http-transport.test.ts(v0.22.7) — 8 cases against real Postgres coveringgbrain serve --httpend-to-end: bearer auth round-trip,last_used_atSQL-level debounce semantics,mcp_request_logrow insertion on success and auth_failed paths,/healthDB-down → 503 (DB-probing health check), and the F1+F2+F3 dispatch round-trip with a real operation. Skips gracefully whenDATABASE_URLis unset.test/e2e/serve-http-oauth.test.ts(v0.26.0, expanded v0.26.2, expanded v0.26.9) — real-Postgres E2E againstgbrain serve --httpwith full OAuth 2.1. Spawns a subprocess server, registers a client via the CLI, mintsclient_credentialstokens, exercises the/mcpJSON-RPC pipeline. v0.26.2 adds: real DCR/registerHTTP-level response-shape test (assertstypeof body.client_id_issued_at === 'number'over the wire — RFC 7591 §3.2.1 spec compliance, not just internal-store shape); real CLI subprocess test forrevoke-client(registers → mints token → revokes viaexecSync→ asserts token rejected at/mcp→ asserts re-run exits 1); server fixture flips on--enable-dcrso/registeris reachable. bun execSync env-inheritance fix: bun'sexecSyncdoes NOT inherit env mutations done viaprocess.env.X = ..., only OS-level env from before bun started. helpers.ts loads.env.testingand setsDATABASE_URLviaprocess.envmutation, which is invisible to subprocesses unlessenv: { ...process.env }is passed explicitly — every subprocess call in this file passesenv: { ...process.env }for that reason. Reference fix for the next maintainer hitting the same failure mode in sibling sync/cycle/dream/claw-test E2Es.afterAllcleanup is guarded onclientId(won't throw ifbeforeAllfailed before registration); cleanup errors surface to stderr without throwing so real test failures aren't masked. Tracks DCR-registered clients alongside the manual one. v0.26.9 adds 2 regressions for the F7 trust-boundary fix: an HTTP MCPsubmit_jobforname: "shell"MUST reject with a permission error (proving the request handler now setsremote: trueandsubmit_job's protected-name guard fires), and the same guard rejects subagent submission. Closes the OAuth-token-to-RCE escalation path. Skips gracefully whenDATABASE_URLis unset.test/e2e/sync-parallel.test.ts(v0.22.13 PR #490) — DATABASE_URL-gated. T2: 60-file Postgres sync at concurrency=4 imports all + no connection leak (probespg_stat_activitybefore/after to confirm worker engines disconnected). P4: 120-file serial-vs-parallel benchmark printsSYNC_PARALLEL_BENCH N files | serial=Xms | parallel(4)=Yms | speedup=Zxfor CHANGELOG quoting. Asserts parallel ≤ serial × 1.5 (CI-noise tolerant; not a strict speedup gate).test/e2e/multi-source-bug-class.test.ts(v0.32.8, PR #860) — 7-case PGLite in-memory regression suite pinning every bug site fixed in this PR:listAllPageRefsordering by(source_id, slug)(F11),getPagewith sourceId picks the right(source, slug)row (F2),extract-takesprocesses both overlappingpeople/alicerows independently,listPagesfilters correctly withPageFilters.sourceId,addLinksBatchwithfrom/to_source_idtargets the right rows (F4),validateSourceIdrejects path traversal (F6), reverse-write disk layout usesbrainDir/.sources/<id>/<slug>.mdfor non-default sources (F6). No DATABASE_URL needed. Wired intoscripts/e2e-test-map.tsso changes to extract-takes / patterns / synthesize / embed / extract / migrate-engine auto-trigger this test. Companion:test/e2e/integrity-batch.test.ts's "multi-source duplicate slugs scan once" case was pinning the pre-fix bug — assertion flipped in v0.32.8 to expect both batch + sequential paths report 2.test/e2e/source-isolation-pglite.test.ts(v0.34.1.0, #861) — 14-case PGLite in-memory regression suite pinning the source-isolation P0 seal at two layers. Engine layer:searchKeyword/searchVector/searchKeywordChunks/listPages/getPage/traverseGraph/traversePathsapplysourceId(scalar fast path) andsourceIds(array path) correctly across both engines. Op-handler layer: routes throughsourceScopeOpts(ctx)so aread+write-scoped OAuth client bound to--source dept-xcannot see rows from neighboring sources viasearch,query,list_pages,get_page, orfind_experts. Covers bothctx.sourceId(single-source clients) andctx.auth.allowedSources(federated_read clients) precedence; federated array wins over scalar wins over nothing. No DATABASE_URL needed.test/openai-compat-multimodal.test.ts(v0.34.1.0, #875) — 11-case unit suite for the gateway's openai-compatible multimodal path: happy-path single + multi-input embedding, unauthenticated proxy mode, dimension-mismatch guard (D12; throwsAIConfigErrorwith model id + observed + expected pre-storage), default-dim fallback when recipe declaresdefault_dims, HTTP 401 / 400 / malformed-JSON / non-array error paths, plus a regression test that the existing Voyage/multimodalembeddingsrecipe still routes through its dedicated path (not the openai-compatible one). Hermetic via the__setEmbedTransportForTestsseam.test/serve-stdio-lifecycle.test.ts(extended v0.34.1.0, #870) — adds 3 new cases for theMCP_STDIO=1env guard: stdin EOF does NOT trigger shutdown when the env is set, SIGTERM still does (guard scope is correct), unset env preserves the pre-v0.34 CLI lifecycle. Exercises theServeOptions.mcpStdio?: booleantest seam directly so tests don't mutateprocess.env.test/oauth.test.ts(extended v0.34.1.0, #909) — 5 new cases for the PKCE DCR public-client gate:registerClientwithtoken_endpoint_auth_method: "none"returns noclient_secretfield on the public client, defaultclient_secret_postclients still get the one-time-reveal secret,getClientNULL→undefined normalization so the SDK's clientAuth path accepts public clients, full PKCE/authorize→/tokenround-trip against a public client (no client_secret presented), and a regression test that the public-vs-confidential branch doesn't break confidential clientclient_secret_postexchange.- Tier 2 (
skills.test.ts) requires OpenClaw + API keys, runs nightly in CI - If
.env.testingdoesn't exist in this directory, check sibling worktrees for one:find ../ -maxdepth 2 -name .env.testing -print -quitand copy it here if found. - Run E2E tests without asking permission. When you want to verify behavior, there's a relevant E2E test, or you're shipping touching anything covered by an E2E suite — just spin up the test DB, run the tests, and tear down. Don't ask, don't propose it, don't defer. The lifecycle is short (~2-30s startup, sub-minute tests, instant teardown) and the gate value is high. Skipping with "DATABASE_URL unset" is silent regression, not caution.
API keys and running ALL tests
ALWAYS source the user's shell profile before running tests:
source ~/.zshrc 2>/dev/null || true
This loads OPENAI_API_KEY and ANTHROPIC_API_KEY. Without these, Tier 2 tests
skip silently. Do NOT skip Tier 2 tests just because they require API keys — load
the keys and run them.
When asked to "run all E2E tests" or "run tests", that means ALL tiers:
- Tier 1:
bun run test:e2e(mechanical, sync, upgrade — no API keys needed) - Tier 2:
test/e2e/skills.test.ts(requires OpenAI + Anthropic + openclaw CLI) - Always spin up the test DB, source zshrc, run everything, tear down.
E2E test DB lifecycle (ALWAYS follow this)
You are responsible for spinning up and tearing down the test Postgres container. Do not leave containers running after tests. Do not skip E2E tests, do not ask permission to run them — see the "run without asking" rule above.
- Check for
.env.testing— if missing, copy from sibling worktree. Read it to get the DATABASE_URL (it has the port number). - Check if the port is free:
docker ps --filter "publish=PORT"— if another container is on that port, pick a different port (try 5435, 5436, 5437) and start on that one instead. - Start the test DB:
Wait for ready:
docker run -d --name gbrain-test-pg \ -e POSTGRES_USER=postgres -e POSTGRES_PASSWORD=postgres \ -e POSTGRES_DB=gbrain_test \ -p PORT:5432 pgvector/pgvector:pg16docker exec gbrain-test-pg pg_isready -U postgres - Bootstrap the schema (required — fresh containers have no
oauth_clients,mcp_request_log,pagesetc.; tests likeserve-http-oauth.test.tswill fail withrelation "oauth_clients" does not existif you skip this):DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test \ bun run src/cli.ts doctor --json > /dev/null 2>&1gbrain doctortriggersinitSchema()on first connect, which is the canonical way to bring a fresh DB to head.apply-migrations --yesalone does NOT seed the base schema — it runs ALTER-style migrations on top ofinitSchema. Tests that bypass the engine (rawexecSync-spawnedauth register-client) hit the schema directly and need this step to have run first. - Run E2E tests:
DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test bun run test:e2e - Tear down immediately after tests finish (pass or fail):
docker stop gbrain-test-pg && docker rm gbrain-test-pg
Never leave gbrain-test-pg running. If you find a stale one from a previous run,
stop and remove it before starting a new one.
Search Mode (v0.32.3)
GBrain ships three named search modes that bundle the search-lite knobs from
PR #897 into a single config key. Pick one at install time; the rest of the
project resolves through src/core/search/mode.ts.
| Knob | conservative |
balanced |
tokenmax |
|---|---|---|---|
cache.enabled |
true | true | true |
cache.similarity_threshold |
0.92 | 0.92 | 0.92 |
cache.ttl_seconds |
3600 | 3600 | 3600 |
intentWeighting |
true | true | true |
tokenBudget |
4000 | 12000 | off |
expansion (LLM multi-query) |
false | false | true |
searchLimit default |
10 | 25 | 50 |
Cost anchors (downstream agent input cost — gbrain itself is rounding error). The corner-to-corner spread is 25x once you pair mode with downstream model. Chunks ~400 tokens avg. Per-query cost @ 10K queries/month (typical single-user volume), full search payload, no cache savings:
| Mode \ Downstream | Haiku 4.5 ($1/M) | Sonnet 4.6 ($3/M) | Opus 4.7 ($5/M) |
|---|---|---|---|
| conservative (~4K) | $40/mo | $120/mo | $200/mo |
| balanced (~10K) | $100/mo | $300/mo | $500/mo |
| tokenmax (~20K) | $200/mo | $600/mo | $1,000/mo |
Scales linearly: multiply by 10 for 100K/mo (heavy power user / multi-user fleet); divide by 10 for 1K/mo (light usage). Natural pairings span ~4x. Mismatches (tokenmax+Haiku, conservative+Opus) waste capacity differently — too-big payload overwhelms a cheap model; too-small payload starves an expensive one.
tokenmax adds ~$1.50 per 1K queries in Haiku expansion calls on top of
the matrix ($15/mo @ 10K). Cache hits cut all numbers ~50%. The cost
picker copy in gbrain init carries the same matrix verbatim — update
both when refreshing.
Per-query math vs real-world spend. The matrix above is what an
isolated benchmark would measure. Real agent loops with disciplined
Anthropic prompt caching see 50-80% discount on top (cache hits skip
downstream entirely). The realistic-scale anchor in
docs/eval/SEARCH_MODE_METHODOLOGY.md walks the natural pairings at
single-power-user volume (~860 turns/mo): tokenmax+Opus ~$700/mo,
balanced+Sonnet ~$430/mo, conservative+Haiku ~$170/mo. Setups WITHOUT
cache-aware prompt layout (frequent prefix churn) see the per-query
matrix dominate — mode + model choice matters more there.
Resolution chain (matches the v0.31.12 model-tier pattern at
src/core/model-config.ts:resolveModel):
per-call SearchOpts → per-key config (search.cache.enabled, …) →
MODE_BUNDLES[search.mode] → MODE_BUNDLES.balanced (fallback)
Mode resolution lives in bare hybridSearch (NOT just the cached wrapper)
per [CDX-5+6] in ~/.claude/plans/lets-take-a-look-validated-parrot.md — so
gbrain eval replay and gbrain eval longmemeval test the same mode-affected
behavior as the production query op.
Cache-key contamination hotfix [CDX-4]: migration v56 added a
knobs_hash column to query_cache. The lookup filter is now
WHERE source_id = $ AND knobs_hash = $ AND embedding similarity < $ so a
tokenmax write (expansion=on, limit=50) can't be served to a conservative
read.
Three CLI surfaces:
gbrain search modes # what is running, with per-knob attribution
gbrain search modes --reset # clear search.* overrides (mode bundle wins)
gbrain search stats [--days N] # cache hit rate, intent mix, budget drops
gbrain search tune [--apply] # data-driven recommendations
The install picker fires inside gbrain init AFTER engine.initSchema()
(non-TTY auto-selects). The upgrade banner fires once via runPostUpgrade
in src/commands/upgrade.ts, gated by search.mode_upgrade_notice_shown.
Eval discipline (v0.32.3)
Every metric printed by any gbrain eval * or gbrain search stats command
resolves through src/core/eval/metric-glossary.ts so industry terms
(P@k, nDCG@k, MRR, Jaccard@k) carry a plain-English line in human
output and a _meta.metric_glossary block in JSON output (one block per
response per [CDX-25], NOT sibling _gloss fields).
The full methodology — datasets, sample selection, pre-registered
expectations, threats to validity, paired-bootstrap + Bonferroni p-value
discipline [CDX-14] — lives in docs/eval/SEARCH_MODE_METHODOLOGY.md.
Auto-regenerated docs/eval/METRIC_GLOSSARY.md is CI-guarded against
drift (scripts/check-eval-glossary-fresh.sh).
Per-run records land at <repo>/.gbrain-evals/eval-results.jsonl per
[CDX-23]. The user's personal ~/.gbrain brain is NEVER touched —
audit trail lives in the source repo's git history.
Skills
Read the skill files in skills/ before doing brain operations. GBrain ships 29 skills
organized by skills/RESOLVER.md (AGENTS.md is also accepted as of v0.19):
Original 8 (conformance-migrated): ingest (thin router), query, maintain, enrich, briefing, migrate, setup, publish.
Brain skills (ported from an upstream agent fork): signal-detector, brain-ops, idea-ingest, media-ingest, meeting-ingestion, citation-fixer, repo-architecture, skill-creator, daily-task-manager.
Operational + identity: daily-task-prep, cross-modal-review, cron-scheduler, reports,
testing, soul-audit, webhook-transforms, data-research, minion-orchestrator. As of
v0.20.4, minion-orchestrator is the single unified skill for both lanes of background
work (shell jobs via gbrain jobs submit shell, LLM subagents via gbrain agent run) ...
the prior gbrain-jobs skill was merged in, Preconditions are shared, and trigger
routing is narrowed to what the skill actually covers.
Skillify loop (v0.19): skillify (the markdown orchestration), skillpack-check (agent-readable health report).
Routing-table compression (v0.32.3.0): skills/functional-area-resolver/ —
two-layer dispatch pattern for shrinking large AGENTS.md / RESOLVER.md files
(>=12KB) without losing routing accuracy. Replaces one row per skill with one
entry per functional area, where each area declares its sub-skills in a
(dispatcher for: ...) clause. The static-prompt analog of hierarchical agent
routing (AnyTool arXiv:2402.04253, RAG-MCP
arXiv:2505.03275, Anthropic Agent Skills
progressive disclosure). Empirically validated across Opus 4.7 / Sonnet 4.6 /
Haiku 4.5: +13 to +17pp over the verbose baseline at 48% the size (25KB → 13KB
on a real fork). The (dispatcher for: ...) clause is the load-bearing signal
— strip it and lenient accuracy collapses to 41.7% on Sonnet (the
resolver-of-resolvers ablation case). A/B eval surface lives at
evals/functional-area-resolver/ (outside skills/ deliberately so the
skillpack bundler doesn't ship eval infrastructure to downstream installs):
gateway-routed TypeScript harness, 20 training + 5 held-out fixtures, strict +
lenient scoring, three committed cross-model receipts in baseline-runs/.
Receipt header binds (model, prompt_template_hash, fixtures_hash, harness_sha,
ts) so future contributors can verify reproduction. Companion rescore.mjs
re-scores existing JSONL with lenient tolerance for zero API cost. Reproduce
with cd evals/functional-area-resolver && node harness.mjs --model {opus|sonnet|haiku} (~$0.30–1.70 per model). Nine v0.33.x follow-up TODOs
filed for held-out corpus growth, cross-vendor verification, hierarchical
area-of-areas, embedding-based pre-router, and the run-1 vs run-2
prompt-design ablation methodology.
Operational health (v0.19.1): smoke-test (8 post-restart health checks with auto-fix
for Bun, CLI, DB, worker, Zod CJS, gateway, API key, brain repo; user-extensible via
~/.gbrain/smoke-tests.d/*.sh).
Conventions: skills/conventions/ has cross-cutting rules (quality, brain-first,
model-routing, test-before-bulk, cross-modal). skills/_brain-filing-rules.md and
skills/_output-rules.md are shared references.
Bulk-action progress reporting
All bulk commands (doctor, embed, import, export, sync, extract, migrate,
repair-jsonb, orphans, check-backlinks, lint, integrity auto, eval, files
sync, and apply-migrations) stream progress through the shared reporter
at src/core/progress.ts. Agents get heartbeats within 1 second of every
iteration regardless of how slow the underlying work is.
Rules:
- Progress always writes to stderr. Stdout stays clean for data output
(
--jsonpayloads, final summaries, JSON action events fromextract). - Non-TTY default: plain one-line-per-event human text. JSON requires the
explicit
--progress-jsonflag. - Global flags (
--quiet,--progress-json,--progress-interval=<ms>) are parsed bysrc/core/cli-options.tsBEFORE command dispatch. - Phase names are machine-stable
snake_case.dot.path(e.g.doctor.db_checks,sync.imports). Documented indocs/progress-events.md; additive changes only. scripts/check-progress-to-stdout.shis a CI guard that fails the build if any new code writes\rprogress to stdout. Wired intobun run test.- Minion handlers pass
job.updateProgressas theonProgresscallback to core functions (DB-backed primary progress channel); stderr fromjobs workstays coarse for daemon liveness only.
When wiring a new bulk command: import { createProgress } from '../core/progress.ts'
and import { getCliOptions, cliOptsToProgressOptions } from '../core/cli-options.ts'.
Create a reporter with createProgress(cliOptsToProgressOptions(getCliOptions())),
start(phase, total?) before the loop, tick() inside it, finish() after.
For single long-running queries, use startHeartbeat(reporter, note) with a
try/finally to guarantee cleanup. Never call process.stdout.write('\r...')
in bulk paths, the CI guard will fail the build.
Capturing test output (NEVER pipe through tail / head)
Iron rule: when running bun test, bun run test:e2e, bun run typecheck,
or any other test/check command, redirect to a file FIRST, then tail the file
separately:
# RIGHT — full output preserved, real exit code visible
bun test > /tmp/ship_units.txt 2>&1
echo "EXIT=$?"
tail -50 /tmp/ship_units.txt
grep -E '(fail\)|✗|error:' /tmp/ship_units.txt | head -30
# WRONG — exit code is `tail`'s (always 0), failures truncated, ship gates fail open
bun test 2>&1 | tail -10
The pipe form silently breaks /ship Step T1 (test failure ownership triage) and the test verification gate (Step 16) because:
$?after a pipe is the LAST command's exit code (tail→ 0), not bun's- bun prints failure details before the summary line, so
tail -Ndrops them - Step T1 needs the full failure list to classify in-branch vs pre-existing
This bit us during v0.26.2 ship: bun test 2>&1 | tail -10 reported "3911 pass / 23 fail"
but no failure details survived, forcing a 23-minute re-run to triage.
Apply the same pattern to any long-running command whose exit code matters:
bun run typecheck, bun run ci:local, migration runs, eval suites, etc.
For background tasks (run_in_background: true), the harness captures the exit
file separately — use it via the bg task's <id>.exit file, not the streamed
output.
Build
bun build --compile --outfile bin/gbrain src/cli.ts
Version locations (single source of truth: VERSION file)
Every release advances the version in five files at once. Keep these in
sync. /ship enforces this via Step 12's idempotency check (VERSION vs
package.json drift), but the canonical list lives here so future runs and
the auto-update agent know where to look.
Version format is mandatory: MAJOR.MINOR.PATCH.MICRO (four numeric
segments, dot-separated, no leading v). Every new release MUST use the
4-segment form. The .MICRO slot is the dot-suffix follow-up channel: when
a release ships its commit subject ahead of its VERSION bump (e.g. PR #795
landing as v0.31.4 without bumping the file), the corrective ship lands
as 0.31.4.1 rather than churning the patch number to 0.31.5. Suffixes
like -fixwave are still allowed as needed (0.31.1.1-fixwave), but the
four numeric segments are required first. Historical 3-segment versions
(0.31.3, 0.22.1) remain valid in git log and migration filenames
(skills/migrations/v0.21.0.md); do NOT rewrite them. Going forward only.
Required (every release must update all five):
| File | What lives there | Format |
|---|---|---|
VERSION |
The single source of truth. Read first by /ship, the binary, and CI version-gate. |
Bare 4-segment string MAJOR.MINOR.PATCH.MICRO (e.g. 0.31.4.1), no leading v. |
package.json |
Bun/npm package version. gbrain --version reads it via the compiled binary's bundled package metadata. CI version-gate cross-checks this against VERSION and fails if they drift. |
"version": "0.31.4.1" |
CHANGELOG.md |
Top entry header ## [0.31.4.1] - YYYY-MM-DD plus the "To take advantage of v0.31.4.1" block. |
Standard Keep-a-Changelog header. |
TODOS.md |
Any TODO entries that mention "follow-up from vX.Y.Z.W" use the version of the release that filed them. Update only when filing NEW follow-up TODOs. | Inline vX.Y.Z.W references in TODO bodies. |
CLAUDE.md |
The Key Files section's per-file annotations carry vX.Y.Z.W (#NNN) tags noting which release introduced a behavior. Update whenever a wave's annotations get folded in. |
Inline vX.Y.Z.W (#NNN, contributed by @user) references. |
Auto-derived (no manual edit; refreshed by their own commands):
bun.lock— root-package version is auto-pinned frompackage.json. After bumpingpackage.json, runbun installto refresh the lockfile.llms-full.txt/llms.txt— auto-generated documentation bundles. Any CLAUDE.md edit MUST be followed bybun run build:llmsin the same commit (or a follow-up commit before push). The committed bundles are checked against fresh generator output bytest/build-llms.test.ts, which runs in CI shard 1. If you edited CLAUDE.md and didn't regenerate, CI will fail. This has bitten the wave 3 times — every CLAUDE.md edit gets abun run build:llmschaser, no exceptions. (Theverifygate doesn't run this test; only the full unit suite does. Sobun run typecheckclean is NOT enough to know you can push after a CLAUDE.md edit.)
Historical (DO NOT bump on release):
skills/migrations/v0.21.0.md— migration files use the version they shipped FROM as their filename. v0.21.0's migration always says v0.21.0.src/commands/migrations/v0_21_0.ts— same: migration code references the schema version it migrates to.test/migrations-v0_21_0.test.ts,test/migration-orchestrator-v0_21_0.test.ts,test/migrate.test.ts— migration tests reference historical migration versions; these are correct as-is and should not move.src/core/db.ts,src/core/migrate.ts,src/core/import-file.ts,src/commands/reindex-code.ts— code comments cite the release that introduced a feature. Once written, these are historical record.README.md— references the latest published feature names by version (e.g. "v0.21.0 Code Cathedral"); update only when the README's marketing copy is intentionally being refreshed, NOT on every micro/patch bump.
The /ship workflow's version idempotency check: Step 12 reads
VERSION and package.json, classifies as FRESH / ALREADY_BUMPED /
DRIFT_STALE_PKG / DRIFT_UNEXPECTED, and refuses to proceed on
DRIFT_UNEXPECTED. This is why the two must move together.
The CI version-gate rejects pushes where VERSION and
package.json disagree, OR where VERSION is not strictly greater
than master's VERSION. If a queue collision claims your version on
master before yours lands, /ship's queue-aware allocator (Step 12)
will detect drift and re-bump on the next run.
Mandatory version-consistency audit (run after EVERY merge or commit that touches VERSION, package.json, or CHANGELOG)
The trio MUST agree. Every merge from master will hit conflicts on VERSION + package.json + CHANGELOG.md because master ships its own version bumps. Auto-merge sometimes resolves these silently in unexpected ways. After any merge, branch update, or version-related edit, run this audit. It's three lines and never lies:
echo "VERSION: $(cat VERSION)"
echo "package.json: $(node -e 'process.stdout.write(require("./package.json").version)')"
grep -E "^## \[" CHANGELOG.md | head -1
All three MUST show the same MAJOR.MINOR.PATCH.MICRO. If any one
disagrees, you have not finished the merge. Fix it before pushing or
shipping. There is no situation in which "I'll fix it next push" is OK,
because:
- A green local test run with mismatched VERSION/package.json still fails the CI version-gate.
- A green CHANGELOG entry under the wrong version header silently lies to release-notes consumers.
- /ship's Step 12 idempotency check classifies a mismatch as
DRIFT_UNEXPECTEDand HALTS — but only if you remember to run /ship before pushing. Manualgit pushskips the check.
Merge-conflict recovery procedure (memorize this)
When git merge origin/master reports conflicts on VERSION,
package.json, or CHANGELOG.md, resolve in this exact order:
- VERSION — overwrite with the wave's version (`echo -n "X.Y.Z.W"
VERSION`). Highest semver wins; do NOT take master's lower version.
- package.json — strip the conflict markers, keep the wave's
version line. Sed pattern:
sed -i.bak '/^<<<<<<< HEAD$/d; /^=======$/,/^>>>>>>> /d' package.json && rm package.json.bak(assumes ours is above the=======). - CHANGELOG.md — strip ALL three conflict markers; both your entry
and master's entry stay. Sed pattern:
sed -i.bak '/^<<<<<<< HEAD$/d; /^=======$/d; /^>>>>>>> origin\/master$/d' CHANGELOG.md && rm CHANGELOG.md.bakThen verify your entry is the topmost## [X.Y.Z.W]and master's newer-than-yours entries (if any) sit below. - Run the 3-line audit above. If it doesn't show your version on all three lines, you missed a marker.
- Run
bun installto refreshbun.lockagainst the resolvedpackage.json. Stage and commit if it changed. - Run
bun run typecheckbefore committing the merge. - Only THEN run
git commitfor the merge.
If the audit shows drift after step 4, do NOT proceed to step 5. Re-run steps 1-3 against the actual file content; you missed a marker or resolved one in the wrong direction.
Anti-pattern to avoid: Resolving via git checkout --ours package.json
and git checkout --theirs scripts/test-shard.sh mixed in the same
commit. The selective directional resolution is fine, but on
VERSION/package.json/CHANGELOG specifically, ALWAYS use the explicit
echo > VERSION + sed-strip-markers pattern above. The directional
checkout flags have bitten us when the conflict shape was unexpected
(e.g. master stripped a section we expected to keep).
Pre-push gate (manual; tighten when you remember to)
Before any git push of a merge commit, run the audit one more time:
echo "VERSION: $(cat VERSION)"
echo "package.json: $(node -e 'process.stdout.write(require("./package.json").version)')"
grep -E "^## \[" CHANGELOG.md | head -1
If you've been editing the branch via /ship you can rely on Step 12's
idempotency check. If you've been editing manually (merge resolution,
conflict fix, version bump), the audit is the last line of defense
before CI yells at you.
Pre-ship requirements
Before shipping (/ship) or reviewing (/review), always run the full test suite. Two equivalent paths:
Path A — local CI gate (recommended, v0.23.1+):
bun run ci:localruns the entire stack inside Docker: gitleaks (host), unit tests withDATABASE_URLunset, and all 29 E2E files sequentially against a fresh pgvector container. Stronger than PR CI's 2-file Tier 1 set; closer to what nightly Tier 1 catches. Spins up + tears down postgres automatically viadocker-compose.ci.yml. Override the host port withGBRAIN_CI_PG_PORT=5435 bun run ci:localif 5434 collides.bun run ci:local:diffruns only the E2E files matched by the diff selector (scripts/select-e2e.ts), falling back to all 29 on unmapped src/ paths or schema/skills/package.json changes. Fast iteration during a focused branch.
Path B — manual lifecycle (still supported):
bun test— unit tests (no database required)- Follow the "E2E test DB lifecycle" steps above to spin up the test DB,
run
bun run test:e2e, then tear it down.
Both must pass. Do not ship with failing E2E tests. Do not skip E2E tests.
Always run typecheck before pushing. bun test (the bun runner)
skips TypeScript type checking — it only enforces runtime behavior.
Three ways to actually gate on types:
bun run test(npm script inpackage.json) — includesbun run typecheckplus the four shell pre-checks (check-jsonb-pattern.sh,check-progress-to-stdout.sh,check-trailing-newline.sh,check-wasm-embedded.sh) before the runner. Use this mid-branch.bun run typecheck—tsc --noEmitstandalone. Fast (~5s on this repo).bun run ci:local— the full local CI gate from Path A.
The trap is: writing a new test, running bun test test/foo.test.ts,
seeing it pass, pushing — and CI's separate typecheck stage rejects an
invalid type literal that the runner accepted. Caught one of these
shipping the v0.23.2 round-trip E2E (type: 'reflection' is not a
member of PageType). Run bun run typecheck once before push, even
when only test files changed.
Post-ship requirements (MANDATORY)
After EVERY /ship, you MUST run /document-release. This is NOT optional. Do NOT skip it. Do NOT say "docs look fine" without running it. The skill reads every .md file in the project, cross-references the diff, and updates anything that drifted.
If /ship's Step 8.5 triggers document-release automatically, that counts. But if it gets skipped for ANY reason (timeout, error, oversight), you MUST run it manually before considering the ship complete.
Files that MUST be checked on every ship:
- README.md — does it reflect new features, commands, or setup steps?
- CLAUDE.md — does it reflect new files, test files, or architecture changes?
- CHANGELOG.md — does it cover every commit?
- TODOS.md — are completed items marked done?
- docs/ — do any guides need updating?
A ship without updated docs is an incomplete ship. Period.
CHANGELOG + VERSION are branch-scoped
VERSION and CHANGELOG describe what THIS branch adds vs master, not how we got here. Every feature branch that ships gets its own version bump and CHANGELOG entry. The entry is product release notes for users; it is not a log of internal decisions, review rounds, or codex findings.
Write the CHANGELOG entry at /ship time, not during development. Mid-branch
iterations, review rounds (CEO/Eng/Codex/DX), and implementation detours belong
in the plan file at ~/.claude/plans/, not in the CHANGELOG. One unified entry
per branch, covering what the branch added vs the base branch.
Never edit a CHANGELOG entry that already landed on master. If master has v0.18.2 and your branch adds features, bump to the next version (v0.19.0, not editing master's v0.18.2). When merging master into your branch, master may bring new CHANGELOG entries above yours — push your entry above master's latest and verify:
- Does CHANGELOG have your branch's own entry separate from master's entries?
- Is VERSION higher than master's VERSION?
- Is your entry the topmost
## [X.Y.Z]entry? grep "^## \[" CHANGELOG.mdshows a contiguous version sequence?
If any answer is no, fix it before continuing.
CHANGELOG is for users, not contributors. Write like product release notes:
- Lead with what the user can now do that they couldn't before. Sell the capability.
- Plain language, not implementation details. "You can now..." not "Refactored the..."
- Never mention internal artifacts: plan file IDs, decision tags (D-CX-#, F-ENG-#), review rounds, codex findings, subcontractor credits. These are invisible to users.
- Put contributor-facing changes in a separate
### For contributorssection at the bottom. - Every entry should make someone think "oh nice, I want to try that."
What to omit:
- "Codex caught X that the CEO review missed" — private process detail.
- "D-CX-3 split errors/warnings" — tag is meaningless to users; name the feature instead.
- "Fix-wave PR #N supersedes #M" — supersede chains belong in PR bodies, not release notes.
- "215 new cases, 3 decisions applied, 7 reviews cleared" — these are planning-mode metrics.
What to keep:
- The user-facing change: what commands exist now, what flag was added, what behavior fixed.
- Numbers that mean something to the user: TTHW, commands that timed out before, detection counts.
- Upgrade instructions:
gbrain upgrade+ any manual step if needed. - Credit to external contributors when a community PR was incorporated.
CHANGELOG voice + release-summary format
Every version entry in CHANGELOG.md MUST start with a release-summary section in
the GStack/Garry voice — one viewport's worth of prose + tables that lands like a
verdict, not marketing. The itemized changelog (subsections, bullets, files) goes
BELOW that summary, separated by a ### Itemized changes header.
The release-summary section gets read by humans, by the auto-update agent, and by anyone deciding whether to upgrade. The itemized list is for agents that need to know exactly what changed.
Release-summary template
Iron rule: lead ELI10, get precise after. The first ~150 words of every entry must be readable by someone who does NOT know gbrain's internals. No file paths, no function names, no internal constants, no acronyms (no "RRF", no "knobsHash", no "MODE_BUNDLES", no "CDX-4"), no jargon that requires reading the codebase to parse. Lead with the user-visible behavior change, in everyday English, like you're explaining it to a smart engineer who has never opened the repo.
THEN, once the reader knows what shipped and why they'd care, drill into the precise details: real file paths, real function names, real config keys, real numbers. The precision part is required (the entry is also the technical record of what changed), but it lives AFTER the plain-English lead, never before it.
The shape:
- One-line bold headline. What changed for the user, in human English. No jargon. No internal terms. Example good: "Your search stops boosting weak pages just because they have a lot of links pointing at them." Example bad: "PostFusionOpts gains floorRatio; KNOBS_HASH_VERSION bumped 2→3."
- Plain-English opener (~3-5 sentences). Describe the problem this fixes in everyday terms. Pretend the reader has a brain full of meeting notes and people pages and wants to know if this release helps them. Concrete example beats abstract description.
- A "How to turn it on" or "How to use it" section with paste-ready commands. Real flags, real config keys. This is where precision starts.
- A "What you'd see in a concrete example" or "The X numbers that matter" section with a table. Use everyday-language column headers ("Page", "Match quality", "Has many backlinks?") even when the underlying mechanism is technical. The table teaches what the feature does without requiring the reader to understand how.
- A "What's safe to know about" or "Things to watch" section for caveats, side effects, cache invalidation, mid-deploy notes. Still in plain language.
- A "What we caught and fixed before merging" section if the work went through review (CEO/eng/codex/outside-voice). Translate review findings into plain English. "We caught a stale-cache bug" beats "knobsHash() did not include floorRatio in the v=2 hash input."
### Itemized changes(precision lives here). File paths, function names, types, constants, line numbers. This section is for engineers who need to know exactly what moved.
Voice rules (apply throughout):
- No em dashes (use commas, periods, "...").
- No AI vocabulary (delve, robust, comprehensive, nuanced, fundamental, etc.) or banned phrases ("here's the kicker", "the bottom line", etc.).
- Real numbers, real file names, real commands AFTER the ELI10 lead. Not "fast" but "~30s on 30K pages." In the ELI10 lead, "fast enough that you won't notice" or "~30 seconds even on a big brain."
- Short paragraphs, mix one-sentence punches with 2-3 sentence runs.
- Connect to user outcomes: "the agent does ~3x less reading" beats "improved precision."
- Be direct about quality. "Well-designed" or "this is a mess." No dancing.
The smell test: if someone who has never opened gbrain reads the first 150 words and walks away knowing what shipped and whether they care, the entry passes. If they need to grep the codebase to follow along, rewrite the lead.
Canonical examples in this CHANGELOG: v0.35.6.0 (floor-ratio gate, written ELI10-lead-first), v0.34.4.0 (embed stale fix wave). Use those shapes when in doubt. Avoid the shape of entries that lead with internal constants or release mechanics; those exist in older history but should not be the model for new work.
Source material to pull from:
- CHANGELOG.md previous entry for prior context
- Latest
gbrain-evals/docs/benchmarks/[latest].mdfor headline numbers (sibling repo) - Recent commits (
git log <prev-version>..HEAD --oneline) for what shipped - Don't make up numbers. If a metric isn't in a benchmark or production data, don't include it. Say "no measurement yet" if asked.
Target length: ~250-350 words for the summary. Should render as one viewport.
"To take advantage of v[version]" block (required, v0.13+)
After the release-summary and BEFORE ### Itemized changes, every ## [X.Y.Z]
entry MUST include a human-readable self-repair block under the heading
## To take advantage of v[version].
Why: gbrain upgrade runs gbrain post-upgrade which runs gbrain apply-migrations.
This chain has a known weak link — upgrade.ts catches post-upgrade failures as
best-effort (so the binary still works). When that chain silently fails, users end
up with half-upgraded brains. The self-repair block gives them a paste-ready
recovery path; the v0.13+ ~/.gbrain/upgrade-errors.jsonl trail + gbrain doctor
integration close the loop.
Template (adapt the verify commands per release):
## To take advantage of v[version]
`gbrain upgrade` should do this automatically. If it didn't, or if `gbrain doctor`
warns about a partial migration:
1. **Run the orchestrator manually:**
```bash
gbrain apply-migrations --yes
-
Your agent reads
skills/migrations/v[version].mdthe next time you interact with it. [One sentence on whether headless agents need manual action, or whether the orchestrator already handled the mechanical side.] -
Verify the outcome:
[release-specific verify commands, e.g. `gbrain graph ... --depth 2`] gbrain stats -
If any step fails or the numbers look wrong, please file an issue: https://github.com/garrytan/gbrain/issues with:
- output of
gbrain doctor - contents of
~/.gbrain/upgrade-errors.jsonlif it exists - which step broke
This feedback loop is how the gbrain maintainers find fragile upgrade paths. Thank you.
- output of
**Skip this block** for patches that are pure bug fixes with zero user-facing action
(rare). If the release has a schema migration, data backfill, or new feature the
user needs to verify, the block is required.
The v0.13.0 entry in CHANGELOG.md is the canonical example.
### Itemized changes (the existing rules)
Below the release summary, write `### Itemized changes` and continue with the
detailed subsections (Knowledge Graph Layer, Schema migrations, Security hardening,
Tests, etc.). Same rules as before:
- Lead with what the user can now DO that they couldn't before
- Frame as benefits and capabilities, not files changed or code written
- Make the user think "hell yeah, I want that"
- Bad: "Added GBRAIN_VERIFY.md installation verification runbook"
- Good: "Your agent now verifies the entire GBrain installation end-to-end, catching
silent sync failures and stale embeddings before they bite you"
- Bad: "Setup skill Phase H and Phase I added"
- Good: "New installs automatically set up live sync so your brain never falls behind"
- **Always credit community contributions.** When a CHANGELOG entry includes work from
a community PR, name the contributor with `Contributed by @username`. Contributors
did real work. Thank them publicly every time, no exceptions.
### Reference: v0.12.0 entry as canonical example
The v0.12.0 entry in CHANGELOG.md is the canonical example of the format. Match its
structure for every future version: bold headline, lead paragraph, "numbers that
matter" with BrainBench-style before/after table, "what this means" closer, then
`### Itemized changes` with the detailed sections below.
## Version migrations
Create a migration file at `skills/migrations/v[version].md` when a release
includes changes that existing users need to act on. The auto-update agent
reads these files post-upgrade (Section 17, Step 4) and executes them.
**You need a migration file when:**
- New setup step that existing installs don't have (e.g., v0.5.0 added live sync,
existing users need to set it up, not just new installs)
- New SKILLPACK section with a MUST ADD setup requirement
- Schema changes that require `gbrain init` or manual SQL
- Changed defaults that affect existing behavior
- Deprecated commands or flags that need replacement
- New verification steps that should run on existing installs
- New cron jobs or background processes that should be registered
**You do NOT need a migration file when:**
- Bug fixes with no behavior changes
- Documentation-only improvements (the agent re-reads docs automatically)
- New optional features that don't affect existing setups
- Performance improvements that are transparent
**The key test:** if an existing user upgrades and does nothing else, will their
brain work worse than before? If yes, migration file. If no, skip it.
Write migration files as agent instructions, not technical notes. Tell the agent
what to do, step by step, with exact commands. See `skills/migrations/v0.5.0.md`
for the pattern.
## Migration is canonical, not advisory
GBrain's job is to deliver a canonical, working setup to every user on upgrade.
Anything that looks like a "host-repo change" — AGENTS.md, cron manifests,
launchctl units, config files outside `~/.gbrain/` — is a GBrain migration
step, not a nudge we leave for the host-repo maintainer. Migrations edit host
files (with backups) to make the canonical setup real. Exceptions: changes
that require human judgment (content edits, renames that break semantics,
host-specific handler registration where shell-exec would be an RCE surface).
Everything mechanical ships in the migration.
**Test:** if shipping a feature requires a sentence that starts with "in
your AGENTS.md, add…" or "in your cron/jobs.json, rewrite…", the migration
orchestrator should be doing that edit, not the user.
**The exception is host-specific code.** For custom Minion handlers
(host-specific integrations like inbox sweeps or third-party API scanners), shipping them as a
data file the worker would exec is an RCE surface. Those get registered in
the host's own repo via the plugin contract (`docs/guides/plugin-handlers.md`);
the migration orchestrator emits a structured TODO to
`~/.gbrain/migrations/pending-host-work.jsonl` + the host agent walks the
TODOs using `skills/migrations/v0.11.0.md` — stays host-agnostic, still
canonical.
## Privacy rule: scrub real names from public docs
**Never reference real people, companies, funds, or private agent names in any
public-facing artifact.** Public artifacts include: `CHANGELOG.md`, `README.md`,
`docs/`, `skills/`, PR titles + bodies, commit messages, and comments in checked-in
code. Query examples, benchmark stories, and migration guides MUST use generic
placeholders.
Why: gbrain runs a personal knowledge brain containing notes on real people and
real companies (YC founders, portfolio companies, funds, investors, meeting
attendees). When a doc copies a query like `gbrain graph diana-hu --depth 2` or
names a specific agent fork like `Wintermute`, that real name gets indexed by
search engines, surfaced in cross-references, and distributed with every release.
**Name mapping** to use in examples:
- Agent forks → `your agent fork`, `a downstream agent`, or `agent-fork`
- Example person → `alice-example`, `charlie-example`, or `a-founder`
- Example company → `acme-example`, `widget-co`, or `a-company`
- Example fund → `fund-a`, `fund-b`, `fund-c`
- Example deal → `acme-seed`, `widget-series-a`
- Example meeting → `meetings/2026-04-03` (generic date is fine)
- Example user → `you` or `the user`, never a proper name
**Specific rule: never say `Wintermute` in any CHANGELOG, README, doc, PR, or
commit message.** When the temptation is to illustrate with the real fork name:
- Reader-facing copy → `your OpenClaw` (covers Wintermute, Hermes, AlphaClaw,
and any other downstream OpenClaw deployment in one term the reader already
recognizes).
- First-person / origin-story copy → `Garry's OpenClaw` (honest that this is
the production deployment driving the feature, without exposing the private
agent's name).
`Wintermute` may appear in private artifacts (scratch plans under
`~/.gstack/projects/…`, memory files, conversation transcripts, CEO-review
plans) — those aren't distributed. Anything checked into this repo or shipped
in a release must use the OpenClaw phrasing above. Sweeping a stale reference
is a small clean-up PR, not a debate.
**When in doubt, ask yourself:** "Would this query reveal private information
about the user's contacts, investments, or portfolio if it were read by a
stranger?" If yes, replace with generic placeholders.
**Illustrative API examples with household-brand companies** (Stripe, Brex, OpenAI,
GitHub, etc.) are fine — they're public entities, not contacts in anyone's brain.
Do not confuse illustrative API examples with queries that reveal real
relationships.
## Responsible-disclosure rule: don't broadcast attack surface in release notes
**When a release fixes a security gap or a user-impacting bug, describe the fix
functionally. Do not enumerate the attack surface, quantify the exposure window,
or highlight the most sensitive records by name in public-facing artifacts.**
Public-facing artifacts include: `CHANGELOG.md`, `README.md`, `docs/`, PR titles
and bodies, commit messages, GitHub issue titles and comments, release pages,
tweets, blog posts.
**Don't write:**
- "10 tables were publicly readable by the anon key for months, including X, Y, Z"
- "X and Y are the most sensitive ones"
- "N tables exposed. Fix: enable RLS on these specific tables: ..."
**Do write:**
- "Security hardening pass. Fresh installs secure by default. Existing brains
brought to the same bar automatically on upgrade."
- "If `gbrain doctor` still flags anything after upgrade, the message names each
table and gives the exact fix."
Why: anyone reading the release page before they've upgraded now has a directed
probe list for unpatched installs. The source code ships the specifics anyway
(`src/schema.sql`, `src/core/migrate.ts`, test fixtures) — reverse engineers can
get them. But the release page is a broadcast channel. Don't hand attackers a
curated list with a banner.
**The test:** if a reader with no prior context could read the release note and
walk away knowing "gbrain at version X has table Y readable by anon key until
they patch," the note is too specific. Rewrite until that's no longer possible.
**What IS fine in public artifacts:**
- The mechanism of the fix ("the check now scans every public table instead of
a hardcoded allowlist").
- User-facing operator ergonomics (the escape-hatch SQL template, the upgrade
commands, the breaking-change flag).
- Credit to contributors.
- Generic framing of severity ("security posture tightening pass") without
quantification.
**What stays in private artifacts (plan files, private memories, internal docs):**
- Specific table names, record counts, exposure duration.
- Which records stand out as highest-risk.
- Detailed before/after tables in the "numbers that matter" format.
If the CEO/Eng review of a plan produces a detailed exposure table, keep it in
the plan file under `~/.claude/plans/` or `~/.gstack/projects/`. Don't copy it
into the CHANGELOG or PR body.
Applies retroactively: if you see a prior CHANGELOG entry naming attack-surface
specifics, scrub it as a small cleanup commit, the same way a stale Wintermute
reference gets swept.
## Schema state tracking
`~/.gbrain/update-state.json` tracks which recommended schema directories the user
adopted, declined, or added custom. The auto-update agent (SKILLPACK Section 17)
reads this during upgrades to suggest new schema additions without re-suggesting
things the user already declined. The setup skill writes the initial state during
Phase C/E. Never modify a user's custom directories or re-suggest declined ones.
## GitHub Actions SHA maintenance
All GitHub Actions in `.github/workflows/` are pinned to commit SHAs. Before shipping
(`/ship`) or reviewing (`/review`), check for stale pins and update them:
```bash
for action in actions/checkout oven-sh/setup-bun actions/upload-artifact actions/download-artifact softprops/action-gh-release gitleaks/gitleaks-action; do
tag=$(grep -r "$action@" .github/workflows/ | head -1 | grep -o '#.*' | tr -d '# ')
[ -n "$tag" ] && echo "$action@$tag: $(gh api repos/$action/git/ref/tags/$tag --jq .object.sha 2>/dev/null)"
done
If any SHA differs from what's in the workflow files, update the pin and version comment.
PR descriptions cover the whole branch
Pull request titles and bodies must describe everything in the PR diff against the
base branch, not just the most recent commit you made. When you open or update a
PR, walk the full commit range with git log --oneline <base>..<head> and write the
body to cover all of it. Group by feature area (schema, code, tests, docs) — not
chronologically by commit.
This matters because reviewers read the PR body to understand what's shipping. If the body only covers your last commit, they miss everything else and can't review properly. A 7-commit PR with a body that describes commit 7 is worse than no body at all — it actively misleads.
When in doubt, run gh pr view <N> --json commits --jq '[.commits[].messageHeadline]'
to see what's actually in the PR before writing the body.
Community PR wave process
Never merge external PRs directly into master. Instead, use the "fix wave" workflow:
- Categorize — group PRs by theme (bug fixes, features, infra, docs)
- Deduplicate — if two PRs fix the same thing, pick the one that changes fewer lines. Close the other with a note pointing to the winner.
- Collector branch — create a feature branch (e.g.
garrytan/fix-wave-N), cherry-pick or manually re-implement the best fixes from each PR. Do NOT merge PR branches directly — read the diff, understand the fix, and write it yourself if needed. - Test the wave — verify with
bun test && bun run test:e2e(full E2E lifecycle). Every fix in the wave must have test coverage. - Close with context — every closed PR gets a comment explaining why and what (if anything) supersedes it. Contributors did real work; respect that with clear communication and thank them.
- Ship as one PR — single PR to master with all attributions preserved via
Co-Authored-By:trailers. Include a summary of what merged and what closed.
Community PR guardrails:
- Always AskUserQuestion before accepting commits that touch voice, tone, or promotional material (README intro, CHANGELOG voice, skill templates).
- Never auto-merge PRs that remove YC references or "neutralize" the founder perspective.
- Preserve contributor attribution in commit messages.
Checking out PRs from garrytan-agents
garrytan-agents is the AI-authored PR account and is NOT a collaborator on
this repo. Its PRs live in a fork, so GitHub Actions triggered by
pull_request events on those PRs do not receive base-repo secrets. Any CI
job that needs ANTHROPIC_API_KEY, OPENAI_API_KEY, or similar will fail
with empty-env auth errors, regardless of what's set on the base repo. This
is a GitHub security default, not a config bug.
When the user says "check out " and the PR is from garrytan-agents
(or any other non-collaborator fork), move the branch into the base repo
before running CI:
gh pr checkout <N>— pull down the fork's branch. Note the PR number and head branch name (gh pr view <N> --json headRefName --jq .headRefName).git push origin HEAD:<branch-name>— push the same branch to the base repo (origin points atgarrytan/gbrain, not the fork). This is the move that gives CI access to secrets.gh pr close <N> --comment "moving to base-repo branch for secret access"— close the fork PR so the queue stays clean.gh pr create --base master --head <branch-name>— open the replacement PR from the base-repo branch. Preserve the original PR's title and body verbatim (gh pr view <N> --json title,body); contributor attribution moves to aCo-Authored-By:trailer if needed.
Why this over alternatives: adding garrytan-agents as a collaborator, or
flipping the repo-wide "send secrets to fork PRs" toggle, both broaden
secret distribution to every fork PR from that account or any fork. Moving
the branch keeps secret scope tight to just the one PR being shipped.
Skill routing
When the user's request matches an available skill, ALWAYS invoke it using the Skill tool as your FIRST action. Do NOT answer directly, do NOT use other tools first. The skill has specialized workflows that produce better results than ad-hoc answers.
NEVER hand-roll ship operations. Do not manually run git commit + push + gh pr
create when /ship is available. /ship handles VERSION bump, CHANGELOG, document-release,
pre-landing review, test coverage audit, and adversarial review. Manually creating a PR
skips all of these. If the user says "commit and ship", "push and ship", "bisect and
ship", or any combination that ends with shipping — invoke /ship and let it handle
everything including the commits. If the branch name contains a version (e.g.
v0.5-live-sync), /ship should use that version for the bump.
Key routing rules:
- Product ideas, "is this worth building", brainstorming → invoke office-hours
- Bugs, errors, "why is this broken", 500 errors → invoke investigate
- Ship, deploy, push, create PR, "commit and ship", "push and ship" → invoke ship
- QA, test the site, find bugs → invoke qa
- Code review, check my diff → invoke review
- Update docs after shipping → invoke document-release
- Weekly retro → invoke retro
- Design system, brand → invoke design-consultation
- Visual audit, design polish → invoke design-review
- Architecture review → invoke plan-eng-review
- Save progress, checkpoint, resume → invoke checkpoint
- Code quality, health check → invoke health