Files
Garry Tan 2efaaf8f8a v0.48.4.0 feat(search,eval): ranker wave — autocut off, relational rerank pin, metadata boost gate; LongMemEval harness with strict recall_all@5 + judged answer accuracy (#4946)
* fix(cli): flag registry attributes compound dispatch blocks to their command

The generator's marker regex only matched plain `if (command === 'X')` and
`case 'X':` markers, so every compound pre-dispatch bypass in handleCliOnly
(`if (command === 'eval' && args[0] === 'longmemeval')`, the `--help`
pre-engine branches, the agent-register guard) was attributed to the
preceding marker. The documented `gbrain eval longmemeval <f>
--retrieval-only --by-type --no-trajectory` exited 1 at the unknown-flag
validator because the eval row carried none of its subcommand flags.

- segmentDispatchBlocks(): labels the `command === 'X'` head of any `if (`
  (plain, compound, multi-line) plus `case 'X':`; block-level module scan
  restricted to ./commands/*.ts so pre-connect helper imports no longer
  add phantom flags (the agent row lost --delete-brain and friends).
- Registry regenerated (idempotent); eval gains its subcommand flags,
  dream/status/backup lose misattributed ones.
- Tests: acceptance + rejection (the eval row stays a union across eval
  subcommands by design) + a subprocess smoke of the documented command.
- The contradictions remediation hint no longer suggests `dream --slug`
  (dream never parsed it; the registry-gated remediation test caught it).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(search): role-tagged fusion arms + expansion_variant_budget knob (legacy default, byte-identical)

One composition point for RRF inputs: hybrid.ts builds `vectorArms`
(original | variant | clause | image) via pushVectorList at every
assembly site and hands them to composeFusionLists (new
src/core/search/fusion-lists.ts), which returns the full weighted list
set. rrfFusionWeighted scores weight/(k+rank); weight defaults to 1, so
fusion is byte-identical to before.

`search.expansion_variant_budget` (ModeBundle, config key, per-call seam)
is the pre-registered fix for LLM multi-query expansion diluting small-k
recall: a total RRF weight shared equally by the non-empty variant/clause
lists (weight_i = b / n_voting_arms). `null` = legacy (every list weight
1) and is the default in all three bundles; the flip waits for the
LongMemEval receipt. One range contract (normalizeExpansionVariantBudget)
serves the config parser and both hybrid.ts seams. KNOBS_HASH_VERSION
28 -> 29 (`evb=` part; one-time cache miss). queries[0] is enforced to be
the caller's query after expandFn. The both-mode fell-open corner no
longer mis-tags the last text list as the image branch.

onRerankPool fires with the exact pre-autocut returnPool (post alias-hop /
exact-lookup / adaptive-return) so an offline autocut replay is faithful.
`gbrain search modes` renders a legitimate null as `legacy (null)`.

Tests: pure fusion arithmetic (knife edge at budget 1.0), hermetic
directional PGLite test, CRITICAL regression on the both-mode demotion
gate, six knobs-hash pins + two titles, bundle snapshots, modes report.
TSV ceilings raised to exact counts (hybrid.ts 3177, mode.ts 1577,
config.ts 1766).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(eval): LongMemEval harness scores the official strict metric and is the receipt producer

`gbrain eval longmemeval` now measures what the paper measures and records
enough to reproduce it:

- Raw-id join through a per-question slug -> raw session-id map (gold ids
  are compared as the dataset spells them; slug collisions touching gold
  abort the question with an error row).
- recall_all@k (every gold session among the distinct sessions of the
  top-k chunk rows) as the headline, recall_any@k as the diagnostic;
  abstention (`_abs`) questions excluded from recall denominators unless
  --include-abstention; --by-type-floor gates on the strict rate
  (--by-type-floor-metric recall_any restores the old semantics).
- Pins: --reranker on|off, --autocut on|off, --expansion-variant-budget,
  --mode; --reranker on preflights readiness (exit 2) and any row that fell
  through un-reranked, or whose vector arm silently degraded, fails the run.
- Schema-v2 by_type_summary with run_config (pins, embedder, dataset sha,
  knobs hash, cache stats, degradation counters); per-row retrieved[] with
  scores, search_meta, retrieval_config_hash; --capture-pool records the
  pre-autocut pool for offline floor replay; --expansion records variants
  and --expansion-replay serves them; --question-ids; --record appends a
  redacted EvalRunRecord (redaction now lives in persistRunRecord).
- Content-addressed embedding cache (src/eval/shared/embed-cache.ts,
  bun:sqlite, WAL, dims verification, canonical hash) through the gateway
  transport seam; resume recomputes both metrics and refuses a mixed
  retrieval_config_hash unless --allow-mixed-run-config.
- LME_FLAGS table drives both parseArgs and printHelp (the flag registry
  scans help text). Help fixed: --mode tokenmax does not imply --expansion.
- Metric glossary gains recall_all@k, recall_any@k, qa_accuracy,
  mean_returned_est_tokens, mean_returned_results (new LongMemEval group).
- Receipt tooling: scripts/eval-spend-guard.sh (fail-closed ledger cap),
  scripts/replay-autocut-floor.ts + src/eval/shared/autocut-replay.ts,
  seeded dev/decision/half splits in evals/longmemeval/.
- Fixture test/fixtures/longmemeval-mixedcase.jsonl (placeholder text,
  dataset-shaped ids) pins the join, strict/any split, and abstention.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs: ranker wave reference docs — fusion arms + expansion budget knob, LongMemEval harness as the reproduction path

- KEY_FILES: current-state entries for fusion-lists.ts, the harness modules
  (metrics/run-config/resume), shared embed cache + autocut replay, the
  spend guard, the flag-registry generator's compound-marker segmentation,
  and the seeded splits; mode.ts and eval-longmemeval.ts entries rewritten.
- RETRIEVAL.md: the expansion sentence now states the receipted truth
  (strict recall_all@5 93.19% plain hybrid vs 54.89% equal-weight expansion)
  and the budget-normalized weighted RRF; pool capture for offline autocut
  replay; the verify block shows the like-for-like and default-path runs.
- search-modes.md: expansion_variant_budget knob row + say-line.
- eval-bench.md: in-repo command is the reproduction path (self-check
  caveats removed), full flags table, schema-v2 summary example, recall_hit
  deprecated alias, --by-type-floor gates on recall_all, sessdiv rows marked
  as from the 2026-09-02 run. Measured numbers unchanged.
- README: the two self-check sentences updated; no new claims.
- CLAUDE.md: three knobs-hash history paragraphs collapsed into one
  current-state sentence (58,557 -> 58,086 bytes).
- TODOS: search.dedup_max_per_page re-pointed to knobs-hash v30.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(eval): LongMemEval judged answer-accuracy lane (--judge) with official prompts and complete-judgment headlines

`gbrain eval longmemeval --judge` grades generated answers with a verbatim
port of the official evaluate_qa.py per-type yes/no prompts (standard,
temporal-reasoning off-by-one tolerance, knowledge-update, preference
rubric, abstention), judge gpt-4o at max_tokens 10 / temperature 0, and
publishes a qa_accuracy block whose headline scores every question the
judge could not grade (timeout, refusal, malformed, budget skip) as
INCORRECT, with accuracy_excluding_errors and the judge_errors count
alongside. Gold and hypothesis travel inside the #4338 data boundary.

- src/eval/shared/judge-runner.ts: prompt-agnostic runner (retry+backoff
  on timeout/429, closed judge_error vocabulary, cost via canonicalLookup,
  BudgetLedger soft-stop); src/eval/shared/bootstrap.ts: seeded question
  bootstrap CI (labelled question-sampling only).
- src/eval/longmemeval/{judge,judge-lane,qa-accuracy,reader,emit}.ts:
  prompts + verdict parse + judge_config_hash, preflight/backfill, the
  qa_accuracy block, the peeled reader (system text now carries the
  abstention instruction; reader_prompt_sha and the provider-reported model
  snapshot are recorded per row), the emitter.
- Flags: --judge, --judge-model, --max-usd (judge spend only), --yes,
  --judge-concurrency, --allow-incomplete-judgments; --judge --resume-from
  is a judge-only backfill gated by judge_config_hash; a run with judge
  errors or budget skips is "not publishable" (exit 1).
- gateway: ChatOpts.temperature threads to generateText; ChatResult
  exposes the response model id.
- Glossary: qa_accuracy describes the headline rule.
- Tests: 32 pure + 9 e2e (stubbed reader/judge) + gateway temperature.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(eval): NamedThingBench reranker on/off paired A/B (rule R1 runner)

scripts/r1-namedthing-rerank-ab.ts seeds the NamedThingBench corpus
(peeled from the gate test into test/fixtures/retrieval-quality/namedthing/
corpus.ts, byte-identical) into one in-memory brain with real embeddings,
runs the 12 questions (optionally the 42 relational ones) under pinned OFF
and ON reranker arms, and reports per-query hit@1/hit@3/create_safety
with the R1 verdict (PASS iff 0 hit@1 losses and <= 1 hit@3 loss). Exits
2 before spending if the reranker is not ready or any ON row fell through
un-reranked. --embed-cache shares query vectors across arms; --stub-embed
is the hermetic dry run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(eval): LongMemEval miss diagnostics (Phase B1) — per-arm gold ranks, miss classes, clause counterfactuals

scripts/lme-miss-diagnostics.ts re-creates each strict miss exactly as the
harness does (same pages, embed cache, pins) and reports, per missing gold
session, vector / keyword / title ranks to depth 200, fused and post-rerank
ranks from one hybridSearch under the same pins, and the pre-registered
class (absent from every arm; in the vector pool but fused out; reranked
out; k ceiling), plus honesty classes when a miss does not reproduce.
Hypothesis probes: the second-event starvation signature, the frozen
clause splitter (between/and, first…or, before/after, how many … between;
guardrails: two content tokens per clause, no split inside quotes or a
Capitalized-Bigram, max two) with counterfactual clause embeds served
outside the shared cache, and pool-depth vs reranker-depth attribution.
Rows carry dev/decision/half-split membership so the fix is chosen on
half A and confirmed on half B. Pure core in
src/eval/longmemeval/diagnostics.ts; 26 hermetic tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(r1): move the seed-contract engine into beforeAll (test-isolation rule R3)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(search): relational-arm rows bypass reranker demotion (search.relational_rerank_pin, default 3)

The NamedThingBench reranker A/B (rule R1) showed the shipped `balanced`
default collapsing relational queries: "who invested in X" / "who works at
X" / "what connects A and B" lost hit@1 on 19 of 39 questions with the
cross-encoder on (hit@1 21 -> 3, hit@3 27 -> 5) while the 11 non-relational
core questions were unaffected. The reranker scores chunk text and cannot
see typed-edge answers, so graph-derived investor and employee pages sink
below pages that merely mention the words.

Fix: after applyReranker, relational-arm rows are re-pinned in their fused
(RRF) order ahead of the reranked text rows, bounded by
`search.relational_rerank_pin` (3 in every bundle; 0/off reproduces the old
behavior; per-call `relationalRerankPin`), mirroring the alias and
exact-lookup tiers that already run after rerank. Pinned rows are stamped
`relational_pinned`, preserved through autocut and excluded from its cliff
computation (without that, autocut cut them straight back out). Pure no-op
for non-relational queries, image modality, and reranker-off paths;
`rrp=` folds into the knobs hash under the v29 epoch.

Pure module src/core/search/relational-rerank-pin.ts + a hermetic PGLite
test on the relational corpus through the real gateway rerank path with an
inverted-relevance stub (pin 0 reproduces the loss, pin 3 restores hit@1 on
all eight "who invested in" questions; non-relational and reranker-off
queries byte-identical). TSV ceilings raised to exact counts. Known limit
filed in TODOS (pin only hop-1 rows whose link types match the parsed
relation; gate on seed resolution margin).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(r1): --autocut on|off overlay so the ON arm can run in the exact shipped configuration

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(eval): honest eval scaffolds, canary ledger record, ranker-wave TODO filings

- Delete the two undispatched scaffolds that returned ok:true for work never
  run (eval-markdown-greenfield, eval-extract-atoms); eval-schema-authoring
  keeps its real aggregator/parseArgs and now returns the #4198
  not_implemented envelope (exit 1) instead of an inconclusive verdict.
- .gbrain-evals/eval-results.jsonl: retrieval canary recorded at this head
  (recall@10 1.0, first_relevant 1.0, expected_top1 0.857 vs floor 0.85,
  14/14 queries).
- TODOS: E4 completed; R1 decision recorded (balanced reranker stays ON with
  the relational pin); expansion + autocut entries marked in progress; six
  follow-ups filed (LoCoMo/BEAM lanes, run-all wiring, inline judge
  concurrency, session-diverse arm, named embed-transport hook,
  per-subcommand registry rows).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs: judge lane, miss diagnostics, R1 runner and relational pin in the reference docs; R1 script text + --relational-pin overlay

- KEY_FILES: entries for the judge lane modules (judge, judge-lane,
  qa-accuracy, reader, emit, shared judge-runner + bootstrap), the miss
  diagnostics tool, the R1 runner and its corpus module; eval-longmemeval
  entry extended with the reader and judge-lane contracts; gateway entry
  notes ChatOpts.temperature and ChatResult.responseModel.
- eval-bench.md: "Judged answer accuracy (--judge)" subsection with the
  protocol and every disclosure (official prompts, gpt-4o at temperature 0,
  data boundary, judge_error class, headline rule, question-sampling CIs,
  no SOTA claim), flags, say-line, "Diagnosing misses".
- RETRIEVAL.md: relational re-pin section verified against the code and
  carries the paired R1 counts (19/39 and 22/39 losses without the pin;
  0 losses with pin 3, autocut on).
- FIX_WAVE_BASELINES.md: ranker-wave block with the receipts that exist and
  refresh commands; pending rows named as pending.
- R1 script: usage/header list --autocut and the new --relational-pin N|off
  overlay (so the no-pin cell reproduces from HEAD); the markdown header
  reports the arms' actual autocut and pin state.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(search): arm-confidence-weighted fusion knob (search.keyword_arm_confidence_floor, default off)

Phase E2 mechanism for the Cat 13 conceptual-recall gap (Voyage space, held-out
concepts: bare vector 60.5 nDCG@5 vs gbrain hybrid 53.0; grep-only 52.2 — the
keyword arm's noise on paraphrase probes drags fusion below the vector arm).
When the keyword arm's top-row margin ratio top/(top+second) is below the
floor, the keyword AND title lists fuse at weight 0.5; vector lists untouched;
never on relational queries, keyword-only fallbacks, or an empty keyword arm.
Statistic and decision live in the pure src/core/search/arm-confidence.ts and
compose inside composeFusionLists; hybrid.ts threads the knob and stamps
HybridSearchMeta.keyword_arm_confidence {margin_ratio, top_score, downweighted}
so the Cat 13 runner can calibrate the floor on the tuning concepts.

Default null (off) in all three bundles: the held-out Cat 13 receipt decides
the flip. Config key search.keyword_arm_confidence_floor ((0,1] | off);
kacf= folds into the v29 knobs-hash epoch. Hermetic tests: statistic,
decision gates, an RRF flip on a decoy corpus, byte-identity when off or
strong, keyword-only fallback untouched. TSV ceilings raised to exact counts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(search): metadata boost gate knob (search.metadata_boost_gate: always | lexical; default always)

Cat 13 localization (tuning split, re-simulation validated 359/359 against the
live order): gbrain's own vector arm ranks the gold concept page well (nDCG@5
60.3) but the live hybrid scores 50.6 because post-fusion metadata boosts —
backlink 1.035-1.124x, graph adjacency, recency — promote hub pages above gold
concept pages that carry none; in 73 of 105 gap probes the vector arm was the
only voter. `lexical` skips the backlink, salience, recency (+chronicle),
graph-signal and alias-resolved stages when no strict keyword, title or
relational row reached fusion; supersede downrank, exact-match and
title-phrase boosts stay. `always` (default in all bundles) is byte-identical
to today; the held-out E3 receipt decides the flip. `mbg=` folds into the v29
knobs-hash epoch; per-call `metadataBoostGate`; meta stamps gate /
lexical_voted / boosts_applied so vector-only-voter queries are countable
before any flip. Pure module src/core/search/metadata-boost-gate.ts; hermetic
PGLite test reproduces the hub-over-concept inversion and its fix.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(eval): --search-pin KEY=VALUE on the LongMemEval harness and the R1 runner (generic knob A/B pins, folded into the knobs hash)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(search): flip search.metadata_boost_gate to lexical in every bundle (Cat 13 E3 receipt)

Pre-registered Phase E3 rule (written before the held-out run): gbrain
nDCG@5 on the 10 held-out conceptual-recall concepts >= 57.0 (E0 53.0 + 4)
with NamedThingBench, BrainBench, the retrieval canary and the LongMemEval
dev slice unchanged. Result: 57.8 (off/off) and 57.9 on the shipped default
(was 55.8); tuning 57.3 matched the E1 projection; NamedThingBench 50/50,
BrainBench PASS same-hash, canary PASS, LME dev slice 40/40 all identical.
Stretch (bare vector 60.5) not met and filed.

Under `lexical` the post-fusion metadata boosts (backlink, salience,
recency + chronicle, graph signals, alias resolution) are skipped when the
vector arm was the only voter; `always` restores the pre-wave pipeline.
DEFAULT_METADATA_BOOST_GATE stays `always` so a knobs literal without the
field keeps its pre-wave hash identity. Tests force `always` explicitly for
the pre-wave rows; default rows expect `lexical`. Docs: KEY_FILES entries
for metadata-boost-gate.ts and arm-confidence.ts, search-modes knob rows,
RETRIEVAL.md gate stage, FIX_WAVE_BASELINES Cat 13 outcome, CLAUDE.md
search-mode table row; llms bundle regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(eval): --search-pin flag row, Phase B located-class receipt, R1 closeout, methodology pre-registrations

- eval-bench.md Flags table documents `--search-pin KEY=VALUE` (folds into
  retrieval_config_hash + knobs hash).
- TODOS.md: the temporal-reasoning entry records the Phase B diagnosis
  (misses are the embedding ranking of near-duplicate sessions; fused rank =
  vector rank; no knob landed) and the R1 note records the Cat 13 reranker
  rows + the cat13b re-run left for a follow-up.
- SEARCH_MODE_METHODOLOGY.md: dev-slice vs decision-set discipline (§3),
  three new threats (§5), the ranker wave's ten pre-registrations with the
  outcomes decided so far (§7), LongMemEval-S dataset sha in the footer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(changelog): draft the 0.48.3.0 ranker-wave entry (Phase A/C/D measured rows pending receipts)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(key-files): relational pin entry states the core-query outcome precisely (0 losses, one hard-negative gain)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(search): autocut off in balanced and tokenmax (rule R2 receipt); replay joins gold from the dataset

Pre-registered rule R2 (autocut floor) on the LongMemEval decision set: keep
0.35 iff the shipped default (reranker on, autocut on) scores >= the
reranker-only arm - 2 questions with no type losing > 1. Result: 379/470 vs
449/470, paired +0/-68 on the 430 (multi-session -22, temporal -27,
knowledge-update -19). The replay from the captured post-rerank pool
(validate-live reproduced all 500 live decisions) found no floor in
{0.10 .. 0.80} within the guardrail on either seeded half (0.80 still -9,
all knowledge-update), so autocut is off in balanced and tokenmax per the
pre-registered chain. DEFAULT_AUTOCUT is unchanged for operators who
re-enable it; any-hit stayed >= 99.4% at every floor and the mean returned
window went 3256 -> 1633 estimated tokens at 0.35, which is the trade
documented in the CHANGELOG.

Replay fix that the receipt depended on: harness capture rows carry gold
COUNTS only, so `scripts/replay-autocut-floor.ts --dataset <longmemeval
json>` joins `answer_session_ids` by question_id and a capture with no gold
anywhere is refused instead of scoring 0% at every floor; `buildRow` now
also emits `answer_session_ids` per row. CLI tests cover the refusal, the
join, and a capture question missing from the dataset.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(todos): R2 decided (autocut off), file session-aware autocut, Cat 13 residual and cat13b reranker follow-ups

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(eval): judge-resume rewrites the output atomically; serial tests opt into autocut explicitly

Pre-landing review findings (adversarially verified):

- `--judge --resume-from FILE --output FILE` re-emits every prior row after
  the judge backfill, and the emitter opened the file with 'w' first — a
  kill mid-backfill (Ctrl-C, timeout, OOM) left a 0-byte file and every paid
  reader row was gone. `makeEmitter(..., { atomicRewrite: true })` now writes
  to `<file>.rewrite.tmp` and renames over the original on close, so the
  resume file is intact until the rewrite is complete. Both rewrite call
  sites use it; `test/longmemeval-emit.test.ts` pins the mid-run state, the
  rename, idempotent close, and the summary-after-rename order.
- Two serial-lane tests inherited the pre-R2 `balanced.autocut = true`
  bundle default and failed after the flip; they now turn autocut on per
  call (the relational-pin file tests the pin THROUGH autocut by design; the
  reranker-integration test passed a non-boolean autocut object that
  hybridSearch ignores).
- Captured `rerank_pool` rows carry `relational_pinned` so the offline
  autocut replay can mirror hybrid.ts's preserve/score predicates.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(eval,search): pre-landing review wave — receipt pins, judge gate, replay parity, spend-guard reservations, R1 pin guard, doc truth

Adversarially verified findings from the branch review (67-agent pass: six
read-only reviewers, two refutation lenses per finding, completeness critic),
fixed in disjoint groups and re-verified (typecheck, verify 54/54, touched
suites green):

- Miss diagnostics read the receipt's pins from a nested `run_config.pins`
  block the harness never writes; they now read the flat `run_config`
  (topK → top_k, reranker.{enabled,model}, autocut, expansion, budget,
  embedder; legacy nested block still accepted), with a round-trip test
  through the real `buildRunConfig`. Before this a reranker-off receipt was
  re-diagnosed under the balanced bundle with only a WARN.
- Judged lane: the publishability gate now uses `qa_accuracy.complete`
  (judge_errors, skipped_budget AND unjudged all 0); a judge call that throws
  inside the backfill stamps `judge_error: 'provider_error'` instead of
  leaving the row silently unjudged; `judge_config_hash` / `mixed_judge_config`
  derive from the hashes actually present on rows (a homogeneous backfill
  under a different resolved reader model is no longer "mixed");
  degradation gates fold prior rows on every resume, not only with --by-type;
  `gold_missing_from_haystack` / `slug_collisions` count the same row set on
  a fresh run and a resume.
- Autocut replay mirrors hybrid.ts's predicates for `relational_pinned` rows
  (preserved through the cut, excluded from cliff math and the top-score
  histogram); stale "rows carry only gold counts" comments corrected.
- Spend guard fails closed in time: a `running` reservation row lands before
  the command starts and a `done` reconciliation after (shared run_id; sum =
  reconciled + unreconciled reservations); INT/TERM/HUP reconcile at the
  estimate; an unwritable ledger refuses the launch.
- R1 runner refuses `--search-pin search.reranker.*` overlays (the arm axis)
  at parse time; parseArgs tests cover --autocut / --relational-pin /
  --search-pin acceptance and rejection.
- Prose truth: README + KEY_FILES + FIX_WAVE_BASELINES + RETRIEVAL + CHANGELOG
  no longer call autocut-on "the shipped shape"; the relational pin IS an R1
  script flag; A1 parity row filled; mode.ts / metadata-boost-gate.ts /
  test comments say `lexical`; ModeBundle.reranker_enabled doc matches the
  bundles; eval-bench legacy_rows and Flags-table claims corrected; TODOS R2
  entry no longer both decided and in progress; eval-schema-authoring help
  says the command is not dispatched yet.
- Phase A outcome recorded: CHANGELOG "If you run tokenmax" carries the
  measured numbers (255 → 394 of 470 with the budget knob; plain hybrid 439),
  bundles keep the legacy weighting, the TODOS entry is DECIDED and the
  CRAG-style conditional-expansion follow-up is filed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(eval): Phase A outcome in search-modes, RETRIEVAL, FIX_WAVE_BASELINES, methodology and README (budget knob real but below rule; bundles keep legacy weighting)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(eval): publish the ranker-wave retrieval receipts — release default 449/470 (95.53%), eight arms, per-type table, tokenmax as released

CHANGELOG Measured section, docs/eval-bench.md "Current measured result",
README eval paragraphs, the company-brain tutorial line, FIX_WAVE_BASELINES
(final release-configuration arm + gate D11 PASS on every leg, tokenmax as
released) and the CLAUDE.md search-mode table (autocut row) now carry the
2026-09-06 in-repo harness numbers: A1 439/470, A2 449/470, A3 255/470,
A4 379/470, A3′ 394/470, A3′R 381/470, tokenmax-as-released 436/470, release
default 449/470 (byte-identical per question to A2). The 2026-09-02 sibling
rows stay referenced as the prior receipt. llms bundle regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(eval): the judged lane reads whole sessions and the judge respects the provider's token floor

Two defects the 25-question judged dry run exposed before the full run spent:

- Every judge call failed with provider_error: the official evaluate_qa.py
  max_tokens of 10 is below the OpenAI API's minimum of 16. JUDGE_MAX_TOKENS
  is 16 (a one-token yes/no verdict is unaffected) and the methodology note
  discloses the deviation.
- The reader abstained on 11/25 questions whose gold session was retrieved
  at rank 1: the sanitizer's 4000-char per-session cap (an extractor-era
  default) cut the answer out of most retrieved sessions (LongMemEval gold
  sessions run 5–23K chars). renderChatBlock/sanitizeChatContent take the
  bound as a parameter; the reader passes READER_MAX_SESSION_CHARS (60K, a
  safety bound above the longest session) so it reads whole sessions, and
  every row records reader_context_chars / reader_context_sessions /
  reader_sessions_truncated. READER_PROMPT_VERSION bumps to
  v3-abstention-fullsessions (the system text and its sha are unchanged).

Tests: reader context construction (full session reaches the prompt once per
distinct session; only a >60K session is cut; chunk fallback), the
sanitizer's per-call bound + truncatedCount, judge token floor. Docs:
eval-bench judged-lane protocol paragraph, CHANGELOG judged-lane bullet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(eval-bench): move the judged-lane protocol paragraph out of the command block into its section

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(todos): file the R1 runner explicit-embedder / fixture-header follow-up

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(eval): same-file judge resumes append as rows land and compact at run end; integer gold answers grade as strings; capture mapper peeled

The first judged run exposed two more defects:

- A `--judge --resume-from FILE --output FILE` pass wrote to a temp file and
  renamed on close (the previous fix for the truncation window), so a pass
  killed by its own `timeout` lost EVERY row it had produced — the bounded
  resume loop made no progress across three 900-second passes. A same-file
  resume now always APPENDS: judged backfill rows and retries land the moment
  they finish as newer duplicates, and `compactJsonlByQuestionId` (emit.ts)
  rewrites the file atomically to one row per question_id (last wins,
  first-seen order, stale summary lines dropped, corrupt tail dropped) before
  the summary is emitted. `readJsonlRows` and `loadResumeSet` read appended
  files last-wins (a retry supersedes an error row; a later error row
  re-opens the question). Rewriting into a DIFFERENT file keeps the atomic
  temp-and-rename path.
- 32 of the 500 LongMemEval-S gold answers are integers; the judge's
  data-boundary escaper called `.replace` on them and the question became an
  error row (six multi-session rows in the first pass). Golds are coerced to
  their decimal string at every judge boundary, as the official evaluator's
  f-string does.
- `--capture-pool` row assembly moved to `src/eval/longmemeval/capture.ts`
  (`buildCaptureExtras`, `poolKey`) to keep the harness under its size cap;
  the pool-row shape and key template are unchanged.

Tests: compaction (last wins, order, summary + corrupt tail dropped, atomic,
idempotent, missing file), last-wins loaders, integer gold escaping; the
collision parity test expects the compacted file. KEY_FILES documents the
append-then-compact contract and the new module.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(changelog): harness bullet notes append-then-compact same-file resumes

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(todos): file the harness actual-cost ledger follow-up

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(eval): publish the first judged answer-accuracy number (433/500, 86.6%) with full protocol disclosure; merge the wave's ledger records

CHANGELOG Measured, README "How it measures up", eval-bench "Judged answer
accuracy" (result table, evidence-vs-verdict cross-tab, protocol pins),
FIX_WAVE_BASELINES and the methodology pre-registration outcome carry the
Phase D result: 500/500 judged, 0 judge errors, 86.6% headline (CI 83.6–89.6),
abstention 29/30, per-type breakdown; retrieval on the same rows 449/470. The
pre-registered ≥ 92% prediction was missed and is published as such; no
comparison claim is made against vendor answer-accuracy rows (protocols
differ). `.gbrain-evals/eval-results.jsonl` gains the 15 EvalRunRecords the
wave's `--record` passes wrote from the pinned worktree (A1–A4, dev sweep,
A3′, A3′R, tokenmax-as-released, FINAL incl. its resume-noop passes, D1).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: autocut integration + modes-report tests opt into autocut explicitly (bundle default is off since rule R2)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(eval,search): pre-landing review round two — pin precedence and hashing, resolved-pin reranker gate, atomic summary, image-modality gate exemption, guard hardening

Findings from the /ship review (six specialists, Claude adversarial pass,
coverage and plan audits; Codex could not run in this sandbox), applied in
disjoint groups and re-verified (typecheck, verify 54/54, ~460 targeted tests):

Harness (`gbrain eval longmemeval`):
- `--search-pin` values are written BEFORE the explicit `--mode/--reranker/
  --autocut/--expansion-variant-budget` pins so explicit flags win, matching
  `resolvePins`; the sorted raw pin map folds into `retrieval_config_hash`
  (only when non-empty, so every existing receipt hash is unchanged), so a
  resume can no longer merge differently-pinned runs.
- The reranker readiness preflight and the un-reranked-rows gate key on the
  RESOLVED pin (flag, pin, snapshot or bundle): a reranked mode without its
  provider key refuses to start (exit 2) instead of quietly scoring
  un-reranked rows.
- A run in which every question errored exits 1 and records `failed`.
- `by_type_summary` rewrites are atomic (temp + rename); the dead
  `atomicRewrite` emitter mode is gone (same-file resumes append then compact).
- The inline judge runs in its own try: a thrown judge call stamps
  `judge_error: provider_error` and keeps the paid reader row.
- Pre-v2 resume rows with normalized session ids are mapped back to raw ids
  before re-scoring (they scored as misses before).
- `complete` also requires zero reader errors; `escapeJudgeData` neutralises
  `< /judge_input>`; diagnostics error rows are secret-redacted; one
  `rawSessionId`, one judge-error field builder, `hasJudgeAttempt` used;
  help text interpolates `JUDGE_MAX_TOKENS`; `SEARCH_MODES`, `isAbstentionQuestion`
  and `normalizeExpansionVariantBudget` replace local copies; a second
  positional is a usage error; the reader snapshot compares against the
  normalized model id. Gateway client and trajectory routing peeled into
  `src/eval/longmemeval/{gateway-client,trajectory-route}.ts`.

Search core:
- `search.metadata_boost_gate=lexical` no longer skips metadata boosts for
  image-modality queries (their lexical arms never vote by construction);
  the decision records `image_modality`.
- `parseRelationalQuery` memoizes the default pattern set (it compiled ~10
  RegExp objects per call and ran twice per search).

Scripts and shared eval modules:
- Embedding cache: raw NUL bytes in the canonical-hash literals replaced by
  `\0` escapes (identical hash), canonical and file hashes stream instead of
  materialising the whole cache (776 MB in the wave's cache).
- Autocut replay: refuses when ANY captured row lacks gold without
  `--dataset`, invalid floors are usage errors, non-array gold counts as
  missing; `mulberry32` imported from bootstrap.ts.
- Spend guard: traps armed before the reservation row (the SIGTERM flake's
  root cause), fail-closed when the ledger audit is not numeric, portable
  control-character escaping, TERM then KILL after a grace period.
- Flag registry: `//` and `/* */` comments stripped from dispatch blocks
  before flag extraction (a comment had given `storage` a phantom flag);
  registry regenerated.
- R1 runner: explicit `--autocut` / `--relational-pin` win over a colliding
  `--search-pin`.

Tests added: parse-args table, run-config loaders, capture extras, splits
fixture integrity, relational-intent memoization, search-pin end-to-end,
resolved-pin gate, all-errored exit, atomic summary, inline-judge throw,
legacy-id resume, duplicate question ids, qa rebuild without --judge,
escapeJudgeData variants, reader snapshot, spend-guard audit refusal and
tab escaping, replay refusals, registry comment stripping, R1 overlay
precedence. Structural test manifest regenerated. Docs: eval-bench,
KEY_FILES, CHANGELOG, TODOS (redaction entry closed).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: chunker-version-insert-default pins its 1536-d embedding shape (shard-order dependence exposed by the wave's new test files)

The file hard-codes 1536-d vectors but let initSchema size the vector columns
from whatever the gateway held; in the reshuffled shard 1 it followed a file
that left a 1280-d configuration and every upsert failed with 'expected 1280
dimensions, not 1536'. It passes alone and on master alone. Pin the shape in
beforeAll (the pattern consolidate-valid-until.test.ts uses) and reset after.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore: bump version and changelog (v0.48.4.0)

VERSION, package.json, the three plugin manifests, the BOOTSTRAP runbook
stamp, the regenerated bootstrap template repo and plugin trees move to
0.48.4.0 in lockstep (master shipped 0.48.3.0 while this wave ran). The
CHANGELOG entry for 0.48.4.0 landed with the wave's commits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(todos): file the unit-parallel OOM-rescue detector gap seen in the v0.48.4.0 ship verification

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs: update project documentation for v0.48.4.0

Post-ship documentation sweep for the ranker wave.

- CLAUDE.md: search-mode table gains the relational_rerank_pin row (3 in
  every bundle) beside metadata_boost_gate and autocut; the relational
  retrieval paragraph names the post-reranker pin and its off switch; the
  "Effective cache availability" note moves back above the cache-key notes
  it introduces (a master merge had left it below them).
- docs/architecture/KEY_FILES.md: the judge entry states JUDGE_MAX_TOKENS
  16 (the OpenAI minimum; the official 10 is rejected), matching judge.ts.
- docs/eval-bench.md: the like-for-like command comment names this doc's
  own A1 row (93.40%) and notes the sibling 93.19% receipt reproduces too.
- docs/guides/search-modes.md: the knob list says seven, not five.
- docs/TESTING.md: the LongMemEval harness entry names the two .slow files
  that actually exist (the cited test/eval-longmemeval.test.ts did not) and
  the inventory gains the wave's test files (harness, judge lane, embed
  cache, diagnostics, spend guard, autocut replay, R1 runner, flag registry,
  gateway temperature, and the four search-knob pure + hermetic pairs); two
  long-stale entries fixed (upgrade.test.ts renamed to upgrade.serial in
  v0.36.1.1; skillpack-sync-guard.test.ts deleted in v0.36.0.0).
- llms-full.txt regenerated (bun run build:llms).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs: independent doc-review drifts — eval-bench perf-gate test name, CHANGELOG reranker attribution, RETRIEVAL current-result sentence, two env knobs in KEY_FILES

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(key-files): document GBRAIN_LME_DEBUG and the spend guard's kill-grace knob

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(eval): semgrep pass — argv-form git probe in the harness, clause splitter iterates g-flagged literals instead of constructing RegExps at runtime

The PR scan attributed three new blocking findings to this branch:
detect-child-process on the harness's git-describe helper (execSync with a
string command) and detect-non-literal-regexp twice in the clause splitter
(new RegExp(re.source, flags) to add the g flag). The helper now runs a fixed
git subcommand through execFileSync with an argv array; the splitter's
patterns carry the g flag at their literal source and are iterated with
matchAll (which clones the regex, so shared literals never leak lastIndex).
Behavior is unchanged; the 47 diagnostics/parse tests pass and typecheck is
clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 20:40:29 -07:00
..