v0.41.17.0 feat: --workers N on every bulk command + facts dim doctor parity (#1519)

* feat(worker-pool): shared sliding pool + bounded semaphore + PGLite-clamp wrapper

T1 + T2 of the v0.41.16.0 workers cathedral. New src/core/worker-pool.ts is
the canonical primitive every --workers N bulk command in this wave (and
future bulk commands) builds on. Atomic-claim invariant enforced by
scripts/check-worker-pool-atomicity.sh (wired into bun run verify).
BudgetExhausted bypass + AbortSignal composition baked into the helper so
budget caps are a structural ceiling under concurrency, not a per-caller
convention.

The new resolveWorkersWithClamp wrapper composes existing autoConcurrency
with PGLite-clamp + per-(command, requested) stderr dedup. Deliberately
NOT a modification to shared autoConcurrency (silent today, used by sync
+ import); embed.ts keeps GBRAIN_EMBED_CONCURRENCY || 20 default per
codex #13.

23 + 12 + 9 = 44 hermetic tests pin every contract.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: structural + dim-check regression suites for v0.41.16.0 wave

- test/embed-helper-migration.test.ts (T3): asserts embed.ts's two
  sliding-pool sites are migrated to runSlidingPool, pre-migration
  shapes (let nextIdx = 0, Promise.all(Array.from(...))) are gone,
  GBRAIN_EMBED_CONCURRENCY || 20 default preserved, failureLabel
  threads page.slug. Per codex #16/#17 these are invariant assertions,
  not byte-equality on progress event ORDERING.
- test/embedding-dim-check-facts.test.ts (T6): readFactsEmbeddingDim
  covers vector(N) + halfvec(N), halfvec-before-vector regex ordering
  pinned (codex #19), buildFactsAlterRecipe emits DROP INDEX + ALTER
  USING + CREATE INDEX (codex #18, not bare REINDEX),
  FactsEmbeddingDimMismatchError tagged class shape,
  assertFactsEmbeddingDimMatchesConfig PGLite skip + Postgres absent-
  column skip, doctor check + insert-cast wiring assertions.
- test/extract-conversation-facts-workers.test.ts (T5): helper
  exports (extractConversationFactsLockId, PER_PAGE_LOCK_TTL_MINUTES),
  structural wiring (runSlidingPool, resolveWorkersWithClamp,
  withRefreshingLock, LockUnavailableError, delete-orphans-first
  before segment loop, preflight before pool, exit 3 when lock_skipped
  > 0), Minion handler round-trip.
- test/extract-workers.test.ts (T7): --workers wiring on all 3 inner
  fs-walk loops (extractForSlugs, extractLinksFromDir,
  extractTimelineFromDir) + CLI parse + opts threading through
  runExtractCore.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: rebump v0.41.16.0 → v0.41.17.0 (queue collision with PR #1510)

PR #1510 (garrytan/dynamic-regex-conversation-formats) claimed v0.41.16.0
on master in parallel. Advancing this wave to v0.41.17.0 so both can land
cleanly. Pure mechanical version bump:

- VERSION + package.json → 0.41.17.0
- CHANGELOG.md header + "To take advantage of v0.41.17.0" block
- TODOS.md section header + v0.41.18+ forward references
- CLAUDE.md inline version tags
- Regenerated llms-full.txt / llms.txt

No code changes. The actual workers cathedral feature set is unchanged
from the two prior commits in this branch.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(test): search-image-column probes column dim at runtime

CI shard 5 failed on `searchVector column routing (v0.27.1)` with:
  error: expected 1280 dimensions, not 1536

The test had a hardcoded `fakeText1536` helper that seeded chunks at
1536-d vectors. Master's default embedding model switched from OpenAI
text-embedding-3-large (1536) to ZeroEntropy zembed-1 (1280) so a fresh
PGLite brain on CI now sizes content_chunks.embedding at 1280; the
test's 1536-d INSERT trips pgvector's CheckExpectedDim.

Fix: probe `content_chunks.embedding` width via
`readContentChunksEmbeddingDim(engine)` in `beforeAll`, store in
`TEXT_DIM`, and build `fakeTextDefault(seed)` at that width. The test
now passes regardless of which default ships (the model has flipped
twice and may flip again). Local dev (1536 from older config) and CI
fresh-install (1280 from new default) both pass.

Image-side vectors stay at 1024 (matches Voyage multimodal-3 + the
column's fixed width on the image side).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(test): bump PGLite hook timeout for shard-4 deep-process files

facts-anti-loop.test.ts and ingest-capture.test.ts were timing out in CI
shard 4 with "beforeEach/afterEach hook timed out" after the v0.41.16.0
master merge brought migration count to 99. When these files run deep in
a shard process that has already created ~20 PGLite engines, the WASM
cold-start + 95-migration replay legitimately exceeds bun's 5s default
hook timeout (observed 5.6s and 7.3s locally when reproducing).

Bun's --timeout=60000 from scripts/test-shard.sh covers TEST timeouts
but NOT hook timeouts; those default to 5s and must be set per-hook via
the optional 2nd arg to beforeAll/afterAll.

Reproduced locally by running the first 21 shard-4 files via
  head -21 /tmp/shard4-list.txt | xargs bun test
  → 179 pass, 2 fail (both with hook-timeout error)

After fix:
  → 198 pass, 0 fail (the 4 anti-loop + 15 ingest-capture tests recover)

Full shard 4 with fix:  955 pass, 0 fail.
Full shard 5 with fix:  1261 pass, 0 fail.

Also added a defensive diagnostic to the two put_page tests: if
facts_backstop is missing in the response payload, throw with the full
payload + isError so future failures surface the actual handler error
instead of a bare "expected {...} got undefined" assertion. No-op when
the test passes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-05-26 18:29:03 -07:00
committed by GitHub
parent f702ec053b
commit 8ab733471b
36 changed files with 3690 additions and 542 deletions

View File

@@ -2956,6 +2956,10 @@ export class PostgresEngine implements BrainEngine {
const embedding = input.embedding ?? null;
const embeddedAt = embedding ? new Date() : null;
const embedLit = embedding ? toPgVectorLiteral(embedding) : null;
// v0.41.15.0 (T6, codex #20): match cast to actual column type so
// a halfvec(N) column doesn't pay an implicit-cast round-trip + can
// run on pgvector versions that lack the auto vector→halfvec cast.
const castSuffix = await this.resolveFactsEmbeddingCast();
// v0.35.4 (D-CDX-5) — typed-claim columns. All four nullable.
const claimMetric = input.claim_metric ?? null;
const claimValue = input.claim_value ?? null;
@@ -2978,7 +2982,7 @@ export class PostgresEngine implements BrainEngine {
) VALUES (
${ctx.source_id}, ${entitySlug}, ${input.fact}, ${kind}, ${visibility}, ${notability}, ${context},
${validFrom}, ${validUntil}, ${input.source}, ${sourceSession}, ${confidence},
${embedLit === null ? null : tx.unsafe(`'${embedLit}'::vector`)}, ${embeddedAt},
${embedLit === null ? null : tx.unsafe(`'${embedLit}'${castSuffix}`)}, ${embeddedAt},
${claimMetric}, ${claimValue}, ${claimUnit}, ${claimPeriod}
) RETURNING id
`;
@@ -3004,7 +3008,7 @@ export class PostgresEngine implements BrainEngine {
) VALUES (
${ctx.source_id}, ${entitySlug}, ${input.fact}, ${kind}, ${visibility}, ${notability}, ${context},
${validFrom}, ${validUntil}, ${input.source}, ${sourceSession}, ${confidence},
${embedLit === null ? null : tx.unsafe(`'${embedLit}'::vector`)}, ${embeddedAt},
${embedLit === null ? null : tx.unsafe(`'${embedLit}'${castSuffix}`)}, ${embeddedAt},
${claimMetric}, ${claimValue}, ${claimUnit}, ${claimPeriod}
) RETURNING id
`;
@@ -3024,6 +3028,59 @@ export class PostgresEngine implements BrainEngine {
return (result.count ?? 0) > 0;
}
/**
* v0.41.15.0 (T6, codex #20): per-process cache for the
* `facts.embedding` cast suffix. Migration v40 creates the column as
* `halfvec(N)` on pgvector >= 0.7 but falls back to `vector(N)` on
* older. The pre-v0.41.15 insert path always cast embeddings as
* `::vector`, which works via implicit cast on pgvector >= 0.7 but
* is honest-only when the column actually IS vector. Probing once
* per process + caching the suffix lets the insert match the column
* type exactly. Initialized lazily in `insertFacts`.
*/
private _factsEmbeddingCastSuffix: '::vector' | '::halfvec' | null = null;
/** Test seam: clear the cached cast suffix so tests can re-probe. */
__resetFactsEmbeddingCastCacheForTest(): void {
this._factsEmbeddingCastSuffix = null;
}
private async resolveFactsEmbeddingCast(): Promise<'::vector' | '::halfvec'> {
if (this._factsEmbeddingCastSuffix !== null) return this._factsEmbeddingCastSuffix;
const sql = this.sql;
try {
const rows = await sql<Array<{ formatted: string | null }>>`
SELECT format_type(a.atttypid, a.atttypmod) AS formatted
FROM pg_attribute a
JOIN pg_class c ON c.oid = a.attrelid
JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = 'public'
AND c.relname = 'facts'
AND a.attname = 'embedding'
AND NOT a.attisdropped
`;
const formatted = rows?.[0]?.formatted ?? null;
// halfvec match first — halfvec contains "vec" so a /vector/i
// regex would shadow it. See readFactsEmbeddingDim's identical
// ordering note.
if (formatted && /halfvec\(\d+\)/i.test(formatted)) {
this._factsEmbeddingCastSuffix = '::halfvec';
} else {
// Default to '::vector' (the pre-v0.41.15 behavior). On a brain
// without the facts.embedding column yet (pre-v40), the cast
// suffix is irrelevant — the INSERT would fail elsewhere
// anyway. Caching the default still saves the SELECT on
// subsequent inserts.
this._factsEmbeddingCastSuffix = '::vector';
}
} catch {
// Probe failed — fall back to '::vector' default. Cache so we
// don't re-probe on every insert.
this._factsEmbeddingCastSuffix = '::vector';
}
return this._factsEmbeddingCastSuffix;
}
async insertFacts(
rows: Array<NewFact & { row_num: number; source_markdown_slug: string }>,
ctx: { source_id: string },
@@ -3031,6 +3088,10 @@ export class PostgresEngine implements BrainEngine {
if (rows.length === 0) return { inserted: 0, ids: [] };
const sql = this.sql;
// v0.41.15.0 (T6, codex #20): resolve the embedding-cast suffix
// ONCE per process so the cast matches the actual column type
// (halfvec vs vector). The probe is cached after first call.
const castSuffix = await this.resolveFactsEmbeddingCast();
// Single transaction so the v51 partial UNIQUE index can roll back
// the whole batch on constraint violation. Per-row INSERTs (not
// multi-row VALUES) keep the embedding-vs-no-embedding branching
@@ -3071,7 +3132,7 @@ export class PostgresEngine implements BrainEngine {
) VALUES (
${ctx.source_id}, ${entitySlug}, ${input.fact}, ${kind}, ${visibility}, ${notability}, ${context},
${validFrom}, ${validUntil}, ${input.source}, ${sourceSession}, ${confidence},
${embedLit === null ? null : tx.unsafe(`'${embedLit}'::vector`)}, ${embeddedAt},
${embedLit === null ? null : tx.unsafe(`'${embedLit}'${castSuffix}`)}, ${embeddedAt},
${input.row_num}, ${input.source_markdown_slug},
${claimMetric}, ${claimValue}, ${claimUnit}, ${claimPeriod},
${eventType}