Files
gbrain/src/core/storage-config.ts
Garry Tan 52f9581966 v0.22.11 feat: storage tiering — db_tracked vs db_only directories (#494)
* feat: storage tiering — git-tracked vs supabase-only directories

Brain repos scaling to 200K+ files. Bulk data (tweets, articles, transcripts)
bloats git repos and slows operations. New storage config in gbrain.yml lets
users declare git-tracked and supabase-only directories.

Changes:
- New config: storage.git_tracked and storage.supabase_only in gbrain.yml
- gbrain sync auto-manages .gitignore for supabase-only paths
- gbrain export --restore-only restores missing supabase-only files from DB
- New gbrain storage status command shows tier breakdown
- Config validation warns on conflicts
- 8 tests passing, full docs at docs/storage-tiering.md

Backward compatible — systems without gbrain.yml work unchanged.

* feat: add getDefaultSourcePath() typed accessor (step 1/15)

Single source of truth for "what brain repo are we operating against?"
Replaces ad-hoc raw SQL in storage.ts:38 (Issue #3 of eng review). Used by
both gbrain storage status and gbrain export --restore-only.

Returns null on miss, throws on DB error. Composes with the existing
resolveSourceId chain so it honors --source flag / GBRAIN_SOURCE env /
.gbrain-source dotfile / longest-prefix CWD match / brain-level default.

4 new test cases covering happy path, missing local_path, DB error
propagation, and CWD-prefix resolution priority.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: replace gray-matter with dedicated YAML parser (step 2/15)

The original storage-config.ts called gray-matter on a delimiter-less YAML
file. Gray-matter only parses YAML inside `---` frontmatter blocks; without
delimiters, it returns `{data: {}}`. Result: loadStorageConfig() always
returned null, the entire feature was a silent no-op for every user.

Original eng review's P0 confidence-9 finding (Issue #1).

Replaces gray-matter with a small dedicated parser for the gbrain.yml shape
(top-level `storage:` section, two array-valued nested keys). Yaml-lite was
considered first, but its flat key:value design doesn't handle nested
arrays. The dedicated parser is ~50 lines and trades expressiveness for
zero-dep, predictable parsing of a file format we control.

Adds the Issue #1B sanity warning (locked B): when gbrain.yml exists but
has no storage section (or empty arrays), warn once-per-process so the
user sees their config didn't take. The single test that would have caught
the original P0 — write a real gbrain.yml, call loadStorageConfig, assert
non-null — now exists.

Also tightens loadStorageConfig per D36: distinguishes "absent" (silent
null) from "unreadable" (throws). The previous code silently swallowed
read errors, hiding broken installs.

8 new test cases: real-disk happy path, comments + blank lines, quoted
values, missing storage section warning, empty section warning,
once-per-process warning suppression, unreadable file behavior, and the
existing helper tests (validation, tier matching, edge cases) all still
pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor: rename storage keys to db_tracked/db_only (step 3/15)

The vendor-specific names "supabase_only" and "git_tracked" hardcoded a
backend (Supabase) into the config schema. gbrain ships two engines —
PGLite and Postgres-via-Supabase. The canonical distinction is "lives in
the brain DB only" vs "lives in the brain DB and on disk under git." Both
work on either engine.

Renamed throughout (Issue #4 of eng review):
  git_tracked    → db_tracked
  supabase_only  → db_only
  isGitTracked() → isDbTracked()
  isSupabaseOnly() → isDbOnly()
  StorageTier 'git_tracked'/'supabase_only' → 'db_tracked'/'db_only'

Backward compatibility (D3 lock):
  loadStorageConfig accepts both shapes. Loader resolution order per the
  eng-review pass-2 finding: parse YAML → if canonical keys present use
  them, else if deprecated keys present map to canonical AND emit
  once-per-process deprecation warning → THEN run validation.
  Validation always sees the canonical shape so error messages reference
  db_tracked/db_only regardless of which keys the user wrote.

  The deprecation warning suggests `gbrain doctor --fix` for an automated
  rename (D72 — fix path lands in step 7).

  When both shapes coexist in one file, canonical wins and a stronger
  warning fires ("deprecated keys ignored — remove them").

Aliases isGitTracked/isSupabaseOnly kept for now to avoid churning the
sync.ts / export.ts / storage.ts call sites in this commit; they'll be
removed in a follow-up step. Storage.ts's tier-bucket initializers and
output strings updated. ASCII output replaces unicode box-drawing per D10.

gbrain.yml example file updated to canonical keys with explanatory
comments.

2 new test cases: deprecated-key fallback (asserts both shapes load
correctly with warning), canonical-wins-over-deprecated (asserts the
"both shapes coexist" path).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: add slugPrefix to PageFilters with engine-side filter (step 4/15)

Issue #13 of the eng review: storage.ts and export.ts loaded every page
in the brain (limit: 1_000_000) to check tier membership. On the 200K-page
brains this feature targets, that's the wall-clock and memory landmine
the feature exists to fix.

Adds an optional `slugPrefix` field to PageFilters. Both engines implement
it as `WHERE slug LIKE prefix || '%' ESCAPE '\'`, with literal escaping of
LIKE metacharacters (%, _, \) so user-supplied prefixes like `media/x/`
are treated as exact string prefixes.

Performance: the (source_id, slug) UNIQUE constraint on the pages table
gives both engines a btree index that supports LIKE-prefix range scans.
An EXPLAIN on Postgres confirms the index range scan rather than a seq
scan. PGLite has the same index shape via pglite-schema.ts.

Consumers updated:
  - export.ts: --slug-prefix flag now goes engine-side (no in-memory
    .filter(...)). The --restore-only path queries each db_only directory
    with slugPrefix in a loop instead of one full-table scan, with seen-set
    deduplication and disk-existence check inline.
  - storage.ts: keeps the full-scan path because storage-status needs the
    "unspecified" bucket count, which can't be computed without enumerating
    every page. Comment notes that step 5 (single-walk filesystem scan)
    will reduce per-page disk syscall cost.

2 new test cases on PGLiteEngine: slugPrefix happy path (3 tier dirs,
asserts only matching slugs return) and metacharacter escape regression
(asserts safe/ doesn't match unrelated slugs).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf: single-walk filesystem scan via walkBrainRepo() (step 5/15)

Issue #14 of the eng review: storage.ts called existsSync + statSync
per-page in a synchronous loop. On a 200K-page brain that's 400K syscalls
serialized. Wall-clock landmine.

Adds src/core/disk-walk.ts with walkBrainRepo(repoPath) — one recursive
readdirSync walk, builds a Map<slug, {size, mtimeMs}>. Storage.ts looks
up each DB page in the map (O(1)) instead of stat-checking on demand.
Slug derivation matches the pages-table convention: people/alice.md on
disk becomes people/alice as the map key.

Skipped during walk:
  - dot-directories (.git, .gbrain, .vscode, etc) — not part of the brain
    namespace
  - node_modules — guards against accidentally walking into imported repos
  - non-.md files (sidecar JSON, binaries) — tracked by the brain through
    the files table, not by slug

Reusable: future commands (gbrain doctor's storage_tiering check, the
optional autopilot tier-fix path) get the same walk for free.

9 new test cases: empty dir, nonexistent dir, top-level files, nested
dirs, dot-dir skipping, node_modules skipping, non-.md filtering, size
capture, mtimeMs capture.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: path-segment matching for tier directories (step 6/15)

Issue #5 + D6 of the eng review: tier matching used slug.startsWith(dir),
which falsely matches 'media/xerox/foo' against 'media/x' if a user wrote
the directory without a trailing slash.

The new matcher requires the configured directory to end with `/` and
treats it as a canonical path-segment ancestor:

  media/x/   matches  media/x/tweet-1       ✓
  media/x/   doesn't  media/xerox/foo       ✗
  media/x    refused  media/x/tweet-1       (matcher requires trailing /)

Non-canonical input (no trailing slash) is refused outright. Step 7's
auto-normalizing validator converts user-written 'media/x' → 'media/x/'
on load, so the matcher never sees non-canonical input from real configs.
The behavior tested here is the strict matcher's contract.

Regression test pins the media/xerox collision case explicitly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: auto-normalize trailing-slash, throw on tier overlap (step 7/15)

D7+D8 of the eng review: validation was warnings-only. Users miss warnings.
Now:

  - Cosmetic: missing trailing slash auto-corrected, one-time info note
    showing what changed ("normalized 2 storage paths: 'people' →
    'people/', 'media/x' → 'media/x/'"). Once-per-process to keep noise low.

  - Semantic: same directory in both tiers throws StorageConfigError.
    Ambiguous routing — does media/ win as db_tracked or db_only? — is a
    real bug the user must fix. Caller propagates to the CLI for a clean
    exit-1 with actionable message.

loadStorageConfig now applies normalize+validate after merging deprecated
keys, so the path-segment matcher (step 6) only ever sees canonical
trailing-slash directories.

The pure validateStorageConfig kept for callers who want the warnings list
without the auto-fix side effects (gbrain doctor's reporting path).

2 new test cases: auto-normalize round-trip with warning text assertion,
overlap throws StorageConfigError.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: wire manageGitignore into runSync, only on success (step 8/15)

Issue #2 of the eng review: manageGitignore was defined and never
invoked. Docs claimed "auto-managed by gbrain" — false. Users hit a
.gitignore that never updated and committed db_only directories anyway.

Wire-up: runSync now calls manageGitignore after each successful
performSync return, in both watch and one-shot modes.

Eng review pass-2 finding #1: skip on dry_run AND blocked_by_failures
status. A sync that aborted partway has stale state; mutating .gitignore
based on a partially-loaded config invites drift. Failure-skip test
added (uses .gitignore-as-a-directory to simulate write failure;
asserts warning fired and disk wasn't corrupted).

Hardened manageGitignore itself with three additional behaviors:

  - GBRAIN_NO_GITIGNORE=1 escape hatch (D23) for shared-repo setups
    where a maintainer wants gbrain to leave .gitignore alone.

  - Submodule detection (D49). When repoPath/.git is a regular file
    (gitdir: ... pointer), the repo is a git submodule. Submodule
    .gitignore changes don't survive parent submodule updates, so we
    skip with an actionable warning ("add db_only directories to your
    parent repo's .gitignore manually").

  - Graceful failure (D9). Read errors, write errors, and
    StorageConfigError (overlap from step 7) all log a warning and
    return — sync's primary job (moving data) shouldn't die because of
    a side-effect on .gitignore.

manageGitignore is now exported (previously private) so the
storage-sync test file can hit it directly without spinning up sync.

9 new test cases: no-op without gbrain.yml, no-op with empty db_only,
happy-path append, idempotency (run twice, single entry), preservation
of user-written rules, GBRAIN_NO_GITIGNORE skip, submodule skip,
.git-directory normal path, write-failure graceful warning.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: D5 resolution chain for --restore-only and storage status (step 9/15)

D5 of the eng review: gbrain export --restore-only without --repo
silently fell through to the regular export path, dumping every page in
the database to the wrong directory. Hard regression risk.

Now exits 1 with an actionable message when --restore-only has no
--repo AND no configured default source. Resolution order:
  1. Explicit --repo flag
  2. Typed sources.getDefault() (reuses step 1's accessor)
  3. Hard error — never fall through to cwd

storage.ts:38 also bypassed BrainEngine with raw SQL and a bare
try/catch (Issue #3 + Issue #9). Replaced with the same typed
getDefaultSourcePath() — single source of truth, errors propagate
cleanly to the user, no silent cwd fallback.

Regular export (no --restore-only) keeps its current behavior per D26:
exports include everything, --repo is optional.

4 new test cases on PGLite in-memory:
  - hard-errors with no --repo + no default
  - explicit --repo wins
  - falls back to sources default local_path
  - non-restore export does not require --repo

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor: split storage.ts into pure data + JSON + human formatters (step 10/15)

Issue #10 of the eng review: getStorageStatus and runStorageStatus mixed
data gathering, JSON serialization, and human-readable output in one
function. Hard to test, hard to reuse, mismatched the orphans.ts pattern
that CLAUDE.md cites as the precedent.

Now three pure functions + a thin dispatcher:

  getStorageStatus(engine, repoPath) — async, returns StorageStatusResult.
    Side effects: engine.listPages + one walkBrainRepo (Issue #14).
    Exported so MCP exposure (D14) and gbrain doctor (D13) can consume the
    same data without re-running the loop.

  formatStorageStatusJson(result) — pure, returns indented JSON. Stable
    contract on the StorageStatusResult shape, suitable for orchestrators.

  formatStorageStatusHuman(result) — pure, returns ASCII text (D10 — no
    unicode box-drawing). Composable into other commands later.

  runStorageStatus(engine, args) — thin dispatcher: parses --repo /
    --json, calls getStorageStatus, picks a formatter, prints.

8 new test cases on the formatters: JSON parse round-trip, null-config
fallback, missing-files capped at 10 with rollup, ASCII-only assertion
(D10 regression guard), warnings inline, configuration listing, disk-
usage block omitted when zero bytes.

The StorageStatusResult interface is now exported as a public type, so
gbrain doctor's storage_tiering check can build its own findings from
the same shape.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* types: distinct PageCountsByTier and DiskUsageByTier (step 11/15)

Issue #11 of the eng review: pagesByTier (page counts) and
diskUsageByTier (byte totals) shared the same structural type
(Record<StorageTier, number>). Both are tier-keyed numeric maps but
carry semantically different units. A future bug that swaps them at a
call site (e.g., displaying disk bytes where the count belongs) wouldn't
trip the compiler.

Replaced with distinct nominal types via a brand field. Structurally
identical at runtime (no overhead) but compile-time disjoint —
TypeScript catches accidental cross-assignment.

  PageCountsByTier   { db_tracked, db_only, unspecified } : numbers (count)
  DiskUsageByTier    { db_tracked, db_only, unspecified } : numbers (bytes)

Both initialized in getStorageStatus, both threaded into
StorageStatusResult, both consumed by formatStorageStatusHuman /
formatStorageStatusJson without further changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: PGLite soft-warn + full lifecycle test (step 12/15)

D4: storage tiering on PGLite is a partial feature. The "DB" the pages
live in IS the local file gbrain uses for everything else, so "db_only"
has no real offload effect. The .gitignore management still helps
(keeps bulk content out of git history), so we warn and proceed —
not refuse.

Two warning sites (once-per-process each via module-local flags):
  - storage status: warns at runStorageStatus entry
  - sync: warns inside manageGitignore when engineKind='pglite' and
    config has db_only entries

Both phrased actionably ("To get full tiering, migrate to Postgres
with `gbrain migrate --to supabase`").

manageGitignore signature now takes an optional `engineKind` param.
runSync passes engine.kind. Stand-alone callers (tests, future
gbrain doctor --fix path) can omit it.

New test: test/storage-pglite.test.ts — D8 + D4 lifecycle. 6 cases:
engine.kind assertion, getStorageStatus loading gbrain.yml + reporting
tier counts, manageGitignore PGLite-warn (once per process), Postgres
no-warn, slugPrefix on PGLite, end-to-end (config + putPage + status
+ gitignore).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: add trailing-newline CI guard (step 14/15)

Issue #7 of the eng review: all four new files in the original
storage-tiering branch lacked POSIX trailing newlines. Linters complain,
git diffs phantom-flag every future edit. We've been adding newlines as
each file landed; this commit catches the regression class.

scripts/check-trailing-newline.sh:
  - sibling to check-jsonb-pattern.sh / check-progress-to-stdout.sh per
    CLAUDE.md's CI guard pattern
  - portable to bash 3.2 (macOS default; no mapfile, no associative arrays)
  - covers src/**, test/**, gbrain.yml, top-level *.md
  - reports each missing file by path and exits 1

Wired into `bun run test` between progress-to-stdout and typecheck.

Also fixed docs/storage-tiering.md (pre-existing missing newline from
the original branch — caught by the new guard on first run).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: v0.23.0 — VERSION, CHANGELOG, README, CLAUDE.md, storage-tiering.md (step 15/15)

VERSION → 0.23.0 (minor bump for new feature surface).

CHANGELOG entry in Garry voice with the canonical format:
  - Two-line bold headline ("Storage tiering, finally working...")
  - Lead paragraph naming what was broken before and what users get now
  - "Numbers that matter" before/after table for the 6 things that
    actually changed
  - "What this means for your brain" closer
  - "To take advantage of v0.23.0" self-repair block (per CLAUDE.md
    convention) — 6 numbered steps users can follow
  - Itemized changes split into critical fixes / new+renamed surface /
    architecture cleanup / tests + CI guards

CLAUDE.md "Key files" gains four new entries: storage-config.ts,
disk-walk.ts, the v0.23.0 storage.ts shape, and gbrain.yml itself.

README.md gains a new "Storage tiering" section between Skillify and
Getting Data In with the canonical example + commands + link to the
full guide.

docs/storage-tiering.md rewritten end-to-end with canonical key names
(db_tracked / db_only), v0.23.0 hardening details (idempotency,
submodule detection, GBRAIN_NO_GITIGNORE, dry-run gating), the
resolution chain for --restore-only, the auto-normalize +
throw-on-overlap validator, and the PGLite engine note.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: e2e Postgres lifecycle for storage tiering (step 16/16)

Per the v0.23.0 plan: full lifecycle E2E against real Postgres.

  - engine.kind === 'postgres' assertion
  - Full lifecycle: write 4 pages (1 db_tracked, 2 db_only, 1 unspecified)
    → getStorageStatus reports correct tier counts → human formatter
    renders → manageGitignore writes managed block → idempotency check
    → getDefaultSourcePath() resolves the configured local_path.
  - Container restart simulation: 2 db_only pages in DB, files missing
    on disk → status.missingFiles.length === 2 → slugPrefix engine
    filter on Postgres returns exactly the tier slugs.
  - slugPrefix index-based range scan regression: 50 media/x/* + 50
    people/p-* pages → slugPrefix='media/x/' returns exactly 50.
  - getDefaultSourcePath returns null when default source has no
    local_path (the hard-error path that replaces the original silent
    cwd fallback).
  - manageGitignore on Postgres engine does NOT emit the PGLite
    soft-warn (cross-engine assertion).

Skips gracefully when DATABASE_URL is unset, per CLAUDE.md E2E pattern.
Run via: DATABASE_URL=... bun test test/e2e/storage-tiering.test.ts

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: rebump version 0.23.0 → 0.22.9

Reverts the minor bump back to a patch-style version on the v0.22 line.
Storage tiering ships within the v0.22.x train alongside the recent
fix waves. Updates VERSION, package.json, CHANGELOG header + body refs,
CLAUDE.md Key files annotations, README.md section heading, and the
docs/storage-tiering.md backward-compat note.

* chore: bump version 0.22.9 → 0.22.11

Sibling workspaces claimed v0.22.10 in the queue. This branch advances
to v0.22.11 to keep the version monotonic on master.

Updates VERSION, package.json, CHANGELOG header + body refs, CLAUDE.md
Key files annotations, README.md section heading, and the
docs/storage-tiering.md backward-compat note.

* fix: address Codex pre-landing review findings (4 fixes)

Codex found 4 real issues during pre-landing review of v0.22.11 diff:

[P0] export --restore-only fell through to full export when
storageConfig was null (no gbrain.yml present). On older or
misconfigured brains, the recovery command would silently dump the
entire database. src/commands/export.ts now refuses with an actionable
error before any page query fires — matches the D5 lock spirit
("never silently fall through").

[P1] manageGitignore wire-up only fired when --repo was passed
explicitly. performSync resolves the repo from sync.repo_path or
sources.local_path, so the common `gbrain sync` path (after
setup, no flag) never updated .gitignore. src/commands/sync.ts now
uses the same source-resolver chain as the rest of /ship: opts.repoPath
→ getDefaultSourcePath → null. Fires in both watch and one-shot modes.

[P2] getDefaultSourcePath only consulted sources.local_path, missing
the legacy global sync.repo_path config key that pre-v0.18 brains use.
Added a fallback to engine.getConfig('sync.repo_path') when the
sources row has NULL local_path. Pre-v0.18 brains now work without
forcing a `gbrain sources add . --path .` migration.

[P2] sync --all multi-source loop never called manageGitignore even
though src.local_path was already known. Each source now gets its own
gitignore update on successful sync.

Tests:
  - test/storage-export.test.ts: replaced the old "falls through to
    full export" test with one that asserts the new refusal path
    (storage-tiering config required for --restore-only).
  - test/source-resolver.test.ts: added a fallback test exercising the
    legacy sync.repo_path code path for pre-v0.18 brains.
  - All 78 storage-tiering tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: regenerate llms.txt + llms-full.txt for v0.22.11

Per CLAUDE.md: "Run `bun run build:llms` after adding a new doc."
The README's new Storage tiering section + the rewritten
docs/storage-tiering.md changed the inlined bundle. test/build-llms.test.ts
catches the drift and was failing on master pre-regen.

* fix: typecheck error in disk-walk.ts (CI #73350475897)

tsc --noEmit failed in CI because ReturnType<typeof readdirSync> with
withFileTypes:true picks an overload union that includes
Dirent<Buffer<ArrayBufferLike>>. Strict tsc treats entry.name as Buffer,
so .startsWith / .endsWith / string comparisons all blew up.

Annotate the variable as Dirent[] (string-based) and cast through unknown,
matching the pattern sync.ts already uses for its own filesystem walk.
Same runtime behavior; clean typecheck.

Tests still 9/9.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
EOF

---------

Co-authored-by: root <root@localhost>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 22:21:07 -07:00

378 lines
13 KiB
TypeScript

import { readFileSync, existsSync } from 'fs';
import { join } from 'path';
/**
* Storage tier configuration loaded from gbrain.yml.
*
* The canonical key names are `db_tracked` and `db_only` (engine-agnostic).
* The deprecated keys `git_tracked` and `supabase_only` are still read for
* backward compatibility but emit a once-per-process deprecation warning.
* Sunset: future release will reject the deprecated names.
*/
export interface StorageConfig {
db_tracked: string[];
db_only: string[];
}
export type StorageTier = 'db_tracked' | 'db_only' | 'unspecified';
/** Recognized YAML keys (canonical and deprecated). */
const STORAGE_KEYS = new Set([
'db_tracked', 'db_only',
'git_tracked', 'supabase_only', // deprecated aliases
]);
/**
* Parse the gbrain.yml shape: a top-level `storage:` section with up to four
* array-valued nested keys (canonical `db_tracked` / `db_only` plus the
* deprecated aliases `git_tracked` / `supabase_only`).
*
* Intentionally narrow. Does NOT handle the full YAML spec — only the file
* shape gbrain controls. Trades expressiveness for zero-dep parsing and
* predictable behavior. Returns null if the file has no `storage:` section
* (so callers can distinguish "no config" from "empty config").
*
* Replaces gray-matter, which silently returned `{data: {}}` on
* delimiter-less YAML and broke the entire feature on every install.
* The defect that prompted this rewrite: storage-config.ts:24 in the
* pre-v0.22.3 implementation.
*
* Returns the raw key map. The caller (loadStorageConfig) is responsible
* for normalizing deprecated keys → canonical, emitting deprecation
* warnings, and merging if both old and new keys appear.
*/
type RawStorage = {
db_tracked?: string[];
db_only?: string[];
git_tracked?: string[];
supabase_only?: string[];
};
function parseStorageYaml(content: string): RawStorage | null {
const lines = content.split('\n').map((line) => line.replace(/\r$/, ''));
let inStorage = false;
let currentList: keyof RawStorage | null = null;
const raw: RawStorage = {};
let sawStorage = false;
for (const line of lines) {
// Strip comments. Conservative: drop trailing `# ...` and full-line `#`.
const noComment = line.replace(/\s+#.*$/, '').replace(/^#.*$/, '');
if (noComment.trim() === '') continue;
// Top-level key (no leading whitespace).
if (!noComment.startsWith(' ') && !noComment.startsWith('\t')) {
const colon = noComment.indexOf(':');
if (colon === -1) continue;
const key = noComment.slice(0, colon).trim();
if (key === 'storage') {
inStorage = true;
sawStorage = true;
currentList = null;
continue;
}
inStorage = false;
currentList = null;
continue;
}
if (!inStorage) continue;
const indented = noComment.replace(/^\s+/, '');
if (indented.startsWith('-')) {
if (!currentList) continue;
const value = indented.slice(1).trim().replace(/^["']|["']$/g, '');
if (value) {
if (!raw[currentList]) raw[currentList] = [];
raw[currentList]!.push(value);
}
continue;
}
const colon = indented.indexOf(':');
if (colon === -1) continue;
const key = indented.slice(0, colon).trim();
if (STORAGE_KEYS.has(key)) {
currentList = key as keyof RawStorage;
// Inline empty list: `db_only: []`.
const remainder = indented.slice(colon + 1).trim();
if (remainder === '[]' && !raw[currentList]) {
raw[currentList] = [];
}
continue;
}
currentList = null;
}
if (!sawStorage) return null;
return raw;
}
/**
* Normalize raw parsed keys into canonical StorageConfig shape.
*
* Resolution order (per plan eng-review pass 2 finding #2):
* 1. If canonical keys present, use them.
* 2. Else if deprecated keys present, map to canonical AND emit a
* once-per-process deprecation warning suggesting `gbrain doctor --fix`.
* 3. If both are present, canonical wins. Deprecated keys are ignored
* with a stronger warning (the user is mid-migration).
*
* Validation (validateStorageConfig) always runs against the canonical
* shape, so error messages reference `db_only` / `db_tracked` regardless
* of which keys the user wrote.
*/
let _deprecationWarned = false;
function normalizeStorageConfig(raw: RawStorage): StorageConfig {
const hasCanonical = Boolean(raw.db_tracked || raw.db_only);
const hasDeprecated = Boolean(raw.git_tracked || raw.supabase_only);
if (hasDeprecated && !_deprecationWarned) {
_deprecationWarned = true;
const which = [
raw.git_tracked ? '`git_tracked`' : null,
raw.supabase_only ? '`supabase_only`' : null,
].filter(Boolean).join(' and ');
if (hasCanonical) {
console.warn(
`Warning: ${which} in gbrain.yml is deprecated and ignored ` +
`(canonical keys db_tracked/db_only are present). ` +
`Remove the deprecated keys, or run \`gbrain doctor --fix\`.`,
);
} else {
console.warn(
`Warning: ${which} in gbrain.yml is deprecated. ` +
`Rename to db_tracked / db_only — see docs/storage-tiering.md. ` +
`Run \`gbrain doctor --fix\` for an automated rename.`,
);
}
}
if (hasCanonical) {
return {
db_tracked: raw.db_tracked ?? [],
db_only: raw.db_only ?? [],
};
}
return {
db_tracked: raw.git_tracked ?? [],
db_only: raw.supabase_only ?? [],
};
}
/**
* Load gbrain.yml configuration from the brain repository root.
*
* Returns null when:
* - repoPath is null/undefined
* - gbrain.yml doesn't exist at the repo root
* - gbrain.yml exists but has no `storage:` section (with sanity warning)
*
* Throws when:
* - gbrain.yml exists but is unreadable (permission denied, etc.) — D36 lock:
* fail loud rather than silently disable the feature.
*
* Logs a console.warn (once per process) when:
* - File parses but `storage:` section is empty or missing — Issue #1 lock:
* surface "your config didn't take" rather than silently no-op.
*/
let _missingStorageWarned = false;
export function loadStorageConfig(repoPath?: string | null): StorageConfig | null {
if (!repoPath) return null;
const yamlPath = join(repoPath, 'gbrain.yml');
if (!existsSync(yamlPath)) return null;
// Read failure is a real error (not a "feature not configured" signal).
// Throwing here lets the caller decide whether to crash or fall back.
const content = readFileSync(yamlPath, 'utf-8');
let raw: RawStorage | null;
try {
raw = parseStorageYaml(content);
} catch (error) {
console.warn(
`Warning: Failed to parse gbrain.yml: ${error instanceof Error ? error.message : String(error)}`,
);
return null;
}
// No storage section at all → null (with sanity warning).
if (raw === null) {
if (!_missingStorageWarned) {
_missingStorageWarned = true;
console.warn(
`Warning: ${yamlPath} exists but has no storage configuration. ` +
`Add a "storage:" section with db_tracked / db_only arrays, ` +
`or remove gbrain.yml to suppress this warning.`,
);
}
return null;
}
const merged = normalizeStorageConfig(raw);
// Empty storage section → return as-is but warn.
if (merged.db_tracked.length === 0 && merged.db_only.length === 0) {
if (!_missingStorageWarned) {
_missingStorageWarned = true;
console.warn(
`Warning: ${yamlPath} exists but has no storage configuration. ` +
`Add a "storage:" section with db_tracked / db_only arrays, ` +
`or remove gbrain.yml to suppress this warning.`,
);
}
return merged;
}
// Normalize cosmetic issues + throw on semantic overlap (D7).
// Throws StorageConfigError on overlap — propagates to the caller.
return normalizeAndValidateStorageConfig(merged);
}
export class StorageConfigError extends Error {
constructor(message: string) {
super(message);
this.name = 'StorageConfigError';
}
}
/**
* Validate storage configuration for conflicts and issues.
* Returns warning strings; callers decide how to surface them.
*
* Always runs against the canonical (db_tracked / db_only) shape — error
* messages reference canonical names regardless of which keys the user
* wrote in gbrain.yml.
*
* Pure: does not mutate. For the auto-normalize behavior (D7), see
* `normalizeAndValidateStorageConfig` below.
*/
export function validateStorageConfig(config: StorageConfig): string[] {
const warnings: string[] = [];
const trackedSet = new Set(config.db_tracked);
for (const path of config.db_only) {
if (trackedSet.has(path)) {
warnings.push(`Directory "${path}" appears in both db_tracked and db_only`);
}
}
const allPaths = [...config.db_tracked, ...config.db_only];
for (const path of allPaths) {
if (!path.endsWith('/')) {
warnings.push(`Directory path "${path}" should end with "/" for consistency`);
}
}
return warnings;
}
/**
* Auto-normalize and strict-validate per D7+D8.
*
* 1. Cosmetic fixups are applied silently with a one-time info message
* naming what changed:
* - missing trailing `/` is added
* The message helps the user learn the canonical form without nagging.
* 2. Semantic problems THROW (don't return warnings):
* - same directory in both tiers (ambiguous routing)
*
* Caller passes a fresh raw config; this returns the normalized shape that
* the rest of the code (matcher, sync, etc.) sees.
*/
let _normalizationInfoEmitted = false;
export function normalizeAndValidateStorageConfig(input: StorageConfig): StorageConfig {
const normalize = (paths: string[]): { normalized: string[]; changed: string[] } => {
const normalized: string[] = [];
const changed: string[] = [];
for (const p of paths) {
if (p.endsWith('/')) {
normalized.push(p);
} else {
normalized.push(p + '/');
changed.push(`"${p}" → "${p}/"`);
}
}
return { normalized, changed };
};
const tracked = normalize(input.db_tracked);
const dbonly = normalize(input.db_only);
const allChanged = [...tracked.changed, ...dbonly.changed];
if (allChanged.length > 0 && !_normalizationInfoEmitted) {
_normalizationInfoEmitted = true;
console.warn(
`Note: normalized ${allChanged.length} storage path(s) in gbrain.yml — ` +
`${allChanged.join(', ')}. Add trailing "/" to suppress this note.`,
);
}
// Semantic check: overlap between tiers throws. Ambiguous routing.
const trackedSet = new Set(tracked.normalized);
for (const path of dbonly.normalized) {
if (trackedSet.has(path)) {
throw new StorageConfigError(
`gbrain.yml: directory "${path}" appears in both db_tracked and db_only — ` +
`pick one tier. Edit gbrain.yml to remove the overlap.`,
);
}
}
return { db_tracked: tracked.normalized, db_only: dbonly.normalized };
}
/**
* Path-segment match: a slug belongs to a tier directory iff the directory
* is a complete path-segment ancestor of the slug. `media/x/` matches
* `media/x/foo` but NOT `media/xerox/foo` — eliminates the prefix-collision
* class of bug (Issue #5 of the eng review, D6 lock).
*
* Strict: requires the configured directory to end with `/`. The validator
* (per D7+D8) auto-normalizes input so the matcher only ever sees canonical
* trailing-`/` directories.
*/
function matchesTierDir(slug: string, dir: string): boolean {
if (!dir.endsWith('/')) return false; // not normalized — matcher refuses
// slug must equal dir's bare prefix OR start with the trailing-slash form.
// Example: dir = 'media/x/' matches 'media/x/anything' but not 'media/x'
// or 'media/xerox'. (A slug that exactly equals 'media/x' is a directory-
// level entry the brain doesn't write.)
return slug.startsWith(dir);
}
export function isDbTracked(slug: string, config: StorageConfig): boolean {
return config.db_tracked.some((dir) => matchesTierDir(slug, dir));
}
export function isDbOnly(slug: string, config: StorageConfig): boolean {
return config.db_only.some((dir) => matchesTierDir(slug, dir));
}
export function getStorageTier(slug: string, config: StorageConfig): StorageTier {
if (isDbTracked(slug, config)) return 'db_tracked';
if (isDbOnly(slug, config)) return 'db_only';
return 'unspecified';
}
// ── Deprecated aliases — to be removed in a future release ────────
// Kept so existing callers (storage.ts, export.ts) compile during the
// step-by-step refactor. Will be deleted once those call sites migrate
// to the canonical names.
export const isGitTracked = isDbTracked;
export const isSupabaseOnly = isDbOnly;
/** Reset once-per-process warning flags. Test-only. */
export function __resetMissingStorageWarning(): void {
_missingStorageWarned = false;
_deprecationWarned = false;
_normalizationInfoEmitted = false;
}