Files
gbrain/skills/migrate/SKILL.md
Garry Tan c0b621923b fix: JSONB double-encode + splitBody wiki + parseEmbedding (v0.12.1) (#196)
* fix: splitBody and inferType for wiki-style markdown content

- splitBody now requires explicit timeline sentinel (<!-- timeline -->,
  --- timeline ---, or --- directly before ## Timeline / ## History).
  A bare --- in body text is a markdown horizontal rule, not a separator.
  This fixes the 83% content truncation @knee5 reported on a 1,991-article
  wiki where 4,856 of 6,680 wikilinks were lost.

- serializeMarkdown emits <!-- timeline --> sentinel for round-trip stability.

- inferType extended with /writing/, /wiki/analysis/, /wiki/guides/,
  /wiki/hardware/, /wiki/architecture/, /wiki/concepts/. Path order is
  most-specific-first so projects/blog/writing/essay.md → writing,
  not project.

- PageType union extended: writing, analysis, guide, hardware, architecture.

Updates test/import-file.test.ts to use the new sentinel.

Co-Authored-By: @knee5 (PR #187)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: JSONB double-encode bug on Postgres + parseEmbedding NaN scores

Two related Postgres-string-typed-data bugs that PGLite hid:

1. JSONB double-encode (postgres-engine.ts:107,668,846 + files.ts:254):
   ${JSON.stringify(value)}::jsonb in postgres.js v3 stringified again
   on the wire, storing JSONB columns as quoted string literals. Every
   frontmatter->>'key' returned NULL on Postgres-backed brains; GIN
   indexes were inert. Switched to sql.json(value), which is the
   postgres.js-native JSONB encoder (Parameter with OID 3802).
   Affected columns: pages.frontmatter, raw_data.data,
   ingest_log.pages_updated, files.metadata. page_versions.frontmatter
   is downstream via INSERT...SELECT and propagates the fix.

2. pgvector embeddings returning as strings (utils.ts):
   getEmbeddingsByChunkIds returned "[0.1,0.2,...]" instead of
   Float32Array on Supabase, producing [NaN] cosine scores.
   Adds parseEmbedding() helper handling Float32Array, numeric arrays,
   and pgvector string format. Throws loud on malformed vectors
   (per Codex's no-silent-NaN requirement); returns null for
   non-vector strings (treated as "no embedding here"). rowToChunk
   delegates to parseEmbedding.

E2E regression test at test/e2e/postgres-jsonb.test.ts asserts
jsonb_typeof = 'object' AND col->>'k' returns expected scalar across
all 5 affected columns — the test that should have caught the original
bug. Runs in CI via the existing pgvector service.

Co-Authored-By: @knee5 (PR #187 — JSONB triple-fix)
Co-Authored-By: @leonardsellem (PR #175 — parseEmbedding)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: extract wikilink syntax with ancestor-search slug resolution

extractMarkdownLinks now handles [[page]] and [[page|Display Text]]
alongside standard [text](page.md). For wiki KBs where authors omit
leading ../ (thinking in wiki-root-relative terms), resolveSlug
walks ancestor directories until it finds a matching slug.

Without this, wikilinks under tech/wiki/analysis/ targeting
[[../../finance/wiki/concepts/foo]] silently dangled when the
correct relative depth was 3 × ../ instead of 2.

Co-Authored-By: @knee5 (PR #187)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: gbrain repair-jsonb + v0.12.1 migration + CI grep guard

- New gbrain repair-jsonb command. Detects rows where
  jsonb_typeof(col) = 'string' and rewrites them via
  (col #>> '{}')::jsonb across 5 affected columns:
  pages.frontmatter, raw_data.data, ingest_log.pages_updated,
  files.metadata, page_versions.frontmatter. Idempotent — re-running
  is a no-op. PGLite engines short-circuit cleanly (the bug never
  affected the parameterized encode path PGLite uses). --dry-run
  shows what would be repaired; --json for scripting.

- New v0_12_1.ts migration orchestrator. Phases: schema → repair → verify.
  Modeled on v0_12_0 pattern, registered in migrations/index.ts.
  Runs automatically via gbrain upgrade / apply-migrations.

- CI grep guard at scripts/check-jsonb-pattern.sh fails the build if
  anyone reintroduces the ${JSON.stringify(x)}::jsonb interpolation
  pattern. Wired into bun test via package.json. Best-effort static
  analysis (multi-line and helper-wrapped variants are caught by the
  E2E round-trip test instead).

- Updates apply-migrations.test.ts expectations to account for the new
  v0.12.1 entry in the registry.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.12.1)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: update project documentation for v0.12.1

- CLAUDE.md: document repair-jsonb command, v0_12_1 migration,
  splitBody sentinel contract, inferType wiki subtypes, CI grep
  guard, new test files (repair-jsonb, migrations-v0_12_1, markdown)
- README.md: add gbrain repair-jsonb to ADMIN command reference
- INSTALL_FOR_AGENTS.md: fix verification count (6 -> 7), add
  v0.12.1 upgrade guidance for Postgres brains
- docs/GBRAIN_VERIFY.md: add check #8 for JSONB integrity on
  Postgres-backed brains
- docs/UPGRADING_DOWNSTREAM_AGENTS.md: add v0.12.1 section with
  migration steps, splitBody contract, wiki subtype inference
- skills/migrate/SKILL.md: document native wikilink extraction
  via gbrain extract links (v0.12.1+)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 07:14:24 +08:00

4.9 KiB

name, description, triggers, tools, mutating
name description triggers tools mutating
migrate Universal migration from Obsidian, Notion, Logseq, markdown, CSV, JSON, Roam
migrate from
import from obsidian
import from notion
put_page
search
add_link
add_tag
sync_brain
true

Migrate Skill

Universal migration from any wiki, note tool, or brain system into GBrain.

Contract

  • Source data is never modified or deleted; migration is additive only.
  • Every migrated page is verified round-trip: written to gbrain, read back, spot-checked.
  • Cross-references from the source system (wikilinks, block refs, tags) are converted to gbrain equivalents.
  • Migration is tested on a sample (5-10 files) before bulk execution.
  • Post-migration health check confirms page count, link integrity, and embedding coverage.

Supported Sources

Source Format Strategy
Obsidian Markdown + [[wikilinks]] Direct import, convert wikilinks to gbrain links
Notion Exported markdown or CSV Parse Notion's export structure
Logseq Markdown with ((block refs)) Convert block refs to page links
Plain markdown Any .md directory Import directory into gbrain directly
CSV Tabular data Map columns to frontmatter fields
JSON Structured data Map keys to page fields
Roam JSON export Convert block structure to pages

Phases

  1. Assess the source. What format? How many files? What structure?
  2. Plan the mapping. How do source fields map to gbrain fields (type, title, tags, compiled_truth, timeline)?
  3. Test with a sample. Import 5-10 files, verify by reading them back from gbrain and exporting.
  4. Bulk import. Import the full directory into gbrain.
  5. Verify. Check gbrain health and statistics, spot-check pages.
  6. Build links. Extract cross-references from content and create typed links in gbrain.

Obsidian Migration

  1. Import the vault directory into gbrain (Obsidian vaults are markdown directories)

  2. Wire the graph with native wikilink support (v0.12.1+):

    gbrain extract links --source db --dry-run | head -20    # preview
    gbrain extract links --source db                         # commit
    

    extract links natively parses [[relative/path]] and [[relative/path|Display Text]] alongside standard [text](page.md) markdown syntax. Ancestor-search resolution handles wiki KBs where authors omit one or more leading ../ prefixes. The .md suffix is inferred automatically for wikilinks.

Obsidian-specific:

  • Tags (#tag) become gbrain tags
  • Frontmatter properties map to gbrain frontmatter
  • Attachments (images, PDFs) are noted but handled separately via file storage

Notion Migration

  1. Export from Notion: Settings > Export > Markdown & CSV
  2. Notion exports nested directories with UUIDs in filenames
  3. Strip UUIDs from filenames for clean slugs
  4. Map Notion's database properties to frontmatter
  5. Import the cleaned directory into gbrain

CSV Migration

For tabular data (e.g., CRM exports, contact lists):

  1. For each row in the CSV, create a page with column values as frontmatter
  2. Use a designated column as the slug (e.g., name)
  3. Use another column as compiled_truth (e.g., notes)
  4. Store each page in gbrain

Verification

After any migration:

  1. Check gbrain statistics to verify page count matches source
  2. Check gbrain health for orphans and missing embeddings
  3. Export pages from gbrain for round-trip verification
  4. Spot-check 5-10 pages by reading them from gbrain
  5. Test search: search gbrain for "someone you know is in the data"

Anti-Patterns

  • Bulk import without sample test. Never import the full dataset before verifying with 5-10 files. The cost of cleaning up hundreds of bad pages is enormous.
  • Destroying source data. Migration is additive. Never modify, move, or delete the source files.
  • Ignoring cross-references. Wikilinks, block refs, and tags from the source system must be converted to gbrain equivalents. Dropping them loses the knowledge graph.
  • Skipping verification. A migration without post-import health check, page count comparison, and spot-check reads is incomplete.

Output Format

MIGRATION REPORT -- [source] -> GBrain
=======================================

Source: [format] ([file count] files, [size])
Mapping: [field mapping summary]

Sample Test (N files):
- Imported: N/N
- Round-trip verified: N/N
- Cross-refs converted: N

Bulk Import:
- Total imported: N
- Skipped (duplicates/errors): N
- Links created: N
- Tags migrated: N

Verification:
- Page count match: [yes/no]
- Health check: [pass/fail]
- Search test: [query] -> [result count] hits

Tools Used

  • Store/update pages in gbrain (put_page)
  • Read pages from gbrain (get_page)
  • Link entities in gbrain (add_link)
  • Tag pages in gbrain (add_tag)
  • Get gbrain statistics (get_stats)
  • Check gbrain health (get_health)
  • Search gbrain (query)