Files
gbrain/docs/guides/chat-connectors.md
Garry Tan 80a899031d v0.46.31.0 feat(connectors): live ChatGPT/Claude chat-history sync into the brain (#4584)
* feat(connectors): live ChatGPT/Claude history sync into the brain

Provider-pluggable chat connectors that fetch a user's own conversation
history via their session credential and ingest it through the existing
runTranscriptsIngest pipeline (redaction, slugging, splitting, idempotency
reused). New src/core/connectors/ module: ChatHistoryProvider seam +
registry (chatgpt/claude), GitHubClient-style HTTP client with off-origin
guard + refresh-on-401, pure classify.ts (Cloudflare-403 vs auth-403),
file-plane credential store (0600), dependency-free OAuth PKCE loopback,
batched spool, and the runConnectorSync orchestrator with a config-scalar
watermark (not op_checkpoint, whose 7-day GC would wipe it) and an
engine-branched embed kickoff (runEmbedCore inline on PGLite, never
submitEmbedBackfill). Surfaces: connectors_status/connector_sync ops
(localOnly), the gbrain connectors CLI family, the connector-sync minion
handler + opt-in autopilot staleness dispatch, and a doctor health check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(connectors): unit + e2e suites against a Bun.serve fixture backend

76 tests: pure classify/config-keys/credentials/pkce/providers units, an
embed-kickoff engine-branch unit, and PGLite e2e suites that run the real
ConnectorClient against a scriptable Bun.serve fixture (keyless-in-CI:
retrieval via getPage + searchKeyword). Covers full pipeline, incremental
watermark incl. the >30-day-gap-survival regression, idempotency/in-place
updates, adapter edge cases, redaction, real HTTP behavior (401-refresh,
403-fingerprint vs 403-auth, drift), dry-run/limit, receipt, spool 0600/
prune, handler lock/terminal-status, and the doctor conditions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(connectors): chat-connectors skill + guide + reference updates

New chat-connectors skill (live account lane; dedup boundaries vs
conversation-archive) with manifest/RESOLVER/lock/plugin registration.
Guide at docs/guides/chat-connectors.md; KEY_FILES + CLAUDE.md reference-map
row; README section; 7 deferred follow-ups filed in TODOS.md; llms bundle
regenerated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v0.46.31.0 feat(connectors): live ChatGPT/Claude chat-history sync

Version bump + CHANGELOG for the chat-connectors feature. Regenerates the
plugin trees + template repo (chat-connectors skill + conversation-archive
dedup edit propagate into the codex/claude plugin trees).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(connectors): post-ship doc-release pass

/document-release audit: add a "bring in your chat history" task to
AGENTS.md's operating protocol (discoverability for non-Claude harnesses),
and disambiguate inbound MCP "Connectors" from outbound `gbrain connectors`
in the README (the word is overloaded on the same page). Regenerate llms.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(connectors): resolve CI failures (verify + test lanes)

- test-isolation: connectors-credentials.test.ts mutates process.env
  (GBRAIN_HOME + env cookies) → rename to *.serial.test.ts (the test is
  process-global-state-bound, so serial is also more correct).
- skill-brain-first: add the brain-first convention callout to the
  chat-connectors skill (a writing skill needs the lookup-before-write
  discipline). Regenerate skills.lock + the plugin tree.
- exit-verdict ownership: route the 9 raw `process.exitCode = 1` writes in
  the connectors command family through setCliExitVerdict().
- doctor-categories drift: categorize the new `connectors` doctor check
  under OPS_CHECK_NAMES.
- remediation text (#3697): teach the flag-registry generator that
  connectors/ is a peeled command dir (facadeExpansion) so the safety flag
  `--dry-run`, consumed in sync.ts, carries consumption evidence into
  depth-zero and registers; split the first-run `config set` / `autopilot
  --install` hint onto separate lines so `--install` isn't misattributed to
  `config`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: add "Say to your agent" end-user phrasebook convention

New CLAUDE.md IRON RULE: every public feature doc carries a common
"**Say to your agent:**" block — the natural-language phrases an END USER
types into their harness, not just the operator CLI. Phrases must come from
(or contain) the backing skill's frontmatter triggers so the doc and the
router never drift; CLI-only features name the command honestly.

Applied throughout README (query, capture, transcripts, connectors, schema,
graph, contradictions, search, minions, FTS, eval, integrations, the loop,
troubleshooting) + the skills phrasebook pointer to RESOLVER.md, the
v0.46.31.0 CHANGELOG "To take advantage" block, and the chat-connectors
guide. Phrases verified against real skill triggers (adversarial two-lens
review). Regenerate llms.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 08:14:26 -07:00

7.2 KiB
Raw Permalink Blame History

Chat Connectors — live sync of ChatGPT + Claude history

Chat connectors sync your own AI-assistant conversation history into the brain using your own session credential, incrementally and (opt-in) on a schedule. They are the LIVE front-end to the export-file lane the conversation-archive skill already documents: fetch replaces the manual download, and everything downstream (redaction, slugging, part-splitting, idempotency) is the exact gbrain transcripts ingest pipeline.

Providers in v1: ChatGPT and Claude (both live). Perplexity has no live connector yet (no transcript adapter) — use the conversation-archive manual conversion for it.

Quick start

Say to your agent: "Connect my chatgpt account and pull my whole history into the brain" — "Connect my claude account" — "Keep my conversations synced automatically." The chat-connectors skill walks the whole flow: cookie capture, dry-run → sample → full backfill, and the opt-in schedule. The commands below are the manual path.

# 1. Connect (cookie paste-in is the primary lane; kept out of argv via stdin)
gbrain connectors auth chatgpt --cookie -      # paste the Cookie header, Ctrl-D
gbrain connectors auth claude  --cookie -      # paste `sessionKey=<value>`, Ctrl-D

# 2. First sync (preview → sample → full)
gbrain connectors sync chatgpt --dry-run
gbrain connectors sync chatgpt --limit 5
gbrain connectors sync chatgpt --full

# 3. (optional) keep it synced automatically — opt-in, daily
gbrain config set connectors.chatgpt.auto_sync true
gbrain autopilot --install

gbrain connectors status shows credential provenance/expiry and sync state (never the secret). gbrain connectors logout <provider> removes a credential.

How it works

  cookie/token (~/.gbrain/connectors/<p>.json, 0600)
        │
        ▼
  ConnectorClient ── list (metadata, newest-first, stop at watermark−7d)
        │                 └─ archived second pass (ChatGPT)
        ▼
  fetch each new conversation ──▶ spool (native-export shape, 0600, batched)
        │                              │
        │                              ▼
        │                    runTranscriptsIngest  (redact → slug → split → import)
        ▼                              │
  watermark (config scalar) ◀──────────┘  advance ONLY on a fully clean run
  connectors.<p>.watermark_iso           receipt → ingest_log; stamp last_sync_at

Incremental sync + gap-heal

Each provider keeps a watermark in the config table (connectors.<provider>.watermark_iso) — the newest conversation update-time imported. Later runs list newest-first and stop at watermark − windowDays (default 7), so:

  • only genuinely new conversations are fetched (detail fetches are the expensive part; the metadata list is cheap and paginated), and
  • a conversation edited just behind the watermark (within the trailing window) is re-listed and re-imported in place — no silent gap.

The watermark advances only on a fully clean run (no fetch errors, no --limit cap, clean ingest). A partial run leaves it untouched so the next run heals. Re-imports are free (content-hash idempotency), so re-running is always safe.

The watermark is deliberately a config scalar, not op_checkpoint: op_checkpoint stores a completed-key set (no scalar timestamp) and GCs rows after 7 days, which would wipe the watermark on any gap longer than a week and trigger a full re-fetch of your entire history — the exact traffic pattern most likely to trip a provider's anti-abuse. The config table is durable and never GC'd.

Automation lanes

Scheduled sync is opt-in per provider and daily by default. It polls your account on a cadence — that is your account making automated requests, so it is off until you enable it.

  • Autopilot (preferred, harness-agnostic): gbrain autopilot --install installs the right OS tick (launchd / systemd / crontab / container start script) and runs the dispatch. It is credential-gated and auto_sync-gated, and a dead cookie stops it (and surfaces in gbrain doctor).
  • Host cron (daemonless): 0 6 * * * gbrain connectors sync --all (daily). Tune the floor with gbrain config set connectors.sync_floor_min <minutes> (default 1440).

On PGLite (the default engine, no worker daemon) sync runs inline; on Postgres, --background submits a connector-sync minion job (single-flight per provider).

Config keys

Key Default Meaning
connectors.source_id default Source the pages land in.
connectors.sync_floor_min 1440 Scheduled-sync cadence floor (minutes).
connectors.embed_kickoff_min_pages 25 Embed backfill after a run importing ≥ this many pages.
connectors.doctor_stale_hours 72 gbrain doctor flags a stalled auto-sync past this.
connectors.<p>.auto_sync off Opt-in scheduled sync for a provider.
connectors.<p>.last_sync_at — Stamped each run (staleness gate).
connectors.<p>.auth_error_at — Stamped on a dead credential.
connectors.<p>.watermark_iso — Incremental watermark.

Env override for a credential (incident escape hatch): GBRAIN_CONNECTOR_<PROVIDER>_COOKIE / _TOKEN.

Security & posture

  • Credentials are session cookies / tokens — password-equivalent. They live file-plane at ~/.gbrain/connectors/<provider>.json (0600, dir 0700), never in the DB, sources.config, the config planes, or any op payload. The only network egress is to the provider's own host.
  • Transcripts are redacted (secret patterns) before any page is written, exactly as the export-file lane does; the spool is 0600 and pruned after ingest.
  • These are ops-facing, local-only operations (the connectors_status / connector_sync ops are localOnly and never expose a credential over MCP).
  • You are syncing your own conversation data with your own account — the same data the provider's official export contains. Keep the cadence polite (daily default) so automated polling doesn't risk your account.

Feasibility caveat (Cloudflare)

chatgpt.com / claude.ai sit behind bot-management that fingerprints the TLS/HTTP2 handshake, and cf_clearance is bound to a real browser. A server-side fetch with a valid cookie may still draw a 403 challenge. When that happens the connector reports forbidden and points you at the official export lane; it never loop-retries. If server-side fetch is reliably blocked in your environment, prefer the export-file lane (conversation-archive) — it always works.

Troubleshooting

Symptom Cause Fix
forbidden Cloudflare/bot challenge on server-side fetch Use the official export + gbrain transcripts ingest
auth_required cookie expired/invalid Re-copy a fresh Cookie header, gbrain connectors auth
partial some fetches failed Watermark not advanced; just re-run
receipt shows drift provider API shape changed Affected threads skipped (not lost); export lane still works

v2 roadmap

Perplexity live client (+ a native perplexity transcript adapter), an advisor collector for connector health, multi-account per provider, attachment/image capture, export-ZIP auto-unwrap, and a nightly spec-target drift probe.