* feat: add `gbrain jobs supervisor` — self-healing worker process manager Adds a first-class supervisor command that: - Spawns `gbrain jobs work` as a child process - Restarts on crash with exponential backoff (1s→60s cap) - Resets crash counter after 5min of stable operation - PID file locking prevents duplicate supervisors - Periodic health checks (stalled jobs, completion gaps) - Graceful shutdown (SIGTERM→35s→SIGKILL) Usage: gbrain jobs supervisor --concurrency 4 Replaces ad-hoc nohup patterns in bootstrap scripts. The autopilot command's internal supervisor can be migrated to use this in a follow-up. Tests: 7 pass (backoff calc, PID management, crash tracking) * supervisor: atomic PID lock, queue-scoped health, env safety, unified exit Lane A of PR #364 review fixes (20-item multi-lane plan). Addresses the codex-tier + CEO + Eng findings on src/core/minions/supervisor.ts: Safety + correctness: - Atomic O_CREAT|O_EXCL PID lock via openSync('wx') with stale-file liveness check. Prevents two supervisors racing on the same PID file. (codex #1) - Health check now queries status='active' AND lock_until < now() matching queue.ts:848's authoritative stalled definition. The prior `status = 'stalled'` predicate returned zero rows forever because 'stalled' is not a persisted value in the schema. (codex #2) - All health queries scoped to WHERE queue = $1 via opts.queue binding. Multi-queue installs no longer see cross-queue false positives. (codex #3) - Class default allowShellJobs flipped true→false AND explicit `delete env.GBRAIN_ALLOW_SHELL_JOBS` when false, so child workers don't silently inherit the var from the parent shell. (eng #8, codex #9) - Unified shutdown(reason, exitCode) — max-crashes now routes through the same drain path as SIGTERM. Single source of truth for lifecycle cleanup; prerequisite for trustworthy audit events (Lane C). (eng #1) - Default PID path moves from /tmp to ~/.gbrain/supervisor.pid with mkdirSync recursive + GBRAIN_SUPERVISOR_PID_FILE env override. Matches the rest of the product's ~/.gbrain/ convention; fresh installs no longer hit ENOENT. (CEO #2 + codex #6) Refinements: - crashCount = 1 after 5-min stable-run reset (was 0, produced calculateBackoffMs(-1) = 500ms by accident). Now reads as 'first crash of a new cycle' with a clean 1s backoff. (Nit 1) - Top-of-file POSTGRES-ONLY docstring documenting why the supervisor can't run against PGLite. (Nit 2) - inBackoff flag suppresses 'worker not alive' warn during the expected null-child window (crash → sleep → next spawn). (eng #2) - Tracked listener refs for SIGTERM/SIGINT removed in shutdown() so integration tests spinning up/tearing down multiple supervisors on one process don't leak handlers. (eng #3) - Single FILTER query replaces two SELECT counts — one round-trip instead of two, three metrics in one pass. (eng #10) - child.on('error') listener emits worker_spawn_failed event for ENOENT/EACCES; exit handler still increments crashCount as usual so max-crashes bounds permanent misconfigurations. (codex #7) - healthInFlight boolean guard with try/finally prevents overlapping health checks from stacking on a hung DB. (codex #8) Documented exit codes (ExitCodes const): 0 CLEAN, 1 MAX_CRASHES, 2 LOCK_HELD, 3 PID_UNWRITABLE Agent can branch on exit=2 ('another supervisor, I'm fine') vs exit=1 ('escalate to human'). Event emitter surface: - started / worker_spawned / worker_exited / worker_spawn_failed - backoff / health_warn / health_error / max_crashes_exceeded - shutting_down / stopped Plumbed through emit() with an onEvent callback hook for Lane C's audit writer. json:false is the default; Lane C's --json mode flips it and writes JSONL to stderr. CLI changes (src/commands/jobs.ts): - `gbrain jobs supervisor` gains --allow-shell-jobs (explicit opt-in mirroring the env-var gate), --cli-path (override auto-resolution for exotic setups), and --json (JSONL lifecycle events on stderr). - Expanded --help body with description, 3 examples, and exit-code table. (DX Fix A per review) - Three-tier PID path resolution: --pid-file > GBRAIN_SUPERVISOR_PID_FILE > ~/.gbrain/supervisor.pid (via exported DEFAULT_PID_FILE). - Removed the catch-fallback to process.argv[1] — resolveGbrainCliPath() throws its own actionable install-hint error, which is what dev users need instead of a cryptic spawn failure on a .ts path. (codex #5) Tests: existing 7 supervisor.test.ts cases continue to pass. Integration tests (crash-restart, max-crashes, SIGTERM-during-backoff, env-inheritance regression) land in Lane E. Out of scope for this lane (tracked in follow-up lanes): - Audit file writer at ~/.gbrain/audit/supervisor-YYYY-Www.jsonl (Lane C) - Documentation pass (Lane B) - supervisor start/status/stop subcommands (Lane C) - gbrain doctor supervisor check (Lane D) - /ship release hygiene (Lane F) - autopilot.ts migration to MinionSupervisor (deferred to follow-up PR per codex — requires non-blocking start() API redesign, not ~30 lines) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: supervisor as canonical worker deployment pattern Lane B of PR #364 review fixes. Reframes docs/guides/minions-deployment.md around `gbrain jobs supervisor` as the default answer (blocker 7), deletes the 68-line legacy bash watchdog (F10), and updates README + deployment snippets to match. docs/guides/minions-deployment.md: - New 'Worker supervision' section at the top with the canonical 3-command agent pattern (start --detach / status --json / stop) and a documented exit-code table (0 clean, 1 max-crashes, 2 lock-held, 3 PID-unwritable). - 'Which supervisor when?' decision table: container = supervisor as PID 1, Linux VM = systemd-over-supervisor, dev laptop = bare terminal. - New 'Agent usage' section for OpenClaw / Hermes / Cursor / Codex — the 3-turn discover-start-maintain workflow that replaces shell archaeology with machine-parseable JSON events + an audit file at ~/.gbrain/audit/supervisor-YYYY-Www.jsonl. - Demoted the 'Option 1: watchdog cron' path entirely; replaced with a straightforward upgrade migration block (stop script, remove cron line, start supervisor, verify via doctor). - Preconditions now check Postgres connectivity directly (supervisor is Postgres-only; the CLI rejects PGLite with a clear error). Snippets: - systemd.service: ExecStart now invokes `gbrain jobs supervisor` instead of raw `gbrain jobs work`. Two-layer supervision (systemd → supervisor → worker) buys automatic restart on reboot plus fast crash recovery. ReadWritePaths expanded to cover $HOME/.gbrain (supervisor PID + audit). - Procfile + fly.toml.partial: same change — platform restarts the container on host events, supervisor restarts the worker on crashes. - minion-watchdog.sh: deleted (git history retains it for anyone in an exotic deployment). Supervisor subsumes every capability it had plus atomic PID locking, structured audit events, queue-scoped health checks, and graceful drain on SIGTERM. README.md: - Added a paragraph under the Minions section pointing `gbrain jobs supervisor` as canonical, noting the --detach / status / stop surface and the audit file path, with a link to the full deployment guide. Kept `gbrain jobs work` documented for direct raw invocation but flagged 'prefer supervisor' for any long-running use. The supervisor `--help` body itself (3 examples + exit-code table in src/commands/jobs.ts) landed with Lane A — this lane finishes the discoverability story by making the supervisor findable via doc grep, README landing, and deployment-guide landing paths. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * supervisor: daemon-manager subcommands + JSONL audit writer Lane C of PR #364 review fixes. Adds the daemon-manager CLI surface so agents can drive `gbrain jobs supervisor` in 3 turns instead of 10, and the audit writer that makes lifecycle events inspectable across process restarts. (Blocker 8, closes DX Fix A/B/C.) New: src/core/minions/handlers/supervisor-audit.ts - writeSupervisorEvent(emission, supervisorPid) appends JSONL to `${GBRAIN_AUDIT_DIR:-~/.gbrain/audit}/supervisor-YYYY-Www.jsonl`. ISO-week rotation via a `computeSupervisorAuditFilename()` helper that mirrors `shell-audit.ts` exactly (year-boundary ISO week math, Thursday anchor, etc). - readSupervisorEvents({sinceMs}) returns parsed events from the current week's file, oldest-first, for Lane D's doctor check. Malformed lines are skipped silently (disk-full truncation is already best-effort at write time). - Reuses `resolveAuditDir()` from shell-audit.ts so the `GBRAIN_AUDIT_DIR` env var override works identically across all gbrain audit trails. src/commands/jobs.ts: supervisor subcommand dispatcher - `gbrain jobs supervisor [start] [--detach] [--json] ...` — default subcommand. Without --detach, runs foreground as before. With --detach, forks a background child (inheriting stderr so the caller can still tail JSONL events), writes a stdout payload: {"event":"started","supervisor_pid":N,"pid_file":"...","detached":true} and exits 0. Stdin/stdout on the detached child are /dev/null so the parent shell isn't held open. - `gbrain jobs supervisor status [--json]` — reads the PID file, checks liveness via `kill -0`, then reads the last 24h from the supervisor audit file to compute crashes_24h / last_start / max_crashes_exceeded. Exits 0 if running, 1 if not. JSON output is machine-parseable; human output is a 5-line ASCII report. - `gbrain jobs supervisor stop [--json]` — reads PID, sends SIGTERM, polls `kill -0` every 250ms for up to 40s (supervisor's own 35s worker-drain + 5s slack). Reports outcome: drained / timeout_40s / pid_file_missing / pid_file_corrupt / process_gone. Exit 0 on clean stop. - `--json` flag is already plumbed through to the supervisor opts from Lane A — this lane adds the onEvent audit-writer callback so every supervisor emission (started, worker_spawned, worker_exited, worker_spawn_failed, backoff, health_warn, health_error, max_crashes_exceeded, shutting_down, stopped) lands in the JSONL file with the supervisor's PID attached. --help body updated: - Three separate usage lines (start / status / stop). - SUBCOMMANDS block with one-line summaries each. - EXIT CODES block (unchanged from Lane A, moved under SUBCOMMANDS). - EXAMPLES block updated with status --json + stop + --detach forms. Tests: existing 127 supervisor + minions tests continue to pass. Integration tests for the new subcommands + audit writer land with Lane E. Follow-up (Lane D): `gbrain doctor` will read readSupervisorEvents() from this module to surface a `supervisor` health check alongside its existing checks (DB connectivity, schema version, queue health). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * doctor: add supervisor health check Lane D of PR #364 review fixes. Closes the observability loop: now that Lane C writes supervisor lifecycle events to `${GBRAIN_AUDIT_DIR:-~/.gbrain/audit}/supervisor-YYYY-Www.jsonl`, `gbrain doctor` surfaces a `supervisor` check alongside its existing health indicators. Implementation (src/commands/doctor.ts, filesystem-only block 3b-bis): - Resolves DEFAULT_PID_FILE via the same three-tier logic as the start path (--pid-file > GBRAIN_SUPERVISOR_PID_FILE > ~/.gbrain/supervisor.pid). - Reads the PID file + `kill -0 <pid>` for liveness. - Calls readSupervisorEvents({sinceMs: 24h}) from the audit module to derive last_start / crashes_24h / max_crashes_exceeded. - Suppresses the check entirely when the user has never invoked the supervisor (no PID file AND no audit events) — avoids noise on installs that don't use the feature. Status thresholds: fail max_crashes_exceeded event seen in last 24h (supervisor gave up; operator needs to restart or triage) warn supervisor not running but audit shows prior use (unexpected stop — likely crash or manual kill) warn running but > 3 crashes in last 24h (supervisor recovering but worker is unstable) ok running + ≤ 3 crashes + no max_crashes event All failure paths emit a paste-ready recovery command. Read/import errors are swallowed (best-effort like the other doctor checks). Tests: all 127 supervisor + minions tests still green; 13 existing doctor tests unaffected. F3 done. All four lanes A/B/C/D are now committed; Lane E (integration tests) and Lane F (/ship v0.20.2) remain. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test: 4 critical integration tests for supervisor lifecycle Lane E of PR #364 review fixes (blocker 10). Fills the ~15% coverage gap flagged in the eng review by actually exercising the code paths that will break in production — crash-restart loop, max-crashes exit, SIGTERM-during-backoff, env-var inheritance — via real spawn() calls against fake shell-script workers. No mocks: real fork, real signals, real env propagation, real audit file writes. test/fixtures/supervisor-runner.ts (new, 55 lines): A standalone bun script that constructs a MinionSupervisor from env vars (SUP_PID_FILE / SUP_CLI_PATH / SUP_MAX_CRASHES / SUP_BACKOFF_FLOOR_MS / SUP_HEALTH_INTERVAL_MS / SUP_ALLOW_SHELL_JOBS / SUP_AUDIT_DIR) and calls start(). Mock engine returns empty rows for executeRaw (health check path still exercised without Postgres). Tests spawn this as a subprocess because MinionSupervisor.start() calls process.exit() on shutdown — can't run it in the test runner's own process. test/supervisor.test.ts (existing; 91 → 300 lines): - Added IntegrationHarness helper: creates a unique tmpdir per test, a fake worker shell script, a PID-file path, and an audit-dir path; cleanup runs in finally. - spawnSupervisor() forks bun on the runner with env vars set. - readAudit() reads the supervisor-YYYY-Www.jsonl file via the existing readSupervisorEvents() helper (Lane C), threading GBRAIN_AUDIT_DIR through so tests don't collide on ~/.gbrain. - waitFor(pred, timeoutMs) polls helper for event-driven tests. Four integration tests (with _backoffFloorMs=5 for <1s suite runs): 1. "respawns the worker after a crash and eventually exits with max-crashes code=1" Worker always `exit 1`. maxCrashes=3. Asserts: exit code 1, PID file cleaned up, audit contains started + 3x worker_spawned + 3x worker_exited + max_crashes_exceeded + shutting_down + stopped, and the stopped event carries {reason:'max_crashes', exit_code:1}. Locks in blockers 1 (PID lock), 2+3+6 (health SQL doesn't 500), 5 (unified shutdown emits right events), F8 (spawn errors counted). 2. "receives SIGTERM while sleeping between crashes and exits 0 cleanly" Worker always `exit 1`, backoff floor 800ms to catch the sleep. Asserts: SIGTERM during backoff → exit code 0 (not 1) in <5s, no signal kill (process.exit via shutdown), audit contains shutting_down {reason:'SIGTERM'} + stopped, PID file cleaned up. Locks in eng Issue 1 (unified exit path), eng Issue 3 (signal handlers don't accumulate across shutdowns). 3. "strips inherited GBRAIN_ALLOW_SHELL_JOBS when allowShellJobs=false, even if parent has it set" ⚠ CRITICAL regression test Parent env has GBRAIN_ALLOW_SHELL_JOBS=1. SUP_ALLOW_SHELL_JOBS=0. Worker writes $GBRAIN_ALLOW_SHELL_JOBS (or 'UNSET' if absent) to an OUT_FILE. Asserts child sees 'UNSET'. Locks in codex #9 + eng #8: the `else delete env.GBRAIN_ALLOW_SHELL_JOBS` branch from Lane A is load-bearing for the supervisor's security posture; this test prevents a future refactor silently re-opening the inheritance hole. 4. "DOES pass GBRAIN_ALLOW_SHELL_JOBS to child when allowShellJobs=true" Positive-path companion to #3. SUP_ALLOW_SHELL_JOBS=1 → worker sees '1'. Confirms the else-branch doesn't over-strip and that operators who explicitly opt in still get shell-exec enabled. Plus two audit-format unit tests: - computeSupervisorAuditFilename format (regex match) - Year-boundary ISO week: 2027-01-01 → supervisor-2026-W53.jsonl (matches the shell-audit.ts pattern exactly) Before: 7 tests covering backoff math + PID helpers (~15% behavioral coverage per eng review). After: 13 tests across all critical lifecycle paths (crash-restart, max-crashes, SIGTERM, env-inheritance, audit rotation). All 146 tests in supervisor + minions + doctor suites green in ~8s. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: bump version and changelog (v0.20.2) Lane F of PR #364 review fixes. Closes the multi-lane plan with release hygiene: VERSION bump 0.19.0 → 0.20.2, package.json sync, CHANGELOG entry in GStack voice with release summary + "numbers that matter" table + "To take advantage of v0.20.2" migration block + itemized changes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: escape template-literal interpolation in supervisor --help The --help body in src/commands/jobs.ts is one big backtick template literal. The supervisor subcommand description I added in Lane B used both `${GBRAIN_AUDIT_DIR:-~/.gbrain/audit}` (parsed as a template interpolation into an undefined variable) and inline `code` backticks (parsed as nested template literals). CI caught it with ~200 tsc parse errors across the file. Fix: - Escape `${...}` → `\${...}` so the audit-file path renders literally. - Replace prose inline-code backticks with plain single-quote fences (`gbrain jobs work` → 'gbrain jobs work', `~/.gbrain/supervisor.pid` → ~/.gbrain/supervisor.pid). `--help` output is human prose; the single-quote form reads cleanly in a terminal without needing to smuggle nested backticks through a template literal. `bunx tsc --noEmit` is clean. 146 tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: regenerate llms-full.txt after Lane B doc rewrite CI drift guard caught that `llms-full.txt` didn't match the current generator output. Root cause: the Lane B rewrite of `docs/guides/minions-deployment.md` (supervisor as canonical, watchdog deleted) changed content that gets inlined into `llms-full.txt`, but I didn't run `bun run build:llms` to regenerate. `bun test test/build-llms.test.ts` now clean (7/7 pass). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: root <root@localhost> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
333 lines
13 KiB
Markdown
333 lines
13 KiB
Markdown
# Minions Worker Deployment Guide
|
|
|
|
Keep `gbrain jobs work` running across crashes, reboots, and Postgres
|
|
connection blips. Written for agents to execute line-by-line.
|
|
|
|
## The problem
|
|
|
|
The persistent worker can die silently from:
|
|
|
|
- Database connection drops (Supabase/Postgres maintenance or network blips).
|
|
- Lock-renewal failures → the stall detector eventually dead-letters jobs.
|
|
- Bun process crashes with no automatic restart.
|
|
- Internal event-loop death (PID alive, worker loop stopped).
|
|
|
|
When the worker dies, submitted jobs sit in `waiting` forever. The
|
|
canonical answer is `gbrain jobs supervisor` — a first-class CLI that
|
|
spawns `gbrain jobs work` as a child and auto-restarts it on crash.
|
|
|
|
## Worker supervision
|
|
|
|
### The canonical pattern
|
|
|
|
`gbrain jobs supervisor` is an auto-restarting wrapper around
|
|
`gbrain jobs work`. It writes a PID file, restarts the worker on crash
|
|
with exponential backoff (1s → 60s cap), emits lifecycle events to an
|
|
audit file, and drains gracefully on SIGTERM (35s worker-drain window
|
|
before SIGKILL). Exit codes are documented so agents can branch on them.
|
|
|
|
**Typical commands:**
|
|
|
|
```bash
|
|
# Start in the foreground (blocks; Ctrl-C to stop).
|
|
gbrain jobs supervisor --concurrency 4
|
|
|
|
# Start detached — returns {"event":"started","supervisor_pid":…} on stdout.
|
|
gbrain jobs supervisor start --detach --json
|
|
|
|
# Check liveness without reading log files.
|
|
gbrain jobs supervisor status --json
|
|
|
|
# Graceful stop (SIGTERM + drain wait + SIGKILL fallback).
|
|
gbrain jobs supervisor stop
|
|
```
|
|
|
|
**Exit codes:**
|
|
|
|
| Code | Meaning |
|
|
|---|---|
|
|
| 0 | Clean shutdown (SIGTERM/SIGINT received, worker drained) |
|
|
| 1 | Max crashes exceeded (worker kept dying) |
|
|
| 2 | Another supervisor holds the PID lock |
|
|
| 3 | PID file unwritable (permission / path error) |
|
|
|
|
An agent seeing exit=2 can safely treat it as "one is already running";
|
|
exit=1 should page a human.
|
|
|
|
### Which supervisor when?
|
|
|
|
The supervisor solves in-process crash recovery. Platform-level
|
|
supervision (systemd, Fly, Render) handles host-level failures. You
|
|
usually want both.
|
|
|
|
| Environment | Recommendation |
|
|
|---|---|
|
|
| **Container (Fly / Railway / Render / Heroku)** | `gbrain jobs supervisor` runs as PID 1. The platform restarts the container on OOM / host loss; supervisor restarts the worker on crash. See [Fly.io](#flyio) / [Render / Railway / Heroku](#render--railway--heroku). |
|
|
| **Linux VM with systemd** | Two-layer recommended: systemd supervises `gbrain jobs supervisor`, which in turn supervises `gbrain jobs work`. Buys you automatic restart on reboot (systemd) plus fast crash recovery (supervisor). See [systemd](#systemd). |
|
|
| **Dev laptop / macOS** | `gbrain jobs supervisor` in a terminal. Ctrl-C stops it. No system-level setup needed. |
|
|
|
|
### Variables used in this guide
|
|
|
|
Substitute these once before copy-pasting any snippet.
|
|
|
|
| Variable | Meaning | Typical value |
|
|
|---|---|---|
|
|
| `$GBRAIN_BIN` | Absolute path to the `gbrain` binary | `$(command -v gbrain)` — often `/usr/local/bin/gbrain` or `~/.bun/bin/gbrain` |
|
|
| `$GBRAIN_WORKER_USER` | OS user that owns the worker process | the same user that ran `gbrain init`; never `root` |
|
|
| `$GBRAIN_WORKSPACE` | `cwd` for shell jobs submitted by this deployment | absolute path, e.g. `/srv/my-brain` |
|
|
| `$GBRAIN_ENV_FILE` | Secrets file sourced by systemd / shell | `/etc/gbrain.env` (mode 600) |
|
|
|
|
### Preconditions
|
|
|
|
Run these before any deployment step.
|
|
|
|
```bash
|
|
# 1. gbrain is on PATH and resolves to an absolute location.
|
|
command -v gbrain || { echo "gbrain not on PATH. Install, then retry."; exit 1; }
|
|
|
|
# 2. DATABASE_URL points at reachable Postgres.
|
|
# (Supervisor is Postgres-only. PGLite's exclusive file lock blocks the
|
|
# separate worker process. If `config.engine === 'pglite'` the CLI rejects
|
|
# with a clear error.)
|
|
gbrain doctor --fast --json | jq '.checks[] | select(.name=="db_connectivity")'
|
|
|
|
# 3. Schema is up to date. If version=0 or status=="fail":
|
|
# gbrain apply-migrations --yes
|
|
gbrain doctor --fast --json | jq '.checks[] | select(.name=="schema_version")'
|
|
|
|
# 4. If you plan to submit `shell` jobs, pass --allow-shell-jobs to the
|
|
# supervisor (or export GBRAIN_ALLOW_SHELL_JOBS=1 before starting).
|
|
# Without the flag, the shell handler is disabled at worker startup.
|
|
```
|
|
|
|
## Agent usage (OpenClaw / Hermes / Cursor / Codex)
|
|
|
|
Three-command pattern an agent can drive without shell archaeology:
|
|
|
|
```bash
|
|
# Start (returns PIDs + pid_file on stdout as JSON, then detaches)
|
|
gbrain jobs supervisor start --detach --json
|
|
# → {"event":"started","supervisor_pid":1234,"worker_pid":1235,"pid_file":"/Users/you/.gbrain/supervisor.pid"}
|
|
|
|
# Check health (machine-parseable JSON, no log scraping)
|
|
gbrain jobs supervisor status --json
|
|
# → {"running":true,"supervisor_pid":1234,"last_start":"2026-04-23T15:30:22Z","crashes_24h":0, ...}
|
|
|
|
# Stop cleanly (SIGTERM + 35s drain + SIGKILL fallback)
|
|
gbrain jobs supervisor stop
|
|
```
|
|
|
|
Every lifecycle event (spawn, crash, backoff, health warning, max-crashes,
|
|
shutdown) is also written to `${GBRAIN_AUDIT_DIR:-~/.gbrain/audit}/supervisor-YYYY-Www.jsonl`
|
|
for historical inspection. `gbrain doctor` reads that file and surfaces
|
|
a `supervisor` check in its health report.
|
|
|
|
## Deployment: systemd
|
|
|
|
For long-running Linux VMs with shell access.
|
|
|
|
```bash
|
|
# Create the worker user if it doesn't exist.
|
|
sudo useradd --system --home "$GBRAIN_WORKSPACE" --shell /usr/sbin/nologin gbrain \
|
|
2>/dev/null || true
|
|
sudo mkdir -p "$GBRAIN_WORKSPACE" && sudo chown gbrain:gbrain "$GBRAIN_WORKSPACE"
|
|
|
|
# Install the env file (secrets stay out of the unit file).
|
|
sudo install -m 600 -o gbrain -g gbrain \
|
|
docs/guides/minions-deployment-snippets/gbrain.env.example /etc/gbrain.env
|
|
sudoedit /etc/gbrain.env
|
|
# Fill in DATABASE_URL, optional GBRAIN_ALLOW_SHELL_JOBS=1.
|
|
|
|
# Install the unit file, substituting /srv/gbrain → your workspace path.
|
|
sudo install -m 644 docs/guides/minions-deployment-snippets/systemd.service \
|
|
/etc/systemd/system/gbrain-worker.service
|
|
sudo sed -i "s|/srv/gbrain|$GBRAIN_WORKSPACE|g" \
|
|
/etc/systemd/system/gbrain-worker.service
|
|
|
|
sudo systemctl daemon-reload
|
|
sudo systemctl enable --now gbrain-worker
|
|
sudo systemctl status gbrain-worker
|
|
journalctl -u gbrain-worker -n 50
|
|
```
|
|
|
|
The shipped unit file invokes `gbrain jobs supervisor` (not `gbrain jobs work`
|
|
directly) so you get two-layer supervision: systemd restarts the supervisor
|
|
on host reboot, supervisor restarts the worker on in-process crash.
|
|
|
|
`Restart=always` + `RestartSec=10s` handle the supervisor-level recovery.
|
|
The unit runs as unprivileged `gbrain` with `PrivateTmp`, `ProtectSystem=strict`,
|
|
and `ReadWritePaths=$GBRAIN_WORKSPACE,$HOME/.gbrain` (for the PID file and
|
|
audit log). `LimitNOFILE=65535` covers Bun + Postgres pool + concurrent
|
|
LLM subagent calls without hitting the default 1024 cap.
|
|
|
|
## Deployment: Fly.io
|
|
|
|
```bash
|
|
# Merge the [processes] block from fly.toml.partial into your fly.toml.
|
|
cat docs/guides/minions-deployment-snippets/fly.toml.partial >> fly.toml
|
|
# Review + edit as needed.
|
|
|
|
# Set secrets (Fly handles restart on crash).
|
|
fly secrets set DATABASE_URL='postgres://…' GBRAIN_ALLOW_SHELL_JOBS=1
|
|
```
|
|
|
|
The `[processes]` block runs `gbrain jobs supervisor` as PID 1. Fly
|
|
restarts the container on host failure; the supervisor restarts the
|
|
worker on in-process crash.
|
|
|
|
## Deployment: Render / Railway / Heroku
|
|
|
|
Drop [`Procfile`](./minions-deployment-snippets/Procfile) at the repo
|
|
root. The shipped Procfile calls `gbrain jobs supervisor`. Set
|
|
`DATABASE_URL` + optional `GBRAIN_ALLOW_SHELL_JOBS=1` via the platform's
|
|
env UI or CLI.
|
|
|
|
## Deployment: inline `--follow` (no persistent worker)
|
|
|
|
For short deterministic scripts on a fixed schedule where you don't need
|
|
a persistent worker between runs. Each cron run brings its own temporary
|
|
worker. `--follow` starts one on the queue and blocks until the
|
|
just-submitted job reaches a terminal state (`completed` / `failed` /
|
|
`dead` / `cancelled`). 2-3 s startup overhead per job; negligible vs job
|
|
duration for scheduled work.
|
|
|
|
```bash
|
|
GBRAIN_ALLOW_SHELL_JOBS=1 gbrain jobs submit shell \
|
|
--queue nightly-enrich \
|
|
--params "{\"cmd\":\"$GBRAIN_BIN embed --stale\",\"cwd\":\"$GBRAIN_WORKSPACE\"}" \
|
|
--follow \
|
|
--timeout-ms 600000
|
|
```
|
|
|
|
Replace `gbrain embed --stale` with whichever gbrain subcommand you're
|
|
scheduling (`sync`, `extract`, `orphans`, `doctor`, `check-backlinks`,
|
|
`lint`, `autopilot`). For strict single-job semantics on shared queues,
|
|
use a dedicated queue name like `nightly-enrich` above.
|
|
|
|
## Upgrading from an older deployment
|
|
|
|
### From `minion-watchdog.sh` (pre-v0.20)
|
|
|
|
Earlier versions of this guide shipped a 68-line bash watchdog
|
|
(`minion-watchdog.sh`). It's been replaced by `gbrain jobs supervisor`
|
|
which handles everything the script did, plus atomic PID locking,
|
|
structured audit events, queue-scoped health checks, and graceful
|
|
drain on SIGTERM.
|
|
|
|
**Migration:**
|
|
|
|
```bash
|
|
# 1. Stop and remove the old watchdog.
|
|
sudo kill $(head -n1 /tmp/gbrain-worker.pid) 2>/dev/null
|
|
sudo rm -f /usr/local/bin/minion-watchdog.sh /tmp/gbrain-worker.pid \
|
|
/tmp/gbrain-worker.log
|
|
crontab -e # delete the "*/5 * * * * /usr/local/bin/minion-watchdog.sh" line
|
|
|
|
# 2. Start the supervisor (systemd users: reinstall the unit from
|
|
# docs/guides/minions-deployment-snippets/systemd.service, which
|
|
# now calls `gbrain jobs supervisor`).
|
|
gbrain jobs supervisor start --detach --json
|
|
# Or: sudo systemctl restart gbrain-worker
|
|
|
|
# 3. Verify.
|
|
gbrain jobs supervisor status --json
|
|
gbrain doctor # 'supervisor' check should report running=true
|
|
```
|
|
|
|
### Schema / migration hygiene
|
|
|
|
Regardless of which deployment path you're upgrading from:
|
|
|
|
1. **Stop the worker before upgrading.** `gbrain jobs supervisor stop`
|
|
(or `sudo systemctl stop gbrain-worker`). Skipping this risks an
|
|
in-flight job landing partial schema.
|
|
2. **Run `gbrain upgrade`**. Then `gbrain apply-migrations --yes` if
|
|
`gbrain doctor` reports any migration as `partial` or `pending`.
|
|
3. **If you run shell jobs:** from v0.14 onward, pass
|
|
`--allow-shell-jobs` to the supervisor (or keep
|
|
`GBRAIN_ALLOW_SHELL_JOBS=1` in `/etc/gbrain.env`). Submitters don't
|
|
need the flag; only the worker does.
|
|
4. **Verify.** `gbrain doctor` should report zero `pending` or `partial`
|
|
migrations plus a healthy `supervisor` check. `gbrain jobs stats`
|
|
should show no unexplained growth in `dead` between pre- and
|
|
post-upgrade.
|
|
|
|
## Known issues
|
|
|
|
### Supabase connection drops
|
|
|
|
The worker uses a single Postgres connection. If Supabase drops it
|
|
(maintenance, connection limits, network blip), lock renewal fails
|
|
silently. The stall detector then dead-letters the job after
|
|
`max_stalled` misses.
|
|
|
|
**Current defaults that make this worse:**
|
|
|
|
- `lockDuration: 30000` (30 s) — too short for long jobs during
|
|
connection blips.
|
|
- `max_stalled: 5` (schema column default — see `src/schema.sql` and
|
|
`src/core/pglite-schema.ts`). Five missed heartbeats before dead-letter.
|
|
- `stalledInterval: 30000` (30 s) — checks too aggressively.
|
|
|
|
**Tune per-job today.** `gbrain jobs submit` accepts `--max-stalled N`,
|
|
`--backoff-type fixed|exponential`, `--backoff-delay <ms>`,
|
|
`--backoff-jitter 0..1`, and `--timeout-ms N` as first-class flags
|
|
(since v0.13.1). These write onto the job row at submit time — which is
|
|
what `handleStalled()` reads — so per-job tuning is the real knob today.
|
|
|
|
### DO NOT pass `maxStalledCount` to `MinionWorker`
|
|
|
|
It's a no-op. The stall detector reads the row's `max_stalled` column
|
|
(set at submit time), not the worker opt in `src/core/minions/worker.ts:74`.
|
|
Use `gbrain jobs submit --max-stalled N` per-job instead.
|
|
|
|
### Zombie shell children
|
|
|
|
When the Bun worker crashes hard, child processes from shell jobs can
|
|
become zombies. The supervisor's SIGTERM → 35s drain → SIGKILL window
|
|
covers the shell handler's 5 s child-kill grace (`KILL_GRACE_MS`). For
|
|
long-running shell jobs, prefer timeouts via `--timeout-ms` on submit
|
|
over relying on hard kills.
|
|
|
|
## Smoke test
|
|
|
|
```bash
|
|
# Supervisor alive?
|
|
gbrain jobs supervisor status --json | jq .running
|
|
|
|
# Aggregate queue health.
|
|
gbrain jobs stats
|
|
|
|
# Jobs currently stalled (still `active` with expired lock_until, pre-requeue).
|
|
gbrain jobs list --status active --limit 10
|
|
|
|
# Dead-lettered jobs.
|
|
gbrain jobs list --status dead --limit 10
|
|
|
|
# Shell handler registered? (check supervisor audit log or worker stderr.)
|
|
gbrain jobs supervisor status --json | jq '.worker_config.allow_shell_jobs'
|
|
```
|
|
|
|
## Uninstall
|
|
|
|
**`gbrain jobs supervisor`** (foreground or `--detach`):
|
|
|
|
```bash
|
|
gbrain jobs supervisor stop
|
|
```
|
|
|
|
**systemd:**
|
|
|
|
```bash
|
|
sudo systemctl disable --now gbrain-worker
|
|
sudo rm /etc/systemd/system/gbrain-worker.service /etc/gbrain.env
|
|
sudo systemctl daemon-reload
|
|
```
|
|
|
|
**Fly / Render / Railway:** delete the `worker` process from `fly.toml`
|
|
/ `Procfile` and redeploy. Secrets set via `fly secrets` persist until
|
|
`fly secrets unset`.
|
|
|
|
**Inline `--follow`:** remove the cron entry. Nothing else to clean up
|
|
— temporary workers exit with their jobs.
|