DEV Community

Hiroshi Toyama
Hiroshi Toyama

Posted on

GitHub Issues as the Only Database: Running Coding CLIs Unattended Without a Server

Every headless coding CLI has the same shape: you hand it a prompt, it works, it exits, and it forgets everything. claude -p, codex exec, cursor-agent -p — all one-shot. Which means the moment you want one to carry a feature from "here's a spec" to "merged to main," you have to own the state: which task is next, which branch, which PR, whether CI went green, whether a human approved.

The usual answer is a server. Webhook receiver, a queue, Postgres to hold run state, a reconciler. Then your laptop sleeps mid-PR and you find out your run state and GitHub's reality have diverged, and there is no cheap way to tell which one is lying.

ghswarm makes the opposite bet: GitHub already is that database. Labels are a mutex. The Issue body is a state blob. Checkboxes are a task queue. The PR carries CI status and review decisions. If all the state lives there, the orchestrator can be a stateless CLI that you can kill -9 at any moment.

The state has nowhere else to live

Two HTML comments at the top of the Issue body hold everything:

<!-- GHSWARM_STATE_START
{ "next_action": "wait_ci", "branch_name": "issue-42", "pr_number": 118,
  "ci_fix_retries": 1, "total_agent_runs": 4, "busy_owner": "mbp.local:9182" }
GHSWARM_STATE_END -->

<!-- GHSWARM_VERIFY_START
verify:
  - uv run --extra dev ruff check .
  - uv run --extra dev python -m pytest -q
GHSWARM_VERIFY_END -->
Enter fullscreen mode Exit fullscreen mode

Invisible in the web UI, parsed by regex, rewritten on every transition. The separation matters: state gets overwritten constantly, verify steps must survive every rewrite, so they live in their own block.

Recovery is then not a feature you build — it's the absence of one. Process dies during wait_ci? Next cycle reads the Issue, sees next_action: wait_ci, polls the PR again. Nothing to reconcile, because there was never a second copy.

Labels as a distributed mutex with a lease

Only one coding CLI may touch a repo at a time, and that has to hold across processes and machines. So the lock is a label: status: busy-claude, status: idle, status: blocked, status: completed.

A lock without a lease is a deadlock waiting to happen, so the holder writes busy_owner: "host:pid" and busy_since into the state block, and staleness is decided without side effects:

def is_stale(lock, labels, state, ttl, now, host) -> bool:
    if not lock.startswith(labels.busy_prefix):
        return False
    owner_host, owner_pid = parse_owner(state.busy_owner)
    if owner_host == host and owner_pid and not pid_alive(owner_pid):
        return True          # same machine, process is gone -> definitely dead
    since = parse_ts(state.busy_since)
    if since is None:
        return True
    return (now - since).total_seconds() > ttl
Enter fullscreen mode Exit fullscreen mode

Same host gets the fast path: os.kill(pid, 0) tells you immediately the holder is gone. Another host only gets the TTL (4h default), because you cannot probe its PIDs. Reclaiming posts a comment on the Issue naming the owner and the acquisition time, so a stolen lock is never silent.

The label transition itself has a detail that only shows up after it bites you:

def set_status(gh, issue, labels, status, agent_names) -> None:
    gh.add_label(issue.number, status)
    current = set(gh.get_issue(issue.number).labels)
    for l in labels.all_status(agent_names):
        if l != status and l in current:
            gh.remove_label(issue.number, l)
Enter fullscreen mode Exit fullscreen mode

Add first, then remove. The status label is what marks an Issue as managed, so if a transient API error lands between a remove and an add, the Issue silently drops out of the swarm forever. Adding first means the worst case is two status labels coexisting for a second — which the stale-reclaim path already handles. And the removal list comes from a fresh get_issue, not the snapshot taken at the top of the cycle, because that snapshot predates the busy label this very run added.

No spec, no run

The presence of the GHSWARM_VERIFY block is the start gate. No block means the Issue blocks as spec_missing; malformed YAML blocks as verify_invalid. An agent is never allowed to start on an Issue where nobody wrote down what "done" means.

Task progress is just markdown checkboxes in the body, worked down and flipped to - [x]. One design choice worth calling out — all unchecked tasks go into a single CLI run, not one run per task:

# Pass all unfinished tasks in a single CLI run. Because each CLI invocation is
# a headless one-shot that loses its context, launching one per task would make
# it re-read the same codebase over and over, and later tasks could not carry
# over the intent of earlier ones.
task_list = "\n".join(f"  {i}. {t.text}" for i, t in enumerate(tasks, 1))
Enter fullscreen mode Exit fullscreen mode

This is the whole reason the granularity is what it is. Per-task runs look cleaner on paper and are strictly worse in practice: you pay full codebase re-reading on every task, and task 3 cannot know why task 1 chose a particular abstraction.

The loop runs all the way to merged

implement → (simplify) → ai_review → create_pr → wait_ci → verify_merge → done, with blocked reachable from anywhere. What makes it actually unattended is that each failure mode has a named recovery rather than a stop:

Failure Recovery Cap
Verify fails after a run Re-prompt with the test log attached max_retries (in-run)
resource_exhausted / 429 / 503 Treated as transient, retried across loops transient_max_retries: 5
PR is CONFLICTING Merge base into the worktree, resolve, push, re-run CI conflict_max_retries: 3
PR CI red Pull the logs, hand them to implement, push a fix ci_fix_max_retries: 3
Spec ambiguous Agent writes .agent_question.md and exits; ghswarm comments and blocks —

Every counter lives in the Issue state and is reset on success, so a retry budget survives a crash instead of resetting to zero and letting the same failure burn tokens forever. On top of that, issue_max_agent_runs (default 10) caps cumulative agent runs per Issue — the backstop against an Issue that is subtly impossible and would otherwise spin all night.

Two gates are worth their own note. require_approval is three-valued, not boolean:

  • false — merge on green CI alone
  • true — GitHub's overall review decision must be APPROVED
  • human — additionally requires a non-bot APPROVED review (user.type == "Bot" and *[bot] logins are excluded)

That third value exists because once review bots are in the loop, "approved" stops meaning "a person looked at this."

And merging is not the end. With post_merge_ci on, the Issue stays open until the CI of the merge commit on the base branch is green:

wait_ci --> verify_merge: CI green + approve
verify_merge --> done: post-merge CI green
verify_merge --> blocked: post-merge CI failed
Enter fullscreen mode Exit fullscreen mode

PR CI passes on a merge preview; regressions from interleaved merges show up only on the base branch. Closing the Issue at merge time would mean the swarm declares victory exactly where the interesting failures start.

Multiple CLIs, pinned per phase, with fallback

Routing is explicit — no LLM decides which model runs. You pin a command per phase, and a list is a fallback chain:

agents:
  implement:
    command:
      - "cursor-agent -p {prompt} --model auto"
      - "claude -p {prompt} --model opus --dangerously-skip-permissions"
  review:
    command: "claude -p {prompt} --model sonnet --dangerously-skip-permissions"
Enter fullscreen mode Exit fullscreen mode

{prompt} is shlex.quoted in. Non-zero exit falls through to the next command, which is how a provider outage degrades into "the run took longer" instead of "the swarm stopped." One subtle bit in the fallback path: the question file is deleted before retrying, so a clarification written by the CLI that just failed cannot be misread as coming from the fallback.

Cheap-model-for-review is a real cost lever here: implement on Opus, review on Sonnet, and the review phase — which runs on every Issue — stops dominating the bill.

Isolation: worktree per Issue, optional Docker per directory

Each Issue gets its own git worktree (../<repo>-worktrees/issue-N), so a run never touches your working tree and parallel repos never collide. worktree_setup runs once on creation (uv venv && uv pip install -e '.[dev]' -q).

Verify steps can then run in Docker, and the config declares only where each step runs — never what it runs, since that belongs to the Issue:

verify:
  - path: terraform
    sandbox: { driver: none }
  - path: backend
    sandbox:
      driver: docker
      image: python:3.12
      isolate_dirs: [.venv]
Enter fullscreen mode Exit fullscreen mode

That split is what makes monorepos work: one Issue and one PR can touch both directories while each verifies in its own environment. The whole worktree is always mounted (only the working directory changes per step) so cross-directory imports still resolve, and isolate_dirs gives .venv / node_modules a container-only writable mount so the container's platform-specific build artifacts never land in the host worktree.

Gotchas

macOS sleep is the one that gets people. A locked screen or running screensaver is not sleep — while the display is on, powerd holds a PreventUserIdleSystemSleep assertion and the loop keeps going. Real system sleep does stop it. Nothing is lost (state is on GitHub), but nothing progresses either. Wrap it:

ghswarm loop -d
nohup caffeinate -i -s -w "$(cat ~/.ghswarm/ghswarm.pid)" >/dev/null 2>&1 &
Enter fullscreen mode Exit fullscreen mode

caffeinate ... ghswarm loop -d does not work — -d double-forks, the foreground process exits instantly, and caffeinate drops the assertion with it. Point -w at the daemon PID instead.

The event DB is derived data. ~/.ghswarm/events.db records each step for observation, and it is explicitly not a source of truth — deleting it costs you nothing but history. Don't build state on it. (There's deliberately no metrics command; the intent is raw events plus an AI writing ad-hoc SQL.)

Headless auto-approval is yours to own. --dangerously-skip-permissions and friends are required for unattended runs, and the worktree + Docker isolation is what makes that a considered risk rather than a blind one.

Why this shape holds up

The real payoff isn't "no server to run," though uv tool install ghswarm and a YAML file is a pleasant setup story. It's that there is exactly one copy of the truth, and it's the copy you were already going to look at. When a run goes sideways you open the Issue: the labels tell you who holds the lock, the checkboxes tell you how far it got, the state block tells you what it was about to do, and the comments tell you why it stopped. No log correlation, no separate dashboard, no "the DB says running but the process is gone."

Everything else — lease-based locks, capped retries, the post-merge CI gate — is just what it takes to be honest about a system where every worker is stateless and every dependency is flaky.

uv tool install ghswarm
ghswarm init && $EDITOR ~/.ghswarm.yaml
ghswarm skills install     # Claude Code skills for the spec side
ghswarm loop               # resident, parallel across repos
Enter fullscreen mode Exit fullscreen mode

Top comments (0)