DEV Community

Cover image for Why AI Coding Agents Need Work Attempts, Leases, and Checkpoints
Egor Komarov
Egor Komarov

Posted on

Why AI Coding Agents Need Work Attempts, Leases, and Checkpoints

We gave AI agents the ability to write code, run tests, open pull requests, and hand work to each other. Then we asked them to coordinate through a TODO.md file and a shared chat window.

That mismatch produces a whole category of bugs that have nothing to do with the model's intelligence.

An agent picks up a task, gets halfway through, and its context window fills up. Another agent, running in a different terminal, grabs the same task because nothing told it the work was taken. A third crashes mid-edit and leaves a task marked "in progress" forever — a tombstone no one can clear.

None of these are reasoning failures. They're coordination failures, and they appear the moment you run more than one agent, or one agent for longer than a single conversation.

This article is about a data model that addresses them. It's drawn from rhizome-mcp, an open-source MCP server I built for this problem, but the ideas are portable. Work attempts, leases, and checkpoints are useful whether or not you ever install it.


Agents are unreliable workers. Model them that way.

Distributed systems settled this decades ago: never assume a worker will finish what it starts. Workers die. Networks partition. Processes get OOM-killed. The queue has to stay correct anyway.

AI coding agents are unreliable workers with three recurring failure modes:

  1. They crash. The process is killed, the terminal closes, the laptop sleeps, the API call times out. Whatever the agent was holding is abandoned.
  2. They forget. Context windows are finite. A long task eventually pushes its own starting instructions out of the window, and the agent loses the thread mid-flight — a failure mode that is unusually frequent and intrinsic to context-window-based agents.
  3. They collide. Two agents — or two runs of one agent, or agents from two different products — work the same repository and pick the same next task, because nothing shared tells them what is already in flight.

The classic tracker model (Jira, Linear, GitHub Issues) was built for human teams, and it encodes two assumptions that quietly break here.

Permanent assignee. You assign a ticket to a person and it stays theirs until reassigned. An agent has no permanent identity — it's a process that exists for one connection and is then gone. "Assigned to agent-7" means nothing an hour later, when agent-7 no longer exists.

In Progress as a stored state. A human moves a card to "In Progress" and moves it back when they're done. An agent that dies in "In Progress" never moves it back. The status becomes a lie the moment the process exits.

Both assumptions share a root flaw: they treat possession of work as durable state that the worker is trusted to clean up. Unreliable workers don't clean up.

So don't store possession as a fact — derive it from a lease that expires.


Work attempts: possession as a time-boxed lease

An issue is never "assigned." An agent claims it, and the server returns a work attempt: one leased execution of that issue by one session.

// claim_issue -> response (trimmed)
{
  "attempt": {
    "id": "01KY7M1XVRE9SJ5H9JTNRRG5XG",
    "issue_id": "01KY7M1QMCGG3TEJ4WRSH3QVQN",
    "kind": "work",
    "status": "active",
    "lease_expires_at": "2026-07-23T14:08:47Z"
  },
  "lease_token": "OwM1IEE3YD8XUqBwKKl6_D_IJxi9KKGq9f10a0NfxtU",
  "lease_expires_at": "2026-07-23T14:08:47Z"
}
Enter fullscreen mode Exit fullscreen mode

Three properties separate this from an assignment.

The lease expires. The attempt is valid only until lease_expires_at. To keep working, the agent calls renew_attempt — a heartbeat. Stop heartbeating, because you crashed or hung or were killed, and the lease lapses. No cleanup is required from the dead worker. The server can reclaim the attempt from elapsed time alone.

The token is a capability, not an identity. The raw lease_token is returned exactly once, and only its hash is stored server-side. Every mutating call on that attempt must present attempt_id + lease_token. Whoever holds the token owns the attempt — no login, no registration. A brand-new session can resume an attempt if it kept the token, which is exactly what handoffs need, and the server never has to establish who the caller is.

Possession is enforced by the database, not by convention. A partial unique index allows at most one active attempt row per issue:

CREATE UNIQUE INDEX idx_one_active_attempt_per_issue
ON work_attempts(issue_id)
WHERE status = 'active';
Enter fullscreen mode Exit fullscreen mode

Two agents racing to claim the same issue: one INSERT succeeds, the other hits a constraint violation and is told the issue is taken. The invariant doesn't depend on an application-level check with a TOCTOU gap — the table itself cannot hold two active attempts for one issue.

One subtlety matters here. A lease can lapse in wall-clock time while its row still reads status = 'active'. Before inserting a new attempt, the claim transaction materializes any lapsed active attempt as expired. The partial unique index then protects the remaining race window: after expiry processing, only one concurrent insert can succeed.

The attempt lifecycle

[*]      -- claim_issue ------------------------> [active]
[active] -- renew_attempt (heartbeat) ----------> [active]
[active] -- finish_attempt (done / review) -----> [completed]   --> [*]
[active] -- finish_attempt (failed) ------------> [failed]      --> [*]
[active] -- finish_attempt (handoff) -----------> [interrupted] --> [*]
[active] -- lease lapses (no heartbeat) --------> [expired]     --> [*]
Enter fullscreen mode Exit fullscreen mode

expired is the only terminal state reached without the agent doing anything. That's the point.


The key insight: in_progress is not a status

This is the decision everything else hangs on:

in_progress is not a stored issue status. It is computed.

An issue's stored status is one of open, ready, blocked, review, done, cancelled. There is no in_progress value in that column. The server derives an effective status at read time:

Issue stored status: ready
        |
        +-- no active leased attempt
        |       -> effective status: ready
        |
        +-- active leased attempt
                -> effective status: in_progress
                         |
                         +-- lease expires
                                 -> attempt: expired
                                 -> effective status: ready
Enter fullscreen mode Exit fullscreen mode

Read that bottom branch again, because it's the whole argument. When the lease expires, nothing has to reset the issue. The stored status was ready the entire time; only the derived view changed, and it changes back on its own.

Under this model, a crashed worker cannot leave a durable in_progress status behind. The stuck-ticket problem is eliminated by the data model rather than repaired by a separate status-reset workflow.

Compare with the stored-status approach, where a crash leaves a real in_progress value that something must notice and reset. That "something" is where the bugs live: the cleanup never runs, or runs too aggressively, or races with a live worker.


Lazy expiry and background cleanup

Leases lapse in wall-clock time, but a database row doesn't update itself. Expiry is materialized two ways:

  1. Lazy expiry. Operations that care about claimability — claim_issue, list_issues, get_work_context — evaluate lease validity as they run and treat a lapsed lease as expired on the spot. The paths that matter are correct without waiting for a cleanup cycle.
  2. A background sweep. A periodic loop finds active attempts past their lease, marks them expired, appends an event, and releases the uniqueness constraint — preserving the attempt's notes and results. Nothing is deleted; the failed attempt stays queryable as execution history.

Expiry never rewrites the issue's stored status. It removes the active lock, nothing more. Whether the issue becomes workable again is decided by its own stored state and blockers, exactly as if the attempt had never happened.

An expired attempt cannot be resumed — its lease is dead. A fresh claim starts a new attempt, and thanks to checkpoints, it doesn't start from zero.


Checkpoints: so "resume" doesn't mean "restart"

Recovering the slot is half the problem. The other half is recovering the work.

An agent that spent forty minutes and half its context window exploring a codebase has built expensive state. If the next attempt rediscovers all of it, you pay for the same work twice.

During an attempt, an agent writes attempt notes in four kinds:

  • progress — an ordinary status update
  • finding — a notable technical discovery
  • warning — a risk or problem worth flagging
  • checkpoint — a restartable summary

The checkpoint is the one that matters. It's the agent writing a brief for its own successor.

// save_attempt_note
{
  "kind": "checkpoint",
  "content": "Repository layer is implemented and tested. Claim transaction is stubbed.",
  "next_steps": [
    "Implement the claim transaction with BEGIN IMMEDIATE",
    "Add the one-active-attempt concurrency test"
  ],
  "important": true
}
Enter fullscreen mode Exit fullscreen mode

When any agent later requests context for that issue, the server assembles a bounded package that prefers the most recent checkpoint over the full note history, plus the previous attempt's result summary and next steps. A new session — possibly a different agent product — starts from a tight brief instead of a cold read.

This also makes deliberate handoffs first-class. An agent that finishes with outcome: interrupted and interruption_reason_code: context_limit isn't failing. It's tapping out cleanly and leaving the baton where the next runner can find it.


What this looks like end to end

14:00 — Agent A claims ISSUE-42.
14:05 — Agent A saves a checkpoint after completing repository analysis.
14:08 — Agent A crashes and stops renewing the lease.
14:13 — The lease expires.
14:14 — Agent B claims ISSUE-42.
14:14 — Agent B receives Agent A's latest checkpoint and next steps.
14:30 — Agent B completes the issue.
Enter fullscreen mode Exit fullscreen mode

The issue was never manually reassigned and never stored a stale in_progress status. The first attempt became historical execution data, while the second attempt continued from its last useful checkpoint.


Token-efficient context

Traditional APIs can afford to return generous JSON responses. Agent-facing APIs often cannot. Every unnecessary field consumes tokens, costs money, and competes with the model's working context. A sufficiently bloated response can contribute directly to the same failure mode the tracker is trying to prevent: the agent loses the thread.

So the tool contract is built to return the least data that is still sufficient, with budgets that are enforced and tested rather than aspirational.

  • Compact list projections. list_issues returns identifiers, title, classification, and computed status fields — but not the free-text description and acceptance_criteria bodies. In the project's own test fixture, a 100-issue page measures ≈46 KB compact versus ≈582 KB full: a ~12× difference over the same issues, driven entirely by whether unbounded bodies are included. An integration test asserts the compact response stays under 64 KB regardless of body size.
  • Graphs exclude free text at the SQL layer. Graph queries select NULL AS description, NULL AS acceptance_criteria, dropping the unbounded fields before traversal. Response size scales with a max_nodes bound you control, not with how verbose your issue bodies are.
  • Search returns snippets, never documents. Full-text search (SQLite FTS5) returns a bounded snippet per hit, truncated in SQL. You get a pointer and a preview, then fetch the one thing you need.
  • Delta sync instead of re-reading state. Every mutation appends to an event log with a monotonic ID. get_changes(since_event_id) returns only what happened since you last looked.
  • One bounded context package. get_work_context returns a single curated bundle: the issue body, unresolved blockers, active decision summaries, the latest checkpoint, the previous attempt's next steps.

The lesson generalizes past this one server. When your API consumer is an LLM, response shape is part of your performance budget. Omit bodies by default, make expensive fields opt-in, ship a delta feed.


Why SQLite fits this workload

A local SQLite database may sound modest compared with a networked database, but it matches this workload unusually well.

The workload is one developer's fleet of agents on one machine: no multi-node requirement, no cross-region replication, no thousands of concurrent writers. What it does demand is atomic, crash-safe multi-table writes. Claiming an issue must, in one indivisible step, create the attempt, record the issue version and event position it started from, and append the attempt_started event. A partial write there would corrupt the coordination model.

SQLite serializes writers and provides the transactional guarantees this coordination model needs, without requiring a separate database service.

PRAGMA journal_mode = WAL;   -- readers don't block the writer
PRAGMA busy_timeout = 5000;  -- wait for the lock instead of erroring
PRAGMA foreign_keys = ON;
PRAGMA synchronous = NORMAL; -- strong process-crash safety with a good speed/durability trade-off
Enter fullscreen mode Exit fullscreen mode
  • WAL mode lets readers proceed while a write is in flight, so list_issues and get_work_context reads never block the agent committing a claim.
  • Serialized writes are a feature here: SQLite gives you exactly the guarantee the one-active-attempt invariant needs. Write transactions stay short — no model calls, no filesystem scans, no graph building inside a transaction — and BEGIN IMMEDIATE takes the writer lock up front for known-write operations, avoiding mid-transaction upgrade deadlocks.
  • Bounded retry handles genuine contention: on SQLITE_BUSY/SQLITE_LOCKED, retry the whole transaction with backoff (≈25 ms, 75 ms, 200 ms). Domain and constraint failures are never retried.

NORMAL is appropriate for this local developer-tool workload and protects well against application crashes. Deployments that prioritize the strongest available power-loss guarantees can choose FULL instead.

The pure-Go driver (modernc.org/sqlite, no CGO) keeps the whole thing a single static binary with no system dependencies.


Optimistic concurrency and idempotency

Physical write safety isn't enough; concurrent agents also produce logical races. Two mechanisms cover them.

Optimistic concurrency. Meaningful issue mutations require an expected_version. If the issue changed since you read it, the write fails with VERSION_CONFLICT instead of silently clobbering someone else's change. No pessimistic locks, no transactions held open across think-time.

Idempotency keys. Mutations accept an optional idempotency key. Replaying the same key with the same request returns the original stored result instead of performing the work twice; replaying it with a different request is rejected with IDEMPOTENCY_CONFLICT. Agents retry constantly — flaky networks, ambiguous timeouts, an unclear response — so retries have to be safe by construction.


Dogfooding

The project runs on this model. rhizome-mcp's own backlog isn't in GitHub Issues or a Markdown file — it lives in a rhizome-mcp database, and every unit of work was created, claimed, executed, and finished through the MCP server by coding agents.

At the time of writing, the project database contains 100 issues, 94 of them completed — including the VS Code extension, the npm and Open VSX publishing pipelines, the SQLite concurrency layer, and the features described in this article. Agents claimed work, hit context limits, wrote checkpoints, handed off, occasionally crashed and had leases reclaimed. The backlog stayed coherent.

That loop surfaced real bugs, too: a bundled-binary resolution path that silently reported success instead of failing hard, validation errors that named the wrong field, a status board that tried to open a temp file the editor's trust rules forbade. Dogfooding turns "it should work" into "it did, and here's the commit from when it didn't."


How this compares to other approaches

Each alternative below is good at the problem it was designed for. The question is which problem that is.

Markdown / TODO.md handoff. Zero setup and human-readable, but it offers no atomic claim, no lease, no crash recovery — and merge conflicts arrive as soon as two agents edit it. Workable for a single agent in a single session.

Hosted-tracker MCP wrappers (Linear, Jira, GitHub Issues). Genuinely useful when humans and agents share one project-management tool. They inherit the human-team model, though: permanent assignees, stored statuses, no lease or heartbeat primitive, plus network latency and rate limits on every call. There's no built-in notion of "this agent holds this task for fifteen minutes and auto-releases if it dies." You'd build the lease model on top, against an API that wasn't designed for it.

Knowledge-graph "memory" servers. These give agents durable facts — entities and relationships that survive a session. Valuable and complementary, but they model memory, not execution: no claiming, no lease, no attempt lifecycle. Memory and coordination are different problems.

These aren't mutually exclusive. A hosted tracker may remain the right source of truth for product planning and human collaboration. A lease-based execution layer solves a narrower problem: coordinating unreliable agent workers while they perform the work.


How rhizome-mcp implements the model

  • Runtime. Go, single static binary, CGO-free. Startup resolves the project, opens SQLite, applies embedded checksummed migrations, verifies WAL and foreign keys, then serves.
  • Protocol. Model Context Protocol over JSON-RPC 2.0, via the official MCP Go SDK. Stdio is the primary transport; an opt-in loopback-only HTTP transport (strict Host/Origin validation, no auth, not for remote exposure) serves local clients.
  • Storage. One SQLite database per project, STRICT tables, CHECK-constrained enums, an append-only event log as the audit and delta-sync backbone, FTS5 as a rebuildable search projection. The database lives outside the repository; the repo holds only a two-field pointer file.
  • Execution model. Issues, work attempts, attempt notes and checkpoints, bounded work context, and a review workflow pinned to an exact issue version.
  • Determinism. All time comes from an injected clock, so lease expiry, heartbeats, and cleanup are tested with a controllable fake clock rather than sleep(). The architecture is hexagonal: domain and application logic behind ports, with MCP, CLI, and SQLite as adapters.

Try it

# zero-install trial
npx rhizome-mcp
Enter fullscreen mode Exit fullscreen mode

Or install the binary and wire up a client:

curl -fsSL https://raw.githubusercontent.com/Odrin/rhizome-mcp/main/scripts/install.sh | sh
rhizome-mcp init
rhizome-mcp connect claude   # or: codex | vscode | json
Enter fullscreen mode Exit fullscreen mode

A VS Code extension provides one-click MCP registration and a status board; other install channels are listed in the repository.

Links


Takeaways, even if you never install it

  1. Treat agents as unreliable workers. Any coordination layer that assumes otherwise will surface those failures as bugs.
  2. Don't store possession — lease it. A renewable lease with a heartbeat makes reclamation automatic.
  3. Derive in_progress, never store it. Possession becomes a computed property of a live lease, so stale state has nowhere to persist.
  4. Checkpoint restartable state. Recovering the slot is cheap; recovering the work is what saves time and tokens.
  5. Response shape is a performance budget when the client is an LLM. Omit bodies by default, make heavy fields opt-in, ship a delta feed.

Give your agents a queue that expects them to fail, and they'll coordinate like a system instead of colliding like a crowd.


I'm especially interested in feedback from developers running multiple coding agents against the same repository. Compatibility reports from Copilot, Claude Code, Codex, and other MCP clients are particularly useful, as are examples of coordination failures this model does not yet handle well.

The project is open source under Apache-2.0, and bug reports, design discussions, and contributions are welcome: github.com/Odrin/rhizome-mcp.

Top comments (0)