DEV Community

Cover image for Your agent is a long-running process with side effects. Kill it and watch what happens.
Harish Kotra (he/him)
Harish Kotra (he/him)

Posted on AI-assisted

Your agent is a long-running process with side effects. Kill it and watch what happens.

A teardown of **Checkpoint: a 30-step agent run on Temporal that survives kill -9, with an exactly-once ledger that proves it in the UI.


Every agent tutorial ends at the same place: a while loop that calls a model, calls a tool, writes something to the world, and repeats. It works beautifully until the process dies at step 27 of 30 and then you learn what your loop actually was: a record of everything it had already done,
held in memory that no longer exists.

The instinctive fix is a checkpoint object — serialize the loop index to disk, restore it on boot. That works right up until the crash lands in the two lines between "the model answered" and "the email was sent", at which point you are writing a distributed database and calling it a bug fix.

Checkpoint is the version where the loop's position is not state you maintain — it is a log someone else already stores durably, and every observable side effect is gated by a key that cannot change across retries. Then I built a button that sends a real SIGKILL to the worker process, and a test suite that fails if anything is done twice.

Here's how it works, and the three bugs I found by actually pressing the button.


The claim being tested

Not "the app is resilient" (unfalsifiable). Instead:

SELECT COUNT(*) - COUNT(DISTINCT id) FROM effects;   -- must equal 0
Enter fullscreen mode Exit fullscreen mode

That number sits in the ledger panel's header, green when it's 0, and the app computes it on every refresh. Kill the worker at step 7 of 16 and the run completes with 48 ledger rows and duplicates: 0.

Everything below exists to make that sentence defensible.

Three processes, two stores

 panel :5173 ──SSE──► Temporal :7233 ─── EVENT HISTORY ───►  the run's position
     │                    ▲   ▲
     │ /api/*             │   │ poll tasks
     ▼                    │   │
 API :3001 ──spawn/kill──►│   └──► WORKER (a pid you can murder)
     │                          workflows/agent.ts   ← the loop
     │                          activities/{llm,tool,effect}.ts
     ▼
 SQLite checkpoint.db  effects · steps · runs · faults · worker_events · notes
     └── artifacts/<runId>/step-007.md · outbox.log
Enter fullscreen mode Exit fullscreen mode

The rule I held myself to: anything the UI presents as evidence must be readable from SQLite or Temporal history after both processes are killed. That single constraint killed most of the ways I would have cheated.

The loop is a plain async function — and that's the trick

// server/src/workflows/agent.ts
export async function agentLoop(input: AgentInput): Promise<AgentSummary> {
  const { workflowId } = workflowInfo();
  for (let index = 1; index <= input.total; index++) {
    await book.markStepRunning({ runId: input.runId, step: index });
    const outcome = await step.llmActivity({ runId: input.runId, step: index, total: input.total, previousNote });

    const toolResult = outcome.decision.action === 'compute'
      ? (await step.toolActivity({ runId: input.runId, step: index, value: outcome.decision.args.value })).result
      : null;

    await step.effectActivity({ runId: input.runId, step: index, workflowId, decision: outcome.decision, toolResult });
    await book.markStepCompleted({ runId: input.runId, step: index });

    await sleep(300 + ((index * 137) % 501));   // 300–800ms, derived from the index
  }
}
Enter fullscreen mode Exit fullscreen mode

On every task, the SDK re-runs this function from the top. Activity calls don't re-execute — their recorded results are fed back from history, so the loop is effectively "seeking" to the frontier. That's why the sleep is a formula rather than Math.random(): replay must schedule identical timers. There is no clock, no fetch, no fs anywhere in workflow code; timestamps come from activities.

SDK trivia: the docs still say workflow.now(). It doesn't exist in 1.24 — the sandbox patches Date/setTimeout inside workflows, and the unsafe escape hatch moved to workflowInfo().unsafe.now(). This loop needs neither.

Where the crash-safety actually lives: the idempotency key

Durable execution gets you re-execution. It does not get you don't-send-two-emails — an activity whose worker died is re-run from the beginning, by design. That's exactly-once processing, not exactly-once effects. The bridge is a key that cannot change:

effectId = `${workflowId}:${stepIndex}:${kind}`     // kind ∈ file | db | email
Enter fullscreen mode Exit fullscreen mode

workflowId is the run id — stable across activity retries, worker respawns, API restarts. Then the
write itself is the lock:

// server/src/ledger.ts
export function applyEffectOnce(db, spec, physical) {
  const claim = db.transaction(() => {
    const res = db
      .prepare('INSERT OR IGNORE INTO effects(id, step, kind, payload, applied_at, attempt) VALUES(?,?,?,?,NULL,?)')
      .run(spec.id, spec.step, spec.kind, spec.payload, spec.attempt);
    return { claimed: res.changes === 1, existing: db.prepare('SELECT * FROM effects WHERE id = ?').get(spec.id) };
  })();

  if (claim.claimed) { physical(); stamp(spec.id); return { applied: true }; }        // we own it
  if (claim.existing.applied_at) return { applied: false };                            // no-op: already done
  physical(); stamp(spec.id); return { applied: false, recovered: true };              // finish an interrupted write
}
Enter fullscreen mode Exit fullscreen mode

Three branches, three crash stories:

situation what happens result
first attempt claim → write → stamp applied_at applied: true
retry, or re-dispatch after a kill PK conflict → read row → skip the write applied: false
killed between claim and stamp row exists with applied_at IS NULL → redo the deterministic write recovered: true

The third row forces a design rule that pays off everywhere else: payloads are built only from activity arguments, because those arguments are already in history. A replayed attempt produces byte-identical file content, so re-running an interrupted write is boring — and boring is the goal. The email stub appends to outbox.log; the claim row is what stops the second append.

Note the deliberate trade-off: this is at-most-once per claim, so a crash between claiming and writing leaves a row that lies. The applied_at IS NULL marker turns that silent corruption into a detectable, repairable state. If your side effect isn't deterministic (a payment, an email with a timestamp in it), you'd instead store the rendered payload on the claim row and hand the same string to the transport on every attempt.

Bug #1 — a "fast" resume took three minutes

First version. Kill the worker, watch the panel: the replacement came up in ~1.5s and… nothing. Step 7 stayed running for 180 seconds, then resumed correctly. The data was never at risk; the latency was a lie I was telling myself about how crash recovery works.

Cause: the activity task was dispatched to worker A. A died without acknowledging. Temporal doesn't guess — for an in-flight activity it waits for startToCloseTimeout, which I'd set to 3 minutes to accommodate slow models.

Fix: heartbeats. The activity proves liveness every 2s; the server notices the silence at
heartbeatTimeout and re-dispatches.

// server/src/activities/heartbeat.ts
export function beat(details: unknown): () => void {
  let stopped = false;
  const once = () => {
    if (stopped) return;
    try { heartbeat(details); } catch { stopped = true; clearInterval(timer); }
  };
  once();                       // never-beaten = can't be declared dead early; announce yourself
  const timer = setInterval(once, 1000);
  timer.unref();
  return () => { stopped = true; clearInterval(timer); };
}

// activity options
{ startToCloseTimeout: '3 minutes', heartbeatTimeout: '5 seconds',
  retry: { initialInterval: '400ms', backoffCoefficient: 2, maximumInterval: '5 seconds', maximumAttempts: 8 } }
Enter fullscreen mode Exit fullscreen mode

That moved kill→resume from minutes to ~12 seconds — and then I stopped pretending to understand
the rest of it, measured it, and was wrong about the cause. Three samples, three builds:

build heartbeatTimeout workflowTaskTimeout measured kill → next step starts
original (no beats) — default 10s >180s — stalled for the whole startToCloseTimeout; verify's 300s budget expired
+ heartbeats 12s default 10s 11.4 – 12.0s
+ tighter beats 5s default 10s 11.9 – 12.5s
+ tighter beats 5s 2s 12.0s, 12.0s, 18.0s

The heartbeat change was worth ~3 minutes. Everything after that was noise: the last ~10 seconds is the Temporal dev server re-delivering the outstanding workflow task, and none of the knobs I had access to moved it. I reverted the workflowTaskTimeout experiment — an unearned deviation from a default is worse than a slow demo — and wrote the number down as 12 seconds instead of the "2–4 seconds" I had confidently put in the README first.

Two lessons worth keeping: measure before you name the mechanism (I blamed heartbeats for a deadline that was never binding), and a negative result belongs in the docs — it's the part that saves the next person an afternoon.

Cancellation gets the same treatment: an activity can only receive it if it beats, so the 1s cadence buys a clean shutdown too.

Bug #2 — retries were invisible, because every activity has its own counter

Verification scenario B injects two transient failures at step 8 and asserts the timeline shows attempt 2. It failed: attempt=1. Injection clearly worked (the step retried and completed) but the ledger said one attempt.

Cause: steps is upserted by several activities — llmActivity logs attempt 2 for the model call, then markStepCompleted writes… attempt 1, because its attempt counter is 1. Last writer won and erased the evidence.

-- server/src/ledger.ts, upsertStep
ON CONFLICT(run_id, step) DO UPDATE SET attempt = MAX(steps.attempt, excluded.attempt), ...
Enter fullscreen mode Exit fullscreen mode

The generic lesson is about observability, not SQL: a field written by more than one process should have an explicit merge rule. MAX was the right semantics for "how many times did this step be worked on"; it would have been the wrong one for, say, ended_at.

Bug #3 — a model that rambles is a config bug, not a crash

The prompt demands JSON. Local models sometimes answer in prose. Per the spec, a parse failure is a non-retryable guardrail — so a 32-step demo died at step 2 because a 4-billion-parameter model decided to explain itself:

throw new ApplicationFailure(`Guardrail: step ${step} returned non-JSON output`, 'GuardrailError', true, [{ raw }]);
Enter fullscreen mode Exit fullscreen mode

Two changes, in order of honesty:

  1. A tolerant extractor: fenced blocks, then a hand-rolled balanced-brace scan for prose-wrapped or truncated JSON (with finish_reason === 'length' recorded, so the UI can say "truncated").
  2. One corrective re-prompt inside the activity — "your previous reply was rejected; return only the JSON object, strings under 15 words" — before the guardrail fires. Persistent garbage still fails fast and non-retryably, with the raw 500 chars stored on the step row and shown in a red banner. The injected garble fault bypasses repair entirely, which is how the test still proves the guardrail path rather than assuming it.

I also stopped trusting preset model ids: local servers answer for whatever is loaded, so the fallback chain now asks /v1/models first. Without that, a silent downgrade to the offline provider was producing a "working" demo that wasn't calling any model at all.

The kill button is not a simulation

// server/src/supervisor.ts
async kill(timeoutMs = 8000): Promise<KillReport> {
  const pid = this.pid ?? readPidFile();
  const atStep = this.getActiveStep();
  const observed = this.waitForExit(timeoutMs);
  process.kill(pid, 'SIGKILL');                     // ← undeliverable, uncatchable, real
  const { signal, exitCode } = await observed;       // resolves from the child's 'exit' event
  return { pid, signal: signal ?? 'SIGKILL', exitCode, confirmed: !pidAlive(pid), atStep };
}
Enter fullscreen mode Exit fullscreen mode

A SIGKILL leaves no exit code, so the panel says SIGKILL · no exit code instead of inventing one.

The worker is spawned detached: true + child.unref(), which enables the best test in the suite:
SIGKILL the API process mid-run. The worker survives (its parent dying is not its problem), a
fresh API reads run/worker.pid, sees the pid alive, and adopts it rather than starting a second
consumer:

PASS  5a. the api process is gone but the worker survived     api pid 96049 alive=false; worker pid 96449 alive=true
PASS  5b. the restarted api adopted the same worker           new api pid 98642 adopted worker 96449; completedSteps already 6
Enter fullscreen mode Exit fullscreen mode

The UI is a thin client over durable data

The panel holds no run state either — that would be embarrassing for an app about statelessness. It subscribes to an SSE stream where the server diffs SQLite each tick:

// server/src/api.ts
for (const s of next.steps) {
  const before = prev?.steps.find((p) => p.step === s.step);
  if (!before || before.status !== s.status || before.attempt !== s.attempt || before.worker_pid !== s.worker_pid) {
    send('step', { step: s.step, status: s.status, attempt: s.attempt, tokens: s.tokens, cost: s.cost_usd ?? 0, durationMs, provider: s.provider, workerPid: s.worker_pid, resumed: next.resumeStep === s.step });
  }
}
Enter fullscreen mode Exit fullscreen mode

resumeStep is derived, never remembered: the first step whose worker_pid differs from the step before it. So the ⚑ resumed here chip is a query result, not a UI memory — reload mid-crash and
it lands in the same place.

The history panel is the raw event log (handle.fetchHistory(), EventType names from @temporalio/proto), which is the honest place to look when the timeline and the ledger disagree.

Proof, not prose

npm run verify runs four scenarios and prints observed numbers:

PASS  1. worker really dies — API reports the observed signal + exit code   signal=SIGKILL exitCode=null pid=65590 confirmed=true observedAtStep=6
PASS  1b. the killed pid is gone from the process table                     previous pid 65590 alive=false; status=worker-dead
PASS  2. run completes after resume (workflow COMPLETED)                    status=COMPLETED completedSteps=16/16 resumedAtStep=7
PASS  2b. step log has no gaps                                              steps 1..16 missing=[]
PASS  2c. a different worker pid finished the run                           worker pids: 65590, 66292
PASS  3. ledger proves exactly-once after a mid-run kill                    COUNT(*) - COUNT(DISTINCT id) = 0 (141 rows)
PASS  4. retried step shows attempt >= 2 and wrote no duplicate effect      step 8: attempt=3 effects=3 uniqueIds=3
PASS  5. no run state was lost: the run finishes after an api restart       completedSteps=16/16 duplicates=0
PASS  6. completed step count equals TARGET_STEPS                           A=16/16  B=12/12  C=16/16
ALL CHECKS PASS — 14/14
Enter fullscreen mode Exit fullscreen mode

Each assertion prints the value it looked at, so a pass is auditable and a failure is a diagnosis:

function check(name, ok, observed) {
  results.push({ name, ok, observed });
  console.log(`${ok ? 'PASS' : 'FAIL'}  ${pad(name, 58)} ${observed}`);
  return ok;
}
Enter fullscreen mode Exit fullscreen mode

I also drove it in a browser and recorded the DOM, because "the state exists in the reducer" is not
the same as "the user sees it": click Kill worker →

worker process killed — a replacement worker is polling the queue now · observed signal SIGKILL
  · exit code none            → pill: worker-dead · RUNNING
(none)                        → pill: running · RUNNING      # resumed on the new pid
run completed — 24 steps, 72 ledger rows, duplicates 0 · applied exactly once across 2 worker processes
Enter fullscreen mode Exit fullscreen mode

And once from outside the app entirely: kill -9 $(cat run/worker.pid). The panel noticed, respawned, the run finished, duplicates stayed 0.

What this does not claim

  • Not exactly-once delivery. Temporal gives at-least-once activity execution; the ledger makes effects idempotent. If you can't render the effect deterministically, you need a transactional outbox, not a cleverer key.
  • Not a crash-proof clock. steps.started_at is stamped by whatever worker ran the attempt, and the resume marker uses worker-identity changes — a worker that restarts for an unrelated reason also flips that flag. It's honest evidence of "a different process did this step", which is the claim I want, but it isn't a general-purpose audit clock.
  • Not production. Single namespace, one task queue, no auth, no multi-tenancy, SQLite on a local disk, dev server with file persistence. Non-goals, deliberately.
  • The verified run above answered on the offline deterministic provider — no real API key in that session (placeholder → HTTP 401) and a hanging local server. Crash-safety claims don't depend on the model, but the "thoughts" in those runs are synthetic. With a key or a healthy local server the same scenarios run model-backed; the UI labels which provider answered every step.

Steal these four ideas

  1. Give every side effect a key derived only from durable identity — (workflowId, step, kind) in my case. If the key can change between retries, you don't have an idempotency key.
  2. Make the write itself the lock (INSERT OR IGNORE + changes === 1). No advisory locks, no separate state machine.
  3. Stamp completion last. Claim with applied_at NULL, do the work, then set the timestamp — and you get a detectable "interrupted" state instead of silent loss.
  4. Assert on numbers the app computed. If your resilience demo's proof is a console.log you wrote by hand, you built a slideshow.

Run it: clone the repo, npm install && npm run temporal:install && npm run dev, start a 32-step run, press Kill worker around step 7, and watch step 7 get claimed twice and applied once.

How it works

Code & more: https://www.dailybuild.xyz/project/273-checkpoint

Top comments (0)