DEV Community

siddarthareddy chinthala
siddarthareddy chinthala

Posted on

I Killed My AI Agent 14 Times. It Kept Working.

Every kill was real. kill -9, no warning, mid-run. Here's the durability layer that made it boring.


The problem nobody demos

Watch any agent framework demo: the agent runs, it finishes, everyone claps. Now ask the question nobody asks on stage:

What happens when the worker dies at 90%?

OOM-killed. Spot instance reclaimed. Someone's finger slips on kill -9. The container gets evicted.

I surveyed the landscape — LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK. The answer is the same everywhere: nothing. No death detection. No scheduled resume. Your three-hour run dies at hour 2:59, and you start over.

This isn't a corner case. If you run agents in production — long research tasks, multi-step coding, data pipelines — worker death is a when, not an if.

What I built

DHP — the Durable Handoff Protocol. It's the durability layer that sits underneath your agent framework:

  • MCP handles agent ↔ tool.
  • A2A handles agent ↔ agent.
  • DHP handles durable execution state: checkpointing, fencing, failover, portable suspended work.

Five guarantees: no lost results, no crossed inputs, no untyped results, no silent budget overrun, no runner lock-in.

The core ideas are old and proven — leases, heartbeats, atomic compare-and-swap ownership, monotonic fencing tokens — applied to a problem the agent world hasn't solved yet.

Kill #1: the one that convinced me

60 real Wikipedia pages to fetch. Worker A starts. At page 20, I kill -9 it. No warning, no graceful shutdown.

What happened:

  1. Worker's heartbeats stop.
  2. Supervisor notices the missed lease after ~6.5 seconds.
  3. Supervisor marks the handoff orphaned.
  4. A standby worker claims it, reads the last checkpoint (page 20, shipped over TCP before the kill).
  5. Standby resumes at page 21.

Result: 60/60 pages fetched. Zero refetched. The kill was invisible in the output — just a 6.5-second pause.

Kills #2–4: the network chaos run

I wasn't satisfied with one kill. The next test: four execution rounds, three real SIGKILLs delivered mid-TCP-shipment — the worst possible moment, when a checkpoint is halfway across the wire.

  • Orphan times: 7.0s, 6.5s, 6.0s.
  • Peer received 24/24 checkpoints.
  • Zero rework.

The transport protocol verifies envelope identity and per-checkpoint hashes on receipt, so a torn write can't corrupt the resume point.

Kills #5–18: randomized chaos

Then the fault-injection harness: randomized SIGKILLs against workers and supervisors. 14 worker kills, 4 supervisor kills, random timing.

The interesting one: I killed the worker and the lead supervisor in the same instant. A follower supervisor promoted via leader election, detected the orphan in 6.6 seconds, and a standby recovered via atomic CAS from the last checkpoint. 1,600 checkpoints, zero rework, result verified correct.

Every invariant held: no lost results, no crossed inputs, exactly one owner at all times (10 simultaneous recovery threads → exactly one winner, verified).

What I was honest about

Sub-second failover is a no-go. I investigated it seriously — for LLM workloads, the reasoning layer is too slow and unpredictable for sub-second death detection without false positives. The honest number is seconds, not milliseconds: ~7s kill-to-resume. That's still 2x faster than the best I could find (Temporal's ~12s for stateful failover), and it's real, measured, not projected.

How it works (the 30-second version)

import dhp

store = dhp.Store("/var/lib/dhp")

# Dispatch durable work
hid = dhp.dispatch(store,
    task_kind="fetch_pages",
    inputs={"urls": [...]},
    output_schema={"type": "array", "items": {"type": "string"}})

# Claim it — heartbeats run automatically in a background thread
with dhp.claim(store, hid, "worker-1") as ctx:
    for page in pages:
        data = fetch(page)
        ctx.checkpoint({"page": page, "data": data})  # resume point
    ctx.complete({"pages": all_data})  # schema-validated
Enter fullscreen mode Exit fullscreen mode

If worker-1 dies inside that with block, the supervisor orphans the handoff and any standby picks it up from the last checkpoint(). The complete() validates the result against the schema — no untyped results sneaking through.

For MCP users, there's an 8-tool server:

claude mcp add dhp -- dhp-mcp --root ~/.dhp
Enter fullscreen mode Exit fullscreen mode

And an A2A bridge — the A2A task ID becomes the stable identity, DHP handoffs are execution attempts underneath. The A2A client sees WORKING → COMPLETED; the worker swap is invisible.

The architecture

Key design decisions:

  • Content-addressed envelopes: identity covers intent, not mutable state. A tampered checkpoint can't be mistaken for a valid one.
  • Fencing tokens: monotonic counters reject stale workers. A worker that was declared dead can't come back and commit.
  • Atomic CAS ownership: exactly one recoverer wins, even under contention.
  • No shared database needed: the transport log (append-only JSON-lines) lets a peer rebuild state from shipped checkpoints alone. I proved this by destroying an entire host — disk and all — mid-run.

Try it

pip install dhp-protocol
# or one line:
curl -fsSL https://raw.githubusercontent.com/SIDDARTHAREDDY8/dhp/main/install.sh | bash
Enter fullscreen mode Exit fullscreen mode

The repo has a runnable kill demo (demo/run_mad_demo.py) — it fetches real pages, kills a real worker, and shows the recovery. Run it yourself; don't take my word for it.

Links:


DHP is MIT licensed. 45 tests green, 10/10 conformance, chaos-tested with real kills. If you run agents in production, this is the layer you're missing.

Top comments (0)