Every kill was real. kill -9, no warning, mid-run. Here's the durability layer that made it boring.
The problem nobody demos
Watch any agent framework demo: the agent runs, it finishes, everyone claps. Now ask the question nobody asks on stage:
What happens when the worker dies at 90%?
OOM-killed. Spot instance reclaimed. Someone's finger slips on kill -9. The container gets evicted.
I surveyed the landscape — LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK. The answer is the same everywhere: nothing. No death detection. No scheduled resume. Your three-hour run dies at hour 2:59, and you start over.
This isn't a corner case. If you run agents in production — long research tasks, multi-step coding, data pipelines — worker death is a when, not an if.
What I built
DHP — the Durable Handoff Protocol. It's the durability layer that sits underneath your agent framework:
- MCP handles agent ↔ tool.
- A2A handles agent ↔ agent.
- DHP handles durable execution state: checkpointing, fencing, failover, portable suspended work.
Five guarantees: no lost results, no crossed inputs, no untyped results, no silent budget overrun, no runner lock-in.
The core ideas are old and proven — leases, heartbeats, atomic compare-and-swap ownership, monotonic fencing tokens — applied to a problem the agent world hasn't solved yet.
Kill #1: the one that convinced me
60 real Wikipedia pages to fetch. Worker A starts. At page 20, I kill -9 it. No warning, no graceful shutdown.
What happened:
- Worker's heartbeats stop.
- Supervisor notices the missed lease after ~6.5 seconds.
- Supervisor marks the handoff orphaned.
- A standby worker claims it, reads the last checkpoint (page 20, shipped over TCP before the kill).
- Standby resumes at page 21.
Result: 60/60 pages fetched. Zero refetched. The kill was invisible in the output — just a 6.5-second pause.
Kills #2–4: the network chaos run
I wasn't satisfied with one kill. The next test: four execution rounds, three real SIGKILLs delivered mid-TCP-shipment — the worst possible moment, when a checkpoint is halfway across the wire.
- Orphan times: 7.0s, 6.5s, 6.0s.
- Peer received 24/24 checkpoints.
- Zero rework.
The transport protocol verifies envelope identity and per-checkpoint hashes on receipt, so a torn write can't corrupt the resume point.
Kills #5–18: randomized chaos
Then the fault-injection harness: randomized SIGKILLs against workers and supervisors. 14 worker kills, 4 supervisor kills, random timing.
The interesting one: I killed the worker and the lead supervisor in the same instant. A follower supervisor promoted via leader election, detected the orphan in 6.6 seconds, and a standby recovered via atomic CAS from the last checkpoint. 1,600 checkpoints, zero rework, result verified correct.
Every invariant held: no lost results, no crossed inputs, exactly one owner at all times (10 simultaneous recovery threads → exactly one winner, verified).
What I was honest about
Sub-second failover is a no-go. I investigated it seriously — for LLM workloads, the reasoning layer is too slow and unpredictable for sub-second death detection without false positives. The honest number is seconds, not milliseconds: ~7s kill-to-resume. That's still 2x faster than the best I could find (Temporal's ~12s for stateful failover), and it's real, measured, not projected.
How it works (the 30-second version)
import dhp
store = dhp.Store("/var/lib/dhp")
# Dispatch durable work
hid = dhp.dispatch(store,
task_kind="fetch_pages",
inputs={"urls": [...]},
output_schema={"type": "array", "items": {"type": "string"}})
# Claim it — heartbeats run automatically in a background thread
with dhp.claim(store, hid, "worker-1") as ctx:
for page in pages:
data = fetch(page)
ctx.checkpoint({"page": page, "data": data}) # resume point
ctx.complete({"pages": all_data}) # schema-validated
If worker-1 dies inside that with block, the supervisor orphans the handoff and any standby picks it up from the last checkpoint(). The complete() validates the result against the schema — no untyped results sneaking through.
For MCP users, there's an 8-tool server:
claude mcp add dhp -- dhp-mcp --root ~/.dhp
And an A2A bridge — the A2A task ID becomes the stable identity, DHP handoffs are execution attempts underneath. The A2A client sees WORKING → COMPLETED; the worker swap is invisible.
The architecture
Key design decisions:
- Content-addressed envelopes: identity covers intent, not mutable state. A tampered checkpoint can't be mistaken for a valid one.
- Fencing tokens: monotonic counters reject stale workers. A worker that was declared dead can't come back and commit.
- Atomic CAS ownership: exactly one recoverer wins, even under contention.
- No shared database needed: the transport log (append-only JSON-lines) lets a peer rebuild state from shipped checkpoints alone. I proved this by destroying an entire host — disk and all — mid-run.
Try it
pip install dhp-protocol
# or one line:
curl -fsSL https://raw.githubusercontent.com/SIDDARTHAREDDY8/dhp/main/install.sh | bash
The repo has a runnable kill demo (demo/run_mad_demo.py) — it fetches real pages, kills a real worker, and shows the recovery. Run it yourself; don't take my word for it.
Links:
DHP is MIT licensed. 45 tests green, 10/10 conformance, chaos-tested with real kills. If you run agents in production, this is the layer you're missing.
Top comments (0)