DEV Community

Cover image for Agents that forget where they are will repeat work forever — checkpointing is the save-point pattern you need — Checkpointing and Resume
Alex Aslam
Alex Aslam

Posted on

Agents that forget where they are will repeat work forever — checkpointing is the save-point pattern you need — Checkpointing and Resume

I watched a research pipeline die at step nine of ten. Forty-three minutes of work, sixty sources fetched, four analyses complete, one synthesis in progress. The worker crashed—OOM, not a logic bug. When it came back, the orchestrator did what I'd told it to do: it started over.

Forty-three minutes became eighty-six. Every source re-fetched. Every analysis re-run. And three of those re-runs wrote to a downstream system that didn't deduplicate, which is how I learned that "just retry the whole thing" is not a recovery strategy. It's a duplicate-generation engine with a nice error message.

That was the week I stopped thinking about agent memory and started thinking about agent position.

Memory Was Never the Problem

I'd spent months on memory. Vector stores, retrieval strategies, context compression. My agents knew a lot. What they didn't know was where they were.

Ask one of them "what have you already done?" and it would tell you about its accumulated context—sources read, conclusions drawn. Ask it "what's left?" and it had nothing. It could describe its knowledge. It couldn't describe its progress.

That distinction is the whole game. Memory answers "what do I know?" Checkpointing answers "where am I?" Most agent systems have the first and not the second, which is why a crash at step nine costs you the entire run instead of the last step.

The failure modes multiply from there. An agent that doesn't know what it already did will happily redo it. An agent that doesn't know which side effects already landed will happily re-fire them. The orchestrator retries, the agent re-reasons, and the external world absorbs the duplication. Nobody throws an error, because from every individual component's perspective, nothing went wrong.

What Checkpointing Actually Is

The pattern is decades old. Durable execution systems have been solving this since before anyone called them agents. The core idea is that a workflow's progress is recorded as a durable log, and recovery means replaying that log rather than re-executing from the start.

Temporal's model is the clearest articulation. Workflow code must be deterministic—no Date.now(), no random, no direct I/O. Activities are where side effects live. When a worker crashes and the workflow resumes on another worker, the workflow code re-executes from the beginning, but activities don't run again. Their results come from the event history. The code re-derives its position by replaying decisions that were already made.

AWS shipped the same concept into Lambda as durable execution, letting a function checkpoint at each step and resume from the last recorded point, with execution windows up to a year. DBOS does it with Postgres as the log and @DBOS.step() decorating each checkpointed unit. Restate journals every handler invocation. Resonate makes every spawn and await a durable promise, so a crash mid-batch re-runs only what was in flight.

LangGraph brings it into the agent world directly. Passing a checkpointer to compile() makes every superstep persist. The thread_id in your config is the identity of the run. A crash mid-graph, a resumed invocation with the same thread_id, and the graph picks up at the last completed node rather than at START.

checkpointer = PostgresSaver.from_conn_string(DB_URI)
graph = workflow.compile(checkpointer=checkpointer)

config = {"configurable": {"thread_id": "session-42"}}
result = graph.invoke({"query": "..."}, config=config)

# after a crash, same thread_id, no new input:
result = graph.invoke(None, config=config)
Enter fullscreen mode Exit fullscreen mode

That invoke(None, ...) is the entire recovery primitive. The graph reads its last checkpoint and continues. Same thread_id, same state, same position.

LangGraph's newer durability modes let you tune how eagerly that happens. "sync" checkpoints before the next step begins—safest, slowest. "async" writes in the background—faster, with a small window where a crash can lose the last step. "exit" only persists at the end—cheapest, useless for recovery. Most production systems land on "async" and accept the small exposure.

The Trap That Bites Everyone

Checkpointing and replay have a failure mode that will absolutely find you: non-determinism.

If your workflow code produces different results on replay than it did the first time, replay diverges. Temporal enforces determinism hard—you use Workflow.now() and Workflow.random() instead of the standard library, and violations are caught by a determinism checker. LangGraph is more permissive, which means it's more dangerous. If a node does something nondeterministic and the checkpoint lands after it, replay may reconstruct a different state than the original run.

The practical rule: anything with side effects or non-determinism belongs behind the checkpoint boundary, not inside the replay path. LLM calls are the obvious case. An LLM call is non-deterministic by definition. If a resumed run re-invokes the model instead of reading the recorded result, you've lost the guarantee entirely.

This is where most agent frameworks quietly fall short. A graph node that calls an LLM and writes a summary is two things: a decision and a side effect. Only the second one should be replayable. Frameworks that checkpoint at the node level rather than the step level conflate them.

Idempotency Is the Other Half

Checkpointing tells your agent where it is. It does nothing about what already happened in the outside world.

The classic failure: the agent writes a record, then crashes before the checkpoint lands. On resume, it replays from before the write and writes again. Now you have two records and no error.

The fix is idempotency keys—derived from the operation's identity, not the timestamp. A refund is keyed by refund:{order_id}:{amount}. A notification is keyed by notify:{customer_id}:{event_id}. A write is keyed by the checkpoint sequence number of the step that produced it. The receiving system either deduplicates or the operation is naturally idempotent because it's a keyed upsert rather than an append.

Without this, checkpointing turns a single failure into a single failure plus duplicated side effects. With it, checkpointing is actually safe.

The pairing is non-negotiable. I learned this the expensive way: my pipeline had checkpointing and no idempotency, so the resume at step nine re-sent three downstream writes that had already succeeded. One of them triggered a customer email. Twice.

The Guardrails That Make It Survivable

Five failure modes matter in production.

Checkpoint bloat. Every checkpoint stores the full state. If your state includes fifty documents and the entire conversation history, you're writing megabytes per step. Storage grows, resume time grows, and eventually you're paying more for persistence than for inference. The fix is compaction—store references to large artifacts rather than the artifacts themselves, and periodically summarize accumulated context into a compact form.

Topology skew. A checkpoint written by graph version 3 may not be resumable by graph version 4. If you deployed a change between the crash and the resume, the replay can hit a node that no longer exists or an edge that now routes elsewhere. Version your graph. Refuse to resume a checkpoint whose version doesn't match, or provide explicit migration paths.

Coarse checkpoints. If you only checkpoint every ten steps, a crash costs you nine steps of re-run work. If you checkpoint every step, you pay write overhead constantly. The right granularity is per-meaningful-side-effect, not per-node.

Checkpoint everywhere, resume nowhere. I've seen teams invest heavily in persistence and never build the resume path. The checkpoint exists, the recovery code doesn't. A crash still costs the full run. Persistence without recovery is a log, not a safety net.

Unbounded retries against a poisoned state. If a checkpoint contains bad data—a malformed artifact, an impossible intermediate value—resuming will fail the same way every time. The system needs a poison-checkpoint detector and an escalation path. Otherwise the workflow loops forever, crashing and resuming in the same place.

What Production Teams Are Actually Running

Lyft's LangGraph-based support system uses thread-level state and checkpointers to manage conversation history across riders and drivers, with the router holding state and re-routing mid-chat as intent shifts. Their agent development timeline dropped from roughly six months to a few weeks, and the state management is what makes multi-turn support workflows survivable in production.

Temporal is the workhorse behind a large share of durable agent infrastructure. The event-history model means a workflow can run for months, survive repeated worker failures, and resume on any available worker without losing position. The determinism constraint is the price of admission.

LangGraph's human-in-the-loop pattern depends entirely on checkpointing. interrupt() pauses a node, the graph checkpoints, and the run stays suspended indefinitely. When a human responds, Command(resume=value) continues from exactly where it stopped. Without a checkpointer, the interrupt has nothing to resume from.

AWS Lambda durable execution extends the same model to serverless functions, checkpointing at each step and resuming across cold starts with execution windows measured in months. DBOS and Restate offer Postgres- and journal-based alternatives with the same guarantee.

When Not to Reach for Checkpointing

Checkpointing is not free, and it's not always justified.

Short, single-shot workflows. If your agent completes in under thirty seconds and a failure just means a retry, checkpointing adds infrastructure for a recovery cost you barely notice.

Workflows with no durable side effects. If everything the agent does is read-only and the output is ephemeral, there's nothing to protect against. Re-run it.

Prototypes. Adding a checkpoint store, a schema, a versioning strategy, and idempotency keys to a workflow you haven't validated is premature. Build the thing first. Checkpoint it when a crash actually costs you something.

When the state can't be serialized. Some agents hold live connections, handles, or in-process state that can't be persisted cleanly. Either the design changes to isolate serializable state, or checkpointing isn't the right tool.

The Trade-Off You're Accepting

Checkpointing buys you recovery. It costs you determinism discipline, storage, and a serialization boundary in your architecture.

You're accepting that some of your code must be pure and replayable. You're accepting that large artifacts need to live outside the state and be referenced. You're accepting that a version change to your graph is now a version change to your persistence format. And you're accepting that idempotency is not optional—if you don't build it, checkpointing will turn single failures into compound ones.

But here's what I've learned from every recovery story that went badly: the teams that suffer aren't the ones without checkpointing. They're the ones with checkpointing and no discipline about what lives inside it. The save point isn't a magic eraser. It's a contract about what your system can and cannot redo safely.

So here's my question: If your agent crashed right now, mid-task, would it know exactly what it already did—or would it start over and hope the side effects were idempotent?

I'd love to hear where you've landed. LangGraph checkpointers, Temporal's event history, a homegrown state table, or a retry-everything approach you haven't been burned by yet—and what finally made you change?

Top comments (0)