DEV Community

Nature (Zhang Chenxi)
Nature (Zhang Chenxi)

Posted on

Four Things LangGraph Checkpoints Won't Do for Multi-Day Agents

A longer version of this article is on towow.ai.

If your agents lose the thread on day three, a different runtime usually will not fix it. LangChain's own comparison page puts LangGraph, Temporal and Inngest in the same group, runtimes, whose job is durable execution, streaming, human-in-the-loop and persistence (docs). A LangGraph checkpointer saves a snapshot of graph state at each super-step, per thread (docs). What goes into that state, and whether it is true, is up to you.

Below are four gaps we hit running agent work over days. Each has the symptom, a fix, and what our own records show; the records come from Flowness, the agent harness we use in our own delivery work. The code blocks are sketches with made-up task and file names, not a library API or Flowness code. The replay sketch follows the pseudo-code in our cursor write-up, and the deploy snippet is LangGraph's own example.

1. Handoffs

Symptom. A new session starts and either continues from a stale picture or duplicates work the old session still holds.

Fix. Keep a handoff record in graph state or a store (persistence docs), and call a check like this at the start of whichever node runs first when a new session picks up the thread. A sketch, not a library API:

handoff = {
    "in_flight": [{"task": "t-17", "owner": "session-a", "step": "tests",
                   "artifact": "branch/perm-tests", "blocked_on": None}],
    "next_waits_on": "t-17 review",
    "refs": ["task_packet@v3", "branch/perm-tests", "report/latest"],
}

def resume_gate(handoff):
    for ref in handoff["refs"]:
        if not exists_and_current(ref):   # packet version, branch, readable report
            raise NeedsRepair(ref)        # fix the sheet; don't continue from memory
    for t in handoff["in_flight"]:
        if activity_belongs_to(t["task"], t["owner"]):
            skip_or_attach(t)             # owner still working: don't start a twin
Enter fullscreen mode Exit fullscreen mode

Check ownership per task, not per project line. "Someone is on the permissions work" does not tell you whether this test task is taken.

Our record. Our handoff sheet lists in-flight tasks, owners, artifacts, blockers and what the next step waits on, and the receiver checks every reference first (write-up, in Chinese). In a 2 September 2026 review of one rebuild, 29 of 32 tasks were recorded as successful while one requirement had no task carrying it all the way (write-up, in Chinese). The sheet keeps the work; the goal needs its own object, inherited by every stage.

2. "Done" claims

Symptom. A step prints success, exits 0, and the target is unchanged.

Fix. A checkpoint records that a node returned. It cannot see the repository, service or file the node claims to have changed. For each step with an outside effect, declare what to observe and where, then read it back from the target, not from the process that made the change:

effect = Effect(target="repo:main", expect="perms.py contains check_access")

result = run_step()                       # Attempt
observed = target_readback(effect)        # fresh read from the target: Effect
record(attempt=result, effect=observed)   # Adoption and Acceptance get their own labels
Enter fullscreen mode Exit fullscreen mode

Report Attempt, Effect, Adoption and Acceptance separately instead of one green tick.

Our record. In 17 real Flowness scenarios a naive terminal label was wrong 10 times. In 9 selected operations, stdout and exit code each matched the true state in 4; effect contract plus read-back matched in 9 (study). These are selected sets, not a general failure rate.

3. Replay

Symptom. After a crash, a restart double-counts or double-acts.

Fix. Assume every step runs twice. LangGraph's docs say replay re-executes nodes after the chosen checkpoint, so LLM calls and API requests fire again (time travel). Interrupts re-run their node, so side effects before an interrupt should be idempotent (interrupts). In exit durability mode, intermediate state is not saved, so a mid-run crash is not recoverable (checkpointers).

For your own side effects, save the result and the dedup evidence in one atomic write, and advance the progress marker last:

pending, end = read_after_cursor()
with lock():
    state = load()
    for offset, key in pending:
        if offset > state.high_water:     # already-counted offsets are skipped
            state.counts[key] = state.counts.get(key, 0) + 1
    state.high_water = max([state.high_water] + [o for o, _ in pending])
    atomic_replace(state)                 # counts and high-water mark together
advance_cursor(end)                       # last
Enter fullscreen mode Exit fullscreen mode

Our record. With synthetic offsets 40, 80 and 120, a crash before the cursor moves leaves counts of 2; the restart rereads three signals and ends at 2 + 0 + 0 + 1 = 3, where naive re-adding gives 2 + 3 = 5. The write-up covers a fixed, append-only local file, a lock among cooperating writers, and recovery from process exit, not power loss (write-up, in Chinese).

4. Pauses and deploys

Symptom. You pause the system and work keeps arriving. Or you ship a fix and in-flight runs change behavior.

Fix. Write down what a pause stops: new dispatch, in-flight work, review and fix lanes, and coordinator patrols are set separately. On resume, read what arrived during the pause before restarting any dispatch; resuming everything at once can restart work someone already holds.

For deploys, LangGraph's backward-compatibility guide says the latest graph is applied to every thread, including threads resuming from a checkpoint, whereas some workflow engines pin a run to its starting code version. Renaming or removing a node while threads are paused at it breaks the resume. The guide recommends stamping a behavioral version on state at thread start and branching on it (docs):

def intake(state):
    return {"flow_version": state.get("flow_version", 2)}  # new threads are stamped 2

def after_triage(state):
    return "policy_check" if state.get("flow_version", 1) >= 2 else "respond"  # old threads default to 1
Enter fullscreen mode Exit fullscreen mode

Our record. In a 15 September pause, two in-flight executors finished their current section, new dispatch went to zero, and review and fix lanes stayed on (write-up, in Chinese).

Which tool for which gap

LangChain's page lists Temporal and Inngest beside LangGraph as runtimes, and the Deep Agents SDK and Claude Agent SDK as harnesses. That grouping is LangChain's; this post does not test them. The four fixes above are about how you design state, handoffs and checks, so they carry to any runtime.


Drafted with AI assistance, based on our project records, fact-checked before publishing.

Top comments (0)