DEV Community

Cover image for Your LangGraph Node Runs Twice After interrupt(). That's by Design.
Priyansh Singhal
Priyansh Singhal

Posted on Originally published at Medium

Your LangGraph Node Runs Twice After interrupt(). That's by Design.

Checkpoints are save points at super-step boundaries. Resume, durability and subgraph state all follow from that.

Put a print on the first line of a LangGraph node, call interrupt() a few lines later, resume the graph with a human's answer, and the print fires twice. The first time you see that you assume you wired something wrong. You didn't. LangGraph does not resume a paused node from the line it stopped on. It goes back to the last saved checkpoint and runs the whole node again from the top, and this time interrupt() returns the answer instead of pausing.

Once that clicks, three things tutorials teach as separate features (human-in-the-loop, crash recovery, time travel) turn out to be one mechanism seen from three angles. A checkpoint is a save point. Everything else is what you can do with a save point. This is the explanation I wanted while working through human-in-the-loop in my own agent notes, so here it is, with the 20-line script that reproduces the double print. Every behaviour below was run on langgraph 1.2.12, not just read in the docs; the one place they disagree is called out.

I wrote this with AI assistance, and I fact-checked and edited every claim myself.

What a checkpoint actually is

LangGraph runs a graph in rounds it calls super-steps. In one super-step, every node that is scheduled runs (in parallel if there are several), their state updates are merged, and the runtime works out which nodes fire next. A plain START -> a -> b -> END graph is three super-steps: one for the input, one for a, one for b.

The docs put it in one line: LangGraph creates a checkpoint at each super-step boundary. Attach a checkpointer, run that three-node graph on a thread, and get_state_history hands back four snapshots: an empty one before the input landed, then one after each step. Each snapshot holds the full state values at that moment, the nodes due to run next, a step number, and a link to its parent snapshot. That is the entire save file.

Two details matter for everything that follows. First, checkpoints are keyed by thread_id, passed in the config on every call. A new thread starts with empty state and never sees another thread's history. Second, the docs are precise about scope: checkpointers are thread-scoped, short-term memory (conversation continuity, human-in-the-loop, time travel, fault tolerance), and a separate store handles cross-thread, long-term memory (user preferences, facts). A checkpointer remembers a thread. It does not remember a user.

Why the node runs twice

Here is the smallest reproduction. No model, no API key, just a node that asks a human before it commits.

from typing import TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.types import interrupt, Command

class State(TypedDict):
    decision: str

def approve(state: State) -> State:
    print("approve node entered")          # runs before the pause
    answer = interrupt("Ship it?")         # pause here, wait for a human
    print(f"resumed with {answer!r}")
    return {"decision": answer}

graph = (
    StateGraph(State)
    .add_node("approve", approve)
    .add_edge(START, "approve")
    .add_edge("approve", END)
    .compile(checkpointer=InMemorySaver())
)

config = {"configurable": {"thread_id": "review-1"}}
graph.invoke({"decision": ""}, config)          # pauses at interrupt
graph.invoke(Command(resume="yes"), config)     # human answered
Enter fullscreen mode Exit fullscreen mode

On langgraph 1.2.12 the output is three lines: approve node entered, approve node entered, resumed with 'yes'.

Walk through what happened. The first invoke enters the node, prints, reaches interrupt("Ship it?"), and stops. The runtime writes a checkpoint for the thread, and the return value of invoke carries the question under a key called __interrupt__. Nothing after the interrupt line ran, so decision in state is still the empty string you seeded.

The second invoke passes Command(resume="yes") on the same thread. This is the part the docs state plainly and most tutorials skip: the runtime restarts the entire node from the beginning. It does not resume from the exact line where interrupt was called. So the print fires again, execution reaches interrupt() again, and this time the call does not pause. The value you passed in Command(resume=...) becomes its return value. The node carries on to the second print and returns.

My read on why it works this way: the alternative is freezing a Python call stack mid-function, and nothing serializes that safely across a restart, a different worker, or a deploy. Saving state at a boundary and replaying is what lets a pause be resumed by something other than the exact process that paused. The docs describe the model in one sentence: execution returns to a checkpoint boundary, and the workflow replays forward until it reaches the pause again.

Three consequences follow, and each one is easy to trip over.

Anything before interrupt() must be safe to run twice. The docs say side effects called before interrupt must be idempotent, and their own example is the database row: inserting one before the pause inserts it again on resume. Put the write after the interrupt, or make it an upsert, or use an idempotency key. Do not put the email send on line one and the approval on line three.

No checkpointer, no resume. Run the snippet above without checkpointer=InMemorySaver() and the first call still returns __interrupt__, which is misleading, because there is nowhere to resume from. The second call raises RuntimeError: Cannot use Command(resume=...) without checkpointer. The docs list a checkpointer as requirement number one for interrupt, and the error message is the runtime agreeing.

Order matters if a node interrupts more than once. Resume values are matched to interrupt calls strictly by index, in the order the calls happen during the replay. A node with two interrupts takes two resumes, and the first resume value lands in the first call. Add a conditional interrupt before an existing one and the wrong answer lands in the wrong call. Keep interrupts in a node few and their order fixed.

Durability: what survives a crash

If checkpoints are save points, the next question is how often the game saves. LangGraph exposes that as a durability argument on invoke and stream, with three modes. Leave it unset and 1.2.12 uses "async"; the default is in langgraph/pregel/main.py, the docs page names none.

"exit" persists only when the run exits, whether by finishing, erroring, or hitting an interrupt. Fastest, and intermediate state is not saved, so a process crash mid-run loses everything since the start. "async" writes each checkpoint while the next step is already executing, which the docs describe as good performance and durability with a small risk of a missed checkpoint if the process dies at the wrong moment. "sync" writes every checkpoint before the next step starts, highest durability, some overhead.

You can see the difference without a crash. The three-node graph from earlier leaves four snapshots in history under "async" or "sync", and exactly one under "exit". That one is the finished state. Time travel to the point before b is impossible on that thread, because the save point was never written.

My view: "exit" is right for a cheap graph where a retry from scratch costs nothing, and wrong the moment a node does anything you would not want repeated. And no mode helps if the checkpointer itself is InMemorySaver. It keeps checkpoints in RAM, and the docs say it plainly: when the process restarts, all checkpoints are lost. Most examples, the docs' own included, use it because it needs no setup. Production needs PostgresSaver or at least SqliteSaver, or your carefully durable checkpoints evaporate on the next deploy.

Subgraphs and time travel are the same save point

The two features that look most separate are the ones that fall out of the model most directly.

Subgraphs. Compile a subgraph without a checkpointer and add it as a node in a parent that has one, and the subgraph inherits the parent's checkpointer. An interrupt() inside the subgraph surfaces at the parent's invoke exactly like a top-level one, and Command(resume=...) on the parent resumes it. Its checkpoints live in the same store under a namespace of the form child_node_name:uuid, nested subgraphs joined with |. You can read them with get_state(config, subgraphs=True).

Where the save points sit decides what replays. From the parent's history the whole subgraph is one step: START -> child -> END leaves three snapshots however many nodes the child has. The child's own save points live under its namespace, so when a two-node subgraph pauses in its second node, resume re-runs only that node; the first one's result was already saved. checkpointer=True drops the uuid from the namespace and, per the docs, makes the subgraph per-thread: its state carries across calls on the same thread. checkpointer=False saves nothing inside. The docs say no interrupts or durable execution. On 1.2.12 it was one step short of that, and it is the place the docs and the run disagree: the interrupt still surfaced, the resume still completed, but both nodes ran again, the first one included, because there was no save point inside to replay from. Same rule, one level up: everything in a stateless subgraph before the pause must be safe to run twice.

Time travel. This is loading an older save. Pull the history, pick the snapshot whose next is the node you want to re-run, call update_state on that snapshot's config with new values, then invoke(None, fork_config). On the three-node graph, forking the snapshot before b with n=100 returns n=1000, because b multiplies by ten and never knew the original run existed. The fork is a new branch on the same thread; the original checkpoints stay in history, which is what makes this safe for debugging a bad agent turn.

Seen this way, human-in-the-loop, fault tolerance, subgraph state and time travel are one feature with four names. The runtime writes a snapshot at every boundary, and resuming from any snapshot replays forward.

How to actually use this

A short checklist, in the order the mistakes usually happen:

  1. Decide what a thread_id means in your app before writing a node. One per conversation is the usual answer. One per user is a store's job, not a checkpointer's.
  2. In any node that calls interrupt(), move side effects below it or make them idempotent. Read the node top to bottom and ask what happens if this runs twice.
  3. Pick durability on purpose per graph instead of taking the "async" default. Long, expensive, or side-effecting runs want "sync"; cheap replayable ones can take "exit".
  4. Swap InMemorySaver for a persistent saver before the first deploy, not after the first lost approval.
  5. For each subgraph, choose: inherit (default), carry state across calls on the thread (checkpointer=True), or stateless (checkpointer=False). Stateless means the whole subgraph replays on resume, so rule 2 applies to every node in it.

If you are working through this stage yourself and want to talk through any of it live, I keep a slot open on Topmate.


More from me: the notes this came out of are at notes.priyanshsinghal.com, part of the full roadmap at notes.priyanshsinghal.com, and I post shorter breakdowns of what I'm learning on LinkedIn.

Top comments (0)