DEV Community

siddhesh kabra
siddhesh kabra

Posted on Originally published at aidevinsider.com

LangGraph Tutorial: 5 Proven Steps to Fix a Fragile Agent

LangGraph Tutorial: 5 Proven Steps to Fix a Fragile Agent

I had an agent that worked in the demo and fell over in production. It was a linear prompt chain, and the moment a tool call failed or a user came back an hour later, it lost the thread. I rewrote it as a LangGraph StateGraph over an afternoon, and this LangGraph tutorial is the result, and the class of bugs that had been eating my week just stopped happening. This LangGraph tutorial walks that rewrite end to end: the mental model, working tool-calling code, memory with a checkpointer, and the failure modes that cost me time so they don't cost you any.

I am assuming you are comfortable with Python and have touched LangChain once. You do not need prior LangGraph knowledge.

Table of contents

What LangGraph is, explained without the hype

LangGraph is a low-level orchestration framework for building stateful, long-running agents. Instead of modelling your app as a pipeline (do this, then this, then this), it models it as a directed graph: each step is a node, the connections are edges, and all shared data flows through one central state object. That one design choice is why loops, conditional branching, retries, and multi-step memory stop being things you bolt on and start being things the structure gives you.

Its GitHub repo passed 30,000 stars and it is among the most active agent frameworks in 2026. You can use it with LangChain or entirely on its own. The reason it matters: a prompt chain has no memory of where it is, so a failed tool call or a resumed conversation breaks it. A graph carries explicit state, so it can loop back, retry a failed step, and then pick up exactly where it stopped.

The graph mental model: state, nodes, edges

Four primitives carry the whole model (state, nodes, edges, and the checkpointer you will add for memory):

  • State is a typed dictionary that every node reads from and writes to. It is the single source of truth for the run.

  • A node is a function. It takes the state, does something (calls the LLM, runs a tool), and returns an update to the state.

  • An edge connects nodes. A normal edge always goes A to B. A conditional edge picks the next node based on the current state, and that is what gives you branching and loops.

  • A checkpointer persists the state per thread, which is what turns a one-shot run into a conversation that survives restarts.

Primitive
What it is
Example

State
Typed shared data
messages, iteration count

Node
A function that acts
reasoning_node, tool_node

Normal edge
Always A to B
START -> reasoning

Conditional edge
Routes on state
reasoning to tools OR output

Checkpointer
Persists state per thread
SqliteSaver

Step 1-5: build a tool-calling agent

Below is a complete tool-calling loop, the pattern you will reuse most. It reasons, decides whether to call a tool, runs it, loops back with the result, and exits when the model has an answer.

Step 1, define the state. A TypedDict with the message list and a safety counter:

from typing import Annotated, TypedDict
from langgraph.graph.message import add_messages

class AgentState(TypedDict):
    messages: Annotated[list, add_messages]  # add_messages appends, never overwrites
    iterations: int

Enter fullscreen mode Exit fullscreen mode

The add_messages reducer is the small detail that matters: it appends new messages to the list instead of replacing it, so history accumulates correctly across loops.

Step 2, write the nodes. Each is a plain function returning a state update:

def reasoning_node(state: AgentState):
    response = llm_with_tools.invoke(state["messages"])
    return {"messages": [response], "iterations": state["iterations"] + 1}

def tool_node(state: AgentState):
    # execute whatever tool the last message asked for, append the result
    results = run_pending_tools(state["messages"][-1])
    return {"messages": results}

Enter fullscreen mode Exit fullscreen mode

Step 3, the router. A conditional edge needs a function that returns the name of the next node:

MAX_ITERATIONS = 10

def should_continue(state: AgentState) -> str:
    last = state["messages"][-1]
    if state["iterations"] >= MAX_ITERATIONS:
        return "output"                      # safety valve, never loop forever
    if getattr(last, "tool_calls", None):
        return "tools"                       # model wants a tool
    return "output"                          # model gave a final answer

Enter fullscreen mode Exit fullscreen mode

Step 4, wire the graph. This is where nodes and edges become a runnable thing:

from langgraph.graph import StateGraph, START, END

graph = StateGraph(AgentState)
graph.add_node("reasoning", reasoning_node)
graph.add_node("tools", tool_node)
graph.add_edge(START, "reasoning")
graph.add_conditional_edges("reasoning", should_continue,
                            {"tools": "tools", "output": END})
graph.add_edge("tools", "reasoning")         # loop back after a tool runs

Enter fullscreen mode Exit fullscreen mode

Step 5, compile and run. compile() turns the builder into an executable graph:

app = graph.compile()
result = app.invoke({"messages": [("user", "What's the weather in Pune?")],
                     "iterations": 0})

Enter fullscreen mode Exit fullscreen mode

That is a working agent. The loop reasoning -> tools -> reasoning runs until should_continue routes to END, and the iteration cap means a confused model degrades to a normal answer instead of spinning.

Adding memory with a checkpointer

The single upgrade that fixed my production problem: a checkpointer. It persists the full state after every step, keyed by a thread id, so a conversation survives a crash, a restart, or a user returning an hour later.

from langgraph.checkpoint.sqlite import SqliteSaver

memory = SqliteSaver.from_conn_string("checkpoints.db")
app = graph.compile(checkpointer=memory)

# every call for the same thread_id resumes that conversation's state
config = {"configurable": {"thread_id": "user-42"}}
app.invoke({"messages": [("user", "and tomorrow?")], "iterations": 0}, config)

Enter fullscreen mode Exit fullscreen mode

Because state is persisted per thread, "and tomorrow?" works: the graph reloads the prior messages for user-42 and continues. No manual session store, no re-sending the whole history yourself. This is also the foundation for human-in-the-loop approval, where the graph pauses at an approval node and resumes when you send a decision.

LangGraph vs a plain prompt chain: when it's worth it

Honestly, LangGraph is not always the right tool, and reaching for it too early is its own mistake. A quick way to decide:

  • Use a plain chain when the flow is genuinely linear and stateless: one prompt, one answer, no tools, no memory. The graph is overhead you don't need.

  • Use LangGraph the moment you have any of these: a tool-calling loop; branching on what the model decides; retries; or memory that must survive a restart. That is exactly where chains turn fragile.

This LangGraph tutorial's rewrite paid off because my agent had all four. If yours is a single summarisation call, skip the graph.

LangGraph best practices and setup notes

LangGraph setup is one install: pip install langgraph langgraph-checkpoint-sqlite. From there, the practices that saved me the most:

  • Always cap your loops. A MAX_ITERATIONS counter in state is not optional. Without it, a model that keeps requesting tools will run until your API bill notices.

  • Keep nodes small and single-purpose. One node reasons, one runs tools, one formats output. Fat nodes that do three things are the hardest to debug because you cannot see which step wrote which state.

  • Let the state be the only shared data. Do not stash things in globals between nodes. If a later node needs it, put it in the state. This is what makes the checkpointer able to fully restore a run.

  • Compile once, invoke many. Building the graph is setup; do it at startup, not per request.

The failure modes that cost me time

Every hour I lost came from one of these:

  • Overwriting messages instead of appending. If you return {"messages": [response]} without the add_messages reducer on the field, you clobber history. Annotate the field.

  • A conditional edge with no exit. If should_continue can never return the END branch, you built an infinite loop. The iteration cap is your seatbelt.

  • Forgetting the checkpointer needs a thread_id. No configurable.thread_id in the config means no memory, silently. The call works; it just never resumes.

  • Reaching for LangGraph on a linear task. If there is no loop, branch, or memory, the graph is pure ceremony.

FAQ

What is LangGraph in one sentence?
A low-level framework for building stateful AI agents as a directed graph of nodes (actions) and edges (transitions) that share one typed state object, which makes loops, branching, and memory natural.

Do I need LangChain to use LangGraph?
No. LangGraph works standalone, though the two are often used together because LangChain provides convenient model and tool wrappers.

How does LangGraph remember conversations?
Through a checkpointer (for example SqliteSaver) that persists the full state after each step, keyed by a thread_id. Any later call with the same thread id resumes that conversation's state.

When should I NOT use LangGraph?
When your workflow is a single, linear, stateless call. If there is no tool loop, no branching, and no memory to preserve, a plain chain is simpler and the graph is just overhead.

Related reading: Your coding agent has a kill switch in its permissions, and mine was off, What an LLM gateway actually does for a coding agent, Claude Code Hooks: 7 Fixes for a Painful Setup.

Related reading: PSSA: 7 Proven Reasons This Rust LM Is Painful to Ignore.

Sources: LangGraph official quickstart and Graph API docs; Real Python: Build Stateful AI Agents with LangGraph; Hostinger LangGraph tutorial (approval + streaming patterns). Code reflects a tool-calling agent I rebuilt from a prompt chain on my own machine in 2026; the checkpointer fix is from that migration.


Originally published at aidevinsider.com.

Top comments (0)