DEV Community

Cover image for Building a Ride-Share Zone-Balancing Agent with LangGraph — Part 4: Letting a Human Step In
Ebrahim Arian
Ebrahim Arian

Posted on

Building a Ride-Share Zone-Balancing Agent with LangGraph — Part 4: Letting a Human Step In

This is Part 4 of a 5-part series. Part 3 gave the agent memory. It can now run for hours on its own, cycle after cycle, without forgetting what it already tried.

That's exactly the problem. A genuinely severe deficit might call for an aggressive surge multiplier, or a costly driver bonus. Right now, the agent applies whatever it decides immediately. Nobody has looked at it first. Nothing pauses for a second opinion, no matter how expensive the call.

Part 4 adds that pause. It uses LangGraph's dynamic interrupt():

  • A node calls interrupt(payload) — but only conditionally, when it actually wants to pause.
  • Execution stops at that exact point.
  • app.invoke(Command(resume=answer), config=thread) continues from there. answer becomes interrupt()'s return value, right at the call site.

Two gated pauses, not one:

  • Early — request_data_edit. Gated on cycle_number == 1. A human validates or corrects the raw starting snapshot once, before detect_imbalance even runs. Later cycles use the agent's own carried-forward numbers, not a fresh unverified reading — so re-confirming every cycle would just be noise.
  • Late — request_approval. Gated on severity == "critical". A wrong automatic call is most expensive in a critical deficit, so that's where a human checkpoint earns its cost. Mild, moderate, balanced, and surplus cycles still run fully autonomously, exactly like Part 3.

One property worth naming: no new fields got added to the state. Every interrupt here only reads state Part 3 already has, and writes back a field Part 3 already has — zone, recommended_policy, explanation. Human review is a new control-flow capability. It's not new data.

The Graph

start_cycle ──▶ apply_scheduled_conditions ──▶ request_data_edit ──▶ detect_imbalance ──▶ classify_severity ──▶ set_candidates
   (gate: cycle_number == 1)                                                                                          │
                                                                                                           ┌──────────┴──────────┐
                                                                                                           ▼ (balanced)          ▼ (deficit/surplus)
                                                                                                   trivial_do_nothing      reconcile_inputs  ← LLM #1
                                                                                                           │                     ▼
                                                                                                           │              resolved_imbalance
                                                                                                           │                     ▼
                                                                                                           │              choose_best_policy
                                                                                                           │                     ▼
                                                                                                           │              generate_explanation ← LLM #2
                                                                                                           └──────────┬──────────┘
                                                                                                                       ▼
                                                                                                             request_approval
                                                                                                        (gate: severity == "critical")
                                                                                                                       ▼
                                                                                                              simulate_and_report
Enter fullscreen mode Exit fullscreen mode

Everything from detect_imbalance through generate_explanation is unchanged from Part 3 — same nodes, reused directly. Two new nodes bracket that core: request_data_edit right after the cycle starts, request_approval right before the outcome is simulated.

Both are plain Python. And both return {} on any cycle where their gate doesn't apply — so an ungated cycle looks exactly like Part 3.

The Nodes That Pause

Each one is a plain function with an if guard at the top. No LangGraph machinery decides whether to pause. The node itself does, then calls interrupt(payload) when it wants to:

def request_data_edit(state: AgentState) -> AgentState:
    if state["cycle_number"] != 1:
        return {}

    zone = state["zone"]
    answer = interrupt({
        "kind": "data_edit",
        "question": (
            f"Review the starting snapshot for {zone['zone_name']}. "
            "Resume with {'corrections': {...}} to fix fields, or {'corrections': {}} to accept it."
        ),
        "zone": zone,
    })
    corrections = answer.get("corrections", {})
    if not corrections:
        return {}
    return {"zone": {**zone, **corrections}}
Enter fullscreen mode Exit fullscreen mode

interrupt()'s return value becomes whatever gets passed to Command(resume=...) — here, a corrections dict.

One real gotcha, worth knowing before you use this: Command(resume={}) is falsy in Python. LangGraph treats a falsy resume as if no answer was given at all. It just re-fires the same interrupt, instead of continuing. Always resume with a non-empty dict — {"corrections": {}} is how you accept the snapshot as-is.

The late pause has the same shape, but a real decision to make:

def request_approval(state: AgentState) -> AgentState:
    if state["severity"] != "critical":
        return {}

    answer = interrupt({
        "kind": "approval",
        "question": (
            f"Critical imbalance in {state['zone']['zone_name']} (cycle {state['cycle_number']}). "
            "Approve, reject, or override the recommended policy."
        ),
        "zone_name": state["zone"]["zone_name"],
        "cycle": state["cycle_number"],
        "imbalance_ratio": state["imbalance_ratio"],
        "recommended_policy": state["recommended_policy"],
        "explanation": state["explanation"],
        "policy_evaluations": state["policy_evaluations"],
        "candidate_policies": state["candidate_policies"],
    })
    action = (answer or {}).get("action", "approve")

    if action == "reject":
        return {
            "recommended_policy": "do_nothing",
            "explanation": "Rejected by human reviewer — falling back to do_nothing.",
        }
    if action == "override":
        chosen = answer["policy"]
        return {
            "recommended_policy": chosen,
            "explanation": f"Overridden by human reviewer: {chosen} chosen instead of "
                            f"{state['recommended_policy']}.",
        }
    return {}  # approve — no change
Enter fullscreen mode Exit fullscreen mode

Three possible resume values:

  • {"action": "approve"} — keep the recommended policy.
  • {"action": "reject"} — fall back to do_nothing.
  • {"action": "override", "policy": "<name>"} — force a specific policy.

Building It

Same graph as Part 3, with these two nodes inserted at the right points. It's wrapped in a build(checkpointer) function, rather than built inline once — the next section needs a second graph object sharing the exact same checkpointer. The structure never changes. Only whether it's the same Python object does:

def build(checkpointer):
    llm = ChatOllama(model="qwen2.5:14b", temperature=0)
    llm_with_reconcile_tool = llm.bind_tools([report_context_and_schedule])

    def _reconcile(state):
        return reconcile_inputs(state, llm_with_reconcile_tool)

    def _explain(state):
        return generate_explanation(state, llm)

    g = StateGraph(AgentState)
    g.add_node("start_cycle",                start_cycle)
    g.add_node("apply_scheduled_conditions", apply_scheduled_conditions)
    g.add_node("request_data_edit",          request_data_edit)
    g.add_node("detect_imbalance",           detect_imbalance)
    g.add_node("classify_severity",          classify_severity)
    g.add_node("set_candidates",             set_candidates)
    g.add_node("trivial_do_nothing",         trivial_do_nothing)
    g.add_node("reconcile_inputs",           _reconcile)
    g.add_node("resolved_imbalance",         resolved_imbalance)
    g.add_node("choose_best_policy",         choose_best_policy)
    g.add_node("generate_explanation",       _explain)
    g.add_node("request_approval",           request_approval)
    g.add_node("simulate_and_report",        simulate_and_report)

    g.add_edge(START, "start_cycle")
    g.add_edge("start_cycle", "apply_scheduled_conditions")
    g.add_edge("apply_scheduled_conditions", "request_data_edit")
    g.add_edge("request_data_edit", "detect_imbalance")
    g.add_edge("detect_imbalance",  "classify_severity")
    g.add_edge("classify_severity", "set_candidates")
    g.add_conditional_edges("set_candidates", route_llm_or_skip, {
        "trivial_do_nothing": "trivial_do_nothing",
        "reconcile_inputs":   "reconcile_inputs",
    })
    g.add_edge("reconcile_inputs",     "resolved_imbalance")
    g.add_edge("resolved_imbalance",   "choose_best_policy")
    g.add_edge("choose_best_policy",   "generate_explanation")
    g.add_edge("generate_explanation", "request_approval")
    g.add_edge("trivial_do_nothing",   "request_approval")
    g.add_edge("request_approval",     "simulate_and_report")
    g.add_edge("simulate_and_report",  END)

    return g.compile(checkpointer=checkpointer)


memory = MemorySaver()
app = build(memory)
Enter fullscreen mode Exit fullscreen mode

Correcting the Starting Snapshot

Cycle 1 always pauses at request_data_edit. Downtown Core's raw snapshot understates its driver count here — it's recorded as 9, but it was actually 15.

Reaching the pause:

downtown = get_zone(zones, DEFAULT_ZONE_NAME, driver_count=9)
r1 = app.invoke(make_initial_state(downtown), config=thread_edit)
Enter fullscreen mode Exit fullscreen mode
PAUSED — kind=data_edit
question: Review the starting snapshot for Downtown Core. Resume with {'corrections': {...}} to fix fields, or {'corrections': {}} to accept it.
zone as recorded: driver_count=9, rider_request_count=16
uncorrected ratio: 1.78
Enter fullscreen mode Exit fullscreen mode

At 1.78, that's a deficit. But it's a wrong one — a data problem, not a real supply problem. Correcting it and resuming:

MY_CORRECTION = {"driver_count": 15}
r2 = app.invoke(Command(resume={"corrections": MY_CORRECTION}), config=thread_edit)
Enter fullscreen mode Exit fullscreen mode
(no interrupt — cycle ran straight through)
ratio now: 1.07 (balanced)
[Cycle 1] [Downtown Core] ratio=1.07 | severity=none | policy=do_nothing | wait 3.3min → 4.6min | resolved=N/A
Enter fullscreen mode Exit fullscreen mode

One correction, and the zone flips from deficit to balanced. That's the entire point of this pause — catching a bad reading before it drives a real decision.

request_data_edit is gated on cycle_number == 1, so it doesn't fire again on cycle 2 of the same thread. The edit gate only ever applies once per thread.

Reviewing a Critical Decision

request_approval only fires when severity == "critical" — a ratio bad enough that the wrong automatic call is expensive.

Reaching the pause, on a zone with 4 drivers against 80 rider requests:

PAUSED — kind=approval
question: Critical imbalance in Airport (cycle 1). Approve, reject, or override the recommended policy.
severity=critical | recommended=surge_pricing | candidates=['surge_pricing', 'driver_bonus', 'demand_redirect', 'do_nothing']
explanation: The surge pricing policy was selected for the Airport zone in Cycle 1 because it generated the highest profit of $142.93, even though it did not resolve the imbalance.
Enter fullscreen mode Exit fullscreen mode

Approving and resuming — but against a different graph object this time:

MY_DECISION = {"action": "approve"}
app_resume = build(memory)   # a different graph object, sharing only `memory`
r3 = app_resume.invoke(Command(resume=MY_DECISION), config=thread_approve)
Enter fullscreen mode Exit fullscreen mode
final policy: surge_pricing
[Cycle 1] [Airport] ratio=20.0 | severity=critical | policy=surge_pricing | wait 5.6min → 13.3min | resolved=NO
Enter fullscreen mode Exit fullscreen mode

That app_resume detail is worth pausing on. It's built fresh from the same build() function, and it only shares one thing with the original app: the MemorySaver. The paused state isn't living inside the app Python object — it's living in the checkpointer. Any graph object built the same way, sharing that checkpointer, can pick the pause back up.

A Third Kind of Pause: Debugging

request_data_edit and request_approval are both gated. They only pause under a specific condition, and the resume value changes what happens next.

A debugging pause is different in kind:

  • It's unconditional — it always pauses, whenever it's wired in.
  • It's purely observational — the resume value is discarded, and it never changes state.
def debug_checkpoint(state: AgentState) -> AgentState:
    interrupt({
        "kind": "debug",
        "question": "Inspect state before the approval gate. Resume with anything to continue.",
        "state_snapshot": {
            "cycle_number": state["cycle_number"],
            "zone": state["zone"],
            "imbalance_ratio": state["imbalance_ratio"],
            "severity": state["severity"],
            "candidate_policies": state.get("candidate_policies"),
            "policy_evaluations": state.get("policy_evaluations"),
            "recommended_policy": state.get("recommended_policy"),
            "explanation": state.get("explanation"),
        },
    })
    return {}
Enter fullscreen mode Exit fullscreen mode

It exists to inspect the full in-flight state right before the approval gate acts on it — not to make or revise a decision. It's opt-in via build_graph(debug_mode=True), kept out of the default graph so it doesn't add a pause to every single run.

What We Built

Part 3 Part 4
Scope A sequence of cycles on one zone Same, plus two review checkpoints per thread
Control flow Fully autonomous, no pauses interrupt() pauses execution; Command(resume=...) continues it
Starting snapshot Trusted as generated Reviewable once, on cycle 1
Critical decisions Applied automatically Held for approve / reject / override — your call, not a script
State Unchanged — no new fields

The decision core still hasn't changed. choose_best_policy is exactly Part 1's function, called exactly the same way. What Part 4 adds is a new way for a human to sit inside the loop, instead of only reading a report after the fact — a real pause, resumable from a different Python object entirely, gated so it only interrupts when it's actually worth someone's attention.

What's still missing: every zone here is still evaluated alone. A policy that pulls a driver from a neighboring zone has been mentioned since Part 1, but never actually built. That's where Part 5 picks up.


Code for this series: github.com/ebiarian/zone-balancing-ridesharing-langgraph-agent

Next — Part 5: coordinating two zones at once, so a driver pulled into one zone is a decision that actually accounts for both.

Top comments (0)