Как проверить конфликты состояния в мультиагентном LangGraph.js: a reproducible TypeScript test | Agent Lab Journal
Agent Lab Journal
Home
Guides
Glossary
Как проверить конфликты состояния в мультиагентном LangGraph.js
Level: advanced · Reading time: 45 min · Updated: 8 October 2026
You run three agents in parallel and the final state says the critic approved publication. Or it says the researcher found one source when it found two. Or the same note shows up three times. Nothing crashed. In a LangGraph.js multi-agent system, this is usually not the model's fault. Parallel nodes write to a shared state channel, and the channel's merge semantics decide which updates survive. This guide shows how to make those conflicts visible with a TypeScript test you can run yourself. It then replaces the fragile channels with three mechanisms that the same suite checks: commutative reducers, explicit channel ownership and deterministic routing.
Evidence boundary. We wrote the code and tests in this article to match the documented behavior of @langchain/langgraph (the Annotation, StateGraph and Send APIs). We did not include run output, because we are not publishing results we did not capture alongside this text. Treat every expect(...) as a specification: if your installed version behaves differently, the failing assertion is the finding. Record your versions with npm ls before you draw conclusions.
1. The problem in one superstep
LangGraph runs a graph in supersteps. Every node scheduled for the same step runs concurrently, and their returned updates are applied to the state together when the step ends. Each state key is a channel, and the channel's type decides what happens when more than one update arrives in the same step:
No reducer (Annotation()): a last-value channel. It accepts one write per step. A second concurrent write raises an error. LangGraph.js reports it as an invalid concurrent graph update and points you to a reducer.
With a reducer (Annotation({ reducer, default })): every update is folded into the current value with reducer(current, update). The step never fails, so any information the reducer throws away is lost silently.
That gives you three typical failure modes, from loudest to quietest:
Loud: two agents write to a last-value channel in parallel, and the run throws.
Silent loss: someone "fixes" the error with reducer: (_old, next) => next. One agent's output disappears, and which one depends on application order.
Silent duplication: a node reads the list, appends to it and returns the whole list to an append reducer. This read-modify-write pattern multiplies existing entries once per parallel writer.
All three are a race condition in the general sense: the result depends on how concurrent writes are combined, not only on what each agent did. You can't catch them by reading prompts. You catch them by testing the merge semantics directly.
2. The concrete case
We use a small editorial pipeline as the running example. A planner fans out tasks to three agents that run in parallel:
researcher records sources and votes on a shared verdict;
critic records risks and votes on the same shared verdict;
editor writes two successive versions of a headline.
Our requirements:
Concurrent writes never fail randomly and never disappear silently.
An agent can write only to keys it owns, plus an explicit shared: namespace.
When two agents disagree at the same version, the state records a conflict. A last-write-wins rule must not hide it.
The final state is byte-identical however the tasks are ordered on input and however long each agent takes.
The fixture data below is invented for the test. It is not output from a real model.
3. Project setup
mkdir langgraph-state-conflicts && cd langgraph-state-conflicts
npm init -y
npm pkg set type=module
npm pkg set scripts.test="vitest run" scripts.typecheck="tsc"
npm install @langchain/langgraph @langchain/core
npm install -D typescript vitest tsx @types/node
# Record exactly what you tested against
npm ls @langchain/langgraph @langchain/core
tsconfig.json:
{
"compilerOptions": {
"target": "ES2022",
"module": "NodeNext",
"moduleResolution": "NodeNext",
"strict": true,
"skipLibCheck": true,
"noEmit": true
},
"include": ["src", "tests", "scripts"]
}
Layout:
src/ledger.ts # pure merge logic: no LangGraph import
src/graphs.ts # three broken graphs and one safe graph
tests/state-conflicts.test.ts
scripts/trace.ts # optional: print per-node updates
The merge logic stays in a file with no framework dependency. That way you can test its algebraic properties exhaustively in milliseconds, apart from graph execution.
4. A conflict-preserving reducer
The core idea: don't store "the value" of a shared key. Store every distinct candidate for each key, sorted by a total order. The winner is derived as the first element. A conflict is a tie at the highest version. Set union with a fixed sort order is commutative, associative and idempotent. The reducer's result therefore can't depend on the order in which LangGraph applies the updates, and re-applying the same update changes nothing (idempotency).
src/ledger.ts:
export type AgentId = "researcher" | "critic" | "editor";
export interface Finding {
id: string; // "<owner>:<name>" or "shared:<name>"
owner: AgentId; // stamped by the worker, never by the model
version: number;
text: string;
}
// Every id maps to a sorted, de-duplicated list of candidates.
export type Ledger = Record<string, Finding[]>;
export function compareFindings(a: Finding, b: Finding): number {
if (a.version !== b.version) return b.version - a.version; // newest first
if (a.owner !== b.owner) return a.owner < b.owner ? -1 : 1;
if (a.text !== b.text) return a.text < b.text ? -1 : 1;
return 0;
}
export function assertOwnership(f: Finding): void {
const ns = f.id.split(":", 1)[0];
if (ns !== f.owner && ns !== "shared") {
throw new Error(`Ownership violation: ${f.owner} wrote ${f.id}`);
}
}
export function mergeLedger(left: Ledger, right: Ledger): Ledger {
for (const [id, list] of Object.entries(right)) {
for (const f of list) {
if (f.id !== id) throw new Error(`Key mismatch: ${id} vs ${f.id}`);
assertOwnership(f);
}
}
const ids = [...new Set([...Object.keys(left), ...Object.keys(right)])].sort();
const out: Ledger = {};
for (const id of ids) {
const all = [...(left[id] ?? []), ...(right[id] ?? [])].sort(compareFindings);
out[id] = all.filter((f, i) => i === 0 || compareFindings(all[i - 1], f) !== 0);
}
return out;
}
export const winner = (l: Ledger, id: string): Finding | undefined => l[id]?.[0];
// A conflict is a tie at the highest version, not ordinary version history.
export const conflicts = (l: Ledger): string[] =>
Object.keys(l).filter(
(id) => l[id].length > 1 && l[id][0].version === l[id][1].version,
);
Design decisions worth calling out:
Sorted keys in the output object. JSON.stringify follows insertion order. Rebuilding the object with sorted ids makes snapshots comparable byte for byte.
No Date.now(), no randomness, no I/O in the reducer. A reducer must be a pure function. A timestamp tie-break looks deterministic but isn't.
The tie-break by owner name is arbitrary on purpose. It exists only so that the derived winner is stable. The conflict itself is still reported through conflicts(), and a person or policy should resolve it.
The reducer throws on ownership violations. Failing the step is better than merging a write that shouldn't exist.
5. Three broken graphs and one safe graph
src/graphs.ts. The broken graphs are deliberately minimal, so each test isolates exactly one failure mode.
import { Annotation, END, Send, START, StateGraph } from "@langchain/langgraph";
import { type AgentId, type Finding, type Ledger, mergeLedger } from "./ledger.js";
/* ---------- 1. Loud: two writers, one last-value channel ---------- */
export const NaiveState = Annotation.Root({
topic: Annotation<string>(),
summary: Annotation<string>(),
});
export const buildNaive = () =>
new StateGraph(NaiveState)
.addNode("researcher", async () => ({ summary: "from researcher" }))
.addNode("critic", async () => ({ summary: "from critic" }))
.addEdge(START, "researcher")
.addEdge(START, "critic")
.addEdge("researcher", END)
.addEdge("critic", END)
.compile();
/* ---------- 2. Silent loss: a reducer that hides the conflict ---------- */
export const LossyState = Annotation.Root({
topic: Annotation<string>(),
summary: Annotation<string>({ reducer: (_old, next) => next, default: () => "" }),
});
export const buildLossy = () =>
new StateGraph(LossyState)
.addNode("researcher", async () => ({ summary: "from researcher" }))
.addNode("critic", async () => ({ summary: "from critic" }))
.addEdge(START, "researcher")
.addEdge(START, "critic")
.addEdge("researcher", END)
.addEdge("critic", END)
.compile();
/* ---------- 3. Silent duplication: read-modify-write into an append reducer ---------- */
export const NotesState = Annotation.Root({
notes: Annotation<string[]>({ reducer: (a, b) => a.concat(b), default: () => [] }),
});
export const buildReadModifyWrite = () =>
new StateGraph(NotesState)
.addNode("seed", async () => ({ notes: ["seed"] }))
// Bug: returns the whole list instead of a delta.
.addNode("researcher", async (s: typeof NotesState.State) => ({ notes: [...s.notes, "r"] }))
.addNode("critic", async (s: typeof NotesState.State) => ({ notes: [...s.notes, "c"] }))
.addEdge(START, "seed")
.addEdge("seed", "researcher")
.addEdge("seed", "critic")
.addEdge("researcher", END)
.addEdge("critic", END)
.compile();
/* ---------- 4. Safe: owned channels, ledger reducer, deterministic routing ---------- */
export interface Task {
taskId: string;
agent: AgentId;
delayMs: number; // test hook to shuffle completion order
writes: Array<Omit<Finding, "owner">>; // the agent cannot choose its owner
}
const sortedUnion = (a: string[], b: string[]) => [...new Set([...a, ...b])].sort();
export const SafeState = Annotation.Root({
tasks: Annotation<Task[]>(), // single writer (graph input): last-value is a guard
ledger: Annotation<Ledger>({ reducer: mergeLedger, default: () => ({}) }),
log: Annotation<string[]>({ reducer: sortedUnion, default: () => [] }),
});
// Node-level ownership: a node may only return the channels it owns.
export function owned<I, O extends Record<string, unknown>>(
node: string,
keys: readonly string[],
fn: (input: I) => Promise<O>,
) {
return async (input: I): Promise<O> => {
const out = await fn(input);
const extra = Object.keys(out).filter((k) => !keys.includes(k));
if (extra.length) throw new Error(`${node} wrote unowned channel(s): ${extra.join(", ")}`);
return out;
};
}
async function planner(s: typeof SafeState.State) {
const ids = s.tasks.map((t) => t.taskId);
if (new Set(ids).size !== ids.length) throw new Error("Duplicate taskId in input");
return { log: [`planner:${s.tasks.length}`] };
}
export function routeTasks(state: typeof SafeState.State): Send[] {
return [...state.tasks]
.sort((a, b) => (a.taskId < b.taskId ? -1 : a.taskId > b.taskId ? 1 : 0))
.map((t) => new Send("worker", t));
}
async function worker(task: Task) {
await new Promise((r) => setTimeout(r, task.delayMs));
const update: Ledger = {};
for (const w of task.writes) {
(update[w.id] ??= []).push({ ...w, owner: task.agent }); // owner stamped here
}
return { ledger: update, log: [`${task.agent}:${task.taskId}`] };
}
export const buildSafe = () =>
new StateGraph(SafeState)
.addNode("planner", owned("planner", ["log"], planner))
.addNode("worker", owned("worker", ["ledger", "log"], worker))
.addEdge(START, "planner")
.addConditionalEdges("planner", routeTasks, ["worker"])
.addEdge("worker", END)
.compile();
Here is how the safe graph maps onto the requirements:
ChannelWritersMerge semanticsWhy
tasksgraph input onlylast valueA second writer would be a bug. The built-in error acts as an ownership assertion.
ledgerall workersmergeLedger: union, total order, conflicts keptConcurrent writes are expected, and none may be lost.
logplanner, workerssorted set unionOrder-independent audit trail.
There are three layers of ownership. The channel layer lets only reducer channels accept concurrent writes. The node layer (owned(...)) stops a node from returning keys it doesn't own. The key layer (assertOwnership) stops an agent from writing into another agent's namespace. The worker stamps owner from the task and doesn't take it from model output, so an agent can't claim to be someone else.
Routing is deterministic because routeTasks builds the Send list from state alone and sorts it by a stable key. An LLM-chosen route or the input order of tasks can't change it. Don't assume a sorted Send list also fixes the order in which updates are applied. The safe design doesn't depend on that order, and that is what the tests verify.
6. The reproducible test suite
tests/state-conflicts.test.ts:
import { describe, expect, it } from "vitest";
import {
buildLossy, buildNaive, buildReadModifyWrite, buildSafe, owned, type Task,
} from "../src/graphs.js";
import { conflicts, mergeLedger, winner, type Ledger } from "../src/ledger.js";
/* ---------- fixtures (invented data, not model output) ---------- */
const fixture: Task[] = [
{ taskId: "t1", agent: "researcher", delayMs: 0, writes: [
{ id: "researcher:source-1", version: 1, text: "RFC draft found" },
{ id: "shared:verdict", version: 1, text: "publish" },
] },
{ taskId: "t2", agent: "critic", delayMs: 0, writes: [
{ id: "critic:risk-1", version: 1, text: "claim lacks source" },
{ id: "shared:verdict", version: 1, text: "hold" },
] },
{ taskId: "t3", agent: "editor", delayMs: 0, writes: [
{ id: "editor:headline", version: 1, text: "Draft A" },
{ id: "editor:headline", version: 2, text: "Draft B" },
] },
];
/* ---------- helpers ---------- */
function rng(seed: number) { // mulberry32: small, seedable, reproducible
return () => {
seed = (seed + 0x6d2b79f5) | 0;
let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
function shuffle<T>(xs: T[], r: () => number): T[] {
const a = [...xs];
for (let i = a.length - 1; i > 0; i--) {
const j = Math.floor(r() * (i + 1));
[a[i], a[j]] = [a[j], a[i]];
}
return a;
}
function* permutations<T>(xs: T[]): Generator<T[]> {
if (xs.length <= 1) { yield xs; return; }
for (let i = 0; i < xs.length; i++) {
const rest = [...xs.slice(0, i), ...xs.slice(i + 1)];
for (const p of permutations(rest)) yield [xs[i], ...p];
}
}
const updates: Ledger[] = fixture.flatMap((t) =>
t.writes.map((w) => ({ [w.id]: [{ ...w, owner: t.agent }] })),
);
/* ---------- 1. reproduce the failure modes ---------- */
describe("failure modes", () => {
it("last-value channel rejects two writers in one superstep", async () => {
await expect(buildNaive().invoke({ topic: "x" }))
.rejects.toThrow(/summary|INVALID_CONCURRENT_GRAPH_UPDATE/);
});
it("last-write-wins reducer completes but keeps only one agent's output", async () => {
const out = await buildLossy().invoke({ topic: "x" });
// No error is raised: the loss is silent. Exactly one value survives.
expect(["from researcher", "from critic"]).toContain(out.summary);
});
it("read-modify-write into an append reducer duplicates existing entries", async () => {
const out = await buildReadModifyWrite().invoke({});
expect(out.notes.filter((n) => n === "seed")).toHaveLength(3);
expect(out.notes).toHaveLength(5);
});
});
/* ---------- 2. algebraic properties of the reducer (no LangGraph) ---------- */
describe("mergeLedger", () => {
it("is order-independent across all permutations of the updates", () => {
const seen = new Set<string>();
for (const p of permutations(updates)) {
seen.add(JSON.stringify(p.reduce(mergeLedger, {} as Ledger)));
}
expect(seen.size).toBe(1); // 6 updates -> 720 orders, one result
});
it("is idempotent: re-applying an update changes nothing", () => {
const once = updates.reduce(mergeLedger, {} as Ledger);
const twice = updates.reduce(mergeLedger, once);
expect(twice).toEqual(once);
});
it("reports ties as conflicts and keeps version history out of them", () => {
const l = updates.reduce(mergeLedger, {} as Ledger);
expect(conflicts(l)).toEqual(["shared:verdict"]);
expect(l["shared:verdict"].map((f) => f.text).sort()).toEqual(["hold", "publish"]);
expect(winner(l, "editor:headline")?.text).toBe("Draft B");
});
it("rejects writes outside the owner's namespace", () => {
expect(() => mergeLedger({}, {
"researcher:source-1": [{ id: "researcher:source-1", owner: "critic", version: 9, text: "x" }],
})).toThrow(/Ownership violation/);
});
});
/* ---------- 3. ownership at the node level ---------- */
describe("owned()", () => {
it("fails a node that returns a channel it does not own", async () => {
const node = owned("rogue", ["ledger"], async () => ({ ledger: {}, tasks: [] }));
await expect(node({})).rejects.toThrow(/unowned channel\(s\): tasks/);
});
});
/* ---------- 4. the full graph ---------- */
describe("safe graph", () => {
it("produces identical state across 20 shuffled, re-timed runs", async () => {
const app = buildSafe();
const snapshots = new Set<string>();
for (let seed = 1; seed <= 20; seed++) {
const r = rng(seed);
const tasks = shuffle(fixture, r).map((t) => ({ ...t, delayMs: Math.floor(r() * 40) }));
const out = await app.invoke({ tasks });
snapshots.add(JSON.stringify({ ledger: out.ledger, log: out.log }));
}
expect(snapshots.size).toBe(1);
}, 30_000);
it("surfaces the verdict conflict instead of hiding it", async () => {
const out = await buildSafe().invoke({ tasks: fixture });
expect(conflicts(out.ledger)).toEqual(["shared:verdict"]);
expect(out.log).toEqual(["critic:t2", "editor:t3", "planner:3", "researcher:t1"]);
});
it("fails the run when an agent writes into another agent's namespace", async () => {
const rogue: Task = { taskId: "t9", agent: "critic", delayMs: 0, writes: [
{ id: "researcher:source-1", version: 9, text: "overwritten" },
] };
await expect(buildSafe().invoke({ tasks: [...fixture, rogue] }))
.rejects.toThrow(/Ownership violation/);
});
it("rejects duplicate task ids before fan-out", async () => {
await expect(buildSafe().invoke({ tasks: [...fixture, fixture[0]] }))
.rejects.toThrow(/Duplicate taskId/);
});
});
Run it:
npm run typecheck
npm test
How to read the results
The "failure modes" block should pass. Each of those tests asserts that the bug exists, so a pass means you reproduced it. If the first test fails because no error was raised, your version treats concurrent last-value writes differently. Note that and look at the actual state, since it's a silent-loss case now.
The error text in test 1 is version-dependent. The regex accepts either the key name or the error code. If neither matches, print the error and update the pattern. Don't delete the test.
The mergeLedger block runs without LangGraph. If it fails, the bug is in your merge logic, not in the framework. This is the cheapest place to find it.
The 20-run test checks invariance, not speed. A snapshots.size above 1 means some part of the final state depends on order or timing. Diff two snapshots to find which channel.
7. Tracing which nodes wrote in which step
When a test fails on your own graph, look at the per-node updates before you change any reducer. scripts/trace.ts:
import { buildSafe, type Task } from "../src/graphs.js";
const tasks: Task[] = [
{ taskId: "t1", agent: "researcher", delayMs: 30, writes: [{ id: "shared:verdict", version: 1, text: "publish" }] },
{ taskId: "t2", agent: "critic", delayMs: 0, writes: [{ id: "shared:verdict", version: 1, text: "hold" }] },
];
const stream = await buildSafe().stream({ tasks }, { streamMode: "updates" });
for await (const chunk of stream) {
console.log(JSON.stringify(chunk));
}
npx tsx scripts/trace.ts
Each chunk is keyed by the node that produced it. What to look for in your own graphs:
two different nodes returning the same key in the same step, where that key has no reducer;
a node returning a full collection where you expected a delta, which is the read-modify-write smell;
keys a node returns that it has no business touching. Wrap that node in owned(...).
If you also want to look at state between supersteps, compile with a checkpointer (compile({ checkpointer: new MemorySaver() })), invoke with a thread_id and iterate app.getStateHistory(config). That's useful when the problem only shows up after several rounds.
8. Applying this to an existing graph
Inventory the channels. For every key in your Annotation.Root, write down which nodes return it and whether those nodes can be scheduled in the same superstep: static fan-out edges, Send or conditional edges that return several targets.
Classify each channel. Single writer → keep it a last-value channel; the error is your guard. Multiple writers → it needs a reducer, and you must decide whether conflicts should be kept (ledger), combined (set union, sum, max) or rejected (throw).
Make updates deltas. Nodes return only what they add. Search for patterns like [...state.x, item] or { ...state.x, k: v } in returns to channels with reducers.
Property-test every reducer. Run all permutations of a representative set of updates and check idempotency. If a reducer can't pass these tests, its channel shouldn't receive parallel writes.
Make routing a pure function of state. If an LLM proposes the plan, persist the plan in state first (a single writer), then route from the stored plan with a stable sort.
Add the invariance test. Shuffle input order and inject delays with a seeded generator. Assert that one snapshot results.
9. Failure cases and anti-patterns
"Fixing" the concurrent update error with (_, b) => b. This turns a loud failure into silent loss. If you really want last-write-wins, define "last" by data (a version field), not by arrival.
Object spread as a reducer. (a, b) => ({ ...a, ...b }) loses one side whenever two agents write the same key, and it doesn't merge nested objects.
Plain concat with snapshot tests. Concatenation is not commutative. If application order ever changes, a snapshot test becomes flaky with no logic change. Sort, or use a set.
Side effects inside reducers. Logging to an external service, calling APIs or reading the clock make the reducer impure and break replay. Keep effects in nodes.
Trusting model-reported identity. If the owner field comes from model output, ownership checks are decoration. Stamp it in code.
LLM-chosen fan-out without a stored plan. Two runs can route differently even when your state code is correct. Test the router with fixed state, not with a live model.
Duplicate Send payloads. If a planner emits the same task twice, an append reducer records it twice. Validate task ids before fan-out, as the planner above does, or make the reducer idempotent.
10. Limitations
No published run output. This article specifies the expected behavior and gives you the means to check it. It doesn't report results from our environment. Your version of @langchain/langgraph decides the actual outcome.
Typing of narrowed node inputs. The worker takes a Task (the Send payload) instead of the full state. LangGraph.js documents this pattern, but if your version's types reject it, type the parameter as typeof SafeState.State & Task or add a cast at the addNode call.
Scope is one run. These tests cover writes inside a single graph execution. Two independent invoke calls on the same thread_id with a shared checkpointer are a different concurrency problem (persistence-level). Reducers don't solve it.
Conflict policy is a business decision. The ledger guarantees that conflicts are kept and detected deterministically. It doesn't decide whether "hold" should beat "publish". Route conflicts to a reviewer or an explicit policy node.
Cost. The ledger keeps every distinct candidate and re-sorts on every merge. That's fine for dozens or hundreds of entries per key. For large collections, keep a compact index and move the full history to external storage.
Timing injection is not proof. Seeded delays show that completion order doesn't matter in the cases you ran. The permutation test on the pure reducer is the stronger guarantee. Keep both.
11. Verification checklist
npm ls @langchain/langgraph @langchain/core output is saved next to the test results.
The three failure-mode tests pass. You reproduced the bugs, so the harness can see them.
Every reducer passes the all-permutations and idempotency tests.
Every multi-writer channel has a documented merge policy; every single-writer channel has no reducer.
Every node is wrapped in owned(...), or the ownership is enforced some other way you can test.
Routing depends only on stored state and uses a stable sort.
The 20-run invariance test yields exactly one snapshot.
Conflicts appear in state and are resolved by an explicit step, never by arrival order.
More hands-on material on testing and controlling agents is in the guides. Terms used here are defined in the glossary.
We publish what works for us—and implement the same solutions for your business.
We design AI automation, Telegram bots, chats, and AI agents for real-world processes.
Discuss your project →
© 2026 Agent Lab Journal · Guides · Glossary
Top comments (0)