<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alex Aslam</title>
    <description>The latest articles on DEV Community by Alex Aslam (@alex_aslam).</description>
    <link>https://dev.to/alex_aslam</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2607368%2Fb74d2406-bd46-4f49-a4ac-9ebe867ee219.jpeg</url>
      <title>DEV Community: Alex Aslam</title>
      <link>https://dev.to/alex_aslam</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alex_aslam"/>
    <language>en</language>
    <item>
      <title>Agents that forget where they are will repeat work forever — checkpointing is the save-point pattern you need — Checkpointing and Resume</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:01:15 +0000</pubDate>
      <link>https://dev.to/alex_aslam/agents-that-forget-where-they-are-will-repeat-work-forever-checkpointing-is-the-save-point-28p3</link>
      <guid>https://dev.to/alex_aslam/agents-that-forget-where-they-are-will-repeat-work-forever-checkpointing-is-the-save-point-28p3</guid>
      <description>&lt;p&gt;I watched a research pipeline die at step nine of ten. Forty-three minutes of work, sixty sources fetched, four analyses complete, one synthesis in progress. The worker crashed—OOM, not a logic bug. When it came back, the orchestrator did what I'd told it to do: it started over.&lt;/p&gt;

&lt;p&gt;Forty-three minutes became eighty-six. Every source re-fetched. Every analysis re-run. And three of those re-runs wrote to a downstream system that didn't deduplicate, which is how I learned that "just retry the whole thing" is not a recovery strategy. It's a duplicate-generation engine with a nice error message.&lt;/p&gt;

&lt;p&gt;That was the week I stopped thinking about agent memory and started thinking about agent &lt;em&gt;position&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory Was Never the Problem
&lt;/h2&gt;

&lt;p&gt;I'd spent months on memory. Vector stores, retrieval strategies, context compression. My agents knew a lot. What they didn't know was where they were.&lt;/p&gt;

&lt;p&gt;Ask one of them "what have you already done?" and it would tell you about its accumulated context—sources read, conclusions drawn. Ask it "what's left?" and it had nothing. It could describe its knowledge. It couldn't describe its progress.&lt;/p&gt;

&lt;p&gt;That distinction is the whole game. &lt;strong&gt;Memory answers "what do I know?" Checkpointing answers "where am I?"&lt;/strong&gt; Most agent systems have the first and not the second, which is why a crash at step nine costs you the entire run instead of the last step.&lt;/p&gt;

&lt;p&gt;The failure modes multiply from there. An agent that doesn't know what it already did will happily redo it. An agent that doesn't know which side effects already landed will happily re-fire them. The orchestrator retries, the agent re-reasons, and the external world absorbs the duplication. Nobody throws an error, because from every individual component's perspective, nothing went wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Checkpointing Actually Is
&lt;/h2&gt;

&lt;p&gt;The pattern is decades old. Durable execution systems have been solving this since before anyone called them agents. The core idea is that a workflow's progress is recorded as a durable log, and recovery means replaying that log rather than re-executing from the start.&lt;/p&gt;

&lt;p&gt;Temporal's model is the clearest articulation. Workflow code must be deterministic—no &lt;code&gt;Date.now()&lt;/code&gt;, no random, no direct I/O. Activities are where side effects live. When a worker crashes and the workflow resumes on another worker, the workflow code re-executes from the beginning, but activities don't run again. Their results come from the event history. The code re-derives its position by replaying decisions that were already made.&lt;/p&gt;

&lt;p&gt;AWS shipped the same concept into Lambda as durable execution, letting a function checkpoint at each step and resume from the last recorded point, with execution windows up to a year. DBOS does it with Postgres as the log and &lt;code&gt;@DBOS.step()&lt;/code&gt; decorating each checkpointed unit. Restate journals every handler invocation. Resonate makes every spawn and await a durable promise, so a crash mid-batch re-runs only what was in flight.&lt;/p&gt;

&lt;p&gt;LangGraph brings it into the agent world directly. Passing a &lt;code&gt;checkpointer&lt;/code&gt; to &lt;code&gt;compile()&lt;/code&gt; makes every superstep persist. The &lt;code&gt;thread_id&lt;/code&gt; in your config is the identity of the run. A crash mid-graph, a resumed invocation with the same &lt;code&gt;thread_id&lt;/code&gt;, and the graph picks up at the last completed node rather than at &lt;code&gt;START&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;checkpointer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;PostgresSaver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_conn_string&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DB_URI&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;graph&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;checkpointer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;checkpointer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;configurable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thread_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;session-42&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;invoke&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# after a crash, same thread_id, no new input:
&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;invoke&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;invoke(None, ...)&lt;/code&gt; is the entire recovery primitive. The graph reads its last checkpoint and continues. Same &lt;code&gt;thread_id&lt;/code&gt;, same state, same position.&lt;/p&gt;

&lt;p&gt;LangGraph's newer durability modes let you tune how eagerly that happens. &lt;code&gt;"sync"&lt;/code&gt; checkpoints before the next step begins—safest, slowest. &lt;code&gt;"async"&lt;/code&gt; writes in the background—faster, with a small window where a crash can lose the last step. &lt;code&gt;"exit"&lt;/code&gt; only persists at the end—cheapest, useless for recovery. Most production systems land on &lt;code&gt;"async"&lt;/code&gt; and accept the small exposure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trap That Bites Everyone
&lt;/h2&gt;

&lt;p&gt;Checkpointing and replay have a failure mode that will absolutely find you: &lt;strong&gt;non-determinism&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If your workflow code produces different results on replay than it did the first time, replay diverges. Temporal enforces determinism hard—you use &lt;code&gt;Workflow.now()&lt;/code&gt; and &lt;code&gt;Workflow.random()&lt;/code&gt; instead of the standard library, and violations are caught by a determinism checker. LangGraph is more permissive, which means it's more dangerous. If a node does something nondeterministic &lt;em&gt;and&lt;/em&gt; the checkpoint lands after it, replay may reconstruct a different state than the original run.&lt;/p&gt;

&lt;p&gt;The practical rule: &lt;strong&gt;anything with side effects or non-determinism belongs behind the checkpoint boundary, not inside the replay path.&lt;/strong&gt; LLM calls are the obvious case. An LLM call is non-deterministic by definition. If a resumed run re-invokes the model instead of reading the recorded result, you've lost the guarantee entirely.&lt;/p&gt;

&lt;p&gt;This is where most agent frameworks quietly fall short. A graph node that calls an LLM and writes a summary is two things: a decision and a side effect. Only the second one should be replayable. Frameworks that checkpoint at the node level rather than the step level conflate them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Idempotency Is the Other Half
&lt;/h2&gt;

&lt;p&gt;Checkpointing tells your agent where it is. It does nothing about what already happened in the outside world.&lt;/p&gt;

&lt;p&gt;The classic failure: the agent writes a record, then crashes before the checkpoint lands. On resume, it replays from before the write and writes again. Now you have two records and no error.&lt;/p&gt;

&lt;p&gt;The fix is idempotency keys—derived from the &lt;em&gt;operation's identity&lt;/em&gt;, not the timestamp. A refund is keyed by &lt;code&gt;refund:{order_id}:{amount}&lt;/code&gt;. A notification is keyed by &lt;code&gt;notify:{customer_id}:{event_id}&lt;/code&gt;. A write is keyed by the checkpoint sequence number of the step that produced it. The receiving system either deduplicates or the operation is naturally idempotent because it's a keyed upsert rather than an append.&lt;/p&gt;

&lt;p&gt;Without this, checkpointing turns a single failure into a single failure plus duplicated side effects. With it, checkpointing is actually safe.&lt;/p&gt;

&lt;p&gt;The pairing is non-negotiable. I learned this the expensive way: my pipeline had checkpointing and no idempotency, so the resume at step nine re-sent three downstream writes that had already succeeded. One of them triggered a customer email. Twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Guardrails That Make It Survivable
&lt;/h2&gt;

&lt;p&gt;Five failure modes matter in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Checkpoint bloat.&lt;/strong&gt; Every checkpoint stores the full state. If your state includes fifty documents and the entire conversation history, you're writing megabytes per step. Storage grows, resume time grows, and eventually you're paying more for persistence than for inference. The fix is compaction—store references to large artifacts rather than the artifacts themselves, and periodically summarize accumulated context into a compact form.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Topology skew.&lt;/strong&gt; A checkpoint written by graph version 3 may not be resumable by graph version 4. If you deployed a change between the crash and the resume, the replay can hit a node that no longer exists or an edge that now routes elsewhere. Version your graph. Refuse to resume a checkpoint whose version doesn't match, or provide explicit migration paths.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Coarse checkpoints.&lt;/strong&gt; If you only checkpoint every ten steps, a crash costs you nine steps of re-run work. If you checkpoint every step, you pay write overhead constantly. The right granularity is per-meaningful-side-effect, not per-node.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Checkpoint everywhere, resume nowhere.&lt;/strong&gt; I've seen teams invest heavily in persistence and never build the resume path. The checkpoint exists, the recovery code doesn't. A crash still costs the full run. Persistence without recovery is a log, not a safety net.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unbounded retries against a poisoned state.&lt;/strong&gt; If a checkpoint contains bad data—a malformed artifact, an impossible intermediate value—resuming will fail the same way every time. The system needs a poison-checkpoint detector and an escalation path. Otherwise the workflow loops forever, crashing and resuming in the same place.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Production Teams Are Actually Running
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Lyft's&lt;/strong&gt; LangGraph-based support system uses thread-level state and checkpointers to manage conversation history across riders and drivers, with the router holding state and re-routing mid-chat as intent shifts. Their agent development timeline dropped from roughly six months to a few weeks, and the state management is what makes multi-turn support workflows survivable in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Temporal&lt;/strong&gt; is the workhorse behind a large share of durable agent infrastructure. The event-history model means a workflow can run for months, survive repeated worker failures, and resume on any available worker without losing position. The determinism constraint is the price of admission.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LangGraph's&lt;/strong&gt; human-in-the-loop pattern depends entirely on checkpointing. &lt;code&gt;interrupt()&lt;/code&gt; pauses a node, the graph checkpoints, and the run stays suspended indefinitely. When a human responds, &lt;code&gt;Command(resume=value)&lt;/code&gt; continues from exactly where it stopped. Without a checkpointer, the interrupt has nothing to resume from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AWS Lambda durable execution&lt;/strong&gt; extends the same model to serverless functions, checkpointing at each step and resuming across cold starts with execution windows measured in months. &lt;strong&gt;DBOS&lt;/strong&gt; and &lt;strong&gt;Restate&lt;/strong&gt; offer Postgres- and journal-based alternatives with the same guarantee.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Not to Reach for Checkpointing
&lt;/h2&gt;

&lt;p&gt;Checkpointing is not free, and it's not always justified.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Short, single-shot workflows.&lt;/strong&gt; If your agent completes in under thirty seconds and a failure just means a retry, checkpointing adds infrastructure for a recovery cost you barely notice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Workflows with no durable side effects.&lt;/strong&gt; If everything the agent does is read-only and the output is ephemeral, there's nothing to protect against. Re-run it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prototypes.&lt;/strong&gt; Adding a checkpoint store, a schema, a versioning strategy, and idempotency keys to a workflow you haven't validated is premature. Build the thing first. Checkpoint it when a crash actually costs you something.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When the state can't be serialized.&lt;/strong&gt; Some agents hold live connections, handles, or in-process state that can't be persisted cleanly. Either the design changes to isolate serializable state, or checkpointing isn't the right tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;Checkpointing buys you recovery. It costs you determinism discipline, storage, and a serialization boundary in your architecture.&lt;/p&gt;

&lt;p&gt;You're accepting that some of your code must be pure and replayable. You're accepting that large artifacts need to live outside the state and be referenced. You're accepting that a version change to your graph is now a version change to your persistence format. And you're accepting that idempotency is not optional—if you don't build it, checkpointing will turn single failures into compound ones.&lt;/p&gt;

&lt;p&gt;But here's what I've learned from every recovery story that went badly: the teams that suffer aren't the ones without checkpointing. They're the ones with checkpointing and no discipline about what lives inside it. The save point isn't a magic eraser. It's a contract about what your system can and cannot redo safely.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;If your agent crashed right now, mid-task, would it know exactly what it already did—or would it start over and hope the side effects were idempotent?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. LangGraph checkpointers, Temporal's event history, a homegrown state table, or a retry-everything approach you haven't been burned by yet—and what finally made you change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>Your multi-agent system doesn’t have a memory problem — it has a shared-context problem — Shared Context Layer</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Fri, 02 Oct 2026 11:03:16 +0000</pubDate>
      <link>https://dev.to/alex_aslam/your-multi-agent-system-doesnt-have-a-memory-problem-it-has-a-shared-context-problem-shared-4ih4</link>
      <guid>https://dev.to/alex_aslam/your-multi-agent-system-doesnt-have-a-memory-problem-it-has-a-shared-context-problem-shared-4ih4</guid>
      <description>&lt;p&gt;I spent six weeks convinced my multi-agent system had a memory problem. I added a vector store. Then a second one. Then a knowledge graph, because the research said graphs were the future. Each addition bought me a week of calm and then the same failures crept back: agents contradicting each other, repeating work, and confidently citing facts that no one had actually established.&lt;/p&gt;

&lt;p&gt;The turning point came when I stopped asking "how do I give my agents better memory?" and started asking "why are they each building their own version of reality?"&lt;/p&gt;

&lt;p&gt;That's when I saw it. My agents didn't have a memory problem. They had a &lt;strong&gt;shared-context problem&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Failure Nobody Names Correctly
&lt;/h2&gt;

&lt;p&gt;I'd built a system where each agent accumulated its own context. The researcher remembered sources it had read. The analyst remembered conclusions it had drawn. The writer remembered the draft it had produced. Each agent's context was internally consistent and completely isolated from every other agent's.&lt;/p&gt;

&lt;p&gt;The first time I noticed the cost, a support agent told a customer that a refund was "being processed" while the billing agent had already flagged the request as ineligible. Both agents were reasoning correctly. Neither was wrong. They were operating on different pictures of the same situation, and no one had told them to compare notes.&lt;/p&gt;

&lt;p&gt;The research had already quantified what I was feeling. In large-scale deployments of multi-agent frameworks, the task failure rate can climb to 40–80% without coordination through shared memory. Roughly 36.9% of those failures are attributed to misalignment issues—inconsistent states or goal deviations between agents. The paper's conclusion was blunt: operating independently without a common memory repository, different agents develop conflicting understandings of the environment, leading to contradictory decisions that undermine the stability of the overall system.&lt;/p&gt;

&lt;p&gt;I had been solving the wrong problem. Vector stores and knowledge graphs help a &lt;em&gt;single&lt;/em&gt; agent remember more. They do nothing for the failure mode where two agents disagree because they're reasoning over different fragments of the truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Shared Context Actually Means
&lt;/h2&gt;

&lt;p&gt;The discipline that replaced my patchwork approach has a name now: &lt;strong&gt;context engineering&lt;/strong&gt;. Google's ADK team describes it as treating context as a first-class system with its own architecture, lifecycle, and constraints. The core thesis is that context should be a &lt;em&gt;compiled view&lt;/em&gt; over a richer stateful system—not a mutable string buffer that each agent accumulates independently.&lt;/p&gt;

&lt;p&gt;The mental model shift is this. In a single-agent system, you ask: "What context does this agent need to answer well?" In a multi-agent system, you have to ask a harder question: "How do several agents share the same business understanding while each agent sees only the context its role requires?"&lt;/p&gt;

&lt;p&gt;That second question is what I had never asked. I'd been optimizing retrieval per agent. I hadn't designed the shared substrate that makes those agents' retrievals compatible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture That Fixed It
&lt;/h2&gt;

&lt;p&gt;The fix has three layers. Each one solves a specific failure mode I'd been papering over with prompt engineering.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, a shared context repository that agents read from and write to.&lt;/strong&gt; Not a message queue. Not point-to-point handoffs. A persistent, versioned store of facts, decisions, and artifacts that every agent can query. The Atlan guide calls this the "single trusted source" layer—every agent uses the same approved definitions, metrics, and policies. When my researcher established a fact, it didn't just pass that fact forward. It wrote it to the shared layer. When my analyst needed context, it didn't ask a peer. It queried the shared layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, role-scoped views instead of a shared prompt.&lt;/strong&gt; This is the part I got wrong the first time. I tried sharing everything with everyone. The result was context saturation and signal degradation. The fix is that each agent gets &lt;em&gt;its own view&lt;/em&gt; of the shared context—a projection filtered by role. Atlan's framework calls these "role-scoped context views": each agent gets only the context it needs for its specific job. Tencent's Team Memory does this with visibility tiers—Private, Team, Restricted—so that a Scout agent researching a market doesn't get the code graph a Builder agent needs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, governed retrieval paths.&lt;/strong&gt; Agents don't pull context from random documents or their own private memory. They pull from approved systems through governed paths. This is what prevents one agent's hallucination from spreading to others through shared state. A fact enters the shared layer only if it came from a governed source. A claim is only reusable if it was validated before it was written.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Research Backs Up
&lt;/h2&gt;

&lt;p&gt;The shared-context pattern isn't just a hunch. It's a research-backed architecture that shows up across three distinct lines of work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ContextDB&lt;/strong&gt; is an open-source unified context layer that replaces the patchwork of vector databases, session stores, and glue code with a single memory operating system. Its multi-agent memory sharing protocol includes conflict resolution and role-aware routing—exactly the layer I was missing. The evaluation reports up to 90% token savings compared to full-context baselines while maintaining production-grade latency under 100ms p95.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SE-Blackboard&lt;/strong&gt; applies the classical blackboard architecture—a shared workspace accessible to all agents—to software engineering pipelines. The paper's key metric is "knowledge drift": the cumulative loss or distortion of technical entities as task-relevant content is paraphrased through successive agent handoffs. Their empirical result: the blackboard architecture improves information fidelity by 62% and raises resolve rates by 4 percentage points compared to message-passing architectures.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Context-Graph Shared Memory pattern&lt;/strong&gt; takes a different approach: storing cross-agent state as typed entity-relationship triples rather than vector chunks. The benchmark data is nuanced—graph memory beats vector RAG on multi-hop join queries (80% vs 20% accuracy) but falls behind on general long-conversation recall by roughly 25 accuracy points at ten times the cost. The pattern's own guidance is refreshingly honest: benchmark all three approaches on the queries your system actually runs before defaulting to any one of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Production Teams Are Building
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Tencent's Team Memory&lt;/strong&gt; is the most complete production implementation I've found. It distinguishes itself from RAG in a way that matters: "RAG answers 'what can be found?' Team Memory also answers 'who can use it, which version is valid, and which Agent should receive it.'" The system registers four kinds of reusable assets—Chat Memory, Skills, LLM-Wiki, and Code-Graph—and equips each agent with an "Agent Loadout" of only what it needs. The repo hit number one on GitHub's TypeScript trending list within days of launch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CrewContext&lt;/strong&gt; takes a database-first approach. PostgreSQL is the source of truth—an append-only event log with versioned entity snapshots and causal link tables. A Policy Router evaluates events against composable rules before they enter the shared state. Neo4j is an optional projection for graph queries and lineage visualization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;memX&lt;/strong&gt;, open-sourced by a Microsoft AutoGen contributor, treats the shared memory layer as a semantic store rather than just a coordination bus. The architecture separates two concerns: a pub/sub layer (Redis) for real-time signals like task status and inter-agent pings, and a vector store (LanceDB) for accumulated knowledge with embedding-based retrieval. The result: a memory written by the research agent can be recalled by the marketing agent with a thematic query, with no direct messaging and no schema coupling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Google's ADK&lt;/strong&gt; builds context engineering into the framework itself. Sessions, memory, and artifacts are the sources; flows and processors are the compiler pipeline; the working context is the compiled view shipped to the LLM for a single invocation. The design principle is separating storage from presentation so that schemas and prompt formats can evolve independently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blackboard-Core&lt;/strong&gt; implements the classical blackboard pattern as a Python SDK: a centralized typed state model, a supervisor LLM that decides which worker runs next, and workers that read from and write to the shared state rather than messaging each other directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;A shared context layer isn't free. It adds a dependency that every agent's reasoning now flows through, and that dependency can become a bottleneck if you don't design it carefully.&lt;/p&gt;

&lt;p&gt;The failure mode is what one team called &lt;strong&gt;"agreement manufactured by copying rather than earned by observation."&lt;/strong&gt; Once an agent's context is shared across a whole team, a wrong fact doesn't cost one person a repeated explanation. It costs the whole team. One agent records a wrong value, others copy it into shared memory, and the system displays consensus that was never real.&lt;/p&gt;

&lt;p&gt;The fix is governance. Every write to the shared layer needs provenance—which agent, which source, which timestamp. Every read needs a scope. Conflict resolution can't be an afterthought; it has to be a first-class operation on the shared layer. And someone has to own the schema.&lt;/p&gt;

&lt;p&gt;You're also accepting more infrastructure. A shared context layer is another system to deploy, monitor, and scale. The payoff is that agents stop duplicating each other's work, stop contradicting each other, and stop burning tokens reconstructing state that already exists elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Fits in Your Architecture
&lt;/h2&gt;

&lt;p&gt;Use a shared context layer when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your agents routinely need facts another agent established.&lt;/strong&gt; If you're passing raw outputs between agents and hoping meaning survives the handoff, you have a shared-context problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You've seen contradictory outputs from agents that were each individually correct.&lt;/strong&gt; That's the signature of fragmented context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You're scaling past three or four agents.&lt;/strong&gt; The coordination overhead of ad-hoc context sharing grows faster than the agent count.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don't use it when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your agents are genuinely independent.&lt;/strong&gt; If each agent operates on its own data and produces its own output with no shared reasoning, a shared context layer is unnecessary overhead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your sessions are short and single-purpose.&lt;/strong&gt; A shared context layer earns its keep over sessions with many agents and many interactions. For a single-shot pipeline, it's premature.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You haven't defined what "correct" looks like.&lt;/strong&gt; The shared layer doesn't fix bad reasoning. It surfaces conflicts. If you can't adjudicate a conflict, sharing context will just make the contradictions more visible—which is progress, but it's uncomfortable progress.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Question I Keep Coming Back To
&lt;/h2&gt;

&lt;p&gt;If two of your agents answered the same question right now, would they agree—not because one copied the other, but because they're both reading from the same version of the truth?&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. A shared blackboard, a context graph, a governed memory hub, or a patchwork that mostly works and what finally made you stop treating it as a memory problem?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>Fallback models are not a safety net if quality drops — the same-bar pattern that keeps reliability without regression — Same-Bar Fallback</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:29:25 +0000</pubDate>
      <link>https://dev.to/alex_aslam/fallback-models-are-not-a-safety-net-if-quality-drops-the-same-bar-pattern-that-keeps-reliability-9pg</link>
      <guid>https://dev.to/alex_aslam/fallback-models-are-not-a-safety-net-if-quality-drops-the-same-bar-pattern-that-keeps-reliability-9pg</guid>
      <description>&lt;p&gt;I once watched a fallback model save a request and destroy a product decision.&lt;/p&gt;

&lt;p&gt;It was a support-ticket classifier. The primary model timed out. The fallback model picked up the same prompt, returned valid JSON, and the dashboard turned green. Availability restored. Everyone exhaled. Then someone noticed that a ticket which should have gone to the urgent queue was sitting in the ordinary one. The fallback had returned &lt;code&gt;"priority": "high"&lt;/code&gt; instead of &lt;code&gt;"urgent"&lt;/code&gt; and &lt;code&gt;"requires_human_review": false&lt;/code&gt; instead of &lt;code&gt;true&lt;/code&gt;. Same schema. Different product. No error anywhere.&lt;/p&gt;

&lt;p&gt;That was the day I stopped treating fallback models as spare parts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Lie Your Dashboard Tells You
&lt;/h2&gt;

&lt;p&gt;The standard fallback pattern is seductive in its simplicity. Primary model fails, fallback model takes over, request completes. The graph goes green, the 200s roll in, and you move on with your day.&lt;/p&gt;

&lt;p&gt;But there are three layers of success that a green dashboard blurs together. &lt;strong&gt;Transport success&lt;/strong&gt; means a model returned a response. &lt;strong&gt;Contract success&lt;/strong&gt; means the response has the shape your application expects. &lt;strong&gt;Product success&lt;/strong&gt; means the application did the thing the user actually needed. A fallback can pass the first two and fail the third without a single log entry.&lt;/p&gt;

&lt;p&gt;I learned this the hard way across three separate incidents, each more expensive than the last. The first was the classifier. The second was a personality drift where a DeepSeek fallback silently erased our agent's entire tone and behavioral rules for 79,000 tokens—40% of a session—before a human noticed "this doesn't feel like Joe anymore". The third was a schema integrity collapse where a fallback model received a payload formatted for a different engine and returned structurally broken JSON that the downstream validator couldn't detect. The pipeline reported 100% completion. The data was useless.&lt;/p&gt;

&lt;p&gt;The common thread wasn't model quality. It was that I had never defined what "acceptable" meant for a fallback in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Research Already Knew
&lt;/h2&gt;

&lt;p&gt;The 2026 literature had already named my problem. A Google Developers Blog post on the strongest AI Agents Challenge submissions identified &lt;strong&gt;same-bar fallback validation&lt;/strong&gt; as one of four essential engineering patterns. The warning was explicit: ensuring fallback models meet the same validation standards as primary paths is not optional.&lt;/p&gt;

&lt;p&gt;The model cascade research made the mechanism clear. In a within-task cascade, the escalation gate—not the ladder—decides which cheap answers ship as final. Everything the gate accepts goes unreviewed, so its false-accept rate caps the cascade's quality. A gate that only checks output shape will pass wrong answers straight through. I had been building gates that checked shape and calling it validation.&lt;/p&gt;

&lt;p&gt;The routing-policy evaluation literature added the failure mode I hadn't named: &lt;strong&gt;fallback rot&lt;/strong&gt;. The primary has been stable for ten months. The fallback chain has drifted. The first time it fires under real load is the first time anyone learns it's broken. I had a fallback model I'd never tested under production conditions because the primary had never failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Same-Bar Pattern
&lt;/h2&gt;

&lt;p&gt;The fix isn't complicated to state. &lt;strong&gt;Every fallback model must clear the same quality bar as the primary model on the actual production tasks, or it doesn't belong in the chain.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The implementation has three layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, define the bar explicitly.&lt;/strong&gt; Not "does it return JSON" but "does it classify the ticket the way the primary model would." The bar has to be measured on your real task distribution, not a generic benchmark. A fallback that scores 92% on a public eval and 61% on your specific extraction task is a liability, not an asset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, test the fallback as a production path, not a spare tire.&lt;/strong&gt; The model swap deserves the same regression testing as a prompt change. Hold the prompt, tools, cases, judge, and inference parameters still, and vary only the model. If your fallback can't hit a reasonable percentage of your primary model's score on your actual production tasks, it's not ready to serve your users.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, gate the fallback with the same validation you'd apply to the primary.&lt;/strong&gt; A schema check is necessary but nowhere near sufficient. A response can conform perfectly to a schema and still classify incorrectly, choose the wrong tool, overstate confidence, or take a tone that doesn't belong in your product. The gate has to reject &lt;em&gt;semantically&lt;/em&gt; wrong output, not just malformed output.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Actually Costs
&lt;/h2&gt;

&lt;p&gt;The honest trade-off: same-bar fallback is slower to implement and more expensive to operate than "route to whatever model is available."&lt;/p&gt;

&lt;p&gt;You're accepting that some fallback models won't make the cut. You're accepting that your fallback chain will be shorter than it could be. You're accepting that during a provider outage, your system might return an explicit &lt;code&gt;NotAnswered&lt;/code&gt; rather than a plausible-but-wrong answer.&lt;/p&gt;

&lt;p&gt;The research on model cascades gives the concrete math. The escalation rate has to stay below &lt;code&gt;1 − (cheap cost ÷ flagship cost)&lt;/code&gt; for the cascade to pay off. If your fallback model fails 80% of the time, the fallback chain costs more than going straight to the primary—and the failures ship silently.&lt;/p&gt;

&lt;p&gt;But here's what I've learned from watching a single fallback model corrupt a support queue without anyone noticing for three days: &lt;strong&gt;the cost of a plausible wrong answer is always higher than the cost of an explicit failure&lt;/strong&gt;. A crash forces a fix. Silent degradation slips into your database and compounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Looks Like in Production
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OpenRouter's Jev-verified cascade&lt;/strong&gt; runs a cheap model draft, verifies it against retrieved context, and escalates to the frontier only when the check fails. On a 50-question benchmark, the cascade shipped the same zero wrong answers as running the frontier model on every question, at about 7% of the cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shopify's LLM proxy&lt;/strong&gt; gives every engineer access to multiple providers with automatic failover. When Claude Fable 5 shut down, the proxy shifted traffic to Claude Opus or GPT 5.5 automatically—but the fallback models were pre-validated to meet the same quality bar, not just to return a response.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The SentinelAgent&lt;/strong&gt; project demonstrates a three-tier fallback chain—Claude → OpenAI → Local BiLSTM—with graceful degradation at each tier. The local model has 4.9M parameters and 77.31% accuracy. It's not as good as Claude. But it's tested against the same task distribution, and the system knows exactly what quality it's delivering at each tier.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Question I Keep Coming Back To
&lt;/h2&gt;

&lt;p&gt;If your fallback model answered a production request right now, could you prove its answer was as good as the primary model's—or would you find out three days later when someone noticed the wrong ticket in the wrong queue?&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Same-bar validation on every fallback, a tested chain you trust under load, or a green dashboard you haven't stress-tested yet—and what finally made you change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>Why your agent bill explodes before the model even thinks — tiered routing and the cheap-check-first fix — Tiered Routing</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:55:11 +0000</pubDate>
      <link>https://dev.to/alex_aslam/why-your-agent-bill-explodes-before-the-model-even-thinks-tiered-routing-and-the-4abc</link>
      <guid>https://dev.to/alex_aslam/why-your-agent-bill-explodes-before-the-model-even-thinks-tiered-routing-and-the-4abc</guid>
      <description>&lt;p&gt;I once watched a single agent workflow burn through $340 in a single afternoon. Not because it failed—because it succeeded. Twenty-three steps, every one of them routed to a frontier model, and nineteen of those steps were formatting JSON, extracting dates, and classifying support tickets. I was paying Cadillac prices for a job a bicycle could do.&lt;/p&gt;

&lt;p&gt;That was the day I stopped blaming the model and started looking at the router.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bill Arrives Before the Model Thinks
&lt;/h2&gt;

&lt;p&gt;The math behind agent cost explosion is deceptively simple. A frontier model like Claude Opus runs 5 to 25 times the per-token cost of a small model like Claude Haiku. On a single call, that's a rounding error. On a seven-step workflow where five frontier calls chain together, you're running an order of magnitude past the equivalent small-model pipeline.&lt;/p&gt;

&lt;p&gt;The research confirms what my bill already knew. Enterprise agentic systems that route every trajectory step to a frontier model waste 60–80% of their inference budget on subtasks that smaller models handle equally well. In enterprise deployments spanning document processing, compliance review, and customer interaction, 55–70% of agent trajectory steps require no frontier-model capability at all.&lt;/p&gt;

&lt;p&gt;The kicker is &lt;em&gt;which&lt;/em&gt; steps cost you. A 2026 study found that only &lt;strong&gt;14.2% of steps genuinely require frontier models&lt;/strong&gt;, yet those steps consume &lt;strong&gt;48.3% of the total cost&lt;/strong&gt;. The other 85.8% of steps—the formatting, the extraction, the classification—are burning frontier tokens for work that a 7B model does identically.&lt;/p&gt;

&lt;p&gt;I had been solving the wrong problem. I kept optimizing prompts for a model that shouldn't have been handling the task in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Insight That Changed the Architecture
&lt;/h2&gt;

&lt;p&gt;A colleague looked at my trace and asked a question that reframed everything: &lt;em&gt;"Why is your summarization step calling the same model as your planning step?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The answer was embarrassing. I had one model configured. Every step used it because that was the default. I'd never asked whether each step &lt;em&gt;needed&lt;/em&gt; it.&lt;/p&gt;

&lt;p&gt;That's when I understood that &lt;strong&gt;model choice is a per-step decision&lt;/strong&gt;, not a per-agent decision. A planning step genuinely requires frontier-class reasoning. The formatting step that follows it needs a 7B model and nothing more. Treating them the same is the architectural equivalent of hiring a senior architect to file paperwork.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tiered Routing: Cheap Check First, Escalate on Failure
&lt;/h2&gt;

&lt;p&gt;The pattern that fixed this is called a &lt;strong&gt;model cascade&lt;/strong&gt;—or within-task model cascade—and it's embarrassingly simple: run the step on a cheap model first, and escalate to the flagship only when a gate rejects the output.&lt;/p&gt;

&lt;p&gt;FrugalGPT introduced the technique and demonstrated up to 98% cost reduction while matching or surpassing GPT-4 accuracy on several benchmarks. The Stanford researchers' insight was that most queries don't need the expensive model, and the expensive model only sees the ones that do.&lt;/p&gt;

&lt;p&gt;But the naive cascade has a trap. The gate that decides whether to escalate—that's the load-bearing piece. A gate that only checks output shape will pass wrong answers straight through, because a malformed JSON is easy to catch but a confidently wrong extraction looks exactly like a correct one.&lt;/p&gt;

&lt;p&gt;The three conditions for a cascade to pay off:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The gate rejects semantically wrong output, not just malformed output.&lt;/strong&gt; Everything the gate accepts ships unreviewed. If the gate can't tell "Q3 2024" from "Q3 2025," it's not a gate. It's a rubber stamp.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The escalation rate stays below break-even.&lt;/strong&gt; You pay the cheap attempt on every item, so the cascade beats always-flagship only while the escalation rate stays under &lt;code&gt;1 − (cheap cost ÷ flagship cost)&lt;/code&gt;. If your cheap model fails 80% of the time, the cascade costs more than going straight to the flagship.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cheap rung isn't slower than the flagship, or latency doesn't matter.&lt;/strong&gt; A 4B local model that takes 12 seconds to produce a draft isn't saving you anything if the flagship finishes in 3.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Research Actually Quantifies
&lt;/h2&gt;

&lt;p&gt;The 2026 literature has moved past theory. The numbers are concrete.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AgentRouter&lt;/strong&gt;, a 12M-parameter classifier with under 5ms overhead per step, maps each trajectory step to one of four model tiers using five features extractable at routing time. Trained on 50,000 annotated agent trajectory steps, it achieves &lt;strong&gt;72% cost reduction&lt;/strong&gt; relative to frontier-only baselines while retaining &lt;strong&gt;97.3% of frontier-only quality&lt;/strong&gt;—less than 3% degradation in end-to-end task completion.&lt;/p&gt;

&lt;p&gt;The per-step routing accuracy is telling: 91% on minimal-complexity steps, 85% on efficient-tier steps, and 76–82% on the harder mid-range and frontier tiers. The router is &lt;em&gt;better at identifying cheap work than expensive work&lt;/em&gt;, which is exactly the right failure mode.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FrugalGPT&lt;/strong&gt; (adapted to four tiers) achieves 44.1% cost reduction with 96.8% quality preserved, but with 48ms routing overhead from sequential model probing. &lt;strong&gt;RouteLLM&lt;/strong&gt; performs worst in cost reduction (31.4%) because its preference-based router, trained on single-turn conversations, misroutes agent steps whose complexity depends on accumulated trajectory context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Planner-as-Router&lt;/strong&gt; takes a different approach: instead of a separate router model, the planner assigns each subtask a model tier as it decomposes the query. It cuts cost 44% against all-frontier routing while giving up 2.9 points of accuracy, matching a faithful FrugalGPT cascade at lower cost.&lt;/p&gt;

&lt;p&gt;The consistent finding across all of these: &lt;strong&gt;step-level routing beats query-level routing in agentic workflows&lt;/strong&gt; because subtask complexity varies widely within a single trajectory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cost Paradox Nobody Warns You About
&lt;/h2&gt;

&lt;p&gt;Here's the part that took me three weeks to understand and two months to fix.&lt;/p&gt;

&lt;p&gt;If you implement routing incorrectly in a multi-turn agent, &lt;strong&gt;you can end up paying more than if you'd never routed at all&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The reason is prompt caching. By Turn 3 of an agent session, cached tokens make up the vast majority of your token payload. Anthropic offers up to a 90% discount on cached input tokens. That's the efficiency you're protecting.&lt;/p&gt;

&lt;p&gt;Switching models mid-session &lt;strong&gt;destroys that protection entirely&lt;/strong&gt;. Every provider's cache is model-specific. The new model has no access to the previous model's stored history and must re-read the entire conversation from scratch at full input token cost. For agentic coding workflows where context windows stretch to 50,000 tokens, a single mid-session model switch can spike costs enough to eliminate all the savings routing was supposed to deliver.&lt;/p&gt;

&lt;p&gt;The mistake I made was routing &lt;em&gt;across&lt;/em&gt; a session. The fix is routing &lt;em&gt;within&lt;/em&gt; a step—choose the model for each generation, complete it, and return to the session model. Or better: summarize the context before switching, so the new model starts with a compressed history rather than a cold cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who's Actually Shipping This
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OpenRouter's Jev-verified cascade&lt;/strong&gt; runs a cheap model draft, verifies it against retrieved context using Jev (a verification layer), and escalates to the frontier only when the check fails. On a 50-question benchmark, the cascade shipped the same zero wrong answers as running the frontier model on every question, at about &lt;strong&gt;7% of the cost&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cisco's CAIPE&lt;/strong&gt; uses tiered routing across Argo CD, Kubernetes, and Komodor. Response times dropped from hours to seconds, and MTTR reduced by up to 80%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fireworks&lt;/strong&gt; reports 60–80% cost reduction and a 10–40× increase in inference speed with their multi-tier inference architecture for agentic workflows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The koi model-router&lt;/strong&gt; package implements cascade escalation with a complexity classifier, confidence evaluators, circuit breakers, and per-tier cost tracking. It routes 60–80% of requests to cheap models with no quality loss.&lt;/p&gt;

&lt;p&gt;The pattern is consistent. The teams winning on cost aren't using one model. They're using a ladder—and they've built the gate before they built the ladder.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Fits in Your Architecture
&lt;/h2&gt;

&lt;p&gt;Tiered routing is not a default. It's a response to a specific problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when your agent trajectory has step-level complexity variance.&lt;/strong&gt; Planning and synthesis steps need frontier reasoning. Extraction, formatting, and classification steps don't. If every step genuinely requires frontier capability, routing won't help.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when your token volume is high enough that the savings justify the infrastructure.&lt;/strong&gt; A router that saves 40% on a $50 monthly bill isn't worth building. A router that saves 40% on a $50,000 monthly bill is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it with step-level features, not query-level features.&lt;/strong&gt; The research is clear that single-turn routers misroute agent steps because they miss trajectory context. The features that matter are the ones available at the step: the instruction, the accumulated context length, the prior step's tier, and the dependency structure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't route across sessions without summarization.&lt;/strong&gt; The cache penalty will eat your savings. Route within a step, or summarize before you switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;Tiered routing buys you cost reduction and, often, latency. It costs you determinism and adds a classifier to your stack.&lt;/p&gt;

&lt;p&gt;Every routing decision is a chance to misroute. A complex step sent to a cheap model fails, and you pay the cheap attempt plus the escalation plus the latency of both. The gate's false-accept rate—not the ladder—decides whether the cascade pays. If your gate accepts wrong answers, you ship them at a discount, and the savings are a lie.&lt;/p&gt;

&lt;p&gt;And you're accepting more moving parts. A router, a gate, a tier configuration, per-tier cost tracking. Every one of those is a thing that can fail independently. The koi model-router's failure modes are instructive: low-confidence answers escalate automatically, provider outages trip circuit breakers, budget exhaustion stops escalation and returns the best available response. None of those are free.&lt;/p&gt;

&lt;p&gt;But here's what I've learned from watching my bill drop by 68% without a measurable quality regression: the teams that are winning with agents aren't using the best model for everything. They're using the &lt;strong&gt;right model for each step&lt;/strong&gt;—and they've accepted that "right" is a decision worth making explicitly.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;If you looked at your agent's cost breakdown by step, how many of those steps would you have paid frontier prices for if you'd chosen the model deliberately?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. FrugalGPT-style cascade, a trained router, a planner that assigns tiers up front, or a bill that finally made you look—and what finally made you change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>Sequential agent chains are a latency tax — the event-driven pattern that makes independent agents run in parallel — Event-Driven Concurrency</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Wed, 30 Sep 2026 20:58:23 +0000</pubDate>
      <link>https://dev.to/alex_aslam/sequential-agent-chains-are-a-latency-tax-the-event-driven-pattern-that-makes-independent-agents-b93</link>
      <guid>https://dev.to/alex_aslam/sequential-agent-chains-are-a-latency-tax-the-event-driven-pattern-that-makes-independent-agents-b93</guid>
      <description>&lt;p&gt;I once watched a nine-minute agent workflow spend seven minutes and forty seconds doing nothing. Not failing. Not retrying. Just waiting each agent idle while the previous one finished a task that had no dependency on it whatsoever.&lt;/p&gt;

&lt;p&gt;I'd built the pipeline as a chain. Researcher, then analyst, then writer, then fact-checker, then formatter. Each step consumed the previous step's output. In my head, that was the workflow. In production, it was a latency tax I'd designed myself and never noticed.&lt;/p&gt;

&lt;p&gt;The fix wasn't a better model or a faster prompt. It was admitting that most of those steps were never actually dependent on each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tax Nobody Itemizes
&lt;/h2&gt;

&lt;p&gt;Sequential chains are the easiest multi-agent pattern to reason about, which is why they're the default. Each stage has a clear input, a clear output, and a clear owner. Debugging is trivial because there's only one path. Microsoft's sequential orchestration documentation describes it exactly this way: a predefined linear order where each agent processes the previous agent's output, creating a pipeline of specialized transformations.&lt;/p&gt;

&lt;p&gt;The problem is that "clear" and "correct" are not the same thing. A pipeline implies dependency. Most real workflows don't have the dependency you assumed.&lt;/p&gt;

&lt;p&gt;My researcher fetched six independent sources. Those six fetches could have run in parallel. My analyst produced three independent analyses—market, risk, and competitive positioning—with no shared state between them. My fact-checker verified claims that were already independently verifiable. Only the writer genuinely needed everything upstream.&lt;/p&gt;

&lt;p&gt;The math is unforgiving. If you have N stages with average latency L, your wall-clock time is N × L. If K of those stages are independent, the floor is (N − K) × L + L, not N × L. In my pipeline, K was 4 out of 5. I was paying nearly five times the necessary latency and calling it "the workflow."&lt;/p&gt;

&lt;p&gt;Amdahl's law applies here in a way that's almost embarrassing once you see it. The speedup from parallelizing independent work is bounded by the serial fraction—and in most agent pipelines I've audited, the serial fraction is far smaller than the architecture assumes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Structural Mistake
&lt;/h2&gt;

&lt;p&gt;The chain wasn't wrong because I used the wrong framework. It was wrong because I encoded dependencies that didn't exist.&lt;/p&gt;

&lt;p&gt;A chain says: &lt;em&gt;B cannot begin until A completes.&lt;/em&gt; That's a claim about the world. When it's true, a chain is correct and efficient. When it's false, you've serialized work that should have been concurrent, and you'll never notice because the system still produces correct output—just slowly.&lt;/p&gt;

&lt;p&gt;I'd assumed the analyst needed the researcher's &lt;em&gt;entire&lt;/em&gt; output before starting. In reality, the analyst needed three fields. The researcher was still fetching sources four and five while the analyst could have started on source one.&lt;/p&gt;

&lt;p&gt;The dependency was real at the level of &lt;em&gt;final output&lt;/em&gt;. It was fake at the level of &lt;em&gt;individual facts&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Event-Driven Concurrency: React to Data, Not to Sequence
&lt;/h2&gt;

&lt;p&gt;The pattern that fixed this wasn't a new framework. It was a change in what triggers work.&lt;/p&gt;

&lt;p&gt;In a sequential chain, work starts because the previous step finished. In an event-driven architecture, work starts because &lt;strong&gt;the data it needs has arrived&lt;/strong&gt;. Agents subscribe to events—a source fetched, a claim extracted, a document chunked—and fire when the specific event they care about publishes.&lt;/p&gt;

&lt;p&gt;The shift has three architectural consequences.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, agents become independent consumers.&lt;/strong&gt; Each agent declares what it needs and what it produces. The orchestrator doesn't route tasks; it publishes events. Agents that can act, act. Agents that need more input, wait. No agent is blocked by an unrelated stage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, latency becomes the critical path, not the sum.&lt;/strong&gt; If six fetches run concurrently and three analyses run as soon as their inputs land, wall-clock time collapses to the longest dependency chain plus fan-in. The parallelism isn't a fan-out you designed—it's an emergent property of the dependency graph.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, backpressure becomes visible.&lt;/strong&gt; In a chain, a slow stage just makes everything slow. In an event-driven system, you can see which topics are backing up, which consumers are lagging, and where the actual bottleneck lives. That observability is worth as much as the latency win.&lt;/p&gt;

&lt;p&gt;This is the same shift microservices made a decade ago, and the same discipline applies: publish events, not commands. Let consumers decide what to do with them. Use correlation IDs to trace a logical request across many asynchronous hops. Use idempotency keys so a retried consumer doesn't double-write.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Orchestration Trap in Event-Driven Systems
&lt;/h2&gt;

&lt;p&gt;The failure mode of event-driven agent systems is the opposite of the chain's. Instead of being too slow, it becomes too opaque.&lt;/p&gt;

&lt;p&gt;Choreography—where agents react to each other's events with no central coordinator—scales beautifully and debugs terribly. When a workflow fails, you have no single place to look. The failure is in the emergent behavior of a dozen independent consumers, none of whom saw the whole picture.&lt;/p&gt;

&lt;p&gt;The pragmatic middle ground that has held up in production is &lt;strong&gt;event-driven execution under a thin orchestration layer&lt;/strong&gt;. A coordinator doesn't route individual tasks; it publishes a workflow-started event and then subscribes to completion events, tracking state and enforcing timeouts. The agents still run independently. The coordinator just knows what "done" means and can surface when it isn't.&lt;/p&gt;

&lt;p&gt;This is where durable execution frameworks earn their keep. Temporal's model—workflows as deterministic code, activities as the side-effecting work—maps cleanly onto agent systems. The workflow defines the dependency graph. The activities run concurrently where the graph allows. The coordinator state is durable, so a crash at step seven doesn't lose the run.&lt;/p&gt;

&lt;p&gt;LangGraph's parallel node execution and conditional edges offer the same property within a graph, and the newer event-driven orchestration patterns—where nodes publish and subscribe to channels rather than passing state directly—are the closest thing to true event-driven concurrency in the agent frameworks I've used.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Guardrails That Make It Survivable
&lt;/h2&gt;

&lt;p&gt;Event-driven concurrency introduces failure modes that sequential chains never had. Four of them matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ordering.&lt;/strong&gt; Events can arrive out of order. If an analysis event arrives before the source event it depends on, the analysis either fails or runs on stale data. The fix is either partition keys that preserve order for related events, or a consumer that buffers until dependencies are satisfied. In agent systems, the cleanest answer is usually to make each event self-contained enough that order doesn't matter—carry the data, not just a pointer to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Duplicates.&lt;/strong&gt; At-least-once delivery is the realistic default. Every consumer must be idempotent. An agent that writes a summary twice is fine. An agent that charges a customer twice is not. Idempotency keys derived from the event ID, not the timestamp, are the standard fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Poison events.&lt;/strong&gt; One malformed event can crash a consumer repeatedly, blocking the queue behind it. Dead letter queues aren't optional. Neither is a circuit breaker that stops consuming from a topic after repeated failures and alerts instead of spinning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fan-in reconciliation.&lt;/strong&gt; When three analyses complete asynchronously, something has to decide when all three are ready and what to do with their outputs. That's the same fan-in problem from the concurrency pattern, with the same rules: scoped branch state, structured claims, and an explicit merge that surfaces conflicts rather than smoothing them over.&lt;/p&gt;

&lt;p&gt;None of this is exotic. It's the standard distributed systems toolkit. The mistake is assuming agent systems are exempt from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who's Actually Shipping This
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cisco's CAIPE&lt;/strong&gt; runs multi-agent orchestration across Argo CD, Kubernetes, and Komodor, where agents chain across tools and respond to cluster events rather than polling in sequence. Their reported outcome: response times dropped from hours to seconds, with MTTR reduced by up to 80%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SAP's event-driven multi-agent orchestration&lt;/strong&gt; was evaluated across 208 production-derived enterprise scenarios at Persona, Department, and Enterprise scale. Their finding is the one that should be posted above every pipeline design review: &lt;strong&gt;scale, not task complexity, dominates orchestration performance&lt;/strong&gt;. Both sequential and reactive architectures degraded at enterprise scale, and the Task Manager they introduced to handle priority inference, related-event merging, and preemption cut high-priority queue latency by 14–75% and improved related-event correctness by more than 20 percentage points.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resonate's&lt;/strong&gt; durable execution model treats every spawn and every await as a durable promise. Fan-out branches run concurrently, each checkpointed independently. A crash mid-batch re-runs only what was in flight. That's the property that makes event-driven agent systems survivable rather than merely fast.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AWS's Expansion-Contraction&lt;/strong&gt; pattern, published at ACM CAIS 2026, walks a domain graph with concurrent path analysis. Their reported results: 98.2% accuracy on a production supply chain, 100% on public benchmarks, concurrent path analysis yielding up to 1.43× speedup, and investigation caching reducing token usage by up to 93.9%. The concurrency is the point, but the caching is what makes it affordable.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Not to Reach for Events
&lt;/h2&gt;

&lt;p&gt;Event-driven concurrency is not a default. It's a response to a specific problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stay sequential when the dependency is real.&lt;/strong&gt; If stage B genuinely needs the complete output of stage A, a chain is correct and simpler. Don't parallelize what isn't parallel.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stay sequential when latency isn't the binding constraint.&lt;/strong&gt; If your pipeline finishes in six seconds and users don't care, you don't have a latency problem. You have an architecture preference, and event-driven systems cost more to build and operate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stay sequential when debugging is the scarce resource.&lt;/strong&gt; A chain gives you a single path and a linear trace. An event-driven system gives you a correlation ID and a scatter plot. If your team is small and the workflow is new, that trade is often not worth making yet.&lt;/p&gt;

&lt;p&gt;The honest rule: &lt;strong&gt;start sequential, measure, and parallelize only the stages that measurement proves are independent and expensive.&lt;/strong&gt; The teams I've seen get this right didn't start with events. They started with a chain, watched it in production, and moved to events when the latency numbers justified the operational cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;Event-driven concurrency buys you latency and throughput. It costs you determinism and traceability.&lt;/p&gt;

&lt;p&gt;A sequential chain has one path. An event-driven workflow has as many paths as there are interleavings, and no two runs are identical. Reproducing a bug means replaying an event stream, not re-running a function. Your observability has to be built for it—distributed tracing, correlation IDs, structured logs that tie back to a single logical request.&lt;/p&gt;

&lt;p&gt;And you're accepting more moving parts. A chain has N stages. An event-driven system has N consumers, a broker, a schema registry, dead letter queues, and a coordinator. Every one of those is a thing that can fail independently.&lt;/p&gt;

&lt;p&gt;But here's what I've learned from every pipeline post-mortem I've sat through: the latency was rarely the thing that killed the project. It was the assumption that a chain was the only honest representation of the work. Once I drew the actual dependency graph instead of the one I'd assumed, most of the sequence disappeared—and so did most of the delay.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;If you drew the real dependency graph for your agent pipeline, how many of those sequential steps would survive?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Sequential by choice, event-driven after a painful migration, or somewhere in between and what finally made you look at the graph instead of the chain?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>An agent should be both client and server — the bidirectional MCP pattern that turns reasoning into infrastructure — Bidirectional MCP</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:49:56 +0000</pubDate>
      <link>https://dev.to/alex_aslam/an-agent-should-be-both-client-and-server-the-bidirectional-mcp-pattern-that-turns-reasoning-into-2l52</link>
      <guid>https://dev.to/alex_aslam/an-agent-should-be-both-client-and-server-the-bidirectional-mcp-pattern-that-turns-reasoning-into-2l52</guid>
      <description>&lt;p&gt;I built a refund agent that could fetch policy documents, check account history, and calculate eligibility. It could not, however, ask a human a single question without me tearing down the entire architecture. That's the moment I understood that most MCP servers are deaf, and most agents are talking to themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Server That Could Only Listen
&lt;/h2&gt;

&lt;p&gt;I spent a month building a customer support system with a handful of MCP servers. One wrapped our policy database. One exposed account management APIs. One connected to the ticketing system. Each one was a clean, well-typed, compliant MCP server. Each one was completely one-directional.&lt;/p&gt;

&lt;p&gt;The problem wasn't tool calling. The tools worked. The refund agent could fetch a policy, query an account, calculate a credit. What it couldn't do was hit an edge case—say, a refund outside the standard window—and ask a human for approval mid-operation. The MCP server had no way to reach back to the client. The client could only ask; the server could only answer.&lt;/p&gt;

&lt;p&gt;I kept looking for a workaround. Could I encode the question in the tool result? Could I use a separate API? Could I poll? Every alternative was a workaround, not an answer. The MCP protocol I was using was designed as a one-way street: agents call tools, tools respond, agents reason. Nothing flows the other way.&lt;/p&gt;

&lt;p&gt;That's the ceiling most teams hit and don't name. Your agent can &lt;em&gt;use&lt;/em&gt; tools, but your tools can't &lt;em&gt;collaborate&lt;/em&gt; with your agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reverse Primitives Nobody Explained to Me
&lt;/h2&gt;

&lt;p&gt;The MCP specification had a name for what I was missing: &lt;strong&gt;reverse primitives&lt;/strong&gt;. Server-initiated requests that let an MCP server reach back to the client mid-operation. Three of them, originally: &lt;strong&gt;Sampling&lt;/strong&gt;, &lt;strong&gt;Roots&lt;/strong&gt;, and &lt;strong&gt;Elicitation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;They weren't in the tutorials I read. They weren't in most of the server examples. They were in the specification, described as "client features"—capabilities the client offers to the server, not capabilities the server exposes to the agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sampling&lt;/strong&gt; lets a server request an LLM completion through the client. The server doesn't need its own API keys, doesn't pick its own model, doesn't manage its own costs. It sends a &lt;code&gt;sampling/createMessage&lt;/code&gt; request to the client, and the client—with user approval—runs the completion and returns the result. This is the piece that turns a tool server into something that can reason mid-task. Instead of returning raw data and hoping the agent interprets it correctly, the server can ask the client's LLM to evaluate, summarize, or decide.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Elicitation&lt;/strong&gt; lets a server ask a user for input during an interaction. A structured form, a URL for out-of-band flows, a schema the client renders. This is the piece that solves the refund approval problem. The server hits an edge case, sends an &lt;code&gt;elicitation/create&lt;/code&gt; request, the client surfaces a form to the human, and the server receives the answer and continues.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Roots&lt;/strong&gt; lets a client specify which directories the server should operate in—a scope boundary, a least-privilege constraint. (I'll come back to this one, because the spec's relationship with it changed in ways that matter.)&lt;/p&gt;

&lt;p&gt;The moment I understood these primitives, the architecture changed. My MCP server wasn't a passive tool endpoint anymore. It was a participant in a conversation. It could ask for clarification. It could borrow the client's reasoning. It could pause mid-operation and wait for a human.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Bidirectional Actually Looks Like
&lt;/h2&gt;

&lt;p&gt;The flow is straightforward once you see it. A client sends a &lt;code&gt;tools/call&lt;/code&gt; request to the MCP server. The server starts processing and realizes it needs something—user input, an LLM judgment, a scope decision. Instead of failing or returning incomplete data, the server sends a server-to-client request back through the same connection.&lt;/p&gt;

&lt;p&gt;The client handles the request—surfaces a form, runs a sampling completion, checks the root boundary—and responds. The server receives the response and continues processing the original tool call. The client gets the final result.&lt;/p&gt;

&lt;p&gt;The AWS Bedrock AgentCore Runtime team described it plainly in their stateful MCP client capability announcement: these capabilities "transform one-way tool execution into bidirectional conversations between your MCP server and clients". Before stateful mode, the runtime couldn't handle this. Every HTTP request was independent. The server couldn't maintain a conversation thread, couldn't ask the user for clarification mid-tool-call, couldn't request LLM-generated content. Stateful mode provisions a dedicated microVM per session, maintains continuity through a session ID, and unlocks all three client capabilities.&lt;/p&gt;

&lt;p&gt;The Go SDK's in-process sampling example shows what the code looks like on the server side. You call &lt;code&gt;mcpServer.EnableSampling()&lt;/code&gt; to advertise the capability. Your tool handler calls &lt;code&gt;mcpServer.RequestSampling()&lt;/code&gt; when it needs an LLM completion. The sampling request is handled directly by the client's sampling handler, and the response flows back into your tool handler. The example output is almost boring in how ordinary it looks: a tool result that contains an LLM-generated answer to a question the server couldn't answer itself.&lt;/p&gt;

&lt;p&gt;The TypeScript SDK handles the client side with &lt;code&gt;setRequestHandler&lt;/code&gt;. You register a handler for &lt;code&gt;sampling/createMessage&lt;/code&gt;, and when the server sends a sampling request, your handler receives the messages, runs them through your LLM, and returns the result. The client stays in control of model selection, permissions, and user approval.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Shift That Made Me Rebuild
&lt;/h2&gt;

&lt;p&gt;I rebuilt the refund agent around two changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, the server became a participant, not a lookup table.&lt;/strong&gt; When the refund tool detects an edge case—amount outside policy, missing documentation, a policy version conflict—it doesn't return an error or a partial result. It sends an elicitation request. The client surfaces a form to the support representative. The representative fills it out. The server receives the answer and completes the refund calculation. The tool call doesn't fail. It just asks a question first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, the server started borrowing reasoning.&lt;/strong&gt; For complex cases—disputed transactions, ambiguous policy language, multi-party refunds—the server sends a sampling request to the client. The client's LLM evaluates the situation and returns a structured judgment. The server uses that judgment to decide which policy path to follow. The server doesn't need its own model. It doesn't need its own API keys. It just needs the client to lend its brain for a moment.&lt;/p&gt;

&lt;p&gt;The architectural shift is subtle but consequential. The server isn't a passive endpoint anymore. It's a reasoning participant. The tool call isn't a request-response transaction anymore. It's a conversation that can pause, ask, and continue.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Protocol Moved Under My Feet
&lt;/h2&gt;

&lt;p&gt;Just as I was settling into this architecture, the MCP specification changed. The 2026-07-28 revision—the fifth spec release, and the largest change since launch—made three moves that fundamentally reshaped the bidirectional story.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Roots, Sampling, and Logging were deprecated.&lt;/strong&gt; SEP-2577 formally deprecated all three. The motivation was adoption data and complexity. Sampling, despite being available since November 2024, had low client adoption. It was complex to implement correctly—human-in-the-loop approval, model selection logic, security considerations, tool loop support. Direct LLM provider APIs gave servers more control. Roots had vague semantics, low adoption, and overlapping alternatives. Logging had mature alternatives in stderr and OpenTelemetry.&lt;/p&gt;

&lt;p&gt;The deprecation doesn't remove them immediately. They remain functional during a twelve-month window. But the direction is clear: &lt;strong&gt;sampling and roots are legacy.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The protocol core became stateless.&lt;/strong&gt; The initialize handshake and session identifier were removed. Every request is now self-contained, carrying protocol version, client information, and capabilities in &lt;code&gt;_meta&lt;/code&gt;. Servers can deploy on serverless and edge infrastructure. Any request can be routed to any server instance behind a round-robin load balancer. No sticky sessions, no shared session store.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi Round-Trip Requests replaced server-initiated requests.&lt;/strong&gt; This is the replacement for the old reverse-primitive pattern. Instead of holding an SSE stream open and sending a server-to-client request, the server returns an &lt;code&gt;InputRequiredResult&lt;/code&gt; object containing the questions and a &lt;code&gt;requestState&lt;/code&gt; blob. The client gathers answers and reissues the original call with the responses and echoed state. Because everything the server needs is in the payload, any server instance can pick up the retry.&lt;/p&gt;

&lt;p&gt;The MRTR pattern is the stateless equivalent of bidirectional. It's not a persistent conversation anymore. It's a request that can loop back with additional information. The server asks, the client answers, the original request is retried with the answers attached. The server can ask multiple questions in one round trip. It can request elicitation, sampling, or root listing. But it does so through a stateless retry pattern, not a persistent session.&lt;/p&gt;

&lt;p&gt;The 4sysops writeup described the trade-off cleanly: "In previous versions, this required holding a Server-Sent Events stream open. The new revision replaces that with Multi Round-Trip Requests". The stream is gone. The conversation is now a series of self-contained requests that carry their own state.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means for Your Architecture
&lt;/h2&gt;

&lt;p&gt;The bidirectional MCP pattern is still real. It's just changed shape.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you're on a pre-2026-07-28 spec&lt;/strong&gt;, you have the full reverse-primitive toolkit: sampling for borrowing client reasoning, elicitation for user input, roots for scope boundaries. Your server can be a reasoning participant. Your tool calls can pause and ask questions. Stateful sessions are maintained through the &lt;code&gt;Mcp-Session-Id&lt;/code&gt; header, and the client handles server-initiated requests through registered handlers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you're on the 2026-07-28 spec and later&lt;/strong&gt;, the architecture shifts. Sampling and roots are deprecated. Elicitation survives, but it's delivered through MRTR instead of a persistent session. The server returns an &lt;code&gt;InputRequiredResult&lt;/code&gt; with elicitation requests, the client gathers the answers, and the original tool call is retried with the responses attached. The server can still ask the user for input mid-operation. It just does so through a stateless retry instead of a live back-channel.&lt;/p&gt;

&lt;p&gt;The practical impact is smaller than it sounds for most use cases. If your server was using sampling to borrow the client's LLM, you now integrate directly with an LLM provider API. You get more control over model selection and parameters. You lose the client's permission management and user approval flow. If your server was using roots for scope boundaries, you now pass paths through tool parameters, resource URIs, or configuration. More explicit, less magical.&lt;/p&gt;

&lt;p&gt;If your server was using elicitation, you're fine. Elicitation survives. The delivery mechanism changed. The capability didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who's Actually Running This
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Gong&lt;/strong&gt; turned MCP into a bidirectional protocol for revenue intelligence. Their architecture exposes an MCP Gateway for inbound agent queries—external agents like Microsoft Copilot, HubSpot AI, and Salesforce can query Gong's data—and an MCP Server for outbound context enrichment, where Gong's own AI features pull in data from external tools. Write-back to Salesforce is gated by Gong-side logic. The protocol isn't just a data source anymore. It's a two-way integration layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Amazon Bedrock AgentCore Runtime&lt;/strong&gt; now supports stateful MCP servers with all three client capabilities. The runtime provisions a dedicated microVM per session, maintains continuity through a session ID, and enables interactive, multi-turn workflows. The AWS team's framing is the clearest I've read: these capabilities "transform one-way tool execution into bidirectional conversations between your MCP server and clients".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mcp-data-platform&lt;/strong&gt; project takes a different approach to bidirectional. It's a semantic data platform MCP server that composes multiple data tools with "bidirectional cross-injection"—tool responses automatically include critical context from other services. Trino query results are enriched with DataHub metadata. DataHub searches include query availability from Trino. S3 listings include matching DataHub datasets. The server isn't just answering the question it was asked. It's injecting context the client didn't know to request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Pydantic AI harness&lt;/strong&gt; has an open proposal to let agents expose themselves as MCP servers. The idea is bidirectional MCP at the agent level: agents can both consume MCP servers and &lt;em&gt;serve&lt;/em&gt; as MCP servers. A specialized database agent built with Pydantic AI becomes an MCP tool for a general coding agent. A code review agent is invoked by a CI/CD pipeline agent as an MCP tool. Agents compose through MCP instead of through framework-specific integrations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;Bidirectional MCP gives you conversational tools. It costs you simplicity.&lt;/p&gt;

&lt;p&gt;Every server-initiated request is a decision point. The client has to handle it—surface a form, run a completion, check a root boundary. The server has to handle the response and decide whether to continue or ask another question. The flow is no longer a straight line. It's a loop that can iterate as many times as the server needs.&lt;/p&gt;

&lt;p&gt;The deprecation of sampling and roots is a signal about what the protocol designers think is worth the complexity. Elicitation survived because it's the one that solves a problem no alternative solves cleanly: getting structured user input mid-operation. Sampling and roots had alternatives—direct LLM APIs, explicit path parameters—that gave more control with less protocol surface.&lt;/p&gt;

&lt;p&gt;The MRTR pattern is the compromise. It keeps the stateless core while preserving the ability for a server to ask for input mid-operation. It doesn't preserve the persistent conversation. It replaces it with a retry loop that carries its own state.&lt;/p&gt;

&lt;p&gt;Here's what I've learned: the teams that are getting bidirectional MCP right aren't using it for everything. They're using it for the moments where a tool genuinely needs to ask a question, borrow a judgment, or pause for input. Everything else stays a simple, one-way tool call. The bidirectional capabilities are exceptions, not the default. They're the edge case handling that makes the happy path possible.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;When your MCP server hits an edge case, does it fail silently, return a partial result, or ask for help—and does your architecture let it ask?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Full reverse primitives, MRTR elicitation, or a workaround you built because the spec didn't have what you needed—and what finally made you change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Your agents can’t collaborate if they only speak tool calls — the A2A protocol that fixes peer handoffs — A2A Under the Linux Foundation</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Mon, 28 Sep 2026 22:20:23 +0000</pubDate>
      <link>https://dev.to/alex_aslam/your-agents-cant-collaborate-if-they-only-speak-tool-calls-the-a2a-protocol-that-fixes-peer-4o0d</link>
      <guid>https://dev.to/alex_aslam/your-agents-cant-collaborate-if-they-only-speak-tool-calls-the-a2a-protocol-that-fixes-peer-4o0d</guid>
      <description>&lt;p&gt;I spent three weeks trying to make a billing agent ask a refund agent a question. Not a tool call—an actual question. The billing agent needed to know whether a refund was appropriate, and the refund agent needed to reason about policy, check account history, and possibly refuse. I wired it up as a function call anyway. The billing agent waited fourteen seconds for a "tool" that was actually trying to think.&lt;/p&gt;

&lt;p&gt;That was the day I stopped treating agents like functions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tool-Call Trap
&lt;/h2&gt;

&lt;p&gt;The mistake is almost inevitable. You build your first agents, you wire them together with tool calls, and it works. The tool returns a value. The caller proceeds. Simple, synchronous, easy to debug.&lt;/p&gt;

&lt;p&gt;Then one of your "tools" grows a brain. It needs to ask clarifying questions. It needs to run for six minutes. It needs to decide the request is out of scope and say no. But your architecture treats it as a passive function, so every one of those behaviors looks like a failure—a timeout, a malformed response, an unhandled edge case.&lt;/p&gt;

&lt;p&gt;I kept tuning prompts. The model wasn't the problem. The architecture was. I had confused the vertical with the horizontal.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Tool Calls Can't Do
&lt;/h2&gt;

&lt;p&gt;The Model Context Protocol (MCP) solved a genuinely hard problem. Before it, connecting five agents to ten tools meant up to fifty bespoke integrations, each one a maintenance liability. MCP collapses that into one interface. By mid-2025, the community had built thousands of active MCP servers, and OpenAI, Microsoft, and Google DeepMind all adopted it. MCP won the tool-calling layer, and it deserved to.&lt;/p&gt;

&lt;p&gt;But MCP has no concept of a peer agent with its own goals, its own model, and its own authority to act. An MCP server that wraps &lt;code&gt;kubectl&lt;/code&gt; does not think, negotiate, or push back. It exposes capabilities and waits. The agent stays in charge; the things it touches are passive.&lt;/p&gt;

&lt;p&gt;When my billing agent needed the refund agent to &lt;em&gt;decide&lt;/em&gt; something—not fetch a value, not execute a command, but reason through policy and arrive at a judgment—that was never an MCP-shaped problem. It was a horizontal problem. And horizontal problems need a different protocol.&lt;/p&gt;

&lt;h2&gt;
  
  
  The A2A Protocol: HTTP for Agents
&lt;/h2&gt;

&lt;p&gt;The Agent2Agent (A2A) protocol is an open standard launched by Google in April 2025 and now governed by the Linux Foundation's Agentic AI Foundation. Where MCP is the vertical integration layer connecting agents to internal tools and databases, A2A acts as the horizontal protocol enabling peer-to-peer collaboration. The A2A specification puts the distinction in one sentence: &lt;strong&gt;MCP standardizes how an agent uses a tool or resource; A2A standardizes how one agent delegates work to another&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Think of A2A as HTTP for AI agents. Just as HTTP enables any web browser to communicate with any web server regardless of their underlying technologies, A2A enables a LangGraph agent to delegate work to a CrewAI agent, which can then call a Google ADK agent, without requiring custom integration code for each pairing.&lt;/p&gt;

&lt;p&gt;The architectural shift matters because it changes what the calling agent expects. When you call an API, it simply returns data or fails. When an agent calls an A2A peer, it initiates a collaboration. The receiving agent can understand intent, refine the plan, push back on incomplete requests, and ask clarifying questions if something is off. That's not a tool call. That's a handoff between peers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Under the Hood: Agent Cards and Task Lifecycle
&lt;/h2&gt;

&lt;p&gt;A2A standardizes four objects: an Agent Card, a Task, a Message, and an Artifact.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Agent Card&lt;/strong&gt; is a JSON manifest served at a well-known URI, describing who an agent is, what it can do, where to reach it, and how to authenticate. This is the discovery mechanism. Your agent doesn't need to know about the refund agent at build time. It queries the registry, reads the card, and decides whether to delegate.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Task&lt;/strong&gt; is the unit of work, with an ID, a context ID, a state, a history, and artifacts. Tasks are stateful and progress through a defined lifecycle: submitted, working, input-required, completed, failed, canceled, rejected. The &lt;code&gt;input-required&lt;/code&gt; state is the one that matters most in practice—it's the protocol's way of letting a receiving agent ask a clarifying question without collapsing the whole interaction.&lt;/p&gt;

&lt;p&gt;The design constraint behind all of this is worth stating plainly: &lt;strong&gt;A2A assumes the agent on the other end is opaque&lt;/strong&gt;. You do not get its tools, its memory, its model, or its internal plan. You get a card, a task ID, and the states it passes through. That opacity is a feature, not a limitation. It's what allows enterprise agents to collaborate without exposing proprietary logic or sensitive data.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Google to Linux Foundation: Neutral Governance
&lt;/h2&gt;

&lt;p&gt;The protocol's lineage matters for anyone deciding whether to bet on it. Google published A2A in April 2025 and donated it to the Linux Foundation that June, with AWS, Cisco, Microsoft, Salesforce, SAP, and ServiceNow as founding organizations. By August 2025, IBM's Agent Communication Protocol merged into A2A rather than competing with it.&lt;/p&gt;

&lt;p&gt;In August 2026, A2A was accepted as a Growth Stage project at the Agentic AI Foundation. By operating under the Linux Foundation-directed AAIF alongside sibling projects like MCP, goose, and AGENTS.md, A2A is protected from single-vendor constraints. The Linux Foundation supplies the legal and operational structure but does not control technical decisions—a governance model intended to prevent A2A from being tied to Google's product roadmap.&lt;/p&gt;

&lt;p&gt;That neutrality is why engineering teams can plan a three-to-five-year deployment around it. Founding members now include every major cloud provider and enterprise SaaS platform that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Production Taught the Teams That Shipped It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Eon&lt;/strong&gt; built a multi-agent assistant called Buzz, where about twenty specialists live in the same service and a function call is enough of a boundary. Then three specialists had to leave the process. The triage agent held production credentials Buzz had no business holding. The Eon agent owned the semantic layer and shipped on a different release train. And tenant admins wanted to plug in agents their own teams built, on frameworks Eon doesn't run, whose source they would never see.&lt;/p&gt;

&lt;p&gt;A bespoke HTTP contract for each would have worked for the first two cases. The third made that impossible. You can't import a stranger's agent—you can only call it, and to call it you need a contract neither side wrote for the other. That's what A2A exists to solve. The team's conclusion: "Would we choose it again? Yes. An agent's capabilities reach us as data we read at runtime, so a tenant can add an agent to Buzz without us shipping code".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;pcell.si&lt;/strong&gt; has been running an A2A-based multi-agent system since May 2026, with 135 active agents, 12 content-generation bots, and over 200 automated patrol cycles of autonomous operation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cisco's CAIPE&lt;/strong&gt; uses multi-agent orchestration with tool calling via both MCP and A2A. Response times dropped from hours to seconds, and MTTR reduced by up to 80%.&lt;/p&gt;

&lt;p&gt;The A2A protocol now runs in production across supply chains, financial services, and mobile platforms, with native support built directly into Google Cloud, AWS Bedrock AgentCore Runtime, and Microsoft Azure AI Foundry. More than 150 organizations support the standard, and major frameworks—LangGraph, CrewAI, Pydantic AI, AG2, IBM BeeAI—have all adopted it.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Use A2A (and When MCP Is Enough)
&lt;/h2&gt;

&lt;p&gt;The decision framework that has held up in production:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use MCP when an agent needs to reach down into the world&lt;/strong&gt;. Fast, structured, request-response operations. Bounded competences that don't need their own reasoning loop. A database query, a file read, an API call. The agent stays in charge; the thing it touches stays passive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use A2A when independent agents need to delegate work across a boundary you don't fully control&lt;/strong&gt;. When the remote endpoint is genuinely an autonomous agent capable of reasoning, not just a passive tool. When the task might run for minutes or hours, and the receiving agent might ask clarifying questions or refuse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use both when your system has both vertical and horizontal needs&lt;/strong&gt;, which most production systems do. An A2A agent calls MCP to fetch context. An MCP-hosted tool triggers an A2A delegation. The layers compose.&lt;/p&gt;

&lt;p&gt;The mistake I made wasn't choosing the wrong protocol. It was not recognizing that I was solving two different problems with one mental model. When my billing agent needed to read account status from a database, that was a tool call. When it needed the refund agent to evaluate whether a refund was appropriate—to reason, check policy, and decide—that was a task delegation.&lt;/p&gt;

&lt;p&gt;I had been treating the refund agent as a function. It was a colleague.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;A2A gives you interoperability and delegation. It costs you some determinism and adds latency.&lt;/p&gt;

&lt;p&gt;Every A2A hop is a network round-trip. Every task lifecycle adds state management overhead. The protocol explicitly assumes the remote agent is opaque, which means you can't inspect its reasoning or debug its internals. That opacity is what makes cross-organization collaboration possible, but it also means that when something goes wrong, you're debugging at the boundary, not inside the other agent's brain.&lt;/p&gt;

&lt;p&gt;And A2A isn't free of failure modes. A 2026 post-mortem documented a four-agent LangChain pipeline communicating via A2A that entered a handoff loop and burned through $47,000 in tokens before anyone noticed. The fix, like the fix for any peer-to-peer system, is cycle detection and hop budgets.&lt;/p&gt;

&lt;p&gt;But here's what I've learned: the teams shipping multi-agent systems to production aren't choosing between MCP and A2A. They're drawing the boundary carefully, treating MCP as the tool-access layer and A2A as the coordination layer, and using each where it actually fits.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;When your agent needs something from another agent, does your architecture make it wait like a tool call—or does it hand off the task and move on?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Pure MCP, A2A where it counts, or a hybrid you had to discover the hard way—and what finally made you draw the line?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>MCP vs. A2A: why treating every agent like a tool is slowing your system down (and how to split the two) — MCP vs. A2A</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Sat, 26 Sep 2026 22:14:35 +0000</pubDate>
      <link>https://dev.to/alex_aslam/mcp-vs-a2a-why-treating-every-agent-like-a-tool-is-slowing-your-system-down-and-how-to-split-the-25n2</link>
      <guid>https://dev.to/alex_aslam/mcp-vs-a2a-why-treating-every-agent-like-a-tool-is-slowing-your-system-down-and-how-to-split-the-25n2</guid>
      <description>&lt;p&gt;I spent three weeks trying to fix a latency problem that didn't exist. Our multi-agent customer support system was slow unacceptably slow, with p99 latencies hitting 14 seconds by the third hop. I blamed the models. I blamed the prompts. I blamed the framework, the network, the token counts.&lt;/p&gt;

&lt;p&gt;Then a colleague looked at the architecture diagram, pointed at a box, and asked the question that ended the investigation: &lt;em&gt;"Why is your billing agent calling your refund agent like it's a tool?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I'd built a system where one agent needed to hand off a task to another agent that had its own reasoning, its own tools, and its own authority to act. But I'd wired it up as a tool call—a synchronous, request-response operation where the "tool" was actually an autonomous system. The billing agent was waiting for the refund agent the way you'd wait for a database query, while the refund agent was trying to reason, ask clarifying questions, and decide whether the request was even valid.&lt;/p&gt;

&lt;p&gt;I had confused the vertical with the horizontal. And I was paying for it in every single interaction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Distinction That Changes Everything
&lt;/h2&gt;

&lt;p&gt;The confusion is understandable. Both MCP and A2A arrived in the same eighteen-month window, both carry the word "protocol," and both promise interoperability for AI agents. Most teams still treat them as competing standards, or worse, as interchangeable.&lt;/p&gt;

&lt;p&gt;They are neither. They sit at different layers of the stack, and a serious agent deployment runs both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MCP is vertical.&lt;/strong&gt; It standardizes how a single agent reaches down into the world: calling tools, reading data, invoking APIs. The agent stays in charge, and the things it touches are passive. An MCP server that wraps &lt;code&gt;kubectl&lt;/code&gt; does not think, negotiate, or push back. It exposes capabilities and waits. MCP optimizes for fast, structured, request-and-response tool calls.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A2A is horizontal.&lt;/strong&gt; It standardizes how one agent asks another agent to do work on its behalf. The far side is not a passive tool but an autonomous system with its own model, its own tools, and its own judgment. It can accept the task, ask a clarifying question, run for six hours, or refuse. A2A optimizes for long-running, stateful negotiation between systems that do not trust each other and do not share memory.&lt;/p&gt;

&lt;p&gt;MCP gives your agent hands. A2A gives it colleagues.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Mistake That Costs You
&lt;/h2&gt;

&lt;p&gt;When you treat an agent like a tool, you inherit every assumption of tool-calling. You assume the call will complete quickly. You assume the response will be a structured value. You assume the callee has no state, no goals, and no authority to say no.&lt;/p&gt;

&lt;p&gt;None of that is true for an agent.&lt;/p&gt;

&lt;p&gt;My billing agent would call the refund agent and wait. The refund agent would try to reason through the request, discover it needed account status, and either fail or return a partial result. The billing agent would either hang, retry, or proceed with bad data. The trace showed a tool call that took 14 seconds. It should have been an A2A task delegation where the billing agent hands off work and moves on, and the refund agent gets back to it when it's done.&lt;/p&gt;

&lt;p&gt;The architecture was wrong, and no amount of prompt tuning could fix a structural mismatch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why MCP Can't Make Two Agents Talk
&lt;/h2&gt;

&lt;p&gt;MCP solves a genuinely important problem: before it, connecting five agents to ten tools meant up to fifty bespoke integrations, each one a maintenance liability. MCP collapses that into one interface. By mid-2025, the community had built thousands of active MCP servers, and OpenAI, Microsoft, and Google DeepMind all adopted the protocol. MCP passed roughly 97 million monthly SDK downloads, with public registries indexing close to twenty thousand servers. MCP has won the tool-calling layer.&lt;/p&gt;

&lt;p&gt;But MCP has no concept of a peer agent with its own goals, its own model, and its own authority to act. If your procurement agent needs your finance agent to approve a payment, MCP has nothing to say about that conversation. The finance agent is not a tool to be called; it is an actor with its own reasoning and its own right to refuse.&lt;/p&gt;

&lt;p&gt;That is a horizontal problem, and it needs a different protocol.&lt;/p&gt;

&lt;h2&gt;
  
  
  What A2A Actually Provides
&lt;/h2&gt;

&lt;p&gt;A2A handles the horizontal case: agents discovering each other, exchanging messages, and coordinating tasks across organizational and platform boundaries.&lt;/p&gt;

&lt;p&gt;Each agent publishes an "Agent Card" describing what it can do, and other agents query that card to decide what work to delegate. The protocol defines task lifecycle states, support for synchronous and asynchronous interaction, and a structure for agents that don't share memory or trust each other.&lt;/p&gt;

&lt;p&gt;By April 2026, more than 150 organizations supported A2A, with active production deployments across supply chains, financial services, and mobile platforms. Major cloud providers—including Google Cloud, AWS Bedrock AgentCore, and Microsoft Azure AI Foundry—have built native A2A support directly into their infrastructure. Enterprise SaaS platforms like ServiceNow, Salesforce, Atlassian, and SAP use A2A to connect workflows across their products.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real-World Systems Running Both
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cisco's CAIPE&lt;/strong&gt; (Community AI Platform Engineering) uses multi-agent orchestration with tool calling via both MCP and A2A. Their agents chain across Argo CD, Kubernetes, and Komodor to synthesize actionable answers. The result: response times dropped from hours to seconds, and MTTR reduced by up to 80%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Elastic's LLM newsroom&lt;/strong&gt; demonstrates the hybrid pattern in production code. A Reporter Agent delegates research tasks to a Researcher Agent via A2A, and delegates archive searches to an Archive Agent via A2A. The Archive Agent then uses MCP to access Elasticsearch tools. A2A enables agent collaboration; MCP provides tool access. They run together in the same system, each doing what it's best at.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AWS Bedrock AgentCore Runtime&lt;/strong&gt; now supports both MCP and A2A natively. Agents built using different frameworks—Strands Agents, OpenAI Agents SDK, LangGraph, Google ADK, or Claude Agents SDK—can share context, capabilities, and reasoning in a common, verifiable format. The complete A2A request lifecycle, from agent card discovery to task delegation, is supported out of the box.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Microsoft's Agent Framework team&lt;/strong&gt; ran a controlled comparison: a ski resort advisor built with A2A specialists versus the same application using distributed skills over MCP. The A2A path used six to seven model calls per request. The skills path used three. Mean elapsed time dropped from 15.48 seconds to 6.35 seconds—roughly 60% faster. The trade-off showed up in token count: the skills path consumed about 22% more tokens because the parent agent's context grew as it loaded specialist instructions. The conclusion: A2A handles collaboration between autonomous agents; MCP handles bounded competences that don't need their own reasoning loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem MCP Hasn't Solved Yet
&lt;/h2&gt;

&lt;p&gt;MCP has a scaling problem, and it's structural.&lt;/p&gt;

&lt;p&gt;Every tool definition you load into an agent's context window consumes tokens. When you connect to a handful of MCP servers, that's manageable. When you connect to dozens, loading every tool definition upfront can consume the majority of the context window before the model has even read the user's message.&lt;/p&gt;

&lt;p&gt;The numbers are brutal. Cloudflare found that tool descriptions alone would consume roughly 244,000 tokens before the agent processes a single user message. Claude Code's &lt;code&gt;/context&lt;/code&gt; shows MCP tools alone consuming 97,000 tokens—about 48% of Claude Sonnet 4.5's window. A single MCP server like GitHub can add 15-20k tokens just for tool schemas.&lt;/p&gt;

&lt;p&gt;This is the "tool explosion" problem, and it's the reason teams are moving away from naive MCP adoption. Perplexity's CTO said they're moving away from MCP internally. Cloudflare replaced MCP's tool-calling mechanism with code generation and cut token usage by 244x.&lt;/p&gt;

&lt;p&gt;But the answer isn't abandoning MCP. It's using it correctly.&lt;/p&gt;

&lt;p&gt;The MCP specification itself now recommends &lt;strong&gt;progressive discovery&lt;/strong&gt;: instead of loading every tool definition upfront, the host provides a lightweight &lt;code&gt;search_tools&lt;/code&gt; meta-tool, and the model loads full definitions only as needed. The comparison is stark: loading all tools upfront consumes ~150,000 tokens on definitions alone; progressive discovery uses ~2,000 tokens by loading only what the task requires.&lt;/p&gt;

&lt;p&gt;Microsoft's Agent Skills extension, standardized in September 2026, takes this further. Skills are addressed via a &lt;code&gt;skill://&lt;/code&gt; URI scheme with progressive disclosure: advertise at roughly 100 tokens per skill, load under 5,000 tokens, read resources and run scripts on demand. The parent agent's context window stays lean.&lt;/p&gt;

&lt;p&gt;The DADL paper from April 2026 quantifies the reduction precisely: on a catalog of 1,833 tool definitions across 20 services, Code Mode reduced the LLM context cost of tool advertisement from approximately 142,000 tokens to approximately 1,000—a 142x reduction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Decision Framework
&lt;/h2&gt;

&lt;p&gt;The research is clear about when to use what:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use MCP when an agent needs to discover and call tools at runtime&lt;/strong&gt;. Use it for fast, structured, request-response operations. Use it for bounded competences—procedures with typed operations that don't need their own reasoning loop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use A2A when independent agents need to delegate work across a boundary you don't fully control&lt;/strong&gt;. Use it when the remote endpoint is genuinely an autonomous agent capable of reasoning, not just a passive tool. Use it for long-running tasks with their own lifecycle, where the callee might ask clarifying questions or refuse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use both when your system has both vertical and horizontal needs&lt;/strong&gt;, which most production systems do. An A2A agent calls MCP to fetch context. An MCP-hosted tool triggers an A2A delegation. The layers compose.&lt;/p&gt;

&lt;p&gt;The Atlan decision guide puts it well: MCP, A2A, and ANP are a complementary stack for three jobs. MCP standardizes how an agent reaches a tool or data source. A2A standardizes how one agent delegates structured work to another. ANP standardizes how agents from different organizations authenticate without a shared broker. You rarely pick one instead of another. You adopt each as its job appears.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd Tell My Past Self
&lt;/h2&gt;

&lt;p&gt;The mistake I made wasn't choosing the wrong protocol. It was not recognizing that I was solving two different problems with one mental model.&lt;/p&gt;

&lt;p&gt;When my billing agent needed to read account status from a database, that was a tool call. MCP. Fast, synchronous, structured response.&lt;/p&gt;

&lt;p&gt;When my billing agent needed the refund agent to evaluate whether a refund was appropriate—to reason, check policy, and decide—that was a task delegation. A2A. Potentially long-running, with its own lifecycle and its own authority to say no.&lt;/p&gt;

&lt;p&gt;I had been treating the refund agent as a function. It was a colleague.&lt;/p&gt;

&lt;p&gt;The fix wasn't rewriting the system. It was reclassifying the boundaries. Tool calls stayed synchronous. Agent delegations became asynchronous task handoffs with their own state and lifecycle. Latency dropped because the billing agent stopped waiting for a reasoning process to complete before it could do anything else.&lt;/p&gt;

&lt;p&gt;The teams that are getting this right aren't choosing sides. They're drawing the boundary carefully, treating MCP as the tool-access layer and A2A as the coordination layer, and using each where it actually fits.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;When your agent needs something from another agent, does your architecture make it wait like a tool call—or does it hand off the task and move on?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Pure MCP, A2A where it counts, or a hybrid you had to discover the hard way and what finally made you draw the line?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Hybrid multi-agent systems are the pragmatic middle ground — here’s how to get control without micromanaging — Hybrid MAS</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Fri, 25 Sep 2026 18:56:34 +0000</pubDate>
      <link>https://dev.to/alex_aslam/hybrid-multi-agent-systems-are-the-pragmatic-middle-ground-heres-how-to-get-control-without-2f04</link>
      <guid>https://dev.to/alex_aslam/hybrid-multi-agent-systems-are-the-pragmatic-middle-ground-heres-how-to-get-control-without-2f04</guid>
      <description>&lt;p&gt;I spent a year building multi-agent systems and a year defending them. First against the "just use one good agent" crowd, then against the "you need a supervisor" crowd, and finally the hardest one against my own production incidents. Every architecture I tried worked beautifully in the demo and fell apart somewhere around step 12, agent 8, or the first real user who asked something the graph didn't anticipate.&lt;/p&gt;

&lt;p&gt;I built flat swarms that deadlocked. I built hierarchies that added latency without adding reliability. I built pipelines that hallucinated reconciliations when branches disagreed. Each time, I fixed the failure and created a new one. The pattern was obvious in hindsight: I was treating orchestration as a binary choice centralized or decentralized, supervisor or peer, control or autonomy when the systems that actually survived production were doing something else entirely.&lt;/p&gt;

&lt;p&gt;They were hybrid. Not as a compromise. As a deliberate architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Binary Was Always Wrong
&lt;/h2&gt;

&lt;p&gt;The 2026 research has converged on a taxonomy that formalizes what I kept rediscovering the hard way. A comprehensive survey places multi-agent architectures on a 3×2 grid: topology (centralized, decentralized, hierarchical) crossed with adaptivity (static versus dynamic-adaptive). The frameworks I'd been using all lived in one cell. LangGraph supervisor is centralized-static. CrewAI manager-mode is centralized-static. AutoGen GroupChatManager auto-select is centralized-dynamic. AgentVerse dynamic expert recruitment is decentralized-dynamic.&lt;/p&gt;

&lt;p&gt;But the real systems—the ones shipping to production and handling failures gracefully—were classified as &lt;strong&gt;hybrid&lt;/strong&gt;: systems that cannot be cleanly placed in a single cell because they deliberately combine properties along both axes.&lt;/p&gt;

&lt;p&gt;The canonical example is DyLAN, which exercises two sub-modes at once: membership mutating (dynamically selecting which agents to include) and learned coordination (learning importance scoring for that selection). Along the topology axis, its outer architecture is hierarchical while its in-team coordination is partly decentralized. It's not an anomaly. It's a design choice.&lt;/p&gt;

&lt;p&gt;A separate study on shuttle-based storage systems put the same finding more plainly: hybrid control architectures combine elements of centralized and distributed control to balance global coordination with local autonomy. Some agents, usually at higher levels, maintain centralized oversight, while others operate independently or coordinate directly with peers. The paper's conclusion was unambiguous: hybrid architectures are "particularly promising" because they meet all five improvement needs—distributed decision-making, adaptive traffic management, inter-agent coordination, modular scalability, and Industry 4.0 interoperability. An additional benefit: incremental deployability. You can add autonomy to an existing system without redesigning the entire control architecture.&lt;/p&gt;

&lt;p&gt;That last point is the one I keep coming back to. Every pure architecture demands a full rewrite. Hybrid lets you evolve.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Hybrid Actually Looks Like
&lt;/h2&gt;

&lt;p&gt;I rebuilt my customer support system around three layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A supervisory control plane&lt;/strong&gt; that holds global state, enforces policy, and handles escalation. This is centralized. It's the part that needs consistency: user identity, conversation history, compliance rules, cost budgets. The control plane doesn't route every task. It sets constraints and observes outcomes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distributed execution agents&lt;/strong&gt; that operate independently within those constraints. A billing agent resolves refund requests using its own tools and local state. A technical agent debugs issues without asking the supervisor for permission on every step. They report outcomes upward, not every action. This is the decentralized layer. It's where latency lives and where autonomy pays off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Escalation paths&lt;/strong&gt; that move work between layers when local resolution fails. When a billing agent hits an edge case—say, a refund outside policy—it doesn't hand off to a peer and hope. It escalates to the control plane with structured context. The control plane either resolves it (rare) or routes to a human (common). This is the hybrid mechanism: neither pure hierarchy nor pure peer coordination, but a defined protocol for moving between them.&lt;/p&gt;

&lt;p&gt;The research validates this structure. The Alignment Flywheel, a governance-centric hybrid MAS, specifies the same separation: a Proposer generates candidate trajectories, and a governed Safety Oracle stack returns safety scores, prediction uncertainty, audit coverage, and evidence hooks through a stable interface. The central engineering principle is &lt;strong&gt;patch locality&lt;/strong&gt;—when safety behavior needs updating, you change the oracle, not the decision components. In my system, when compliance rules changed, I updated the control plane. The execution agents kept running.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Control Mechanisms That Make It Work
&lt;/h2&gt;

&lt;p&gt;Hybrid architecture is only half the story. The other half is what you put in the seams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structured handoffs with contracts.&lt;/strong&gt; Every boundary between control plane and execution agent has an explicit schema. The execution agent declares what it will produce. The control plane validates the contract before accepting the result. A 2026 study on hybrid MAS for insurance decision support found that explicit orchestration and verification eliminated common LLM tool-use failures, and completing the inter-agent policy contract increased agreement from 91.8% to &lt;strong&gt;100%&lt;/strong&gt;. That's not incremental. That's the difference between a system you trust and one you audit manually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Human-in-the-loop as a first-class escalation path.&lt;/strong&gt; The Winnex Maestro architecture implements a four-layer validation pipeline: sequential sandbox, checklist, human validation, and partner review, with an AutoRollbackSystem that triggers on 5-second anomaly detection and completes rollback within 30 seconds. The "Strategy Room" is a facilitator-mediated multi-agent collaboration protocol with five formal phases, cryptographic audit, Ed25519 signatures, and automatic escalation on timeout. This is what hybrid control looks like when you're serious about auditability. It's not a supervisor agent hoping for the best. It's a system designed so that when the AI is uncertain, a human—with the right context—makes the call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stateful policy enforcement beyond prompts.&lt;/strong&gt; Databricks' Omnigent, released in alpha under Apache 2.0, sits above existing agents and provides an orchestration layer for teams using multiple models and frameworks. Its control functions include stateful policies that track agent actions and enforce guardrails beyond prompt-based controls, plus cost budgeting and operating system sandboxing. The key word is &lt;em&gt;stateful&lt;/em&gt;. A prompt says "don't exceed the budget." A stateful policy tracks spend in real time and halts execution when the limit is hit. One is a suggestion. The other is a control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Centralized training, decentralized execution (CTDE).&lt;/strong&gt; This is the reinforcement learning pattern that maps directly onto hybrid MAS. Agents learn globally optimal strategies during training, then operate independently during execution, coordinating through local communication and conflict resolution protocols. A study on legacy ERP modernization found that CTDE-based decentralized agent orchestration increased throughput by &lt;strong&gt;6.9%&lt;/strong&gt; over a monolithic setup while offering increased agility without platform replacement risk. The trade-off was a 6.3% error rate from aggressive allocation policies—which is exactly the kind of thing the control plane is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Actually Ships
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Vivix Vidros Planos&lt;/strong&gt;, a glass manufacturer, moved beyond traditional AI copilots to orchestrate a hybrid workforce of people and AI agents, reducing quality complaint resolution times by &lt;strong&gt;75%&lt;/strong&gt;—from 10 days down to 2.5. The architecture isn't pure automation. It's a hybrid workforce where humans and agents each own the parts they're best at.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;EY&lt;/strong&gt; embedded AI agents into its global audit operations with a mandatory human-in-the-loop model. Each auditor using multi-agent tools retains ultimate professional skepticism and legal responsibility for reviewing all AI-generated workpapers. The agents handle the scale. The humans handle the judgment. Neither replaces the other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SAP&lt;/strong&gt; evaluated DAG Plan &amp;amp; Execute and ReAct across 208 production scenarios at Persona, Department, and Enterprise scales. The finding that matters here: &lt;strong&gt;scale, not task complexity, dominates orchestration performance&lt;/strong&gt;. Both architectures degraded at enterprise scale as agent discovery noise became the primary bottleneck. SAP's Task Manager—a hybrid control mechanism—reduced high-priority queue latency by 14-75% and improved related-event correctness by over 20 percentage points at enterprise scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cisco's JARVIS&lt;/strong&gt; (now open-sourced as part of the CNOE Agentic AI Community) uses LangGraph to implement a hierarchical supervisor MAS where distributed agents are connected using the AGNTCY Agent Connect Protocol. The reflection agent determines whether the system has addressed a request, routing queries back to the supervisor or finalizing. It's hierarchical in structure but decentralized in execution—agents run in their own environments, communicating through standardized protocols.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;IBM watsonx Orchestrate&lt;/strong&gt; has evolved into what IBM calls an "agentic control plane" for the multi-agent era, where organizations can deploy agents from any source with consistent policy enforcement and accountability. IBM's positioning is explicit: the control plane should be hybrid to avoid lock-in, and a clear majority of enterprises (51% by end of 2026) expect a hybrid control plane—provider-native plus external orchestration—rather than handing control to a provider-managed service.&lt;/p&gt;

&lt;p&gt;The pattern is consistent. The teams winning with multi-agent systems aren't choosing centralized or decentralized. They're building a control plane that provides consistency and a distributed execution layer that provides speed and autonomy, with defined protocols for moving between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;Hybrid architecture isn't free. It's more complex than either pure approach. You're designing two systems—the control plane and the execution layer—and the protocol that connects them.&lt;/p&gt;

&lt;p&gt;You're accepting that some decisions will be slower because they require escalation. You're accepting that your control plane is a critical dependency. You're accepting that the boundary between "local autonomy" and "central control" is something you'll be tuning for months.&lt;/p&gt;

&lt;p&gt;But here's what I've learned from every production incident, every 3 AM debugging session, and every post-mortem where the root cause was "the architecture didn't let us see what was happening": &lt;strong&gt;the systems that survive are the ones where control and autonomy are deliberately balanced, not ideologically chosen&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A flat swarm can't tell you why it deadlocked. A pure supervisor can't scale past seven agents. A pipeline can't stop a hallucination from cascading. A hybrid system can do all three—because it has a control plane that sees everything, execution agents that act locally, and escalation paths that surface problems before they compound.&lt;/p&gt;

&lt;p&gt;The research is clear that most production systems blend patterns. The survey's conclusion is explicit: the right combination depends on task structure, agent count, fault tolerance requirements, and cost budget. There is no universally correct architecture. There is only the architecture that fits your problem and degrades gracefully when your problem changes.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;When your multi-agent system fails at 2 AM, does your architecture let you see where control broke down and what the execution layer was actually doing—or are you back to reading logs and guessing?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Pure hierarchy, flat swarm, or the hybrid middle ground and what finally made you stop choosing sides?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Your agent org chart is flat, and that’s why it fails at scale — the 3-tier pattern that handles 50+ agents — Paperclip Pattern</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Fri, 25 Sep 2026 16:38:08 +0000</pubDate>
      <link>https://dev.to/alex_aslam/your-agent-org-chart-is-flat-and-thats-why-it-fails-at-scale-the-3-tier-pattern-that-handles-43nh</link>
      <guid>https://dev.to/alex_aslam/your-agent-org-chart-is-flat-and-thats-why-it-fails-at-scale-the-3-tier-pattern-that-handles-43nh</guid>
      <description>&lt;p&gt;I spent three weeks watching a flat swarm of eighteen agents slowly eat itself. Not crash. Not error. Just gradually degrade until every output contradicted the last, and no agent could tell you why. The trace was a beautiful, terrifying mess of peer-to-peer handoffs with no owner, no escalation path, and no single place where anyone—human or agent—could say “this is wrong, stop.”&lt;/p&gt;

&lt;p&gt;That was the week I stopped believing in flat agent org charts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Flat Org Chart Illusion
&lt;/h2&gt;

&lt;p&gt;The pitch for flat multi-agent systems is intoxicating. No supervisor bottleneck. No context saturation at the top. Every agent talks to every other agent, information flows freely, and the system self-organizes. OpenAI’s Swarm library, LangGraph’s handoff primitives, and a dozen frameworks all made peer-to-peer coordination feel like the future.&lt;/p&gt;

&lt;p&gt;It works beautifully for three to five agents. Microsoft’s own orchestration training materials are blunt about what happens when you scale: coordination overhead grows quadratically, errors cascade unpredictably, and subdomains have no clear ownership. With N agents, you’re managing N(N-1)/2 possible relationships. A group chat with ten participants is chaos when nobody respects turn order.&lt;/p&gt;

&lt;p&gt;I hit all three failure modes in the same week. Compliance and risk agents contradicted each other on the same account, and nobody was designated to adjudicate. A billing agent failed silently, and the workflow either aborted entirely or continued with incomplete analysis depending on which path the handoff took. And every agent was nominally responsible for everything, which meant no agent was actually responsible for anything.&lt;/p&gt;

&lt;p&gt;The research confirms what I felt in production. A 2026 paper on multi-agent organization found that hierarchical settings improved performance over flat systems by &lt;strong&gt;102.73%&lt;/strong&gt; while reducing token usage by &lt;strong&gt;74.52%&lt;/strong&gt; on SQuAD 2.0. Another study found that a population-independent “causal floor” on achievable error in flat systems can only be removed by hierarchical organization. The coordination cost of flat topologies eventually eclipses the gains from parallelism.&lt;/p&gt;

&lt;p&gt;The flat org chart isn’t wrong. It’s just not scalable. And the pattern that fixes it has a name that sounds absurd until you’ve lived the alternative.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Paperclip Pattern
&lt;/h2&gt;

&lt;p&gt;In early March 2026, a pseudonymous developer dropped a Node.js server on GitHub with a tagline that sounded like a joke: “open-source orchestration for zero-human companies.” Three weeks later, Paperclip had crossed 44,900 stars. CrewAI took most of a year to reach comparable numbers.&lt;/p&gt;

&lt;p&gt;The architecture is deliberately unglamorous. You configure a CEO agent, give it a high-level goal, and it delegates down through a tree of managers and workers. Every agent reports to exactly one manager. Workers spawn for specific tasks and have no awareness of anything outside their assigned scope. Paperclip calls its coordination layer an org chart—complete with ticketing, role definitions, monthly budgets per agent, and governance gates that prevent agents from hiring subordinates without approval.&lt;/p&gt;

&lt;p&gt;Under the hood, agents don’t run continuously. They fire in short execution windows called “heartbeats” triggered by a scheduler. This keeps costs predictable and prevents runaway token spend, which has been a recurring pain point in persistent agent deployments. Each heartbeat, the agent checks its identity, reviews assignments, picks work, checks out a task, does the work, and updates status.&lt;/p&gt;

&lt;p&gt;The pattern isn’t new. IBM’s enterprise AI teams documented hierarchical decomposition as the standard approach for large-scale deployments since at least 2024: domain-specific agent clusters, each supervised by a mid-tier coordinator, all reporting to a strategic orchestrator at the top. What Paperclip did was make it approachable enough that a developer could have a two-level hierarchy running with budget controls and approval gates in an afternoon.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Research Backing the Pattern
&lt;/h2&gt;

&lt;p&gt;The Paperclip pattern isn’t just a viral framework. It’s the most visible implementation of a three-tier architecture that the research literature has been converging on for two years.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OrgAgent&lt;/strong&gt;, a company-style hierarchical multi-agent framework from researchers at CUHK and IBM, decomposes reasoning into three layers: a &lt;strong&gt;governance layer&lt;/strong&gt; for planning and resource allocation, an &lt;strong&gt;execution layer&lt;/strong&gt; for task solving and review, and a &lt;strong&gt;compliance layer&lt;/strong&gt; for final answer control. The results across reasoning tasks and model sizes were consistent: hierarchical organization outperformed flat collaboration in most settings while reducing token consumption. For GPT-OSS-120B, the hierarchical setting improved performance by 102.73% over flat while reducing token usage by 74.52%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SAP’s enterprise AI research&lt;/strong&gt; evaluated DAG Plan &amp;amp; Execute and ReAct across 208 production scenarios spanning Persona (&amp;lt;10 agents), Department (20-80), and Enterprise (200) scales. The finding was stark: &lt;strong&gt;scale, not task complexity, dominates orchestration performance&lt;/strong&gt;. Both architectures performed well at small scale but degraded at enterprise scale as agent discovery noise became the primary bottleneck. Simple tasks degraded more sharply than complex ones. The Task Manager they introduced reduced high-priority queue latency by 14-75% and improved related-event correctness by over 20 percentage points at enterprise scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agensh&lt;/strong&gt;, a self-organized multi-agent harness, scaled from 1 to 128 agents and raised mean test-pass rates from 19.31% to 28.78%. On pandoc, scaling from 1 to 1,024 agents raised test-pass rates from 33.89% to 55.06%. The key architectural insight: rather than a central orchestrator, concurrent workers execute a cooperation loop, continuously gathering context, claiming sub-tasks, taking action, sharing findings, verifying results, and merging progress asynchronously.&lt;/p&gt;

&lt;p&gt;The consistent finding across all three: &lt;strong&gt;hierarchy helps most when tasks benefit from stable skill assignment, controlled information flow, and layered verification&lt;/strong&gt;. Artificial hierarchies add cost without value. The pattern works when it maps to natural problem decomposition—breaking a product launch into marketing, engineering, and operations domains, for instance. It fails when the tree is imposed on tasks that don’t actually decompose that way.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a 3-Tier Hierarchy Actually Looks Like
&lt;/h2&gt;

&lt;p&gt;The architecture that handles 50+ agents is not three times more complicated than a flat swarm. It’s structurally simpler.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 1: Governance.&lt;/strong&gt; The CEO agent receives the company goal, proposes a strategy for approval, breaks approved goals into tasks, and assigns them to managers based on role and capability. It doesn’t execute work. It plans, allocates, and monitors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 2: Execution.&lt;/strong&gt; Manager agents receive workstreams from the CEO and decompose them into subtasks for their reports. They review worker output, handle escalation, and compress information before sending it upward. A manager doesn’t need to see every token from every worker—it reads structured summaries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 3: Workers.&lt;/strong&gt; Specialist agents execute the actual work. They report to exactly one manager, operate in short heartbeat windows, and escalate blockers through their own chain of command. Cross-team tasks are possible, but the receiving agent’s manager handles escalation if the task gets blocked.&lt;/p&gt;

&lt;p&gt;The magic is what doesn’t happen. The CEO never sees raw worker output. Managers never debug worker prompts. Workers never coordinate with peers outside their team. Local summarization at each layer prevents context saturation at the top. The global N² conflict problem decomposes into multiple local N’² conflicts—and N’ is much smaller.&lt;/p&gt;

&lt;p&gt;SAP’s research quantified the payoff: &lt;strong&gt;scale dominates performance&lt;/strong&gt;, and the hierarchy is what makes scale survivable. A Task Manager handling priority inference, related-event merging, and preemption reduced high-priority queue latency by up to 75% at enterprise scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who’s Actually Shipping This
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Toyota Motor North America&lt;/strong&gt; runs a platform with over 50 agents in production across manufacturing, research, and supply chain operations. ToyotaGPT’s multi-agent system reduced development timelines dramatically, with the architecture organized hierarchically to manage the scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Midea&lt;/strong&gt;, the Chinese manufacturing giant, built what it calls a “factory brain” that orchestrates &lt;strong&gt;14 AI agents across 38 business scenarios&lt;/strong&gt;. The system acts as the plant’s nervous system, built on a distributed, scalable multi-agent architecture with agent-to-agent communication and industrial large model inference engines. Average efficiency improvement exceeded &lt;strong&gt;80%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lyft&lt;/strong&gt; rebuilt their customer support system on a router-based multi-agent architecture using LangGraph. A meta agent acts as a stateful router, dispatching to specialized subgraphs that are themselves full &lt;code&gt;StateGraph&lt;/code&gt; instances. Agent development accelerated from roughly &lt;strong&gt;six months to just a few weeks&lt;/strong&gt;, hallucination rates dropped by &lt;strong&gt;20%&lt;/strong&gt;, and AI resolution rates increased by &lt;strong&gt;16%&lt;/strong&gt;. They handle millions of interactions for riders and drivers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SAP&lt;/strong&gt; deployed event-driven multi-agent orchestration across 208 production-derived enterprise scenarios. Their Task Manager architecture, designed specifically for continuous operation at enterprise scale, reduced high-priority queue latency by 14-75% and improved related-event correctness by over 20 percentage points.&lt;/p&gt;

&lt;p&gt;These aren’t demos. They’re production systems where failure means lost revenue, broken supply chains, or regulatory exposure. And every one of them chose hierarchy over flat coordination.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Use the Paperclip Pattern (and When Not To)
&lt;/h2&gt;

&lt;p&gt;The research is clear about when hierarchy pays off. A 2026 study found that &lt;strong&gt;hierarchical systems with two delegation levels outperform flat architectures by 28% on complex multi-step tasks&lt;/strong&gt;. Adding a third delegation tier provides only &lt;strong&gt;7% additional improvement&lt;/strong&gt; while increasing latency by &lt;strong&gt;40%&lt;/strong&gt;. Most production teams cap hierarchies at two levels for this reason.&lt;/p&gt;

&lt;p&gt;Use the Paperclip pattern when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your task decomposes naturally into domains.&lt;/strong&gt; Product launch → marketing, engineering, operations. Customer support → billing, refunds, technical, account. Supply chain → inventory, procurement, logistics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You’re scaling beyond 10-12 agents.&lt;/strong&gt; The coordination overhead of flat systems grows quadratically. Hierarchy linearizes it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need stable skill assignment and controlled information flow.&lt;/strong&gt; Hierarchy helps most when tasks benefit from layered verification and clear ownership.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fault tolerance matters.&lt;/strong&gt; When a worker fails, the manager can reassign or escalate. When a manager fails, the CEO can redistribute its workstreams. There’s no single point of failure at the system level—only at the individual agent level.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don’t use it when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your task is genuinely flat.&lt;/strong&gt; A research pipeline with six independent sources doesn’t need a CEO and three managers. It needs a fan-out and a fan-in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You’re under five agents.&lt;/strong&gt; The overhead of hierarchy exceeds the coordination savings. Start flat, add structure when flat produces concrete pain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The hierarchy is artificial.&lt;/strong&gt; If you can’t articulate why a domain deserves its own manager, it probably doesn’t. Artificial hierarchies add cost without value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency is the binding constraint.&lt;/strong&gt; Every tier adds a communication hop. Hierarchical patterns trade latency for scalability and fault tolerance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Trade-Off You’re Accepting
&lt;/h2&gt;

&lt;p&gt;Hierarchy gives you scalability, fault tolerance, and debuggability. It costs you latency and some flexibility.&lt;/p&gt;

&lt;p&gt;Every escalation up the chain adds round-trips. Every structured summary at a layer boundary loses some detail. The CEO never sees what the workers actually saw—it sees what the managers chose to summarize. That’s the price of preventing context saturation at the top.&lt;/p&gt;

&lt;p&gt;But here’s what I’ve learned: the teams that are winning with 50+ agents aren’t the ones who found the perfect prompt. They’re the ones who stopped treating agent orchestration as a prompt engineering problem and started treating it as an &lt;strong&gt;organizational design problem&lt;/strong&gt;. They’re writing org charts. They’re defining reporting lines. They’re thinking about escalation paths and information compression at every layer.&lt;/p&gt;

&lt;p&gt;The Paperclip pattern isn’t exciting. It’s not the pattern you demo. It’s the pattern you build when you’ve been burned by flat coordination and you need something that survives contact with real workloads at scale.&lt;/p&gt;

&lt;p&gt;So here’s my question: &lt;strong&gt;When your flat swarm fails at agent number twelve, does your architecture let you escalate the problem or does it just quietly degrade until nobody can tell you what went wrong?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I’d love to hear where you’ve landed. Flat swarm, two-tier hierarchy, or the full CEO/manager/worker stack and what finally made you change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Swarms sound like the future, but most teams can’t debug them — here’s when peer-to-peer actually pays off — Swarms / P2P</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Thu, 24 Sep 2026 22:42:05 +0000</pubDate>
      <link>https://dev.to/alex_aslam/swarms-sound-like-the-future-but-most-teams-cant-debug-them-heres-when-peer-to-peer-actually-318k</link>
      <guid>https://dev.to/alex_aslam/swarms-sound-like-the-future-but-most-teams-cant-debug-them-heres-when-peer-to-peer-actually-318k</guid>
      <description>&lt;p&gt;I spent a week watching a swarm of six agents burn through $400 in tokens, hand off a billing question to a refund specialist, get handed back, and loop for forty minutes without ever answering the customer. Every individual agent was working. The system was deadlocked. Nobody had thrown an error. Nobody had timed out. The trace just showed a beautiful, expensive circle.&lt;/p&gt;

&lt;p&gt;That's the moment I understood why swarms sound like the future and why most teams can't ship them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pitch That Sells Itself
&lt;/h2&gt;

&lt;p&gt;The swarm pattern is intoxicating on paper. No central supervisor. No bottleneck. Each agent owns a domain and decides for itself when to answer or hand off to a peer. OpenAI's Swarm library introduced the primitive in late 2024: agents transfer control to each other by specialization, and the "manager" disappears. LangGraph's swarm library, Microsoft's Agent Framework, and AWS Bedrock AgentCore all implement some version of peer-to-peer handoff. AWS's own documentation describes it as "dynamic and peer-to-peer, based on agent discoveries and needs" with "distributed" decision-making where agents make local choices about task planning and handoffs.&lt;/p&gt;

&lt;p&gt;I bought the pitch. I built a six-agent customer support swarm: triage, billing, refunds, technical, account, and escalation. Each agent had its own tools, its own prompt, its own domain knowledge. No supervisor meant no context saturation. No bottleneck meant horizontal scaling. The demo was flawless.&lt;/p&gt;

&lt;p&gt;Then I turned it loose on real traffic and discovered what the research already knew.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Debugging Wall Nobody Warns You About
&lt;/h2&gt;

&lt;p&gt;The first failure was quiet. A customer asked about a refund. Triage handed to billing. Billing handed to refunds. Refunds handed back to billing because the refund policy check needed account status. Billing handed to account. Account handed back to refunds. Refunds handed to triage. No agent produced a wrong answer. No agent crashed. The system simply spun, and the customer waited.&lt;/p&gt;

&lt;p&gt;The second failure was louder but equally invisible. Two agents returned contradictory information about the same account. The synthesis step, which I'd naively built to "aggregate results," picked one arbitrarily. The final answer was wrong in a way that only made sense if you replayed the entire trace and read every tool output at every hop.&lt;/p&gt;

&lt;p&gt;This is the debugging problem that kills swarms. Sentry's engineering team put it perfectly: "The bug is in the interaction between them, and no single agent's logs will show it". Multi-agent traces are DAGs, not trees. Blame is distributed. The worst failures look fine because an agent returns a plausible-but-thin result, the next agent incorporates it without question, and by the time the output arrives, weak data has been confidently summarized through multiple layers.&lt;/p&gt;

&lt;p&gt;I went looking for tooling. What I found was a research literature that had already named my exact problem: failure attribution in multi-agent systems remains "underexplored and labor-intensive," with verbose system logs turning debugging into a bottleneck. A benchmark paper called TraceElephant found that full observability improves step-level attribution from 16% with output-only inspection to 28-30% with full traces. Even with perfect observability, you're catching fewer than a third of failures.&lt;/p&gt;

&lt;p&gt;That's the wall. Swarms give you distribution and resilience. They take away the single point of observation that made centralized systems debuggable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Research Actually Says About When Swarms Pay Off
&lt;/h2&gt;

&lt;p&gt;The ICML and survey literature is clear about the trade-offs. A 2026 survey of LLM-based multi-agent orchestration frameworks found that decentralized patterns trade &lt;strong&gt;low debugging ease&lt;/strong&gt; for &lt;strong&gt;high scalability and fault tolerance&lt;/strong&gt;. Anthropic's guidance is blunter: start with a single richer agent, add multi-agent only when you can name the specific constraint it relieves. Multi-agent runs measured at roughly 15x the tokens of a chat interaction, and moving from a single agent to a team typically adds another ~4x on top.&lt;/p&gt;

&lt;p&gt;SWARM+ is the clearest research validation that decentralized orchestration can scale if you build the right primitives. A fully decentralized workload management system, it scales coordination to &lt;strong&gt;990 distributed agents&lt;/strong&gt; with approximately 1-second per-job selection time at 110 agents. Under distributed failures, it maintains over &lt;strong&gt;97% job completion&lt;/strong&gt; with graceful degradation. Under correlated site outages—the worst case—completion drops to &lt;strong&gt;95.6%&lt;/strong&gt;, and mean selection time stays under 2.5 seconds. The system doesn't lose jobs when agents die. It autonomously reselects orphaned work through delegation monitoring, triggering additional consensus rounds to recover.&lt;/p&gt;

&lt;p&gt;That's the fault tolerance argument in numbers. But SWARM+ isn't an LLM agent swarm. It's a distributed systems protocol wearing an agent costume. The teams shipping LLM swarms to production are the ones who treat the pattern as a distributed systems problem, not a prompt engineering problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who's Actually Shipping This
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Included Health&lt;/strong&gt; built Dot, a healthcare navigation system on a federated multi-agent architecture using LangGraph. It's not a pure swarm—Dot runs through a main "supergraph" with sub-workflows handling urgent care, scheduling, and behavioral health—but the architecture is deeply peer-oriented. Different product teams own different graph sections. When a member moves between services, Deep Agents' filesystem lets the outgoing agent summarize the conversation and pass both the summary and a file path to the full history, so members don't repeat themselves across team boundaries. Human handoff is built into the graph through durable execution: when Dot reaches uncertainty, it pauses, routes to a human care advocate, and resumes with the context of what the human did. Dot launched with &lt;strong&gt;75% higher chat engagement&lt;/strong&gt; and over &lt;strong&gt;99% high-risk situation detection&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LangChain's agentic engineering report&lt;/strong&gt; describes a pilot of 20+ debugging workflows where coordinated agent execution produced a &lt;strong&gt;93% reduction in time-to-root-cause&lt;/strong&gt; compared to historical baselines. Worker agents retrieve context that extends beyond source code, notify other agents, and trace agentic activity. The swarm isn't the product—it's the debugging infrastructure.&lt;/p&gt;

&lt;p&gt;The pattern that emerges from these deployments is consistent: swarms work when the topology is &lt;strong&gt;explicitly designed&lt;/strong&gt;, not when agents are given a list of peers and told to figure it out.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Guardrails That Make Swarms Survivable
&lt;/h2&gt;

&lt;p&gt;The failure mode I hit—the infinite handoff loop—is the primary one. The literature calls it handoff cycles, and the fix is structural, not prompt-based. A robust swarm &lt;strong&gt;detects handoff cycles&lt;/strong&gt; (a peer already in the trace), &lt;strong&gt;detects dead ends&lt;/strong&gt; (no peer fits), and &lt;strong&gt;escalates to a human&lt;/strong&gt; when either happens or the hop budget runs out. The standard guidance recommends a chain depth limit of three to four handoffs before escalation.&lt;/p&gt;

&lt;p&gt;But cycle detection alone isn't enough. The deeper fix is observability that treats every handoff as a &lt;strong&gt;first-class span&lt;/strong&gt; in a distributed trace. OpenTelemetry-based tooling like llmas-otel combines distributed tracing with fault injection, letting you target specific interaction points and inspect what each agent saw at each hop. SwarmTrace, a time-travel debugger for multi-agent pipelines, builds span trees where every agent action is an OTel graph node with replay from any step.&lt;/p&gt;

&lt;p&gt;The mental model shift is this: &lt;strong&gt;a swarm is a distributed system&lt;/strong&gt;, and distributed systems need distributed tracing. Parent-child spans, state snapshots, checkpoint replay. The same primitives that made microservices debuggable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Fits in Your Architecture
&lt;/h2&gt;

&lt;p&gt;I'm not going to tell you swarms are the answer. The research is clear that most production systems blend patterns. A hierarchical backbone might use decentralized handoffs inside each team. A supervisor might hand off sub-problems to peer groups. The right combination depends on task structure, agent count, fault tolerance requirements, and cost budget.&lt;/p&gt;

&lt;p&gt;Use swarms when your problem is genuinely open-ended, when agents need to negotiate across domains, or when the topology of the work isn't known at design time. Use them when &lt;strong&gt;fault tolerance matters more than debuggability&lt;/strong&gt;. Use them when you've built the tracing infrastructure to see what's happening at every hop.&lt;/p&gt;

&lt;p&gt;Don't use them because they're elegant. Don't use them because the demo looked good. And for the love of everything, don't use them without cycle detection and hop budgets.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;When your swarm deadlocks at hop seven, does your architecture show you the cycle—or do you find out when the token bill arrives?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Peer-to-peer handoffs, a supervisor you never fully escaped, or a hybrid that finally made sense and what finally made you change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Fan-out agents can 10x your throughput or 10x your chaos — the concurrency pattern that keeps both — Concurrent / Fan-Out Pattern</title>
      <dc:creator>Alex Aslam</dc:creator>
      <pubDate>Thu, 24 Sep 2026 13:18:41 +0000</pubDate>
      <link>https://dev.to/alex_aslam/fan-out-agents-can-10x-your-throughput-or-10x-your-chaos-the-concurrency-pattern-that-keeps-both-32ln</link>
      <guid>https://dev.to/alex_aslam/fan-out-agents-can-10x-your-throughput-or-10x-your-chaos-the-concurrency-pattern-that-keeps-both-32ln</guid>
      <description>&lt;p&gt;I once watched a fan-out of 40 parallel agents finish in under three minutes and produce a result so contradictory that the downstream synthesis agent hallucinated a reconciliation just to make the pieces fit. Every individual agent had done its job well. The chaos wasn’t in any single worker. It was in the space between them.&lt;/p&gt;

&lt;p&gt;That was the day I stopped thinking of fan-out as "throughput" and started thinking of it as "concurrency control with an LLM on top."&lt;/p&gt;

&lt;h2&gt;
  
  
  The Efficiency That Bites Back
&lt;/h2&gt;

&lt;p&gt;The fan-out pattern is simple to describe. Take a complex task, split it into independent subtasks, spawn an agent per subtask, run them concurrently, then fan-in to aggregate. Microsoft's Agent Framework documentation describes it exactly this way—fan-out and fan-in edges allow multiple branches to run simultaneously, aggregate their results, and scale workflows efficiently.&lt;/p&gt;

&lt;p&gt;The appeal is obvious. Instead of one agent sequentially reading ten documents, ten agents read one document each. Wall-clock time approaches the slowest branch, not the sum. Resonate's durable fan-in documentation puts the benefit plainly: total wall time approaches the slowest channel, not the sum.&lt;/p&gt;

&lt;p&gt;I built my first serious fan-out for a research pipeline: one question, six sources, six subagents, one synthesis. It was beautiful in the demo. Then production happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Chaos I Didn't See Coming
&lt;/h2&gt;

&lt;p&gt;The first sign was subtle. Two agents returned conflicting numbers for the same metric. Not wrong numbers—conflicting. One had read a cached page, the other a live one. Both were "correct" against their sources. The synthesis agent picked one arbitrarily and moved on.&lt;/p&gt;

&lt;p&gt;The second sign was louder. Three agents wrote to the same shared state object during the fan-out. One wrote a validated result, another overwrote it with a stale one. The fan-in consumed the stale value. No error. No exception. Just a quietly corrupted output.&lt;/p&gt;

&lt;p&gt;I went looking for a framework bug. What I found was a research literature that had already named my problem.&lt;/p&gt;

&lt;p&gt;A 2026 position paper at ICML argues that many multi-agent failures are &lt;strong&gt;fundamentally concurrency control problems&lt;/strong&gt;. Agents concurrently read and write shared state, and long LLM inference windows amplify the risk of stale reads, lost updates, and inconsistent outcomes. The paper maps failure modes commonly attributed to "coordination" or "communication" breakdowns directly onto classical concurrency anomalies.&lt;/p&gt;

&lt;p&gt;That hit hard. I had been thinking about my fan-out as a workflow problem. It was a database problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Anomalies That Kill Fan-Out
&lt;/h2&gt;

&lt;p&gt;The arXiv paper "Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent LLM Systems" formalized four specific anomalies that every fan-out engineer should know by heart:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stale-generation&lt;/strong&gt;: An agent reads shared state, begins a long generation, and commits a result based on a state that has since changed. The paper's travel-booking example: a flight agent reads "trip date = June 14," drafts a reservation for several seconds, and during that window the user updates the date to June 21. The agent commits the original date. No fault occurred at any layer. Yet the system produced an external effect no current state justifies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phantom-tool&lt;/strong&gt;: Two agents discover and register the same tool concurrently. One overwrites the other's registration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Causal-cascade&lt;/strong&gt;: Agent A reads, Agent B writes a correction, Agent A proceeds on the stale read. In message-based multi-agent systems, this maps directly to the write-after-read pattern.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool-effect reordering&lt;/strong&gt;: Concurrent agents invoke tools in an order that produces effects inconsistent with the intended sequence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The paper also found a silent lost update in ByteDance's deer-flow and tool-effect reordering in LangGraph's ToolNode on unmodified output. These aren't exotic. They're in the tools we're using today.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Semantic Problem Nobody Warns You About
&lt;/h2&gt;

&lt;p&gt;Structural concurrency anomalies are the first layer. The second layer is worse: &lt;strong&gt;semantic conflicts&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two agents can both be structurally correct and semantically incompatible. One returns "revenue grew 12% year over year." Another returns "revenue declined 3% quarter-over-quarter." Both are true. Both are correct extractions from valid sources. The synthesis agent has no protocol for deciding which claim should currently be trusted.&lt;/p&gt;

&lt;p&gt;A 2026 paper on Semantic Consensus identifies this as &lt;strong&gt;Semantic Intent Divergence&lt;/strong&gt;—cooperating LLM agents develop inconsistent interpretations of shared objectives due to siloed context. The paper's framework achieved 100% workflow completion by detecting and resolving these conflicts before actions were committed, compared to 25.1% for the next-best baseline.&lt;/p&gt;

&lt;p&gt;I didn't have that framework. I had a synthesis prompt that said "reconcile any conflicts" and hoped for the best.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Works
&lt;/h2&gt;

&lt;p&gt;I rebuilt the fan-out around three structural changes that map directly onto the research.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared state needs isolation levels, not hope.&lt;/strong&gt; The ICML position paper's argument is that concurrency control is the missing discipline. Snapshot isolation adds roughly 8% tokens on one workload; pessimistic locking costs 1.6-2.3×, not the order-of-magnitude penalty commonly assumed. I moved to a model where each fan-out branch writes to its own scoped state, and the fan-in merges explicitly. No branch reads another branch's in-flight writes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fan-in needs adversarial verification, not passive aggregation.&lt;/strong&gt; A July 2026 paper on verify-gated fan-out found that the empirical Pareto frontier is dominated by the cheapest tier at minimum breadth—and that production harnesses increasingly close every fan-out with adversarial verification. The fan-in isn't a summarizer. It's a gate. If two branches disagree, the disagreement is surfaced, not smoothed over.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Durable promises, not best-effort parallelism.&lt;/strong&gt; Resonate's fan-out implementation makes every spawn and every await a durable promise. If one branch fails and retries, the others stay checkpointed and don't re-execute. Crash recovery only re-runs whatever was in flight when the worker died. I had been using &lt;code&gt;Promise.all&lt;/code&gt;, which collapses latency but gives up checkpointing: a crash mid-batch re-runs everything, and a retry blasts the same side effect twice.&lt;/p&gt;

&lt;p&gt;The paper's line is the one I keep coming back to: &lt;strong&gt;the fact that each is durable is what makes this safe—every other implementation has subtle re-execution bugs&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who's Actually Shipping This
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Lyft&lt;/strong&gt; runs a parallel safety fan-out in production. Before any LLM reasoning, malicious-intent and safety-issue detectors run concurrently. The fan-out adds 90ms to each request but lifts block recall to &lt;strong&gt;99.2%&lt;/strong&gt; on red-team probes. Their meta-agent router holds conversation state and re-routes mid-chat when intent shifts, while safety checks fan out in parallel rather than blocking on a single sequential gate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Amazon's Expansion-Contraction&lt;/strong&gt; pattern, published at ACM CAIS 2026, uses concurrent path analysis over a domain graph. The results: &lt;strong&gt;98.2% accuracy&lt;/strong&gt; on a production supply chain, &lt;strong&gt;100% on public benchmarks&lt;/strong&gt;, outperforming single agent baselines by &lt;strong&gt;14+ percentage points&lt;/strong&gt;, with concurrent path analysis yielding up to &lt;strong&gt;1.43× speedup&lt;/strong&gt; and investigation caching reducing token usage by up to &lt;strong&gt;93.9%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;lionagi&lt;/strong&gt;, a governed multi-agent orchestration framework, exposes fan-out as a first-class CLI operation: three workers in parallel, then a synthesis pass. Its architecture keeps branches, sessions, and flows as ordinary Python objects with typed, inspectable state—no opaque blobs, no hidden prompt assembly.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Fan Out (and When Not To)
&lt;/h2&gt;

&lt;p&gt;The Bright Data engineering team's writeup is the most honest guidance I've found. Fan-out is well suited to &lt;strong&gt;reading tasks&lt;/strong&gt;. Actions that &lt;strong&gt;write&lt;/strong&gt; should not be delegated to subagents in the same way. Their rule: fan out when the question has independent parts and the subagents only read. Anything that writes stays on the root agent behind an approval gate.&lt;/p&gt;

&lt;p&gt;Three 2026 papers report limits on when multi-agent systems help. Silo-Bench found that as a system grows, the cost of coordinating agents cancels the gains from running them in parallel. A Nature Machine Intelligence paper reported that the more capable models tested gained less from collaboration. And a matched-budget study from April 2026 found that single agents matched or outperformed multi-agent systems when reasoning tokens were held constant. &lt;strong&gt;Fan-out only becomes competitive when a single agent stops using its context window well, or when more compute is spent&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Anthropic's workflow guidance is equally direct: move to parallel only when latency is the bottleneck and tasks are independent. Start sequential. Default to sequential. Fan out when you can measure the latency gain against the coordination cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-Off You're Accepting
&lt;/h2&gt;

&lt;p&gt;Fan-out gives you throughput. It costs you determinism.&lt;/p&gt;

&lt;p&gt;Every concurrent branch is a place where state can diverge. Every fan-in is a reconciliation you have to design, not hope for. The coordination cost grows with agent count, and the semantic conflict rate grows with the diversity of sources. A benchmark on financial document processing found that hierarchical architectures occupy the most favorable position on the cost-accuracy Pareto frontier, while parallel fan-out with merge sits elsewhere—better latency, different cost profile.&lt;/p&gt;

&lt;p&gt;The teams that are winning with fan-out aren't the ones who spawn the most agents. They're the ones who treat &lt;strong&gt;concurrency control as a first-class design concern&lt;/strong&gt;, isolate state per branch, verify at the fan-in, and use durable execution so a crash at branch seven doesn't cost them the whole run.&lt;/p&gt;

&lt;p&gt;So here's my question: &lt;strong&gt;when your fan-out finishes and the branches disagree, does your architecture surface the conflict or does your synthesis agent quietly pick a winner and move on?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear where you've landed. Adversarial fan-in, durable promises, isolation levels you actually implemented or something simpler that worked because you stayed sequential a little longer than everyone told you to?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
