I once watched a nine-minute agent workflow spend seven minutes and forty seconds doing nothing. Not failing. Not retrying. Just waiting each agent idle while the previous one finished a task that had no dependency on it whatsoever.
I'd built the pipeline as a chain. Researcher, then analyst, then writer, then fact-checker, then formatter. Each step consumed the previous step's output. In my head, that was the workflow. In production, it was a latency tax I'd designed myself and never noticed.
The fix wasn't a better model or a faster prompt. It was admitting that most of those steps were never actually dependent on each other.
The Tax Nobody Itemizes
Sequential chains are the easiest multi-agent pattern to reason about, which is why they're the default. Each stage has a clear input, a clear output, and a clear owner. Debugging is trivial because there's only one path. Microsoft's sequential orchestration documentation describes it exactly this way: a predefined linear order where each agent processes the previous agent's output, creating a pipeline of specialized transformations.
The problem is that "clear" and "correct" are not the same thing. A pipeline implies dependency. Most real workflows don't have the dependency you assumed.
My researcher fetched six independent sources. Those six fetches could have run in parallel. My analyst produced three independent analyses—market, risk, and competitive positioning—with no shared state between them. My fact-checker verified claims that were already independently verifiable. Only the writer genuinely needed everything upstream.
The math is unforgiving. If you have N stages with average latency L, your wall-clock time is N × L. If K of those stages are independent, the floor is (N − K) × L + L, not N × L. In my pipeline, K was 4 out of 5. I was paying nearly five times the necessary latency and calling it "the workflow."
Amdahl's law applies here in a way that's almost embarrassing once you see it. The speedup from parallelizing independent work is bounded by the serial fraction—and in most agent pipelines I've audited, the serial fraction is far smaller than the architecture assumes.
The Structural Mistake
The chain wasn't wrong because I used the wrong framework. It was wrong because I encoded dependencies that didn't exist.
A chain says: B cannot begin until A completes. That's a claim about the world. When it's true, a chain is correct and efficient. When it's false, you've serialized work that should have been concurrent, and you'll never notice because the system still produces correct output—just slowly.
I'd assumed the analyst needed the researcher's entire output before starting. In reality, the analyst needed three fields. The researcher was still fetching sources four and five while the analyst could have started on source one.
The dependency was real at the level of final output. It was fake at the level of individual facts.
Event-Driven Concurrency: React to Data, Not to Sequence
The pattern that fixed this wasn't a new framework. It was a change in what triggers work.
In a sequential chain, work starts because the previous step finished. In an event-driven architecture, work starts because the data it needs has arrived. Agents subscribe to events—a source fetched, a claim extracted, a document chunked—and fire when the specific event they care about publishes.
The shift has three architectural consequences.
First, agents become independent consumers. Each agent declares what it needs and what it produces. The orchestrator doesn't route tasks; it publishes events. Agents that can act, act. Agents that need more input, wait. No agent is blocked by an unrelated stage.
Second, latency becomes the critical path, not the sum. If six fetches run concurrently and three analyses run as soon as their inputs land, wall-clock time collapses to the longest dependency chain plus fan-in. The parallelism isn't a fan-out you designed—it's an emergent property of the dependency graph.
Third, backpressure becomes visible. In a chain, a slow stage just makes everything slow. In an event-driven system, you can see which topics are backing up, which consumers are lagging, and where the actual bottleneck lives. That observability is worth as much as the latency win.
This is the same shift microservices made a decade ago, and the same discipline applies: publish events, not commands. Let consumers decide what to do with them. Use correlation IDs to trace a logical request across many asynchronous hops. Use idempotency keys so a retried consumer doesn't double-write.
The Orchestration Trap in Event-Driven Systems
The failure mode of event-driven agent systems is the opposite of the chain's. Instead of being too slow, it becomes too opaque.
Choreography—where agents react to each other's events with no central coordinator—scales beautifully and debugs terribly. When a workflow fails, you have no single place to look. The failure is in the emergent behavior of a dozen independent consumers, none of whom saw the whole picture.
The pragmatic middle ground that has held up in production is event-driven execution under a thin orchestration layer. A coordinator doesn't route individual tasks; it publishes a workflow-started event and then subscribes to completion events, tracking state and enforcing timeouts. The agents still run independently. The coordinator just knows what "done" means and can surface when it isn't.
This is where durable execution frameworks earn their keep. Temporal's model—workflows as deterministic code, activities as the side-effecting work—maps cleanly onto agent systems. The workflow defines the dependency graph. The activities run concurrently where the graph allows. The coordinator state is durable, so a crash at step seven doesn't lose the run.
LangGraph's parallel node execution and conditional edges offer the same property within a graph, and the newer event-driven orchestration patterns—where nodes publish and subscribe to channels rather than passing state directly—are the closest thing to true event-driven concurrency in the agent frameworks I've used.
The Guardrails That Make It Survivable
Event-driven concurrency introduces failure modes that sequential chains never had. Four of them matter.
Ordering. Events can arrive out of order. If an analysis event arrives before the source event it depends on, the analysis either fails or runs on stale data. The fix is either partition keys that preserve order for related events, or a consumer that buffers until dependencies are satisfied. In agent systems, the cleanest answer is usually to make each event self-contained enough that order doesn't matter—carry the data, not just a pointer to it.
Duplicates. At-least-once delivery is the realistic default. Every consumer must be idempotent. An agent that writes a summary twice is fine. An agent that charges a customer twice is not. Idempotency keys derived from the event ID, not the timestamp, are the standard fix.
Poison events. One malformed event can crash a consumer repeatedly, blocking the queue behind it. Dead letter queues aren't optional. Neither is a circuit breaker that stops consuming from a topic after repeated failures and alerts instead of spinning.
Fan-in reconciliation. When three analyses complete asynchronously, something has to decide when all three are ready and what to do with their outputs. That's the same fan-in problem from the concurrency pattern, with the same rules: scoped branch state, structured claims, and an explicit merge that surfaces conflicts rather than smoothing them over.
None of this is exotic. It's the standard distributed systems toolkit. The mistake is assuming agent systems are exempt from it.
Who's Actually Shipping This
Cisco's CAIPE runs multi-agent orchestration across Argo CD, Kubernetes, and Komodor, where agents chain across tools and respond to cluster events rather than polling in sequence. Their reported outcome: response times dropped from hours to seconds, with MTTR reduced by up to 80%.
SAP's event-driven multi-agent orchestration was evaluated across 208 production-derived enterprise scenarios at Persona, Department, and Enterprise scale. Their finding is the one that should be posted above every pipeline design review: scale, not task complexity, dominates orchestration performance. Both sequential and reactive architectures degraded at enterprise scale, and the Task Manager they introduced to handle priority inference, related-event merging, and preemption cut high-priority queue latency by 14–75% and improved related-event correctness by more than 20 percentage points.
Resonate's durable execution model treats every spawn and every await as a durable promise. Fan-out branches run concurrently, each checkpointed independently. A crash mid-batch re-runs only what was in flight. That's the property that makes event-driven agent systems survivable rather than merely fast.
AWS's Expansion-Contraction pattern, published at ACM CAIS 2026, walks a domain graph with concurrent path analysis. Their reported results: 98.2% accuracy on a production supply chain, 100% on public benchmarks, concurrent path analysis yielding up to 1.43× speedup, and investigation caching reducing token usage by up to 93.9%. The concurrency is the point, but the caching is what makes it affordable.
When Not to Reach for Events
Event-driven concurrency is not a default. It's a response to a specific problem.
Stay sequential when the dependency is real. If stage B genuinely needs the complete output of stage A, a chain is correct and simpler. Don't parallelize what isn't parallel.
Stay sequential when latency isn't the binding constraint. If your pipeline finishes in six seconds and users don't care, you don't have a latency problem. You have an architecture preference, and event-driven systems cost more to build and operate.
Stay sequential when debugging is the scarce resource. A chain gives you a single path and a linear trace. An event-driven system gives you a correlation ID and a scatter plot. If your team is small and the workflow is new, that trade is often not worth making yet.
The honest rule: start sequential, measure, and parallelize only the stages that measurement proves are independent and expensive. The teams I've seen get this right didn't start with events. They started with a chain, watched it in production, and moved to events when the latency numbers justified the operational cost.
The Trade-Off You're Accepting
Event-driven concurrency buys you latency and throughput. It costs you determinism and traceability.
A sequential chain has one path. An event-driven workflow has as many paths as there are interleavings, and no two runs are identical. Reproducing a bug means replaying an event stream, not re-running a function. Your observability has to be built for it—distributed tracing, correlation IDs, structured logs that tie back to a single logical request.
And you're accepting more moving parts. A chain has N stages. An event-driven system has N consumers, a broker, a schema registry, dead letter queues, and a coordinator. Every one of those is a thing that can fail independently.
But here's what I've learned from every pipeline post-mortem I've sat through: the latency was rarely the thing that killed the project. It was the assumption that a chain was the only honest representation of the work. Once I drew the actual dependency graph instead of the one I'd assumed, most of the sequence disappeared—and so did most of the delay.
So here's my question: If you drew the real dependency graph for your agent pipeline, how many of those sequential steps would survive?
I'd love to hear where you've landed. Sequential by choice, event-driven after a painful migration, or somewhere in between and what finally made you look at the graph instead of the chain?
Top comments (0)