TL;DR: Most published multi-agent wins compare a team of agents against a single agent given far less compute. Studies that match the budget find the advantage mostly disappears: automatically generated multi-agent systems underperformed a simple self-consistency baseline while costing up to 10x more, and on a sequential planning benchmark every multi-agent variant tested scored 39–70% worse than a single agent. In a worked five-agent refund pipeline, 64% of the tokens go on duplicated context and handoff notes, and a fact dropped in one handoff produces a wrong refund. Merging the agents into one fixes the lost fact but not the bill. Redesigning the tools does both, cutting the cost to under half. The test for whether you need a second agent: can you say what it would be evaluated on by itself? Parallel research, context isolation, separate permissions and verification pass that test. Splitting work by job title (planner, coder, tester, reviewer) usually doesn't. Before adding an agent, give your single agent the same token budget and compare them on your own labelled cases.
"Add another agent" has become the default answer to an agent that isn't good enough. The planner misses things, so a reviewer agent is added. The reviewer lacks context, so a researcher agent is added to feed it. Six months later there's an org chart of five agents, each with its own prompt, and nobody can say which one is responsible when a customer gets the wrong answer.
The 2026 name for everything around the model is the harness: tools, permissions, loops, memory, verification. Multi-agent design is a harness decision, and usually the most expensive one available. This article makes the case that most multi-agent systems are one agent whose tools are badly designed. It also gives a test for the cases where that isn't true: if you can't say what each agent would be evaluated on by itself, you have one agent.
A support pipeline with five agents
Here's an invented but typical example: a B2B software company handling refund requests. The team built it the way the architecture diagrams suggest:
- A router reads the ticket and decides what kind of request it is.
- An account agent looks up the customer and summarises their account.
- A billing agent pulls invoices and payments and summarises billing history.
- A policy agent reads the refund policy and decides what the customer is owed.
- A writer drafts the reply.
Each agent passes a short written note to the next. It demos well. Every agent has a clear job title, and the diagram has five neat boxes.
Then a ticket comes in from a customer cancelling an annual plan. They paid $4,800 in March. In June they downgraded and received a $1,200 credit. The billing agent's note reads: "Annual plan, $4,800, renewed 3 March, no open disputes." All of that is true. The June credit is in the raw ledger the billing agent read, but it isn't in the note, because the note had a length limit and the agent picked what looked most relevant. The policy agent never sees the ledger, only the note. It calculates a pro-rated refund on the full $4,800. The writer promises it to the customer.
No agent made a mistake given what it was shown. The failure is in the handoff, and debugging it means opening five traces and working out which summary dropped which fact. Anthropic calls this the "telephone game": when agents are split by problem type, they end up passing "information back and forth with each handoff degrading fidelity."
Dat Tran and Douwe Kiela put a formal argument behind the same intuition, using the Data Processing Inequality from information theory. Every agent downstream of a handoff works from a processed version of the original context, and processing can lose information but never add it. A single agent with the whole context in front of it doesn't pay that cost. Their argument also says when multi-agent designs can come out ahead. One case is when a single agent's use of its own context breaks down: "once effective single-agent context utilization deteriorates sufficiently, a well-designed MAS may recover task-relevant information more reliably than a degraded single pass." The other is when the multi-agent system simply gets to spend more compute. That second condition turns out to explain a lot of the published results.
The comparisons that flattered multi-agent systems
Most of the evidence for multi-agent designs compares a multi-agent system against a single agent that was given much less to work with. That's the confound each of the three studies below goes after.
Match the compute and the advantage mostly disappears. Tran and Kiela's starting point is that reported multi-agent gains are often confounded by extra test-time computation. They compared single-agent and multi-agent setups on multi-hop reasoning, across three model families, with the thinking-token budget held equal, and found single agents matched or beat the multi-agent systems once computation was normalised. A typical row from their results: on four-hop MuSiQue questions with DeepSeek-R1-Distill-Llama-70B at a 2,000-token budget, the single agent scored 0.418, a sequential multi-agent pipeline 0.327, and no multi-agent design in that row scored above 0.352. Their error analysis found the sequential pipeline "over-explores and drifts," and their conclusion is that "many reported MAS gains are better explained by compute and context effects than by inherent architectural superiority."
Automatically generated multi-agent systems lost to a simple baseline at a fraction of the cost. The Illusion of Multi-Agent Advantage (Jwalapuram et al., Salesforce Research and collaborators, June 2026) tested six frameworks that design multi-agent systems automatically against one single-agent baseline: chain-of-thought with self-consistency, which just samples the same model five times and takes a majority vote. The authors point out that earlier comparisons "rarely control for inference budgets such as number of LLM calls, total cost, retries, or sampled paths." When they controlled for them, the result was blunt: "automatic MAS consistently underperform CoT-SC despite being up to 10x more expensive."
The paper then took the generated systems apart to see what the extra agents were doing. Mostly, nothing useful:
- In DyLAN, a framework built on the premise that agent diversity drives performance, agents "reach immediate, unanimous consensus in ∼70% of GPT-4o cases and >90% of GPT-5 cases." Agents that agree instantly are one agent's answer, paid for several times over.
- Replacing the task-specific expert roles with generic "assistant" agents scored slightly better (54.4% vs. 53.4%).
- In a framework that searches for the best workflow, 7 of the 14 final workflows it found were a single prompt run three times and aggregated, which is self-consistency under a more complicated name.
The authors' name for agents that cost full price and change nothing is "expensive witnesses."
Sequential and tool-heavy work gets worse, not better. Google Research and MIT's Towards a Science of Scaling Agent Systems (Kim et al.) is the largest study of the three: 260 configurations across six agentic benchmarks, five architectures and three model families, with tools, prompts and compute standardised so that only the architecture varied. Three of its findings map directly onto production systems:
- On a planning benchmark built from sequential constraints, "all multi-agent variants universally degrade performance," by 39% to 70%. A refund decision is a sequential task. Each step depends on what the previous one found.
- Under fixed budgets, "tool-heavy tasks suffer disproportionately from multi-agent inefficiency," because splitting the budget across agents leaves each one less room to use its tools.
- Once a single agent already scored above roughly 45%, adding agents gave negative returns.
The same paper also shows where multi-agent designs do win. Centralised coordination improved performance by 80.8% on a decomposable financial-reasoning benchmark. So "never use multiple agents" is the wrong conclusion. The right one is narrower: the benefit is real for work that splits into genuinely independent pieces, and negative for most other work.
Even Anthropic's widely cited result needs reading carefully. Its multi-agent research system "outperformed single-agent Claude Opus 4 by 90.2% on our internal research eval," on the kind of breadth-first research that parallelises well. The same post reports that "token usage by itself explains 80% of the variance" in performance on the BrowseComp benchmark, and that "multi-agent systems use about 15× more tokens than chats." That 15× is measured against a chat interaction, not against a single agent. Much of the improvement came from spending more tokens, which a single agent can also do. The post is clear about where it doesn't apply: "most coding tasks involve fewer truly parallelizable tasks than research."
Where the tokens go
Back to the support pipeline. Here's the token accounting for one ticket, using made-up but plausible sizes: a 1,200-token ticket, an 800-token system prompt per agent, 1,500 tokens from each account and billing lookup, and 2,000 tokens of retrieved refund policy.
| Agent | System prompt | Ticket | Notes read | Tool results | Input | Output |
|---|---|---|---|---|---|---|
| Router | 800 | 1,200 | — | — | 2,000 | 250 (note) |
| Account | 800 | 1,200 | 250 | 1,500 | 3,750 | 400 (note) |
| Billing | 800 | 1,200 | 650 | 1,500 | 4,150 | 400 (note) |
| Policy | 800 | 1,200 | 1,050 | 2,000 | 5,050 | 400 (note) |
| Writer | 800 | 1,200 | 1,450 | — | 3,450 | 350 (reply) |
| Total | 4,000 | 6,000 | 3,400 | 5,000 | 18,400 | 1,800 |
That's 20,200 tokens per ticket. Now split it by what the tokens were for. The work only one agent had to do is reading the ticket once, one system prompt, the 5,000 tokens of tool results and the 350-token reply: 7,350 tokens, or 36%. The other 64% is the cost of there being five agents. The ticket is read four extra times (4,800), there are four extra system prompts (3,200), and handoff notes are written (1,450) and then re-read by every later agent (3,400). Anthropic's own guidance describes the same breakdown: "In our testing, multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks," from "duplicating context across agents, coordination messages between agents, and summarizing results for handoffs."
Now the part that surprises people. Collapse the five agents into one agent that calls the same three tools, and the saving almost disappears. An agent loop re-reads its growing context on every turn, so four turns (three tool calls and a reply) with a 1,200-token system prompt come to about 19,700 input tokens and 650 output: 20,350 in total. It's no cheaper. It does fix the refund, because the one agent read the raw ledger and saw the June credit. But the token bill is the same.
The saving comes from the tools. The original lookups return raw API output, full invoice objects and the whole policy document, because the tools were written to return whatever the backend returned. Rewrite them around what the agent actually needs to decide:
-
get_billing_context(customer_id)returns one pre-joined record of about 600 tokens: plan, amount paid, and every credit, refund and plan change with its date. -
check_refund_eligibility(customer_id, reason)runs the refund policy as code and returns about 150 tokens: eligible or not, the maximum amount, which rule applied, and why.
The loop becomes three turns (look up, check eligibility, reply): about 8,850 input and 550 output tokens, or 9,400 in total. That's under half the five-agent pipeline, and it gets the June credit right, because the credit is a field in the tool's output and not something a summary might drop. Prompt caching would lower the repeated system-prompt cost further. Your ratios will differ with your sizes, but the pattern holds. Removing agents fixes the lost context. Better tools reduce the cost.
This accounting is only possible if you can see the assembled prompt and the per-step token count for every call, which most orchestration frameworks hide by default. The orchestration lesson in SophiArch's AI Applications with LLMs course starts from that problem: work out what the framework conceals before deciding whether to use one.
Notice what happened to the policy agent. It became a function. That's common. Look closely at a multi-agent diagram and you'll often find a business rule that has been given a prompt and a job title. A rule written as code can be unit-tested, never paraphrases the policy, and costs nothing per call.
The test: what would each agent be evaluated on by itself?
Better tools are now standard advice. Anthropic's guide puts it plainly: "a well-designed single agent with appropriate tools can accomplish far more than many developers expect," and notes that "we've seen teams invest months building elaborate multi-agent architectures only to discover that improved prompting on a single agent achieved equivalent results." The advice exists, and teams still reach for a second agent, so something more concrete is needed at the moment of that decision.
Here's the test. For every agent in the design, write down the evaluation you would run on that agent by itself, with inputs, a pass condition, and a labelled set to run it on, without running the rest of the system. If you can write it, that agent may be worth separating. If the only way to judge the agent is to look at the final output of the whole system, it isn't a separate agent. It's a step inside one.
Apply it to the support pipeline:
| Agent | Independent evaluation? | Verdict |
|---|---|---|
| Router | Yes: routing accuracy against labelled tickets. But it has a fixed label set, no tools and no loop. | A classification call, not an agent |
| Account agent | The lookup can be tested (right customer, right records). The summary can't: "complete enough" depends on what the next agent needs. | A tool |
| Billing agent | Same as above. The note that dropped the June credit had no pass condition it could have failed. | A tool |
| Policy agent | Yes: the correct outcome for each policy case is known in advance. | A function, tested with unit tests |
| Writer | Only by judging the final reply, which is the whole system's output. | The agent |
Five agents turn into one agent, two tools, a function and a classifier. Nothing was lost. Every part that could be tested separately is now easier to test than before, because a function with fixed cases is easier to test than a prompt.
The test also explains the published results. The successful system in The Illusion of Multi-Agent Advantage was the one the authors designed by hand for a task built to reward decomposition: a finance problem with independent per-investor calculations over a large table of prices. Their "Expert-MAS" used an extractor agent that retrieves specific data and a calculator agent that reasons over isolated snippets, both of which can be evaluated on their own. Coordination was done by a deterministic Python orchestrator, and the final comparison was computed in code. With GPT-5 it scored 96.5%, against 57.0% for the self-consistency baseline, "with cost comparable to CoT-SC." The multi-agent system that worked was mostly a program, with the model called only for the pieces that could each be checked.
When a second agent passes the test
There are real cases, and they're narrow enough to name. Each one comes with an evaluation that doesn't need the rest of the system:
- Genuinely parallel work. Five independent research questions, five subagents, each judged on whether it found the facts for its own question. This is the case Anthropic's research system and the 80.8% financial-reasoning result both come from.
- Context protection. A subagent reads 300 pages of logs and returns 20 lines. Evaluation: does the summary contain the line that explains the incident? The main agent's context stays clean. This is also the condition in which Tran and Kiela's argument says a multi-agent design can come out ahead: when one agent's context would otherwise be too long or too noisy to use well.
- Different permissions. A subagent that can write to a production system, and nothing else, evaluated on whether the end state is correct and it took only allowed actions. The separation exists because of what it's allowed to do, not because of its job title.
- Verification. Anthropic's guide calls the verification subagent "one multi-agent pattern that consistently works well across domains." It has the cleanest independent evaluation of all: seed known defects and measure how many it catches. Kim et al.'s error-amplification finding points the same way. Independent agents with no checkpoint amplified errors 17.2×, while centralised coordination with a validation step held it to 4.4×.
What's missing from that list is splitting by job title. Planner, implementer, tester and reviewer is the most common multi-agent design and the one the evidence argues against most directly. Anthropic's guide says "dividing by type of work (one agent writes features, another writes tests, a third reviews code) creates constant coordination overhead," and in their own experiment with those roles "the subagents spent more tokens on coordination than on actual work." A recent preprint, At Equal Inference Cost, Multi-Agent Structure Does Not Beat a Single Frozen Agent (Dylan et al.), tested the planner–executor–critic split directly, holding the number of model calls fixed and optimising each role's prompt. All of the gain came from the executor. The planner and critic prompts "froze empty," and the full team (0.769) was not statistically separable from the single agent (0.754) while using 1.8× the calls. It's a small study, with one 7B model and two benchmarks, but it points the same way. Berkeley's study of multi-agent failures (Cemri et al.) found failure rates of 41% to 86.7% across seven open-source multi-agent frameworks, and grouped 14 failure modes under three headings: system design, misalignment between agents, and task verification. The middle one, misalignment between agents, is a failure a single agent can't have at all.
Run the budget-matched comparison this week
If you already run a multi-agent system, there's a cheap way to find out whether it's earning its cost. It needs one thing you may already have: a set of labelled cases with a checkable pass condition.
def budget_matched(cases, run_multi, run_single):
"""Compare a multi-agent system against one agent given the same token budget."""
rows = []
for case in cases:
multi = run_multi(case.input) # -> answer, tokens_used
single = run_single(case.input, # same model, all the tools,
token_budget=multi.tokens_used) # spend up to the same budget
rows.append({
"multi_pass": case.passes(multi.answer),
"single_pass": case.passes(single.answer),
"multi_tokens": multi.tokens_used,
"single_tokens": single.tokens_used,
})
return rows
How run_single spends the budget is up to you: more reasoning tokens, or several samples with a majority vote, which is exactly the baseline the Illusion paper used. The one rule is that it gets the same budget as the multi-agent run, not a fraction of it. Compare the tokens each side actually used, not the budget you asked for. Tran and Kiela found that with Gemini 2.5, "the visible thought text produced by SAS tends to plateau well below the requested budget, while MAS surfaces more visible thought content under the same requested budget B, due to multiple calls." A single agent that stops early hasn't had a fair comparison. Then read the results in this order:
- If the single agent matches or beats the multi-agent system at the same budget, the extra agents are spending tokens, not adding capability.
- If the multi-agent system wins, look at which cases it wins. If they're the cases that split into independent parts, you've found a real use for a second agent. Keep it, and give that agent its own evaluation.
- Either way, look at the single agent's tool calls on the cases it failed. Tools that return raw API output, and business rules the model has to re-derive from a document every time, are where the next improvement is.
The one-line version
A second agent is worth adding when you can evaluate it without the rest of the system. When you can't, what you actually need is a better tool, a function or a cleaner context, and all three are cheaper, faster and easier to test than another agent.
References
- Jwalapuram, P., Lin, H., Li, C., Jiao, F., Wang, S., Ming, Y., Ke, Z., Qin, C., Carenini, G., & Joty, S. (2026). The Illusion of Multi-Agent Advantage. arXiv:2606.13003. Six automatic multi-agent frameworks against budget-controlled chain-of-thought with self-consistency; source of the "up to 10x more expensive" result, the consensus and role-ablation findings, and the Expert-MAS result.
- Tran, D., & Kiela, D. (2026). Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets. arXiv:2604.02460. Equal-budget comparison, the Data Processing Inequality argument, the MuSiQue results and the budget-tracking caveat.
- Kim, Y., et al. (2025, revised 2026). Towards a Science of Scaling Agent Systems. arXiv:2512.08296v3. 260 controlled configurations across six agentic benchmarks, five architectures and three model families; source of the sequential-planning degradation, tool-coordination trade-off, capability threshold, Finance Agent and error-amplification figures.
- Dylan, D., Brennan, A., Murphy, C., O'Sullivan, N., Kelly, C., & Walsh, S. (2026). At Equal Inference Cost, Multi-Agent Structure Does Not Beat a Single Frozen Agent. arXiv:2609.04217. Planner–executor–critic team against a single agent at a fixed number of model calls.
- Cemri, M., Pan, M. Z., Yang, S., et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657. 1,642 annotated traces from seven frameworks; the MAST taxonomy of 14 failure modes.
- Hadfield, J., Zhang, B., Lien, K., Scholz, F., Fox, J., & Ford, D. (2025, June 13). How we built our multi-agent research system. Anthropic Engineering. Source of the 90.2%, 15×-versus-chat and 80%-of-variance figures.
- Phillips, C. (2026, January 23). Building multi-agent systems: When and how to use them. Claude blog. Source of the 3–10× single-agent comparison, the "telephone game" and type-of-work decomposition warnings, and the verification-subagent pattern.
- SophiArch. AI Applications with LLMs. Course covering the probabilistic contract, context window architecture, output validation layers, evaluation frameworks, orchestration patterns (including what frameworks hide), and cost and latency engineering for LLM systems.
Top comments (2)
The refund case is the cleanest argument in the piece. Each agent stayed truthful to its own slice, and the June credit died in a length-capped handoff note before the policy agent ever saw the ledger. Matching the token budget before adding a second agent would have made the duplicated context visible instead of hiding it behind a new job title. If a split only exists because the planner misses things, what metric would you use to show that a single agent with a better tool beats the reviewer?
iin1007am
The handoff note part matches what happened to me. I split a small support flow into a planner and a responder, and the responder kept missing one detail the planner had summarized away. Folding it back into one agent with a better lookup tool fixed it and the bill dropped. Your test of whether a second agent can be evaluated on its own is a nice filter. Has separate permissions been the most common valid reason you've seen?