GroundedDocs answers questions over a document set with exact citations and refuses when the evidence is missing. On a frozen set of 40 questions (18 answerable, 11 partially answerable, and 11 it should refuse), it still got [N] wrong or only partly right after I added a hybrid search and a reranker.
The obvious next step was the one most agent tutorials suggest: a planner, a retrieval agent and a verifier. Before building them, I asked one question of every failure: was the evidence in the model's context when it answered? For [X] of [N], it was not.
A model that never saw the evidence does not need colleagues. It needs the evidence.
This post is that check, written so you can run it on your own pipeline. It is not a guide to building multi-agent systems.
Why does "add another agent" look like the fix?
The failures cluster on hard questions: multi-part, spread across sections. A wrong answer to a hard question looks like missing intelligence, and the standard response to missing intelligence is more agents.
That is the wrong model. It assumes every failure is a reasoning failure.
A second agent buys exactly one thing: a separate context window. That gives you isolation and the option to run work in parallel. It does not make the model more capable. Here is what it costs:
- Tokens. Anthropic reported its agents use about 4x the tokens of a chat, and its multi-agent system about 15x.
- A lossy handoff. A subagent returns a summary, not what it read. Every decision it made along the way is invisible to the caller. That was Cognition's original argument.
- Compounding errors. A Google and MIT study of 180 agent configurations found independent agents working in parallel amplified errors 17.2x.
- Harder debugging. One failure becomes several traces to join. An analysis of 1,600+ multi-agent traces found 14 distinct failure modes.
If you run backend systems, this is familiar. A subagent is a service boundary: you pay for serialisation (the summary), transport (the tokens) and distributed failure (the joined traces). Same rule as splitting a monolith: don't split until a measured limit forces you to.
What should you ask of every failed run?
One question: was the evidence in the model's context when it answered?
Context is everything the model can read at the moment it answers. If the evidence was not there, no number of agents could fix it. They will fail the same way at a higher price. If it was there, you have a reasoning problem, and more steps might help.
Label each failure. The first three labels are context problems.
| Label | Looks like | Fix |
|---|---|---|
| 01 Evidence never arrived | Half-right answers, or refusals on answerable questions | Retrieval: a second query, wider candidates before reranking, the parent section |
| 02 Tool output floods the window | Fine at step 5, poor at step 20, repeats tool calls | Store large results outside the context; pass an ID and a short summary |
| 03 History never compacted | Forgets an early rule, contradicts an earlier decision | Keep a short record of decisions and open questions; drop the raw transcript |
| 04 Evidence present, still wrong | Right text retrieved, wrong answer | More steps, not more agents |
| 05 Too much independent reading | Window overflows, or too slow in sequence | Orchestrator with subagents |
How do you check it?
Every question in my eval set has hand-labelled evidence chunks. For each failure, I compare them against the chunks that actually went into the prompt: the final prompt after reranking and truncation, not the retriever's top-k.
from collections import Counter
def label_failure(case) -> str:
gold = set(case.gold_chunk_ids) # evidence you labelled by hand
if not gold:
return "refusal-case" # should-refuse questions: check separately
seen = set(case.trace.prompt_chunk_ids) # chunks that actually reached the model
missing = gold - seen # <- the line that matters
if missing == gold:
return "01-no-evidence"
if missing:
return "01-partial-evidence"
return "04-evidence-present"
print(Counter(label_failure(c) for c in failed_cases))
The line that matters is gold - seen. If you can't compute seen because you don't log which chunks reached the prompt, that is your first fix.
For me, [X] of [N] came back as label 01.
Measure recall on the failed questions only. The whole-set average hides exactly the cases you care about.
When is another agent the right call?
Two labels earn more machinery.
04: The evidence was there, and the answer was still wrong. This needs more steps, not more agents: one agent that checks its answer and retries inside a hard step limit. Measure the pass rate and cost per success.
05: too much independent reading for one window. This is where a lead agent with parallel subagents wins. Anthropic reported a 90.2% improvement over a single agent on its internal research eval and was clear that work where agents share context or depend heavily on each other is a poor fit. Even Cognition, which wrote "Don't Build Multi-Agents", has since shipped specific multi-agent flows: one orchestrator owns the state, and short-lived subagents return a compressed result.
If you do split, run both versions on the same frozen question set and set the pass mark before you run it. If the split does not clear it, keep one agent and write down why.
Do this before you add your next agent.
- Pull your last 20 failed runs.
- For each one, log what was actually in the prompt: chunk IDs, tool results, history length.
- Give each run a label from 01 to 05.
- If most are 01 to 03, fix context. Measure recall on the failed set only.
- Only for 04 or 05, try more steps or a split on the same frozen set with the pass mark set first.
In GroundedDocs, deterministic RAG stays the default path. The label-04 failures get a bounded agent path next, with hard limits on steps and cost. The measured results are the next post in this series.
The rule I use now: before adding an agent, check what the first one saw.
What was the failure that finally made you split one agent into two, and was the evidence already in context when it happened?


Top comments (0)