DEV Community

Cover image for My RAG system didn't need more agents. It never saw the evidence
Rakesh Singh
Rakesh Singh

Posted on

My RAG system didn't need more agents. It never saw the evidence

GroundedDocs answers questions over a document set with exact citations and refuses when the evidence is missing. On a frozen set of 40 questions (18 answerable, 11 partially answerable, and 11 it should refuse), it still got [N] wrong or only partly right after I added a hybrid search and a reranker.

The obvious next step was the one most agent tutorials suggest: a planner, a retrieval agent and a verifier. Before building them, I asked one question of every failure: was the evidence in the model's context when it answered? For [X] of [N], it was not.

A model that never saw the evidence does not need colleagues. It needs the evidence.

This post is that check, written so you can run it on your own pipeline. It is not a guide to building multi-agent systems.

Why does "add another agent" look like the fix?

The failures cluster on hard questions: multi-part, spread across sections. A wrong answer to a hard question looks like missing intelligence, and the standard response to missing intelligence is more agents.

That is the wrong model. It assumes every failure is a reasoning failure.

A second agent buys exactly one thing: a separate context window. That gives you isolation and the option to run work in parallel. It does not make the model more capable. Here is what it costs:

A subagent reads many chunks and makes decisions the lead agent never sees; only a short summary crosses back

If you run backend systems, this is familiar. A subagent is a service boundary: you pay for serialisation (the summary), transport (the tokens) and distributed failure (the joined traces). Same rule as splitting a monolith: don't split until a measured limit forces you to.

What should you ask of every failed run?

One question: was the evidence in the model's context when it answered?

Context is everything the model can read at the moment it answers. If the evidence was not there, no number of agents could fix it. They will fail the same way at a higher price. If it was there, you have a reasoning problem, and more steps might help.

Decision fork: if the evidence was not in context, it is a context problem (labels 01 to 03); if it was, it is a reasoning problem (labels 04 and 05)

Label each failure. The first three labels are context problems.

Label Looks like Fix
01 Evidence never arrived Half-right answers, or refusals on answerable questions Retrieval: a second query, wider candidates before reranking, the parent section
02 Tool output floods the window Fine at step 5, poor at step 20, repeats tool calls Store large results outside the context; pass an ID and a short summary
03 History never compacted Forgets an early rule, contradicts an earlier decision Keep a short record of decisions and open questions; drop the raw transcript
04 Evidence present, still wrong Right text retrieved, wrong answer More steps, not more agents
05 Too much independent reading Window overflows, or too slow in sequence Orchestrator with subagents

How do you check it?

Every question in my eval set has hand-labelled evidence chunks. For each failure, I compare them against the chunks that actually went into the prompt: the final prompt after reranking and truncation, not the retriever's top-k.

from collections import Counter

def label_failure(case) -> str:
    gold = set(case.gold_chunk_ids)          # evidence you labelled by hand
    if not gold:
        return "refusal-case"                # should-refuse questions: check separately
    seen = set(case.trace.prompt_chunk_ids)  # chunks that actually reached the model
    missing = gold - seen                    # <- the line that matters
    if missing == gold:
        return "01-no-evidence"
    if missing:
        return "01-partial-evidence"
    return "04-evidence-present"

print(Counter(label_failure(c) for c in failed_cases))
Enter fullscreen mode Exit fullscreen mode

The line that matters is gold - seen. If you can't compute seen because you don't log which chunks reached the prompt, that is your first fix.

For me, [X] of [N] came back as label 01.

Measure recall on the failed questions only. The whole-set average hides exactly the cases you care about.

When is another agent the right call?

Two labels earn more machinery.

04: The evidence was there, and the answer was still wrong. This needs more steps, not more agents: one agent that checks its answer and retries inside a hard step limit. Measure the pass rate and cost per success.

05: too much independent reading for one window. This is where a lead agent with parallel subagents wins. Anthropic reported a 90.2% improvement over a single agent on its internal research eval and was clear that work where agents share context or depend heavily on each other is a poor fit. Even Cognition, which wrote "Don't Build Multi-Agents", has since shipped specific multi-agent flows: one orchestrator owns the state, and short-lived subagents return a compressed result.

If you do split, run both versions on the same frozen question set and set the pass mark before you run it. If the split does not clear it, keep one agent and write down why.

Do this before you add your next agent.

  1. Pull your last 20 failed runs.
  2. For each one, log what was actually in the prompt: chunk IDs, tool results, history length.
  3. Give each run a label from 01 to 05.
  4. If most are 01 to 03, fix context. Measure recall on the failed set only.
  5. Only for 04 or 05, try more steps or a split on the same frozen set with the pass mark set first.

In GroundedDocs, deterministic RAG stays the default path. The label-04 failures get a bounded agent path next, with hard limits on steps and cost. The measured results are the next post in this series.

The rule I use now: before adding an agent, check what the first one saw.

What was the failure that finally made you split one agent into two, and was the evidence already in context when it happened?

Top comments (0)