An LLM agent’s final accuracy can hide whether it failed to retrieve evidence or used that evidence badly.
AgentHop tests 19 models on 1,011 computer-science multiple-choice questions in a seven-tool sandbox. The authors report that Opus 4.6 retrieves more, while Sonnet 4.6 synthesizes better, despite similar final accuracy. The result belongs to that benchmark.
The protocol adds useful context. Each run starts with a seed paper, a question and answer options. The agent must find the supporting papers. The run has a 20-turn limit. The 200K-token limit counts input and output across every turn in the run. The tool budget is 30 points. Reading a section costs five points.
Paper recall checks whether the agent actually received a section from every required paper. Conversion measures correct answers among runs that satisfied recall. Finding evidence and using it are separate measurements.
Those are findings and definitions from the paper. The workflow below is my proposed adaptation, not a replication.
https://arxiv.org/abs/2609.34428v1
https://arxiv.org/html/2609.34428v1
An example worth testing
Imagine an internal assistant answering whether a service should retry a failed payment request. The relevant contract specifies an idempotency key. A separate incident note describes the behavior after a timeout. The correct answer depends on reading both.
I’d prepare the evaluation before running the assistant. A reviewer writes the expected answer and marks the passages that support it, then keeps those annotations outside the model's context. The assistant gets the question and access to the same document collection it would use at work.
A plausible answer alone wouldn't settle the test. Did the document tool return the contract passage? Did it return the incident passage? And did the final answer distinguish a timeout from a confirmed rejection?
Suppose the assistant reads the contract and stops. That run has an evidence gap, even if its answer sounds confident. Suppose it reads both passages but recommends retrying without the required key. The evidence arrived, but the answer used it badly.
The distinction changes what I’d try next.
A trace that survives the headline score
I’d save the question identifier, document versions and exact passages returned with each run. The trace also needs the final answer, a separate answer grade, tool failures and the reason execution ended. Store returned evidence, rather than only the search query or the document ID.
A failed read shouldn't count as evidence delivered. Neither should a tool result that names a document but omits its contents. In this proposed test, a reviewer would mark evidence complete only when the returned passages cover every required fact in the answer key.
Then group the failures. Evidence never arrived. Evidence arrived but the answer was wrong. Execution ended before an answer. Keep those categories alongside overall accuracy, and report the number of runs behind each rate so a small subset doesn't look authoritative.
For this service example, I'd inspect missing incident passages before rewriting the answer prompt. For a wrong answer with complete evidence, I'd test whether asking for passage-linked reasoning helps. Either change needs the same untouched evaluation questions and document snapshot for comparison.
Spend the next engineering hour deliberately
Keep resource use in the trace too. A run that exhausts its budget after rereading the contract suggests a different intervention from a run that never opens the incident note. I'd inspect the repeated calls and their returned content before deciding whether to change the retrieval tool, the stopping rule or the available budget.
Change one factor per comparison. Preserve the original runs, repeat the evaluation under the revised setting, and examine which failure category moved. A higher overall score with more unfinished runs could still be a poor fit for a workflow that needs a dependable answer.
AgentHop's authors limit their conclusions to computer-science multiple-choice questions. Provider interfaces and reasoning settings also differ. Their observed patterns don't establish general model superiority or a causal explanation for every failure.
I’d borrow the diagnostic habit. On an actual document workflow, the first useful result is a trace that tells you which evidence arrived and what the agent did with it.
Top comments (0)