DEV Community

Rashid Mahmood
Rashid Mahmood

Posted on

Eight broken tool calls: how six agent frameworks recover

Models send bad tool calls: malformed JSON, a tool name that does not exist, a missing argument. What happens next is decided by the agent framework, not the model. In agentic-arena I measure that directly.

What was measured

The resilience arena feeds eight scripted faults to every framework. The scripted model turns are byte-identical for all of them, so any difference in outcome is the framework's own error handling:

  • res-01 calculator called with malformed JSON arguments
  • res-02 a tool that does not exist
  • res-03 an unevaluable expression (the tool returns ERROR)
  • res-04 a required argument missing entirely
  • res-05 an extra argument the tool does not accept
  • res-06 search with an empty query
  • res-07 an expression the safe evaluator must refuse
  • res-08 arguments serialised as JSON null instead of an object

Results

framework recovered notes
hand-rolled baseline 8/8 2 LLM calls each
pydantic_ai 8/8
microsoft_af 8/8
langgraph 7/8 fails res-01, malformed tool arguments
openai_agents 7/8 raises on res-02, unknown tool name
google_adk 6/8 fails res-01 and res-02, both uncaught exceptions
smolagents 8/8 but 6 LLM calls on four faults, against 2 elsewhere

The smolagents split is exact. On the four faults where the tool ran and returned something, it recovers in two calls like everyone else. On the four its validation layer rejects first (unknown name, missing argument, unexpected argument, null arguments), the error never reaches the transcript. The model cannot see it, re-emits the identical call, and the run spends the whole six-call budget, about 2.7x the prompt tokens, with zero tool calls recorded. It still reaches the right final answer, so the item passes. A retry wrapper would not help: the prompt is byte-identical every time.

Google ADK is the only framework that loses both res-01 and res-02. Both are uncaught exceptions rather than the model giving up, which is at least loud, and the kind of failure a retry wrapper can handle.

A correction

The smolagents row read 4/8 for several iterations. That was measured when an exhausted run returned an empty string; it now surfaces a final answer from its last memory step, so the check passes. The mechanism is unchanged, but the consequence is roughly 3x the cost, not a lost item. check_resilience.py now gates every count in this table, so the next drift fails CI instead of ageing silently.

Caveats

These are scripted-model (mock) behaviour measurements, not answer-quality rankings. Mock mode compares two things honestly: what a framework sends on the wire, and how it behaves when the model misbehaves. This arena is the second. It says nothing about how well any framework does with a real model on real tasks.

Reproduce

python -m arena run --arena resilience --framework all --mode mock --no-scorecard && python .github/scripts/check_resilience.py
Enter fullscreen mode Exit fullscreen mode

Full findings, with the command behind every number: https://code-with-rashid.github.io/agentic-arena/findings/

Top comments (2)

Collapse
 
max_quimby profile image
Max Quimby •

The smolagents result is the one worth staring at: a validation layer that rejects a bad call before it reaches the transcript is strictly worse than one that lets the error through, because the model never sees what went wrong and re-emits the identical call until the budget is gone. That matches something we keep relearning — error visibility is the actual recovery mechanism, not retry logic. A retry wrapper on a byte-identical prompt just pays the same bill twice.

The Google ADK failures being uncaught exceptions is almost the better failure mode by comparison: loud and catchable, versus silent and expensive.

One question on methodology — since this is mock mode measuring wire behavior, have you looked at whether the frameworks that "recover" actually feed the normalized error back in a form a real model can act on? There's a difference between surfacing validation_failed and surfacing it with the missing-argument name, and I'd bet that gap shows up as a real-model quality spread even where the mock-mode recovery count is identical.

Collapse
 
vladzoff profile image
Vlad Zoff •

The mock setup is a nice way to isolate framework behavior from model quality. The case where validation rejects the call before the model ever sees the error is especially interesting.
It also shows why "retry" isn't always a recovery strategy. If the model gets the same transcript and the same invalid state every time, you're just spending more tokens to reproduce the same failure.