DEV Community

Abhisek Roy
Abhisek Roy

Posted on AI-assisted

AgentToolEval: Grading How LLM Agents Use Tools, Not Just What They Answer

Kaggle Benchmarking Challenge Submission

Written for: dev.to readers and the Kaggle Benchmarking Challenge judges.

I kept your format and headings, shortened the intro, added the two local models and the decision test, and filled the Ollama placeholders with numbers from your latest results.

Before you paste it, two corrections affect what the post can claim:

  1. tev1 is not faster than the cloud models on fresh prompts. The 0.096s in your decision results came from Ollama's prompt cache, because your full run reused the exact prompts from my earlier test run. Measured fresh, it takes about 2.1s per decision, against 0.66–1.43s for the cloud models. I re-ran tev1 fresh and replaced its rows in results/decisions/, so the time chart is now correct. The post says "close to cloud speed while running on a laptop", which is true. "Faster" would not be.
  2. gemma4:31b and gpt-oss didn't run on your machine. They ran on Ollama Cloud (your .env points at ollama.com). Only tev1 and qwen2.5:3b ran locally. The post now says this.

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I built AgentToolEval, a small benchmark that checks whether LLM agents use tools correctly, not just whether they reach the right answer. Did the agent check stock or guess? Did it make five calls when two would do? When a tool failed, did it retry, give up, or make something up? A final-answer check hides all of that.

The model plays a shopping assistant for a fake electronics store with four tools: search_products (prices but not stock), check_stock, place_order, and get_shipping_estimate, a distractor no task needs. The tools are deterministic, so every run is reproducible, and failures are injected on purpose (a timeout once, a service that never recovers, an out-of-stock order).

Scenario The trap
Cheapest in-stock laptop The cheapest laptop overall is out of stock
Laptops under ₹50,000 in stock One laptop costs exactly ₹50,000 ("under" means strictly below)
Order an out-of-stock phone The model must refuse, not claim success
MacBook price The search fails once, so the model must retry
Acer stock check The stock service never recovers, so the model must answer UNKNOWN

Each scenario scores 0 to 1: 60% correct answer, 15% right tools, 15% correct arguments, 10% efficiency (only when the answer is right). The model never sees the answer key; grading happens afterwards, from a log of every tool call.

Models Tested

I ran 8 models on Kaggle and 5 models with Ollama, chosen in groups that each answer a specific question.

On Kaggle ([5] scenarios):

Model Maker Why it's in the lineup
claude-opus-5-5-default Anthropic Newest frontier model, paired with Opus 4.5
claude-opus-4-5-20251101 Anthropic Previous generation, to see whether newer means better at tools
gpt-5.4-2026-03-05 OpenAI Frontier closed model
gpt-oss-20b OpenAI Small open model from the same maker
gemini-3.5-flash Google Fast closed model
gemma-4-26b-a4b-it Google Small open model from the same maker
glm-5 Z.ai Open model from another maker
grok-4.20-0309-non-reasoning xAI Non-reasoning model from a fourth maker

With Ollama (12 scenarios):

Model Where it ran Why
gemma4:31b Ollama Cloud Counterpart to the Gemma model on Kaggle
gpt-oss:120b Ollama Cloud Large vs small from the same family...
gpt-oss:20b Ollama Cloud ...to test whether size improves tool use
qwen2.5:3b My laptop A small general model, as a baseline
tev1:0.8b My laptop A 0.8B decision model built to pick one option from a list, not to run as an agent

The lineup lets me ask:

  • Does a newer generation help? Claude Opus 5.5 vs Opus 4.5.
  • Closed vs open from the same maker: GPT-5.4 vs gpt-oss-20b, and Gemini 3.5 Flash vs Gemma 4.
  • Does size help? gpt-oss 20b vs 120b, and 1–3B local models against cloud ones.
  • Is a small decision model good enough to choose an agent's next step? tev1 vs the general models.

Findings

Kaggle leaderboard

Rank Model Maker Score
1 gemini-3.5-flash Google 0.99
1 claude-opus-5-5-default Anthropic 0.99
3 glm-5 Z.ai 0.86
4 gemma-4-26b-a4b-it Google 0.85
5 gpt-5.4-2026-03-05 OpenAI 0.84
5 claude-opus-4-5-20251101 Anthropic 0.84
7 grok-4.20-0309-non-reasoning xAI 0.71
8 gpt-oss-20b OpenAI 0.69

The top two tied at 0.99, and there's a clear gap after them. Both solved every scenario correctly. The 0.01 they lost comes from [e.g. one extra tool call on one scenario]. The next group (GLM-5, Gemma, GPT-5.4, Opus 4.5) is packed between 0.84 and 0.86, so the order within that group isn't meaningful.

A newer generation made a big difference. Claude Opus 5.5 scored 0.99 and Opus 4.5 scored 0.84, a jump of 15 points from the same maker. [Which scenario Opus 4.5 failed, from its trace.]

A small open model kept up with frontier models. Gemma 4 26B (0.85) scored about the same as GPT-5.4 (0.84) and Opus 4.5 (0.84), and close to GLM-5 (0.86). On structured tool tasks, size wasn't decisive.

Fast models did well. Gemini 3.5 Flash tied for first. Speed and tool discipline don't seem to conflict here.

The bottom two: Grok non-reasoning (0.71) and gpt-oss-20b (0.69). [What went wrong, from the traces.]

Ollama: Gemma was the most efficient

Model Where Correct answers Avg tool calls Tokens per task Time per task
gemma4:31b Cloud 91.7% 2.5 2,073 3.0s
gpt-oss:120b Cloud 91.7% 2.75 2,776 3.0s
gpt-oss:20b Cloud 83.3% 2.75 2,689 4.8s
tev1:0.8b Laptop 41.7% 2.92 3,435 9.5s
qwen2.5:3b Laptop 16.7% 1.0 1,282 4.9s

gemma4:31b tied gpt-oss:120b on correct answers (91.7%) but got there with fewer calls and 25% fewer tokens, so it was the cheapest per correct answer. gpt-oss:120b beat its 20b sibling by 8 points, so size did help within one family.

The small local models struggled as agents. qwen2.5:3b made one call on average and then guessed. tev1 did more work but often skipped the required FINAL: answer line, and twice took actions nobody asked for (below). Fewer tokens didn't mean cheaper: both small models cost about 3.5× more than Gemma per correct answer.

Chart: accuracy by model

Bar chart comparing task accuracy across evaluated models

Chart: tokens per task, input + output

Bar chart comparing average tokens used per task, including input and output tokens

Chart: time per task

Bar chart comparing average execution time per task across the evaluated models, measured in seconds

Chart: cost per task

Cost

The decision test: tev1 is good at choosing, not at acting

tev1 is a decision model, so judging it only as an agent felt unfair. I added a second test: 20 single decisions taken from the same scenarios (which tool first, which tool next, when to stop, retry or give up). Every model saw the same situation and the same five options. tev1 answered through its decision endpoint; the other models answered a plain prompt.

Model Where Correct decisions Input tokens Output tokens Time per decision
gemma4:31b Cloud 100% 232 4 0.66s
gpt-oss:120b Cloud 95% 262 75 0.92s
gpt-oss:20b Cloud 95% 262 78 1.43s
tev1:0.8b Laptop 85% 315 1 2.13s
qwen2.5:3b Laptop 55% 229 3 2.44s
  • Same model, two very different results: tev1 chose correctly 85% of the time, but completed only 42% of tasks as an agent. Choosing the next step works; carrying out a whole task doesn't.
  • It beat a model nearly 4× its size: 85% vs 55% for qwen2.5:3b, both on the same laptop.
  • One output token per decision, against about 75 for gpt-oss. Its total tokens (~316) were in line with the cloud models.
  • Close to cloud speed on a laptop: about 2.1s per decision, against 0.7–1.4s for cloud models running on data-centre hardware.
  • Its three misses were telling: it checked stock before finding the product, chose to order again after an order was already confirmed, and gave up after a "please retry" error.

_Chart: accuracy by mode _
Bar chart comparing task accuracy across evaluated models. The chart shows the percentage of tasks completed correctly for each model.

Chart: tokens per task, input + output
Bar chart comparing average tokens used per task, including input and output tokens, across the evaluated models.

Chart: time per task

Bar chart comparing average execution time per task across the evaluated models, measured in seconds.

Specific failures from the traces (Ollama run)

  • Persistent failure (Acer stock): all three cloud models answered UNKNOWN. tev1 said the laptop was in stock and placed an order nobody asked for.
  • The ₹50,000 boundary: every Ollama model that answered with product IDs included the laptop priced at exactly ₹50,000. This was the hardest trap in the set.
  • Out-of-stock order: no model claimed success. gpt-oss:20b answered UNKNOWN instead of OUT_OF_STOCK.
  • Distractor tool: no model called get_shipping_estimate.
  • Unneeded tools: for "what is 15% of 2000?", tev1 placed an order for 300 laptops, four times. The cloud models answered directly.

What surprised me

I expected the frontier models to be clearly on top. Instead, a fast model tied for first, and a 26B open model scored the same as GPT-5.4 and Claude Opus 4.5. The biggest single difference was between two generations of the same model, not between makers or sizes.

The second surprise was tev1: the same 0.8B model looked weak as an agent and strong as a decision-maker. How a model is used mattered as much as how big it is.

What changed in how I think about these models: grading the path, not just the answer, showed that models with the same final answer can behave very differently. One checks stock properly and the other gets lucky. Defining "correct" also forced decisions I hadn't expected, like whether ₹50,000 counts as "under ₹50,000."

Caveats

  • With [5] scenarios on Kaggle, one scenario is worth up to 20 points, so scores within a few points (0.84–0.86) are effectively tied.
  • Each model ran once. Models can behave differently on repeat runs.
  • Laptop and cloud times run on different hardware, so compare them loosely. I measured local times on fresh prompts, because repeated prompts hit Ollama's cache and look 20× faster.
  • Token counts aren't fully comparable across model families, because each uses its own tokenizer and chat template.

What I'd measure next

  • A small decision model inside an agent: let tev1 pick each next step and a larger model fill in the arguments, to see whether the pair beats either model alone on cost.
  • Reasoning on vs off: Grok's reasoning version alongside the non-reasoning one.
  • Held-out phrasings: paraphrases, typos, and Hinglish ("MacBook Air M2 ka price kya hai?").
  • Repeat runs per model, to measure consistency.

My Benchmark

AgentCallBench: Tool-Calling Reliability on Kaggle

The task code, including the products, scenarios, fake tools, and scoring, is in the benchmark's task notebook.


Top comments (1)

Collapse
 
abhisekroy169 profile image
Abhisek Roy •

Please share thoughts on that Decisions models vs LLM models