Written for: dev.to readers and the Kaggle Benchmarking Challenge judges.
I kept your format and headings, shortened the intro, added the two local models and the decision test, and filled the Ollama placeholders with numbers from your latest results.
Before you paste it, two corrections affect what the post can claim:
-
tev1is not faster than the cloud models on fresh prompts. The 0.096s in your decision results came from Ollama's prompt cache, because your full run reused the exact prompts from my earlier test run. Measured fresh, it takes about 2.1s per decision, against 0.66–1.43s for the cloud models. I re-rantev1fresh and replaced its rows inresults/decisions/, so the time chart is now correct. The post says "close to cloud speed while running on a laptop", which is true. "Faster" would not be. -
gemma4:31bandgpt-ossdidn't run on your machine. They ran on Ollama Cloud (your.envpoints atollama.com). Onlytev1andqwen2.5:3bran locally. The post now says this.
This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I built AgentToolEval, a small benchmark that checks whether LLM agents use tools correctly, not just whether they reach the right answer. Did the agent check stock or guess? Did it make five calls when two would do? When a tool failed, did it retry, give up, or make something up? A final-answer check hides all of that.
The model plays a shopping assistant for a fake electronics store with four tools: search_products (prices but not stock), check_stock, place_order, and get_shipping_estimate, a distractor no task needs. The tools are deterministic, so every run is reproducible, and failures are injected on purpose (a timeout once, a service that never recovers, an out-of-stock order).
| Scenario | The trap |
|---|---|
| Cheapest in-stock laptop | The cheapest laptop overall is out of stock |
| Laptops under ₹50,000 in stock | One laptop costs exactly ₹50,000 ("under" means strictly below) |
| Order an out-of-stock phone | The model must refuse, not claim success |
| MacBook price | The search fails once, so the model must retry |
| Acer stock check | The stock service never recovers, so the model must answer UNKNOWN |
Each scenario scores 0 to 1: 60% correct answer, 15% right tools, 15% correct arguments, 10% efficiency (only when the answer is right). The model never sees the answer key; grading happens afterwards, from a log of every tool call.
Models Tested
I ran 8 models on Kaggle and 5 models with Ollama, chosen in groups that each answer a specific question.
On Kaggle ([5] scenarios):
| Model | Maker | Why it's in the lineup |
|---|---|---|
claude-opus-5-5-default |
Anthropic | Newest frontier model, paired with Opus 4.5 |
claude-opus-4-5-20251101 |
Anthropic | Previous generation, to see whether newer means better at tools |
gpt-5.4-2026-03-05 |
OpenAI | Frontier closed model |
gpt-oss-20b |
OpenAI | Small open model from the same maker |
gemini-3.5-flash |
Fast closed model | |
gemma-4-26b-a4b-it |
Small open model from the same maker | |
glm-5 |
Z.ai | Open model from another maker |
grok-4.20-0309-non-reasoning |
xAI | Non-reasoning model from a fourth maker |
With Ollama (12 scenarios):
| Model | Where it ran | Why |
|---|---|---|
gemma4:31b |
Ollama Cloud | Counterpart to the Gemma model on Kaggle |
gpt-oss:120b |
Ollama Cloud | Large vs small from the same family... |
gpt-oss:20b |
Ollama Cloud | ...to test whether size improves tool use |
qwen2.5:3b |
My laptop | A small general model, as a baseline |
tev1:0.8b |
My laptop | A 0.8B decision model built to pick one option from a list, not to run as an agent |
The lineup lets me ask:
- Does a newer generation help? Claude Opus 5.5 vs Opus 4.5.
- Closed vs open from the same maker: GPT-5.4 vs gpt-oss-20b, and Gemini 3.5 Flash vs Gemma 4.
- Does size help? gpt-oss 20b vs 120b, and 1–3B local models against cloud ones.
- Is a small decision model good enough to choose an agent's next step? tev1 vs the general models.
Findings
Kaggle leaderboard
| Rank | Model | Maker | Score |
|---|---|---|---|
| 1 | gemini-3.5-flash | 0.99 | |
| 1 | claude-opus-5-5-default | Anthropic | 0.99 |
| 3 | glm-5 | Z.ai | 0.86 |
| 4 | gemma-4-26b-a4b-it | 0.85 | |
| 5 | gpt-5.4-2026-03-05 | OpenAI | 0.84 |
| 5 | claude-opus-4-5-20251101 | Anthropic | 0.84 |
| 7 | grok-4.20-0309-non-reasoning | xAI | 0.71 |
| 8 | gpt-oss-20b | OpenAI | 0.69 |
The top two tied at 0.99, and there's a clear gap after them. Both solved every scenario correctly. The 0.01 they lost comes from [e.g. one extra tool call on one scenario]. The next group (GLM-5, Gemma, GPT-5.4, Opus 4.5) is packed between 0.84 and 0.86, so the order within that group isn't meaningful.
A newer generation made a big difference. Claude Opus 5.5 scored 0.99 and Opus 4.5 scored 0.84, a jump of 15 points from the same maker. [Which scenario Opus 4.5 failed, from its trace.]
A small open model kept up with frontier models. Gemma 4 26B (0.85) scored about the same as GPT-5.4 (0.84) and Opus 4.5 (0.84), and close to GLM-5 (0.86). On structured tool tasks, size wasn't decisive.
Fast models did well. Gemini 3.5 Flash tied for first. Speed and tool discipline don't seem to conflict here.
The bottom two: Grok non-reasoning (0.71) and gpt-oss-20b (0.69). [What went wrong, from the traces.]
Ollama: Gemma was the most efficient
| Model | Where | Correct answers | Avg tool calls | Tokens per task | Time per task |
|---|---|---|---|---|---|
| gemma4:31b | Cloud | 91.7% | 2.5 | 2,073 | 3.0s |
| gpt-oss:120b | Cloud | 91.7% | 2.75 | 2,776 | 3.0s |
| gpt-oss:20b | Cloud | 83.3% | 2.75 | 2,689 | 4.8s |
| tev1:0.8b | Laptop | 41.7% | 2.92 | 3,435 | 9.5s |
| qwen2.5:3b | Laptop | 16.7% | 1.0 | 1,282 | 4.9s |
gemma4:31b tied gpt-oss:120b on correct answers (91.7%) but got there with fewer calls and 25% fewer tokens, so it was the cheapest per correct answer. gpt-oss:120b beat its 20b sibling by 8 points, so size did help within one family.
The small local models struggled as agents. qwen2.5:3b made one call on average and then guessed. tev1 did more work but often skipped the required FINAL: answer line, and twice took actions nobody asked for (below). Fewer tokens didn't mean cheaper: both small models cost about 3.5× more than Gemma per correct answer.
Chart: accuracy by model
Chart: tokens per task, input + output
Chart: time per task
Chart: cost per task
The decision test: tev1 is good at choosing, not at acting
tev1 is a decision model, so judging it only as an agent felt unfair. I added a second test: 20 single decisions taken from the same scenarios (which tool first, which tool next, when to stop, retry or give up). Every model saw the same situation and the same five options. tev1 answered through its decision endpoint; the other models answered a plain prompt.
| Model | Where | Correct decisions | Input tokens | Output tokens | Time per decision |
|---|---|---|---|---|---|
| gemma4:31b | Cloud | 100% | 232 | 4 | 0.66s |
| gpt-oss:120b | Cloud | 95% | 262 | 75 | 0.92s |
| gpt-oss:20b | Cloud | 95% | 262 | 78 | 1.43s |
| tev1:0.8b | Laptop | 85% | 315 | 1 | 2.13s |
| qwen2.5:3b | Laptop | 55% | 229 | 3 | 2.44s |
- Same model, two very different results: tev1 chose correctly 85% of the time, but completed only 42% of tasks as an agent. Choosing the next step works; carrying out a whole task doesn't.
- It beat a model nearly 4× its size: 85% vs 55% for qwen2.5:3b, both on the same laptop.
- One output token per decision, against about 75 for gpt-oss. Its total tokens (~316) were in line with the cloud models.
- Close to cloud speed on a laptop: about 2.1s per decision, against 0.7–1.4s for cloud models running on data-centre hardware.
- Its three misses were telling: it checked stock before finding the product, chose to order again after an order was already confirmed, and gave up after a "please retry" error.
Chart: tokens per task, input + output

Chart: time per task
Specific failures from the traces (Ollama run)
- Persistent failure (Acer stock): all three cloud models answered UNKNOWN. tev1 said the laptop was in stock and placed an order nobody asked for.
- The ₹50,000 boundary: every Ollama model that answered with product IDs included the laptop priced at exactly ₹50,000. This was the hardest trap in the set.
- Out-of-stock order: no model claimed success. gpt-oss:20b answered UNKNOWN instead of OUT_OF_STOCK.
-
Distractor tool: no model called
get_shipping_estimate. - Unneeded tools: for "what is 15% of 2000?", tev1 placed an order for 300 laptops, four times. The cloud models answered directly.
What surprised me
I expected the frontier models to be clearly on top. Instead, a fast model tied for first, and a 26B open model scored the same as GPT-5.4 and Claude Opus 4.5. The biggest single difference was between two generations of the same model, not between makers or sizes.
The second surprise was tev1: the same 0.8B model looked weak as an agent and strong as a decision-maker. How a model is used mattered as much as how big it is.
What changed in how I think about these models: grading the path, not just the answer, showed that models with the same final answer can behave very differently. One checks stock properly and the other gets lucky. Defining "correct" also forced decisions I hadn't expected, like whether ₹50,000 counts as "under ₹50,000."
Caveats
- With [5] scenarios on Kaggle, one scenario is worth up to 20 points, so scores within a few points (0.84–0.86) are effectively tied.
- Each model ran once. Models can behave differently on repeat runs.
- Laptop and cloud times run on different hardware, so compare them loosely. I measured local times on fresh prompts, because repeated prompts hit Ollama's cache and look 20× faster.
- Token counts aren't fully comparable across model families, because each uses its own tokenizer and chat template.
What I'd measure next
- A small decision model inside an agent: let tev1 pick each next step and a larger model fill in the arguments, to see whether the pair beats either model alone on cost.
- Reasoning on vs off: Grok's reasoning version alongside the non-reasoning one.
- Held-out phrasings: paraphrases, typos, and Hinglish ("MacBook Air M2 ka price kya hai?").
- Repeat runs per model, to measure consistency.
My Benchmark
AgentCallBench: Tool-Calling Reliability on Kaggle
The task code, including the products, scenarios, fake tools, and scoring, is in the benchmark's task notebook.






Top comments (1)
Please share thoughts on that Decisions models vs LLM models