DEV Community

Cover image for Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens
xbill Subscriber for Google Developer Experts

Posted on

Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

This article reports a Kaggle benchmark for a step every agent performs and few people test, counting what a tool returns, and ends with a step by step guide to reproducing it.

An agent's search, database or API tool usually hands back a list of records. When the user asks how many, the model does the counting. That step passes every quick test with a handful of rows, and nothing flags it when it goes wrong: the model sends the right query and quotes a confident number.

This benchmark asks ten models the same 68 counting questions with two versions of one tool: count_ids returns the exact count, list_ids returns the matching ids.

With the count, every model but one answered every question about 330 ids correctly. With the ids, the models split in two. The ones that spent 2,700 to 8,200 output tokens per question counted 15 to 21 of 21 lists correctly. The ones that answered in under 600 tokens counted 0 to 10.

The surprise was who landed where: Claude Opus 5, the most expensive model in the lineup, answered in 579 tokens and counted 9 of 21. Gemma 4 26B, a 26B open-weight model, spent 8,153 and counted 20.

https://github.com/xbill9/devto-kaggle

https://www.kaggle.com/benchmarks/xbillwork/count-it-or-compute-it


What I Benchmarked

The itch is eleven ids:

0, 2, 3, 20, 21, 22, 23, 10, 11, 12, 13
Enter fullscreen mode Exit fullscreen mode

How many are 10 or more? The answer is 8, and every model here gets it. Eleven ids say nothing about three hundred, and eleven is about the size of the lists most quick tests use.

Language models' trouble with counting is familiar from examples like counting the letters in a word. This benchmark measures it where agents meet it: inside tool use, with every miss traced to the query or the count, and every answer priced in tokens.

The benchmark holds everything fixed except who does the arithmetic:

Task What the model gets Who counts
count-engine count_ids(where), which returns the exact count, minimum and maximum The tool
count-rows-tool list_ids(where), which returns the matching ids The model, from the returned list
count-python-tool Every id in the prompt, plus run_python with ids already defined The model's code, if it writes any

In the first two tasks the ids never appear in the prompt, so the only difference is whether the tool returns a number or a list. Both tasks check every filter the model sends by the ids it selects, so id > 9 counts as right for "10 or more", and every wrong answer is traced to either a wrong query or a wrong count.

Each task asks 68 questions: lists of 11, 110 and 330 ids, seven phrasings of the threshold ("10 or more", "no less than", "under", "between 5 and 9 inclusive" and so on), three seeds each, and the original eleven ids five times. Every threshold is an id in the list, so > and >= always give different answers.


Models Tested

Vendor Models
Google Gemini 2.5 Flash, Gemini 3.7 Flash, Gemini 3.8 Flash, Gemma 4 26B A4B
Anthropic Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5
OpenAI GPT-5.4 nano, GPT-5.4 mini, gpt-oss-20b

The lineup takes a small, a mid-sized and a large model from each vendor, plus the open-weight models from Google and OpenAI, so the comparison covers price, size and what anyone can run. Each model runs with the settings Kaggle's model proxy serves it with. Some of those reason before answering and some answer straight away, and that setting tracks the result more closely than size or price.

GPT-6 Astra is refused function tools by the proxy (Function tools with reasoning_effort are not supported for gpt-6-astra in /v1/chat/completions). Gemini 3.5 Flash-Lite and both Qwen 3 Next 80B models returned 429 or 503 on most calls.


Findings

At 330 ids, the size where the models separate:

Model Rows tool correct Output tokens per question, rows tool Engine correct Rows-tool cost vs engine
Gemma 4 26B A4B 20/21 8,153 21/21 18.6x
Gemini 3.7 Flash 20/21 2,785 21/21 13.5x
Gemini 3.8 Flash 21/21 2,742 21/21 14.3x
gpt-oss-20b 15/21 2,684 19/21 5.2x
Gemini 2.5 Flash 0/21 580 21/21 2.6x
Claude Opus 5 9/21 579 21/21 2.4x
Claude Sonnet 5 10/21 401 21/21 1.9x
Claude Haiku 4.5 5/21 143 21/21 1.5x
GPT-5.4 mini 5/21 39 21/21 1.9x
GPT-5.4 nano 0/21 35 21/21 1.7x

The Claude rows-tool figures come from their 2026-09-25 run, which asked the same 68 questions with the same tool; their 2026-09-28 rows-tool run stopped on the daily quota. Every other figure comes from the 2026-09-28 runs.

1. With the Count in the Tool, Every Model Is Right at a Flat Cost

Every model but one answered all 21 engine questions about 330 ids correctly, at 35 to 451 output tokens per question, and each model's tokens stayed about the same from 11 ids to 330. Turning "no less than 244" into id >= 244 and quoting the number back is something every model here does reliably. The exception, gpt-oss-20b, answered 5 of 68 engine questions with a number other than the count it was given, all of them 0, 1 or 2, and answered once without calling the tool.

2. With the Rows, Quick Tests Pass and Larger Lists Fail

Eight of the ten models counted all 26 questions about 11 ids correctly. At 330 ids, five of those eight counted 10 or fewer of 21 correctly. None of them returned an error or a hedge: each answer was a single confident number.

3. The Models That Count Right Spend the Tokens

The models split into two groups by how many tokens they spend, and model size and price do not predict the split. That was the surprise. The models that count correctly spend more tokens as the list grows: Gemma 4 26B went from 394 output tokens per question at 11 ids to 8,153 at 330. The models that miscount spend about the same at every size: GPT-5.4 nano spent 35 tokens per question at 11, 110 and 330 ids, which leaves no room to count anything. Claude Opus 5, the most expensive model here, spent 579 output tokens per question at 330 ids and counted 9 of 21 correctly; Gemma 4 26B, an open-weight model, spent 8,153 and counted 20.

The same models landed in the same group in every rows-tool run, three or four runs per model between 2026-09-25 and 2026-09-28. The order inside a group moves by a few questions between runs.

Counting right by reasoning costs 5 to 19 times what the engine costs for the same answer. An earlier run with 1,100 ids in the prompt shows the same link from the other side: with output capped at 8,192 tokens, Gemini 3.7 Flash counted 1 of 21 lists correctly, against 19 and 15 of 21 without the cap, and still gave a number every time.

4. The Query Was Right Every Time

No answer on either task rested on a filter that selected the wrong ids. Every miss on the rows tool was the model counting the correct list wrong, apart from 3 answers given without calling the tool and 1 that could not be read as a number. The failure is in the arithmetic.

5. Python Works When the Model Uses It

With every id in the prompt and a Python tool available, 7 of 10 models scored 68 of 68; Gemini 2.5 Flash called the tool on 1 of 42 questions about 110 and 330 ids and scored 37.

What It Changed About How I Think About These Models

In these runs, counting a list was work the model did in tokens, and the models that answered straight away had not done it. A tool that returns rows moves that work onto the model and makes its accuracy depend on a setting the caller may never have looked at. The count belongs in the tool: return the count, the minimum and the maximum, and every model in this lineup answers correctly at a fraction of the tokens.


Compare and Contrast

Engine Rows tool Python tool
Who counts The tool The model, from the returned list The model's code, if it writes any
Output tokens per question at 330 ids 35 to 451 35 to 8,153 —
What went wrong Quoting a number other than the count Miscounting the right list Skipping the tool, no print()
Answers built on a wrong filter 0 0 —

So, Which One?

  • 🟢 Engine — put the count, minimum and maximum in the tool. Every model is right at a flat cost.
  • ⚠️ Python tool — reliable when the model uses it, prints the result and quotes it.
  • ❌ Rows tool — correct only with a model that reasons through the list, at 5 to 19 times the cost per question.

What I'd Measure Next

  • The same model with reasoning on and off. Claude Opus 5 with extended thinking against its default. The runs here show that tokens and accuracy move together across models; this would show whether switching reasoning on is enough to fix one model.
  • Other arithmetic. Sums, averages and the largest value over returned rows, which the same tool-shape question applies to.
  • Harder filters. "At least 10 but under 20" and "outside 5 to 9", where a wrong filter with an exact count would show up.

How to Reproduce It

These steps rebuild the benchmark from the repository; the tasks, the runs and every table above come from these commands.


At This Point You Should Have…

  • A Kaggle account, with the Kaggle CLI installed (pip install kaggle, version 2.2.4 here) and logged in with kaggle auth login
  • Python 3 for the local checks and the report; the kaggle_benchmarks library the tasks import is already installed on Kaggle
  • The repository cloned: git clone https://github.com/xbill9/devto-kaggle and cd devto-kaggle

Step 1 — Check the Questions Locally

Every question is generated in code, and its expected count comes from running a filter over the ids. check.py builds all 68 questions with no model calls and no Kaggle library:

python3 tasks/check.py
Enter fullscreen mode Exit fullscreen mode
ok: 68 rows per task, by size {11: 26, 110: 21, 330: 21}
Enter fullscreen mode Exit fullscreen mode

Step 2 — Push Each Task

Each file in tasks/ is one Kaggle task. The first @kbench.task in the file returns the share of questions answered correctly, because Kaggle names the task and reads its score from the first one it finds. A second task answers one question and runs once per question:

@kbench.task(name="count-engine")
def count_engine(llm) -> float:
    runs = count_engine_row.evaluate(
        llm=[llm], evaluation_data=pd.DataFrame(ROWS), n_jobs=4, on_failure="continue")
    return summarize(runs, len(ROWS), "engine")
Enter fullscreen mode Exit fullscreen mode

Each question's result keeps the answer, every filter the model sent, and the tokens and cost Kaggle's model proxy reported for it. The push slug must match the task name:

kaggle b t push count-engine -f tasks/count_engine.py --wait
kaggle b t status count-engine
Enter fullscreen mode Exit fullscreen mode
Task:     count-engine
Version:  13
Status:   Completed
Created:  2026-09-28 16:11:45
Enter fullscreen mode Exit fullscreen mode

Step 3 — Run the Models

kaggle b auth -y fetches the short-lived key for Kaggle's model proxy, and kaggle b t models lists the model slugs. A run takes one -m per model:

kaggle b t run count-rows-tool -m gemini-2.5-flash -m gemini-3.8-flash -m gemma-4-26b-a4b-it -m gpt-oss-20b -m gpt-5.4-nano-2026-03-17 -m gpt-5.4-mini-2026-03-17 --wait
Enter fullscreen mode Exit fullscreen mode
All runs completed:
  gpt-5.4-mini-2026-03-17: COMPLETED
  gpt-5.4-nano-2026-03-17: COMPLETED
  gpt-oss-20b: COMPLETED
  gemma-4-26b-a4b-it: COMPLETED
  gemini-3.8-flash: COMPLETED
  gemini-2.5-flash: COMPLETED
Enter fullscreen mode Exit fullscreen mode

🔎 Tip: Cap the Output to Stay Inside the Quota

Kaggle's model proxy reserves the worst-case cost of a call, based on the output-token limit, and refuses the call when that exceeds what is left of the day's quota. With no limit set, GPT-6 Astra reserved $6.40 per call and Claude Opus 5 $3.20. Passing a limit keeps the reservation to cents:

llm.prompt(prompt, schema=int, extra_api_params={"max_completion_tokens": 8192})
Enter fullscreen mode Exit fullscreen mode

Step 4 — Download the Runs and Build the Tables

report.py builds every table in this article from the downloaded runs, including the tokens and cost per question:

for t in count-engine count-rows-tool count-python-tool; do kaggle b t download $t -o results; done
python3 tasks/report.py results
Enter fullscreen mode Exit fullscreen mode
| Model | Size | Engine correct | Engine out tokens | Engine $ | Rows correct | Rows out tokens | Rows $ |
|---|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | 330 | 21/21 | 140 | 0.0010 | 21/21 | 2,742 | 0.0146 |
| GPT-5.4 nano | 330 | 21/21 | 35 | 0.0002 | 0/21 | 35 | 0.0003 |
Enter fullscreen mode Exit fullscreen mode

Step 5 — Group the Tasks Into a Benchmark

A benchmark is created in the Kaggle web UI. On the benchmark page, Add Tasks adds the tasks and Add Models the models. Under Settings, set the overall score to Average of task scores: the default, Percentage of tasks passed, ignores tasks that return a number. Setting the visibility to Public is permanent.


🔎 Tip: The Leaderboard Shows Each Model's Latest Run

Each leaderboard cell shows the model's most recent run of that task, so a run that fails on the quota replaces a complete one before it. Check the quota before a run, and rerun the pair if one fails. pending.py lists pairs whose latest run is incomplete.


My Benchmark


Summary

The goal of this article was to measure whether models count correctly what their tools return. The key to the solution was changing only what the tool returns, a count or a list, and recording every filter and every token.

The results were:

  • 🟢 With the count in the tool, every model but gpt-oss-20b answered all 330-id questions correctly.
  • ⚠️ With the rows, the models that spent 2,700 to 8,200 output tokens per question counted 15 to 21 of 21 lists of 330 correctly.
  • ❌ The models that answered in under 600 tokens counted 0 to 10 of 21, Claude Opus 5 among them, with no error to show it.

Each model ran each task on Kaggle's model proxy with its default settings and output capped at 8,192 tokens; the engine and rows-tool figures come from runs on 2026-09-28, apart from the Claude rows-tool figures from 2026-09-25, the Python-tool figures from 2026-09-28, and the token-cap result from earlier in-context runs with lists of 1,100 ids.

The strategy for benchmarking counting across 10 models was validated with an incremental step by step approach.


References

Top comments (2)

Collapse
 
reidmarlow profile image
Reid Marlow •

The 14x to 18x token multiple on reasoning models is the clearest breakdown of the hidden tax on naive tool schemas I've seen.

When developers wrap REST or database endpoints into MCP tools, the default reflex is dumping the raw collection because it feels flexible. But pushing aggregation onto the model doesn't just invite off-by-ten arithmetic misses when output budgets are tight. It also pollutes the entire downstream transcript: once those 330 IDs enter the session history, every subsequent turn pays context cache and latency costs for tokens that were only needed to answer a single 'how many' query.

Returning a structured count alongside a tiny head sample (or separating listing from aggregation tools) keeps the execution cost flat and stops transcript bloat before it starts.

Collapse
 
dexoryn profile image
Dexoryn •

I actually think The 330-ID result makes the tool-design lesson hard to ignore. If the tool already knows the count, making the model reconstruct it from hundreds of rows is just expensive and less reliable. Returning the aggregation directly feels like a much better agent architecture than asking the LLM to do deterministic work.