You run a coding agent against a local model like for example qwen2.5-coder on Ollama, because it's a solid coder and it fits on your GPU. You give it a task, it thinks for a bit, prints a plausible-looking summary of what it did, and stops.
Except it didn't do anything. No files changed. The agent never called a single tool.
This isn't a rare failure, and it isn't only the small models - I saw it on the 7B and the 14B. So I ran an actual benchmark instead of trusting my own impression.
The benchmark
Six small coding tasks (add a CLI subcommand, fix an off-by-one, rename a function across two files, trace a runtime error, etc.), run through four agents against the same local models, 3 trials each, automated pass/fail scoring:
- Polyglot (mine) - parses tool calls out of the model's raw text instead of trusting native function-calling, with a repair pass for the malformed ones
- pi - a genuinely good CLI agent, native function-calling only
- Hermes Agent (Nous Research) - advertises "11 tool-call parsers"
-
Goose - Linux Foundation's Agentic AI Foundation (AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, and Microsoft are platinum members), tested with their own
GOOSE_TOOLSHIMfix for this exact problem explicitly turned on
Why qwen2.5-coder first, and why I didn't stop there. It's not the newest model, but it's still one of the most-run local coding models on Ollama - testing it is testing what people actually have installed today, not chasing a target picked to guarantee a bad number. But I didn't want this to be a "gotcha" on one aging model, so I scaled it on purpose: same six tasks, same scoring, all the way from 7B up through 32B in the same family, then a full newer generation (qwen3-coder, 2025), then the actual current, hyped release (Qwen3.8-27B, out last month). The results below cover the whole range - the pattern either holds or it doesn't, and I'd rather show you both than only the part that makes the point.
Results on qwen2.5-coder:7b
| Agent | Tasks completed | Runs with zero tool calls |
|---|---|---|
| Polyglot | 39% (7/18) | 0/18 |
| pi | 0% (0/18) | 18/18 |
| Hermes | 0% (0/18) | 18/18 |
| Goose | 0% (0/18) | 18/18 |
On the 14B: Polyglot ~79%, Hermes and Goose still 0%.
Scaling up: does it go away at 32B?
Same family, same tasks, just bigger: qwen2.5-coder:32b.
| Agent | Tasks completed |
|---|---|
| Polyglot | 83% (15/18) |
| pi | 0% (0/18) |
| Goose | 0% (0/18) |
| Hermes | 0% (0/18) |
| opencode | 0% (0/18) |
Scale genuinely helps Polyglot here - 39% -> ~79% -> 83% as the model gets bigger, a real, sensible improvement. What doesn't move is the other side: every competitor still made zero tool calls, on every run, at four times the parameters. Whatever's keeping the native tool-calling channel from firing isn't something raw scale fixes on its own within this family - and getting an honest read on this 32B number specifically took a real fix on my end too, the same class of bug that shows up again later in this post: my first pass under-reported it because of a test-harness timeout miscalibrated for a bigger, slower-to-run model, not because of anything the model actually did wrong.
Why it happens
Hosted models (Claude, GPT) emit tool calls through a dedicated, structured channel. Open-weight models are trained to do the same thing, but the training is thinner and less consistent - the smaller the model, the more it wobbles. So qwen2.5-coder:7b, asked to edit a file, will often write this as ordinary text in the middle of its reply instead of through the native channel:
{
"name": "edit",
"arguments": { "path": "math.mjs", "edits": [ "..." ] }
}
Right intent, wrong place. A runtime that only listens on the native tool-call channel sees nothing there and treats the whole reply as a finished answer. Task over. Zero tools called. No error surfaces - it just silently didn't do the work.
The one that surprised me: Goose
Goose ships a real, documented fix for exactly this - GOOSE_TOOLSHIM routes a non-tool-calling model's output through a second, interpretive model to convert it into a proper call. I turned it on and reran everything. Tool-call attempts went up immediately - some runs made 20 of them instead of one - so the mechanism is doing something. Task completion barely moved: 0/18 on the 7B, 1/18 on the 14B, mostly failing on schema mismatches or running out of turns before finishing. A foundation-scale project with a purpose-built fix for this exact failure mode still didn't close the gap in my testing.
(Honest caveat: I substituted llama3.2:3b as the interpreter model since their documented default, mistral-nemo, wasn't already on my machine - worth someone re-running with the real default before treating that specific number as final.)
And Hermes's "11 tool-call parsers"?
Reading the hermes-agent source: those parsers live at the model-training and serving layer, not the runtime. The live agent loop reads the native tool_calls field only, and there's a helper that actively deletes text-emitted tool-call blocks from the model's output as noise - with a comment treating in-context tool-call syntax as "data, do not re-emit it as a tool call." Deliberate design choice, opposite of mine.
One generation newer: qwen3-coder
32B ruled out size within the qwen2.5 family. Next question: does a newer generation, still not the current hyped one, change anything? qwen3-coder (2025) is a real jump - and here, native function-calling is genuinely solid: Polyglot and pi are both strong (~99% vs 89%). This isn't "native function-calling is always bad" - it's specifically a gap on models that wobble, which is most non-flagship open-weight releases and most quantized models people actually run locally.
Not every model has this problem; the point is you can't know which ones do without testing.
What Polyglot does differently
- Teaches the tool-call grammar in the system prompt.
- Parses tool calls out of the streamed text as it arrives.
- Repairs the near-misses - trailing commas, single quotes, a fenced code block wrapped around the arguments, a tool name one character off - and runs the call anyway.
- Shows you every repair, with the raw model output one keypress away. A parser fix should never quietly hide a model getting worse.
Full methodology and more honest caveats in the original post.
The last rung: the model generating buzz right now
qwen3-coder showed the gap isn't universal - some models handle this fine. So the top of the range is the actual current release people are talking about: Qwen3.8-27B, out last month. I benchmarked it the same way as everything above - and got the first clean sweep in this whole series: pi, goose, and opencode all hit 18/18.
That's a real result and worth saying plainly: a well-trained current-generation model can nail tool-calling without any repair layer at all. But the full story turned out to be more useful than the clean number alone - my own first result on this model was wrong (12/18), and digging in found two real bugs: a tool-call closing-tag variant my own parser didn't recognize (a genuine product bug, now fixed), and a test-harness timeout calibrated for smaller/faster models that unfairly cut off a bigger one mid-fix. Fixed both, reran, landed at 18/18 - tied with the field,not ahead, not behind.
The actual takeaway: reliability is genuinely uneven - across models, across configs, even within
one "good" model's own supporting tools (goose-toolshim landed at 16/18 on the same model, hermes
at 17/18) - and that variance is invisible until you actually measure it. That's a better argument
than "every model is broken," honestly, because it doesn't expire as models improve. Full writeup with the debugging details here.
Try it
Free and open source: npm install -g @usepolyglot/cli, point it at whatever's already running on your GPU, see if your agent is doing the work or just narrating it.
Top comments (0)