DEV Community

Giuseppe Sirigu
Giuseppe Sirigu

Posted on Originally published at usepolyglot.dev AI-assisted

Your local coding agent might be doing nothing at all

You run a coding agent against a local model like for example qwen2.5-coder on Ollama, because it's a solid coder and it fits on your GPU. You give it a task, it thinks for a bit, prints a plausible-looking summary of what it did, and stops.

Except it didn't do anything. No files changed. The agent never called a single tool.

This isn't a rare failure, and it isn't only the small models - I saw it on the 7B and the 14B. So I ran an actual benchmark instead of trusting my own impression.

The benchmark

Six small coding tasks (add a CLI subcommand, fix an off-by-one, rename a function across two files, trace a runtime error, etc.), run through four agents against the same local models, 3 trials each, automated pass/fail scoring:

  • Polyglot (mine) - parses tool calls out of the model's raw text instead of trusting native function-calling, with a repair pass for the malformed ones
  • pi - a genuinely good CLI agent, native function-calling only
  • Hermes Agent (Nous Research) - advertises "11 tool-call parsers"
  • Goose - Linux Foundation's Agentic AI Foundation (AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, and Microsoft are platinum members), tested with their own GOOSE_TOOLSHIM fix for this exact problem explicitly turned on

Why qwen2.5-coder first, and why I didn't stop there. It's not the newest model, but it's still one of the most-run local coding models on Ollama - testing it is testing what people actually have installed today, not chasing a target picked to guarantee a bad number. But I didn't want this to be a "gotcha" on one aging model, so I scaled it on purpose: same six tasks, same scoring, all the way from 7B up through 32B in the same family, then a full newer generation (qwen3-coder, 2025), then the actual current, hyped release (Qwen3.8-27B, out last month). The results below cover the whole range - the pattern either holds or it doesn't, and I'd rather show you both than only the part that makes the point.

Results on qwen2.5-coder:7b

Agent Tasks completed Runs with zero tool calls
Polyglot 39% (7/18) 0/18
pi 0% (0/18) 18/18
Hermes 0% (0/18) 18/18
Goose 0% (0/18) 18/18

On the 14B: Polyglot ~79%, Hermes and Goose still 0%.

Scaling up: does it go away at 32B?

Same family, same tasks, just bigger: qwen2.5-coder:32b.

Agent Tasks completed
Polyglot 83% (15/18)
pi 0% (0/18)
Goose 0% (0/18)
Hermes 0% (0/18)
opencode 0% (0/18)

Scale genuinely helps Polyglot here - 39% -> ~79% -> 83% as the model gets bigger, a real, sensible improvement. What doesn't move is the other side: every competitor still made zero tool calls, on every run, at four times the parameters. Whatever's keeping the native tool-calling channel from firing isn't something raw scale fixes on its own within this family - and getting an honest read on this 32B number specifically took a real fix on my end too, the same class of bug that shows up again later in this post: my first pass under-reported it because of a test-harness timeout miscalibrated for a bigger, slower-to-run model, not because of anything the model actually did wrong.

Why it happens

Hosted models (Claude, GPT) emit tool calls through a dedicated, structured channel. Open-weight models are trained to do the same thing, but the training is thinner and less consistent - the smaller the model, the more it wobbles. So qwen2.5-coder:7b, asked to edit a file, will often write this as ordinary text in the middle of its reply instead of through the native channel:

{
  "name": "edit",
  "arguments": { "path": "math.mjs", "edits": [ "..." ] }
}
Enter fullscreen mode Exit fullscreen mode

Right intent, wrong place. A runtime that only listens on the native tool-call channel sees nothing there and treats the whole reply as a finished answer. Task over. Zero tools called. No error surfaces - it just silently didn't do the work.

The one that surprised me: Goose

Goose ships a real, documented fix for exactly this - GOOSE_TOOLSHIM routes a non-tool-calling model's output through a second, interpretive model to convert it into a proper call. I turned it on and reran everything. Tool-call attempts went up immediately - some runs made 20 of them instead of one - so the mechanism is doing something. Task completion barely moved: 0/18 on the 7B, 1/18 on the 14B, mostly failing on schema mismatches or running out of turns before finishing. A foundation-scale project with a purpose-built fix for this exact failure mode still didn't close the gap in my testing.

(Honest caveat: I substituted llama3.2:3b as the interpreter model since their documented default, mistral-nemo, wasn't already on my machine - worth someone re-running with the real default before treating that specific number as final.)

And Hermes's "11 tool-call parsers"?

Reading the hermes-agent source: those parsers live at the model-training and serving layer, not the runtime. The live agent loop reads the native tool_calls field only, and there's a helper that actively deletes text-emitted tool-call blocks from the model's output as noise - with a comment treating in-context tool-call syntax as "data, do not re-emit it as a tool call." Deliberate design choice, opposite of mine.

One generation newer: qwen3-coder

32B ruled out size within the qwen2.5 family. Next question: does a newer generation, still not the current hyped one, change anything? qwen3-coder (2025) is a real jump - and here, native function-calling is genuinely solid: Polyglot and pi are both strong (~99% vs 89%). This isn't "native function-calling is always bad" - it's specifically a gap on models that wobble, which is most non-flagship open-weight releases and most quantized models people actually run locally.
Not every model has this problem; the point is you can't know which ones do without testing.

What Polyglot does differently

  1. Teaches the tool-call grammar in the system prompt.
  2. Parses tool calls out of the streamed text as it arrives.
  3. Repairs the near-misses - trailing commas, single quotes, a fenced code block wrapped around the arguments, a tool name one character off - and runs the call anyway.
  4. Shows you every repair, with the raw model output one keypress away. A parser fix should never quietly hide a model getting worse.

Full methodology and more honest caveats in the original post.

The last rung: the model generating buzz right now

qwen3-coder showed the gap isn't universal - some models handle this fine. So the top of the range is the actual current release people are talking about: Qwen3.8-27B, out last month. I benchmarked it the same way as everything above - and got the first clean sweep in this whole series: pi, goose, and opencode all hit 18/18.

That's a real result and worth saying plainly: a well-trained current-generation model can nail tool-calling without any repair layer at all. But the full story turned out to be more useful than the clean number alone - my own first result on this model was wrong (12/18), and digging in found two real bugs: a tool-call closing-tag variant my own parser didn't recognize (a genuine product bug, now fixed), and a test-harness timeout calibrated for smaller/faster models that unfairly cut off a bigger one mid-fix. Fixed both, reran, landed at 18/18 - tied with the field,not ahead, not behind.

The actual takeaway: reliability is genuinely uneven - across models, across configs, even within
one "good" model's own supporting tools (goose-toolshim landed at 16/18 on the same model, hermes
at 17/18) - and that variance is invisible until you actually measure it. That's a better argument
than "every model is broken," honestly, because it doesn't expire as models improve. Full writeup with the debugging details here.

Try it

Free and open source: npm install -g @usepolyglot/cli, point it at whatever's already running on your GPU, see if your agent is doing the work or just narrating it.

Repo: https://github.com/giuseppe-sirigu/polyglot

Top comments (0)