Last week I posted a small benchmark on whether coding agents still work when the model you run yourself is shaky at tool calls. It used three runs per task, a handful of models, and a harness I kept private. Fair criticism followed, so here's the bigger, stricter version.
Seven models running locally, including the current ones people actually pick today. Six agents. Thirty runs per agent per model. Tiers decided before the first run. And the whole thing - harness, raw results, failed-run transcripts - is public on GitHub.
Disclosure up front: I build Polyglot, one of the six agents. That's exactly why everything is public.
The result
Each cell is how many of 30 runs completed the task, checked automatically against the files the agent left behind.
| Model | Polyglot | pi | goose | goose + toolshim | Hermes | opencode |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | 27 | 28 | 30 | 27 | 28 | 30 |
| gpt-oss 20B | 30 | 29 | 18 | 5 | 18 | 15 |
| Devstral Small 2 24B | 29 | 28 | 24 | 18 | 30 | 23 |
| qwen3-coder 30B | 30 | 26 | 29 | 25 | 29 | 23 |
| qwen2.5-coder 32B | 24 | 0 | 0 | 28 | 0 | 0 |
| qwen2.5-coder 14B | 27 | 0 | 0 | 21 | 0 | 0 |
| qwen2.5-coder 7B | 12 | 0 | 0 | 4 | 0 | 0 |
I set three tiers before running anything: works reliably is 26 or more out of 30, works sometimes is 12 to 25, and fails is under 12. Thirty runs is enough to tell those apart, but not to rank agents inside a tier - 28 out of 30 is consistent with anything from 79% to 98%. So a 30 next to a 28 is a tie, not a win.
Read that way, Polyglot is the only agent that doesn't fail on any of the seven models. Its weakest cells are qwen2.5-coder 32B and 7B, both "works sometimes" - and the 7B sits right on the line at 12.
The newest models don't separate the agents
On Qwen3.8-27B every agent works reliably. On qwen3-coder four of the six do, on Devstral three, and the rest still work sometimes. These models are good at emitting tool calls through the native channel most agents listen on, so the agents mostly differ in how they plan and edit, not in whether they work at all.
If you run one of these models, pick your agent on other grounds.
Where it breaks
It breaks on models that write their tool calls as plain text instead of through the native channel. All three qwen2.5-coder sizes do this through Ollama. On those, pi, Hermes and opencode made no tool call in any of the 30 runs: the model writes a perfectly reasonable {"name": "edit_file", ...} into its reply, the agent treats the whole reply as a final answer, and the task ends with nothing done. goose makes some calls but completes no task.
goose's toolshim is the honest exception. It routes the model's text through a second, small model whose job is to turn it into a proper tool call, and it works: on qwen2.5-coder 32B it scored 28, ahead of Polyglot's 24. The same second model hurts on gpt-oss, where it dropped goose from 18 to 5. Polyglot does the same job with a parser instead of a second model, which is why it holds up on both.
gpt-oss is the other model that separates the field. It's strong, but it emits tool calls in its own trained format, and goose, Hermes and opencode land in "works sometimes" on it while Polyglot and pi stay reliable.
Smaller prompts, too
I also measured what each agent sends to the model, from Ollama's own request log so every agent is counted the same way. Polyglot opens each task with about 1.2k tokens of fixed prompt and pi with 1.2-1.6k. Hermes starts at 3.9-4.7k, and goose and opencode at 5.5-7.6k. Across a whole task, Polyglot used several times fewer prompt tokens per completed task than goose, Hermes and opencode. pi reuses Ollama's prompt cache better than Polyglot does - something for me to fix, not to advertise.
What I got wrong the first time
My earlier numbers for goose, Hermes and opencode were wrong, and not in their favour. Ollama defaults to a 4,096-token context and silently cuts anything longer. Those three agents send several thousand tokens of instructions, so on my first runs they were working from a prompt with most of their instructions missing.
Every model in this grid now runs with a 32k context, and the harness refuses to start below 16k. After each agent's runs, it scans Ollama's log for truncated prompts and marks the results invalid if it finds any. If you benchmark agents on Ollama yourself, check this first.
Where Polyglot lost runs
I read every one of Polyglot's failed transcripts. Outside the 7B, it missed 13 of 150 runs. In up to 5 of those, my parser was at fault: a model stuttering an empty tag before the real one, or running past a broken closing tag into an invented next turn. Those are fixed in Polyglot 0.13.2, along with every other format I found in the reruns. The rest were the model's own mistakes - a wrong edit, a misread file, an answer it made up.
The Polyglot column above is a pre-release build of 0.13.2, the first of three I ran. The two later builds scored 173 and 172 out of 210 against this one's 179 - the same tiers, with the spread coming from the models, not the code. gpt-oss scored 30, 27 and 25 across the three builds without a single parsing error among its misses. That's what run-to-run noise looks like at 30 runs, and it's why I report tiers.
From here on, Polyglot's column comes from the current npm release, run once, whatever it scores. No rerunning a build to replace a number I don't like, and every run stays in the repository.
Methodology
- Tasks. Six small coding jobs: add a CLI subcommand, fix an off-by-one, answer a question from a file without editing it, remove dead code, rename a function across two files, and trace a runtime error to its cause. Each runs in a fresh directory.
- Models. All local, through Ollama, on one RTX 5080 16GB, each with a 32k context: Qwen3.8-27B, gpt-oss 20B, Devstral Small 2 24B, qwen3-coder 30B, and qwen2.5-coder 7B, 14B and 32B.
- Agents. Polyglot, pi 0.85.1, goose 1.52.0 (with and without its toolshim, using llama3.2 3B as the interpreter), Hermes Agent 0.19.0 and opencode 1.18.32, each run headless with tools auto-approved and an isolated config.
- Scoring. The same automated check for every agent: the files must end up right and a verify command must pass. Five trials of each task, so 30 runs per cell, 600 seconds per run.
The README has the commands to reproduce it, and results/GRID.md links every cell to its raw file. If you think I got something wrong, open an issue on the run in question - I'd rather fix the benchmark than defend it.
Try it
Polyglot is free and open source: npm install -g @usepolyglot/cli, point it at whatever's already running on your GPU, and see whether your agent is doing the work or just describing it. Repo: https://github.com/giuseppe-sirigu/polyglot
Top comments (2)
Thirty runs is the right call. Tiny N makes every agent look like a coin flip. Love that you opened the harness up too. Local models still win or lose on tool calling reliability more than raw chat scores and your setup shows that cleanly.
Thanks, Brian. That was the lesson for me too: with only three or five runs, a lucky or unlucky streak moved the numbers more than any real change did. It's also why I set a rule for Polyglot's numbers: they always come from the current npm release, run once, whatever it scores, so we can't pick the best run.
If you test the harness and point it at a model I didn't cover, I'd love to see what you get, and I'm happy to add it to the results.
What do you use for your own coding: Claude Code, Cursor, aider, something local?