DEV Community

Eryk Kubiak
Eryk Kubiak

Posted on

Local Coding Agent: Weak Model? No, Bad Plumbing

In this video:

0:00 Why agents fail around turn ten
0:19 It runs but is useless
1:54 The base-URL swap
3:26 Where compatibility breaks
5:28 From tokens to a tool call
8:31 Context, speed and turn count
12:25 Local vs hosted trade-offs
14:11 Harness, server and model pairings

A local coding agent that falls apart by turn ten is usually blamed on a weak model. Usually it's plumbing. Under 24 GiB of VRAM, Ollama defaults to a 4k context and silently drops history. Three budgets decide whether the agent survives.

An agent that runs but is useless

A fully local coding agent has four parts. The first is an inference server that exposes the API the harness speaks: chat completions, Responses, or Anthropic Messages. The second is a model that emits tool calls. The third is a harness pointed at that endpoint. The fourth is three tools: read a file, write a file, run a command.

Wire those together and the agent starts, accepts a prompt, and answers. By every status check, it works. Then you give it a real task.

A coding task is not one request. It is tens of turns: read the file, edit it, run the tests, read the error, edit again. Any small defect per turn gets multiplied by the turn count. A tool call may carry a malformed argument. A file may silently fall out of the context window.

Take a two percent chance of a bad call on each turn. Over thirty turns, that is about a forty-five percent chance the run hits at least one.

Speed compounds the same way. On a cache miss, the server re-reads about fifteen thousand tokens at roughly five hundred a second. That is thirty seconds, and across forty turns, twenty minutes of staring at a spinner. That is a worst case, and it varies by hardware.

Three budgets that land on you

A hosted model mostly hides these problems behind per-token pricing. The provider handles serving, parsing and throughput. Run the same loop locally and all three land on you: tool-calling reliability, context length, and throughput.

These are the constraints that separate a workable agent from one that only technically works. When one of them fails, nothing crashes. The agent just stops being useful. Very often the cause is the plumbing between the model and the harness, such as a small context default or the wrong tool-call parser, before it is the model's raw ability.

The obvious setup: swap the base URL

The first move is almost always the same, and it is a sensible one. Start an inference server: llama-server, Ollama, vLLM or LM Studio, whichever is already on the machine. Pull a well-regarded chat model, one that ranks high and feels sharp in conversation. Then change a single value in the harness settings, the base URL, so it points at localhost instead of a hosted provider.

This looks like it should work because every one of those servers advertises an OpenAI-compatible endpoint, and the label reads like a single contract. The request shape is the same and so is the response shape. So any model that chats well should be able to drive the loop, and the harness should not care who answers.

The check seems to agree. You send a message, get a fluent reply, maybe a clear explanation of a function. Everything is green.

But look at what that check exercised: one short prompt and one plain-text answer. It involved no structured output, no long prompt and no repeated turns. Those are exactly the three things the loop demands, and a chat check touches none of them.

What the harness can and cannot repair

The harness is the loop around the model: a system prompt, a set of tool definitions, and some context management. That is all. It sends text in and reads text back.

Suppose the server fails to parse the model's tool call. Or it quietly trims the prompt. Or the model was never trained on the call format. In each case the harness receives the broken result and has nothing to repair it with. So the swap succeeds, and proves almost nothing.

Where the swap actually breaks

The label OpenAI-compatible used to mean one thing: the chat completions endpoint. It no longer does. Codex CLI removed the chat completions wire API for custom providers. A local server therefore has to implement the responses endpoint, and a config still set to chat refuses to start.

Claude Code speaks the Anthropic Messages format instead. Ollama added that in version 0.14.0, in January 2026, and llama-server serves it at /v1/messages. The label is the same, but there are three different wire formats. A mismatch here is the loud failure, and the easy one. The quiet ones are worse.

Tool templates and parsers

llama-server does tool use only with the jinja flag on. Without it, the model's tool template is never applied, and what should be a call comes back as ordinary text. Some models need a template file override on top of that.

vLLM needs enable auto tool choice, plus the tool call parser that matches the model family. Pick the wrong parser and the model can emit a perfectly good call that the server cannot read. The harness sees prose where a tool call should be, and the verdict is: this model can't call tools. The weights were fine.

Context defaults

Ollama sets its default context by VRAM: 4k under 24 GiB, 32k from 24 to 48, and 256k at 48 or more. Its own docs say agents need at least 64000.

A system prompt, a set of tool definitions and two source files can fill four thousand tokens quickly. Aider's docs warn that whatever overflows is silently discarded. There is no error and no warning. The agent simply forgets the file it read three turns earlier, and edits it from a guess. That looks like a model with a bad memory.

No grammar behind auto tool choice

Even with the right flags, vLLM's auto tool choice has no grammar behind it. Calls are extracted from raw text. The docs say arguments may occasionally be malformed or violate the schema, and that malformed markup can leak into the response. Most clients never set strict, so nothing stops it.

"Occasionally, over thirty turns" is the arithmetic from the opening. Every one of these failures reads from the outside as bad reasoning. None of them is.

From tokens to a tool call

The model never emits an API object. It emits text. Depending on the model family, that text is JSON inside special tokens, or XML-style tags, or Harmony markup with channels and recipients.

The tool_calls field your harness reads is built afterwards, by the server, in two steps. First, the chat template renders your tool definitions into the prompt in the format the model was trained on. Then a parser reads the generated text and tries to find a call in it.

In llama.cpp, a native handler covers templates it recognises. Anything else falls to a generic fallback, which is a best guess at the format. In vLLM, you choose the parser yourself. Either way, the call is only as real as the parser's ability to find it.

How a missed parse ends the loop

When the parser misses, nothing dramatic happens. The text is not a call, so it becomes the assistant's final answer, and the loop ends.

One LangChain forum user reported exactly this with gpt-oss-120b and the Harmony format. After about five tool calls, the model put a raw call on the analysis channel instead of the commentary channel. The parser ignored it, the harness treated it as the answer, and the run stopped. Earlier calls in the same run had parsed fine. That is a user report, not an official finding. But it shows the shape of the failure: a text-level slip, with no error anywhere.

The fix is to stop relying on the model to get the markup right. Constrained decoding forces the output through a grammar, so the markup must be valid. vLLM supports this with strict tools and an operator flag called tool-strict-level. The catch is that most clients never set strict on their tools. So the server operator has to raise the floor for every request. If you run the server, that is your job.

Quantization: parsed, but with worse values

The common story is that four-bit breaks JSON. The evidence is more nuanced. One small BFCL test of tiny Qwen3 models found schema validity almost unchanged at Q4_K_M: 0.877 at full precision, 0.873 quantized. But argument correctness fell, from 0.605 to 0.575. The call still parsed. It just carried a worse value.

That is one author, with very small models, so hold it loosely. Community long-context reports point the same way. Quantized weights or KV cache can silently shift which tool gets picked, or what arguments it gets, while the prose stays fluent. Fluency tells you nothing here.

Per-call error rates compound

This is arithmetic, not a measurement. Say each call has a two percent chance of going wrong. 0.98 to the twentieth power is about 0.67. Over twenty calls, the run survives only about two-thirds of the time. Small per-call rates are not small per run.

The parser is part of that rate. Unsloth's log for Qwen3-Coder-Next shows a llama.cpp bug fix on February 4 for looping and output problems. It shows another on February 19, after which tool calling should be even better once parsing was fixed. The weights were the same, but a different runtime gave different reliability.

A tool call that parses and carries the right arguments is one budget. The second fails with even less noise: whether the model still has the file it read three turns ago.

The loop budget: context grows every turn

A model does not remember anything between turns. Every turn, the harness sends the entire history again: the system prompt, the tool definitions, the files being edited, every tool output, every earlier reply. So the prompt only grows. By turn twenty it holds everything the agent has read and done, and the KV cache for that context grows with it.

That cache lives in memory next to the weights. On a long session it often hits the memory limit before the weights ever do. A 256K window on a model card is a ceiling, not a promise.

Prefill, decode and the prefix cache

The time cost has two separate parts. Prefill is the model reading the prompt. Decode is the model writing the answer. They run at very different speeds.

One measurement of a Radeon AI PRO R9700 running Qwen3.6-35B-A3B at Q4 under llama.cpp, at 32k tokens of depth, gave 2,113 tokens per second of prefill and 111 of decode. That depth matters. The headline pp512 figure, a short prompt, flatters every machine, because both numbers slow as context grows.

Prefix caching is what keeps the loop affordable. The server keeps the computed state for the unchanged start of the prompt. A normal turn then only prefills the new tokens: the last tool output and the model's last reply.

A cache miss breaks that, and the server must re-prefill everything. One first-hand report from a Mac Studio, a blog post and not a controlled benchmark, measured about 500 tokens per second of prefill. At that rate a 15,000-token context is a 30-second pause before the first new token. Anything that changes the early part of the prompt can cause it.

Turn-time arithmetic, and what thinking does to it

The number that decides usability is turns multiplied by prefill plus decode time. Here is an illustration, and it is arithmetic, not a measurement.

Take the R9700 figures and forty turns. Those rates were taken at 32k depth, so they are pessimistic for the early, short-context turns. Each turn prefills 2,000 new tokens, about one second, and decodes 500 tokens, about four and a half. That is under four minutes for the whole task.

Then let the model think. Some models overthink by default: Simon Willison wrote in August 2026 that Qwen3.8-27B is excellent but defaults to heavy overthinking. Suppose the same Qwen3.6-35B-A3B setup did that, adding 3,000 thinking tokens per turn. Decode alone is about thirty-one seconds. Forty turns then take roughly twenty-two minutes. The hardware, the model and the task are all the same.

A dense model like that 27B would likely decode slower on this card, so its real figure would be worse. Decode is the term that thinking inflates, and it is the one you multiply the most.

How model shape changes the budgets

Hardware and model shape change these terms in different ways. Mixture-of-experts models decode fast because only the active parameters are read for each token. Qwen3-Coder-Next has 80 billion parameters in total but only 3 billion active.

The catch is memory. Every expert still has to be loaded, so the whole 80 billion must fit, and Unsloth lists about 46 gigabytes for a 4-bit build. Fast decode does not make the model small.

The other lever is attention. Only 12 of that model's 48 layers use full attention, each with 2 KV heads. The rest use linear attention. That keeps the KV cache unusually small for a 256K window. The context budget therefore stretches much further than the parameter count suggests.

Put the pieces together. Context sets how much the agent remembers and how much memory the cache eats. Prefill sets the cost of a long prompt, and the cache decides whether you pay it once or every turn. Decode sets the cost of everything the model says. All three get multiplied by the turn count.

These numbers come from specific machines, and a loop only has to fit on one of them.

What local costs you

With a hosted model, the main constraint you manage is cost per token. Context limits, serving speed and most tool-call parsing are handled by the provider, though hosted models have limits and failures of their own.

Going local swaps that bill for three others. The first is hardware, with memory sized for the weights plus the KV cache, which grows with every turn. The second is tuning time: the flags, the parser, the context setting, all the plumbing described above. The third is slower turns, because your machine prefills and decodes at its own pace, not a data center's.

When to use a hosted model instead

Local is the wrong choice when the work is long-horizon. Vendors and StorageReview both say that long-running agentic work lags frontier cloud models. A job that needs fifty turns across a large repository stacks every budget at once.

Context is the clearest case. Raising it raises memory use, and it slows both prefill and decode. Cline's docs say to keep tasks focused and to start a new task when context grows. That is advice for a local setup, and it also describes the work local suits: focused, short-horizon tasks.

Local also wins where hosted cannot go. That includes privacy-bound code that must not leave the machine, and offline use. There, the slower turns are the price of being allowed to do the task at all.

Leaderboards don't answer the question

The model landscape moves monthly. As of October 2026, Qwen3.8-27B, open weights released on August 14, 2026, is the newest Qwen open-weight release. Qwen 4 27B was only named at Apsara on September 22, with no specs and no date.

A leaderboard position cannot tell you whether a model's tool calls parse on your server, or whether its cache fits beside its weights. The three budgets can.

Real setups: harness, server and model that fit together

A working local setup starts with the harness, because the harness decides which wire format the server must speak. Claude Code is pointed at Ollama or llama-server through its Anthropic base URL setting, and both serve the messages endpoint. Codex CLI needs a backend that implements the responses endpoint. GitHub Copilot CLI has run against Ollama, vLLM and Foundry Local since April 2026, and its offline mode needs a model with a 128K context window.

Harness design also changes how things fail. Cline offers a compact prompt built for local models, which spends less of a small context on instructions. Aider goes further. It parses diff or whole-file edits out of plain text, so there is no JSON tool call to mangle. That sidesteps JSON tool-call parsing, but swaps it for edit-format failures, and Aider's own docs say quantized local models are more prone to those.

Then pick the model by memory. VDF's September 2026 shortlist names Devstral Small 2 or Qwen3.8-27B for a 32 GB GPU, and Qwen3-Coder-Next or gpt-oss-120b for 128 GB. OpenHands' docs, in a May 2026 note, recommend Qwen3.6-35B-A3B as the first local model to try.

A four-step check before the first task

Before the first real task, verify in a fixed order. One: the API surface matches what the harness speaks. Two: the jinja flag or the parser flag is set. Three: context is at least 32k to 64k. Four: measure turn time at realistic depth, not on a short prompt.

If all four hold and tool calls still fail constantly, OpenHands says the model may be the problem, not your setup.

That order is the thesis in practice. Tool calls, context and throughput are three budgets, and most of the time the leak is in the plumbing, so check the plumbing before you blame the model.

Sources & further reading — Local Coding Agent: Weak Model? No, Bad Plumbing

  • Tool Calling - vLLM — https://docs.vllm.ai/en/latest/features/tool_calling/ vLLM needs auto tool choice plus the parser matching the model family. Auto tool choice has no grammar behind it, so arguments can be malformed and markup can leak into the response. Strict tools and the operator-level strict setting force valid output.
  • llama.cpp/docs/function-calling.md at master — https://github.com/ggml-org/llama.cpp/blob/master/docs/function-calling.md llama-server does tool use only with the jinja flag. Native handlers cover recognised templates and other templates fall back to a generic format. Some models need a template override.
  • Context length - Ollama — https://docs.ollama.com/context-length Ollama's default context depends on VRAM (4k under 24 GiB, 32k for 24 to 48, 256k at 48 or more), and its docs say agents need at least 64000.
  • Deprecating chat/completions support in Codex · openai/codex · Discussion #7782 (2025) — https://github.com/openai/codex/discussions/7782 Codex CLI dropped the chat completions wire API for custom providers, so a local server must implement the Responses endpoint.
  • Ollama (aider docs) — https://aider.chat/docs/llms/ollama.html A small Ollama context window silently discards whatever overflows, with no error, so the agent forgets files it already read.
  • Qwen3-Coder-Next: Pushing Small Hybrid Models (2026) — https://qwen.ai/blog?id=qwen3-coder-next Qwen3-Coder-Next has 80B total and 3B active parameters, and only 12 of its 48 layers use full attention. This keeps the KV cache small for a 256K window.
  • Qwen3-Coder-Next: How to Run Locally (2026) — https://unsloth.ai/docs/models/qwen3-coder-next A 4-bit build needs about 46 GB of memory. The llama.cpp fixes of February 4 and February 19 for looping and tool-call parsing show that the same weights behave differently on different runtimes.
  • Run Local LLMs with OpenHands - OpenHands Docs (2026) — https://docs.openhands.dev/openhands/usage/llms/local-llms OpenHands recommends Qwen3.6-35B-A3B as the first local model to try, and says that if tool calls still fail constantly the model may be the problem.

Disclosure: this article is the written companion to the video above. Its script was drafted with AI assistance and the narration is an AI voice.

Top comments (0)