DEV Community

Cover image for Local LLM vs Claude Code: 96% of My Requests Failed Locally. Half the Steps Didn't.
Ken Imoto
Ken Imoto

Posted on

Local LLM vs Claude Code: 96% of My Requests Failed Locally. Half the Steps Didn't.

The pitch for running Claude Code on a local model goes like this: point it at Ollama, stop paying per token, keep your code on your machine. I have an RTX 4070 with qwen3.5:4b on it, so I wanted that to be true.

Before swapping anything, I counted. I took 100 requests I had actually sent to Claude Code and asked, for each one, whether a local model could have done it end to end.

96 could not. So much for cancelling the subscription.

Then I counted a different way, and the answer flipped. This post is about the second count and the small config file that came out of it.

Counting whole requests: 96 of 100 go to the frontier

A request to Claude Code is usually one line. "Fix the failing test in the billing module." "Why is this deploy slow?" The request itself contains very little context. The work happens after that, when the agent reads files, runs commands, fetches docs, and decides what to try next.

So when you judge the request as a whole, you are judging the hardest part of it: the planning. That part needs the frontier model almost every time. In my set, 96 of the 100 had to go there.

If you route at the request level, a local model gets almost nothing. Swapping the whole agent onto a local model asks it to do exactly that planning.

Counting single steps: 97 of 200 run fine locally

The agent's work is a chain of steps: tool calls, fetches, summaries, commit messages. I broke Claude Code sessions into 200 of those steps and asked the same question for each one.

97 of 200 were fine on a local model.

Dot chart: 4 of 100 whole requests could run on a local LLM, versus 97 of 200 single steps

The clearest case was WebFetch: fetch a page, pull out the part that answers a question. That is extraction, not planning. I tried it on 20 real fetches, and it was not perfect: some pages never loaded, and some answers came back partly wrong. That is why the example config sends a failed fetch to the frontier (on_fetch_error: frontier). But when the page arrived, the local model usually got the extraction right.

Claude Code already does a version of this. Its own WebFetch tool description says the fetched page is processed with a small, fast model. That makes the tool itself an example of choosing a model per step.

One rule beat a 9B router

My first router asked a local model for a probability: "should this step go local?" It worked, sort of. Then I wrote a one-line rule, "WebFetch goes local", and the rule beat the 9B model's probability.

That changed the design. Rules decide what rules can decide. When enabled, the probability judge only sees what is left. Here is the rules block from the companion repo's example config:

rules:                      # step rules, top to bottom, first match wins
  - match: {tool: WebFetch}
    route: local
  - match: {tool: [Agent, Task]}
    route: frontier
  - match: {kind: commit_message}
    route: frontier         # frontier until a probability judge is added

probability: null           # off until measured with the same judge model
default_route: frontier
Enter fullscreen mode Exit fullscreen mode

The probability judge is off by default. The routing function is short:

def route(self, step: dict) -> Decision:
    for r in self.rules:
        if r.matches(step):
            return Decision(r.route, "rule")
    if self.probability and self.local_threshold is not None:
        p = self.probability(step)
        return Decision("local" if p >= self.local_threshold else "frontier", "probability", p)
    return Decision(self.default, "default")
Enter fullscreen mode Exit fullscreen mode

In my config, sub-agents (Agent, Task) are pinned to the frontier.

Routing table checked top to bottom: WebFetch goes local, Agent and Task go to the frontier, commit messages go to the frontier, the probability judge is skipped, anything else defaults to the frontier

Getting a probability out of Ollama

For the steps rules can't settle, you want a number, not a paragraph. The trick is to make the model answer with one letter (Y or N) and read that letter's probability from logprobs.

This is the request body I send to Ollama's native /api/chat:

{
  "model": "qwen3.5:4b",
  "stream": false,
  "think": false,
  "logprobs": true,
  "top_logprobs": 20,
  "options": {"temperature": 0, "num_predict": 1},
  "messages": [{"role": "user", "content": "... Answer with a single letter, Y or N."}]
}
Enter fullscreen mode Exit fullscreen mode

Three things fail silently here, with no error:

  1. The OpenAI-compatible endpoint drops the probabilities. Use the native API. The repo's scripts expect Ollama 0.12.11 or later.
  2. Thinking models don't answer in the first token. The first token is the start of their reasoning, so the Y/N you are looking for isn't there. think: false and num_predict: 1 pin it to the answer.
  3. A candidate outside the top 20 looks like probability 0. Also, " Y" and "Y" are different tokens.

The fix for the third one is to add up the probability that landed on your candidate letters and refuse to trust the judgment when that total is small:

raw = {lab: 0.0 for lab in labels}
for t in response["logprobs"][0]["top_logprobs"]:
    tok = t["token"].strip()  # " Y" and "Y" are the same answer
    if tok in raw:
        raw[tok] += math.exp(t["logprob"])
mass = sum(raw.values())      # below 0.5: the answer wasn't where we looked
Enter fullscreen mode Exit fullscreen mode

0.5 is a line nobody chose

I used the same kind of judge as a gate in front of Claude Code: "does this request contain a secret?" It read a note with a production database password in it and said it was fine, with 69% confidence. The password was there with 100% confidence.

The ranking was fine. Secrets scored higher than non-secrets. The scale was off: real secrets were landing well below 0.5. Cut at 0.5, the gate missed 20 secrets. When I drew the line from my own labelled examples instead, it missed 2.

A local secrets gate cut at 0.5 missed 20 secrets; with a line drawn from my own labelled examples it missed 2

That is why the gate threshold in the example config looks strange:

judge:
  model: qwen3.5:4b
  threshold: 0.0071    # example only. lowest P among your confidential examples x 0.3
  at_or_above: block
on_timeout: block
Enter fullscreen mode Exit fullscreen mode

0.0071 is only an example from my setup, so don't copy it. The repo has a small CLI (steprouter-threshold) that takes your labelled judgments and prints a line, AUROC, and 5-fold cross-validated misses and false alarms.

A regex pass (gitleaks plus personal-data patterns) runs before the model. It catches secrets that have a recognizable shape. The model is there for the ones that don't.

The free local worker cost the most

One more result surprised me. I set up a strong model to plan and a free local model to do the work. That setup was the most expensive one I tested. Free, the way a puppy is free.

The local model's tokens were free. The orchestrator re-reading everything the worker sent back was not. What comes back from a routed step is part of the routing, so the example config pins the return value to a small typed shape:

return:
  fields: [green, mypy, ruff, pytest_failed, files, summary]
  summary_max_lines: 2
  full_log: file
Enter fullscreen mode Exit fullscreen mode

Pass/fail counts, the files touched, and one or two lines of what was done. The full log is written to a file instead of the reply.

Try it on your own logs

My numbers come from my work. To get yours:

  1. Take a day of Claude Code sessions and list the steps, not the requests.
  2. Mark which step types always go one way. Those become rules.
  3. Leave the probability judge off until you have measured it on your own examples.
  4. Before you add a local worker, decide what it is allowed to send back.

The companion code, including the gate, the logprobs judge, the threshold CLI, the router, and a made-up evaluation set, is on GitHub as local-step-router (MIT). Clone it, run pip install -e ., and start by editing examples/routing.yaml. The README has an English summary at the bottom.


The full measurement (550 of my own requests and steps on one RTX 4070, how big the judge model needs to be from 0.8B to 35B, why none of my six real secrets matched the evaluation set, and the final blueprint) is written up in Local LLM or Claude Code?. Chapters 2 and 3 cover the step count and the rules-first router; chapter 8 is the threshold worksheet.

Top comments (0)