DEV Community

Cover image for A local model is good enough for most of my tooling
Ahmet Zeybek
Ahmet Zeybek

Posted on Originally published at zeybek.dev

A local model is good enough for most of my tooling

In February I wrote down everything in my day that called a hosted model, and the list was longer than I expected. The coding agent was on it. So was a git hook that drafts the commit message, a script that summarises a pull request for the changelog, a tool that reads a stack trace and guesses which file to open, a shell function that turns a sentence into a jq expression, a test runner plugin that names a failing test's likely cause, a thing that rewrites my Slack drafts into shorter Slack drafts, and a small classifier that sorts incoming GitHub notifications into "read now" and "read later".

That came to nine tools. Eight of them sent code, logs or messages to an API for jobs a strong model finishes in under two seconds, and a weak model would finish in under two seconds too. Only the coding agent needed the frontier.

So I moved the eight to a model running on the laptop. Three months later they're still there, and I haven't missed the API for any of them.

The category

The tasks that moved all look alike. The input is small, a few hundred to a few thousand tokens. The output is small and constrained: a commit message, a category, a one paragraph summary, a file path. There's a right answer a reasonable engineer would agree on, or a narrow range of acceptable ones. And a wrong answer costs little, because I see the output straight away and can throw it out.

For tasks like that, I can't see a difference between a frontier model and a good 30 billion parameter open-weight model. I did check. I ran 200 commit diffs through both and had two colleagues rank the messages blind, and they picked the hosted model's message 52 percent of the time, which is a coin flip. On the stack trace to file task I used 100 traces from our error tracker. Both models found the right file 94 percent of the time, and they disagreed with each other on four.

The tasks that stayed look the opposite way: large input, open ended output, many steps, and a wrong answer that costs a lot. The coding agent doing a refactor across twenty files, say, or anything where the model has to plan. The frontier is still clearly better there and I won't pretend it isn't.

The setup

It's a MacBook Pro, M4 Max, 64 GB. The model is a Qwen 3 variant of around 30B parameters in a 4 bit quantisation. It takes about 18 GB of memory and generates roughly 40 tokens a second on this machine. I run it behind a local server that speaks the OpenAI style chat API, because all eight tools already spoke that, so switching meant changing a base URL and a model name.

# ~/.config/tooling/env
LLM_BASE_URL=http://127.0.0.1:11434/v1
LLM_MODEL=qwen3-30b-a3b
LLM_API_KEY=local
Enter fullscreen mode Exit fullscreen mode

For six of the eight tools, that file was the whole migration.1

The server starts at login and sits at about 2 GB until the first request, when it maps the weights. The first request after idle takes about four seconds, and after that a typical commit message comes back in under two. The fan doesn't come on. Over a working day the battery cost is noticeable without being dramatic, maybe an extra ten percent, and on mains I don't care.

One thing about the model itself: with a mixture-of-experts layout, a model this size only activates about 3B parameters per token. That's why it's fast on a laptop with 30B sitting in memory, and that family of models is why this got practical this year.2

The commit message hook, since people ask

#!/usr/bin/env bash
# .git/hooks/prepare-commit-msg
set -euo pipefail
[[ "${2:-}" == "merge" || "${2:-}" == "squash" ]] && exit 0
diff=$(git diff --cached --no-color | head -c 12000)
[[ -z "$diff" ]] && exit 0

msg=$(curl -s "$LLM_BASE_URL/chat/completions" \
  -H "content-type: application/json" \
  -d "$(jq -n --arg d "$diff" '{
    model: env.LLM_MODEL, temperature: 0.2, max_tokens: 120,
    messages: [
      {role:"system", content:"Write a git commit subject line under 72 characters, imperative mood, conventional commits prefix, no trailing period. Output only the line."},
      {role:"user", content:$d}
    ]}')" | jq -r '.choices[0].message.content' | head -1)

# Put the suggestion above whatever git already put in the file.
{ echo "$msg"; echo; cat "$1"; } > "$1.tmp" && mv "$1.tmp" "$1"
Enter fullscreen mode Exit fullscreen mode

The suggestion shows up at the top of the editor and I either take it or rewrite it. About 70 percent go in unchanged. The 30 percent I rewrite are mostly diffs that don't say why the change was made, and no model can guess that from a diff.

Why not cost

People assume I did it for the API bill. For eight tools making a few hundred calls a day, that bill was around 20 dollars a month.3 So cost doesn't explain it.

What does is that these tools see everything. The commit hook sees every diff before it's pushed, including the ones on branches that never will be. The log triage tool sees production stack traces with customer identifiers in them. The Slack rewriter sees drafts I decided not to send, and the notification classifier sees the titles of private repositories.

If the local model does the task just as well, the data has no reason to leave the machine, and I think no reason should win.4 It also ended a compliance conversation with one client. The log triage tool was the only thing on my laptop sending their production data anywhere, and now it sends it nowhere.

Then there's latency and availability. The hook works on a train with no signal. The API had a bad afternoon in April and all eight tools fell over at once, which is when I noticed how many there were.

Where it fell short

Two of the eight tasks needed prompt changes to do as well locally. The PR summariser ran long with the local model in a way the hosted one didn't, and one sentence in the system prompt, "three sentences maximum", fixed it. The jq generator got the syntax right less often, about 85 percent against 96. I added three examples to the prompt and it went up to 93. Small models need examples more than big ones do, and that was all the adjusting I had to do.

Structured output needs some care too. The hosted APIs guarantee valid JSON if you ask for it. The local server does as well, through grammar constrained decoding, but you have to turn it on per request. Before I did, the classifier now and then returned a category with an explanation tacked on the end, which broke the parser. It took one flag.

And there's a ceiling. I tried moving the coding agent's simple mode, the one I use for "rename this and fix the imports", to the local model. The rename worked. As soon as the task meant looking at more than five files it got lost. That's where the line is, and for the 30B class on a laptop it isn't close to moving yet.

The team version

What works on one laptop doesn't carry over to a team on its own. Three colleagues asked for the setup within a month, and we ended up with something a bit different from mine. The differences are worth writing down.

Not everyone has 64 GB. Two people have 16 GB machines, where an 18 GB model doesn't fit. They use a smaller model from the same family, around 8B parameters, which fits in 5 GB and runs the same eight tasks. I ran the same blind comparison on commit messages, and the 8B model lost to the hosted one 61 to 39. That's a real gap. It still wins four times in ten, though, and when it loses the message is still usable. On the classifier I couldn't tell them apart. On the jq generator it was noticeably worse, so that person kept the hosted API for that one tool.

The tools moved into a shared repository along with the prompts, the base URL config and an install script, so everyone runs the same version and a better prompt reaches everyone. Within a week the prompts had already started to drift between machines.

We also added a small shared eval: 50 inputs per tool, each with a known good output, run on your own machine whenever you change the model or a prompt.5 It caught a model update that changed every commit message from "feat: ..." to "feat(scope): ...", which would have annoyed everyone for a week before someone worked out why.

What it costs to keep running

Local models aren't free in the way people imagine. Here's what they actually cost me.

Model updates are your job. A hosted API gets better under you with no change on your side. The local model stays the version you downloaded until you download another. The family I use has shipped three updates since February, every one of them worth taking, and each took an hour to evaluate and roll out. The shared eval is why that's an hour and not a day.

Memory is shared with everything else. On the 64 GB machine that doesn't matter. On the 16 GB ones, running the model next to a browser, an editor and a container or two means something gets swapped out, and now and then it's the model, which turns a two second commit message into a fifteen second one. The people on those machines start the server when they need it instead of at login.

There's no rate limit either. That sounds like a plus, and it also lets you write a tool that hammers the model in a loop and pins the CPU for a minute. The notification classifier did exactly that on its first day: it processed 400 notifications one at a time on startup. A batch endpoint and a small concurrency limit fixed it.

None of this changes my mind. Compared with the hosted API it costs less money and more attention, and for tools that see everything I type, I'll take that.

What I would tell someone

Write the list. You've probably got more of these than you think, and they probably all point at an API because that was the easy path when you wrote them.

Sort it by input size and by what a wrong answer costs. Anything small that's cheap to get wrong is a candidate.

Put the local model behind the API shape your tools already use and change the base URL. Don't rewrite anything.

Check the quality with a blind comparison on a hundred real inputs. Your gut feeling about which model is better is worth less than you think.

Keep the frontier model for the agent and use the laptop for the rest. Most of what I ask a model to do all day is small, and small jobs can stay in the room.


Originally published at zeybek.dev.


  1. The other two had the provider's SDK hard coded and needed a ten line change to use a generic client. ↩

  2. Two years ago a local model was either small and weak or large and slow. ↩

  3. The laptop loses more than that to depreciation every month. ↩

  4. I trust the hosted providers with data more than I trust most companies, and that still isn't my reason. ↩

  5. It's the same idea as the eval sets I use for production features, only much smaller. ↩

Top comments (0)