DEV Community

Todd Sullivan
Todd Sullivan

Posted on

The 8B Local Model Was Worse Than the 4B One

I tried the obvious local-LLM upgrade on a small Mac mini: move from a 4B model to an 8B model.

It was worse.

The machine is an 8GB Apple Silicon mini running a personal assistant stack. The useful path was already working with mlx-community/Qwen3-4B-4bit through Swama's OpenAI-compatible endpoint:

4B weights: ~2.3GB
warm plain reply: ~1s
warm tool-call round: ~3-4s model time
full tool response: ~7-18s, depending on external APIs
Enter fullscreen mode Exit fullscreen mode

So I pulled the 8B quantized model to see if the extra capability was worth it.

The first trivial request was enough:

8B weights: ~4.3GB
"say hello": 14s
throughput: 6.8 tok/s
swap churn: ~1.6GB in/out
free memory: basically zero
Enter fullscreen mode Exit fullscreen mode

That was not an edge case. That was the cheapest possible prompt.

A real assistant request is heavier: system prompt, recent conversation, tool schemas, maybe a tool round, then a final answer. If the 8B model swaps on hello, it is not going to survive a weather/tool/home-status request while the rest of the app is running.

The annoying part is that the model choice looked sensible on paper. 8B should be smarter. It probably is, in isolation. But local AI work is not model leaderboard work. It is memory pressure, cold starts, prompt prefill, tool reliability, and whether the box is still usable after the request.

The fix was not glamorous:

MAX_HISTORY_MESSAGES = 10

@lru_cache(maxsize=1)
def get_ai_handler():
    return SwamaAIHandler()
Enter fullscreen mode Exit fullscreen mode

Cap the prompt history. Reuse the handler. Keep long-term memory as distilled facts, not raw transcript replay. Use the 4B model that stays resident and answers quickly enough.

I also had to add boring production guardrails:

  • run the real Swama binary, not the symlink, because first inference crashed loading MLX's metallib
  • keep the model warm because idle eviction made cold starts painful
  • strip <think> blocks from streamed responses because /no_think was not enough after tool calls
  • bias the prompt toward tool use for live data, because small models will confidently answer from stale weights if you let them

The lesson was simple: on tiny local hardware, the best model is the one that leaves enough machine behind for the product.

The 4B model shipped. The 8B model got deleted.

Top comments (0)