I built a local-LLM coding IDE, and I noticed something odd: giving
the model all the CPU cores made the app feel slower. Tokens
were arriving at the same speed, but the UI lagged behind. So I
measured it — and then, mid-verification, a new Ollama release
changed the whole story.
TL;DR — On Ollama 0.21.0, running inference on all 6 cores vs. 4
produced the same token rate but 90x worse display lag (P95 181ms →
2ms) and 6x worse UI click response (929ms → 158ms). On Ollama
0.40.2 the gap nearly vanished, because the runner internals changed.
I shipped the thread-cap setting anyway — here's why.
What I actually wanted to measure
Not generation speed. API tok/s is easy to measure, but my question
was "how fast does what the API returned get onto the screen?"
Hypothesis: with all cores allocated, the LLM saturates the CPU, the
renderer is starved, and incoming tokens can't be painted in time.
With fewer cores, generation speed is nearly identical but the display
stays smooth — so it feels faster.
To quantify "feel," I split the clocks:
- Wire arrival time — a thin proxy between the IDE and Ollama timestamps every NDJSON chunk. Accurate even when the renderer is starving.
- DOM render time — a MutationObserver records streaming text length over time, giving per-chunk arrival→paint lag.
- Event-loop delay — drift of a 100ms setTimeout.
- Effective FPS — requestAnimationFrame fire count.
- UI click probe — actually clicking a real UI button during generation and timing the response.
- Fixed conditions — temperature 0, fixed seed, num_predict 600, think: false.
Environment: Ryzen 5 4500U (6 logical cores), CPU inference, Windows.
Compared "all cores" vs "4", median over a 60s window. The harness is
e2e/bench-threads.mjs (Playwright + measurement proxy) in the
repo.
Part 1: Ollama 0.21.0 — same speed, 90x more invisible
| Metric | All 6 cores | 4 cores |
|---|---|---|
| Wire-arrival speed (≈ generation) | 4.6 chars/s | 4.7 chars/s |
| Display lag P95 (arrival→DOM) | 181 ms | 2 ms |
| Display lag Max | 553 ms | 6 ms |
| Event-loop delay P95 | 70 ms | 13.6 ms |
| Effective FPS | 55.9 | 60.0 |
| UI click response (real click) | 929 ms | 158 ms |
Three takeaways:
- Generation speed was identical — decode is memory-bandwidth bound, so cutting 6→4 threads didn't slow token production.
- Display lag diverged ~90x at P95 — on all cores, tokens waited up to half a second to appear. Streaming output "backs up."
- One UI click took 929ms vs 158ms (~6x) — a proxy for "can you still do light work while the LLM runs."
On the old engine the runner occupied ~5.7 cores continuously during
generation. "4 cores feels faster" wasn't an illusion — it existed as
real delay in the display pipeline.
Part 2: Ollama 0.40.2 — I upgraded, and the premise changed
Before publishing, I wondered "isn't 0.21.0 old?" and re-ran on
0.40.2. The gap nearly disappeared:
| Metric | num_thread=6 | num_thread=4 |
|---|---|---|
| Wire-arrival speed | 7.6 chars/s | 7.5 chars/s |
| Display lag P95 | 13 ms | 3 ms |
| Event-loop delay P95 | 14.2 ms | 13.9 ms |
| Effective FPS | 59.9 | 59.9 |
| UI click response | 123 ms | 114 ms |
Same story on a 9.6GB model (lagP95: 5ms vs 3ms). Inspecting the
runner process showed why:
| Ollama 0.21.0 | Ollama 0.40.2 | |
|---|---|---|
| Runner |
ollama runner (old engine) |
llama-server (llama.cpp) |
| Threads, unspecified | all cores (=6) | n_threads=3 (measured) |
| Actual CPU at num_thread=6 | ~5.7 cores saturated | ~4.3–4.6 cores, no saturation |
- The default changed: unspecified now means 3 threads — it never grabs all cores in the first place.
- Even explicit num_thread=6 doesn't saturate — decode is memory- bound and the new engine's threads simply wait instead of spinning.
- Headroom exists now: 1.5–2 spare cores at all times, so the renderer/OS/other apps can preempt. UI starvation doesn't happen.
- Prefill behaves the same way — UI response stayed 90–180ms during prompt processing.
The "all cores make it stutter" experience was a real property of the
old engine's design, and the new engine fixed it internally.
So is a thread cap useless? Not quite
Two reasons to keep the setting:
-
Guarantee, not tuning. Engine internals change per version —
the unspecified default moved 6→3 in one release.
num_threadis an explicit ceiling that stays put regardless of version, model, or whatever else is running on the machine. -
API reach. The option only exists on Ollama's native
/api/chat.
Bonus finding: the "thinking" switch
While measuring, I noticed the model (qwen3.5:4b, a thinking model)
produced a long reasoning trace — input analysis, greeting candidates,
style polish — before answering a plain "hello." On a modest PC that
trace alone costs tens of seconds to minutes. Ollama's native API has
a think flag; think:false suppresses it (verified: zero
pre-content chunks in our runs). For practical local-LLM speed it's as
effective a switch as the thread cap.
Where this becomes a tooling difference
Here's the part that turned this into a product feature: these
options only exist on Ollama's native /api/chat.
| What you want | Ollama /api/chat
|
OpenAI-compatible /v1
|
|---|---|---|
| Cap CPU cores | options.num_thread |
No way to pass it |
| Thinking on/off | think |
No way to pass it |
| Output length / seed | options.* |
Partially |
Most VS Code-extension agents don't expose "CPU core count for local
LLM" as a setting — with an OpenAI-compatible client there's simply no
channel to send it through. To run a local LLM without killing the
host PC, you need a design that reaches the native API's options.
I shipped this in Teaspoon IDE v1.2.0 — a free, standalone
Electron AI coding IDE I maintain — as "CPU Threads" and "Thinking"
settings, plus the IDE's Ollama path moved to /api/chat. It's BYOK
(bring your own Gemini key) or fully offline via Ollama, and
source-available under FSL-1.1-MIT.
Product page: https://cuculhart.com/teaspoon.en.html
Lessons
- tok/s explains only half the experience — same generation speed, but starved rendering changed perceived speed by a measured 90x.
- "Use everything" isn't always optimal — and "how much to use" depends on engine internals that change between versions.
- Re-verify on the latest environment — re-testing on the new release showed the phenomenon was engine-dependent, which made the article stronger, not weaker.
- UI responsiveness is measurable — split wire time and DOM time, and "feels slow" becomes a number.
-
Local-LLM features are bounded by API reach — a design limited
to the compatibility API's common denominator can't implement
num_threadorthink.
Methodology: harness e2e/bench-threads.mjs, raw data in
e2e/bench-threads-results.json — reproduce with
BENCH_THREADS=6,4 BENCH_RUNS=3 BENCH_MODEL=qwen3.5:4b node e2e/bench-threads.mjs.
n=2–3 on a single Ryzen 5 4500U, so treat the exact numbers as
one-machine data; the measurement method is the reusable part.
Top comments (0)