DEV Community

cuculhart
cuculhart

Posted on

Tokens arrived on time — the renderer didn't: measuring a 90x display lag in local-LLM streaming (and how Ollama fixed it)

I built a local-LLM coding IDE, and I noticed something odd: giving
the model all the CPU cores made the app feel slower. Tokens
were arriving at the same speed, but the UI lagged behind. So I
measured it — and then, mid-verification, a new Ollama release
changed the whole story.

TL;DR — On Ollama 0.21.0, running inference on all 6 cores vs. 4
produced the same token rate but 90x worse display lag (P95 181ms →
2ms)
and 6x worse UI click response (929ms → 158ms). On Ollama
0.40.2 the gap nearly vanished, because the runner internals changed.
I shipped the thread-cap setting anyway — here's why.

What I actually wanted to measure

Not generation speed. API tok/s is easy to measure, but my question
was "how fast does what the API returned get onto the screen?"

Hypothesis: with all cores allocated, the LLM saturates the CPU, the
renderer is starved, and incoming tokens can't be painted in time.
With fewer cores, generation speed is nearly identical but the display
stays smooth — so it feels faster.

To quantify "feel," I split the clocks:

  • Wire arrival time — a thin proxy between the IDE and Ollama timestamps every NDJSON chunk. Accurate even when the renderer is starving.
  • DOM render time — a MutationObserver records streaming text length over time, giving per-chunk arrival→paint lag.
  • Event-loop delay — drift of a 100ms setTimeout.
  • Effective FPS — requestAnimationFrame fire count.
  • UI click probe — actually clicking a real UI button during generation and timing the response.
  • Fixed conditions — temperature 0, fixed seed, num_predict 600, think: false.

Environment: Ryzen 5 4500U (6 logical cores), CPU inference, Windows.
Compared "all cores" vs "4", median over a 60s window. The harness is
e2e/bench-threads.mjs (Playwright + measurement proxy) in the
repo.

Part 1: Ollama 0.21.0 — same speed, 90x more invisible

Metric All 6 cores 4 cores
Wire-arrival speed (≈ generation) 4.6 chars/s 4.7 chars/s
Display lag P95 (arrival→DOM) 181 ms 2 ms
Display lag Max 553 ms 6 ms
Event-loop delay P95 70 ms 13.6 ms
Effective FPS 55.9 60.0
UI click response (real click) 929 ms 158 ms

Three takeaways:

  1. Generation speed was identical — decode is memory-bandwidth bound, so cutting 6→4 threads didn't slow token production.
  2. Display lag diverged ~90x at P95 — on all cores, tokens waited up to half a second to appear. Streaming output "backs up."
  3. One UI click took 929ms vs 158ms (~6x) — a proxy for "can you still do light work while the LLM runs."

On the old engine the runner occupied ~5.7 cores continuously during
generation. "4 cores feels faster" wasn't an illusion — it existed as
real delay in the display pipeline.

Part 2: Ollama 0.40.2 — I upgraded, and the premise changed

Before publishing, I wondered "isn't 0.21.0 old?" and re-ran on
0.40.2. The gap nearly disappeared:

Metric num_thread=6 num_thread=4
Wire-arrival speed 7.6 chars/s 7.5 chars/s
Display lag P95 13 ms 3 ms
Event-loop delay P95 14.2 ms 13.9 ms
Effective FPS 59.9 59.9
UI click response 123 ms 114 ms

Same story on a 9.6GB model (lagP95: 5ms vs 3ms). Inspecting the
runner process showed why:

Ollama 0.21.0 Ollama 0.40.2
Runner ollama runner (old engine) llama-server (llama.cpp)
Threads, unspecified all cores (=6) n_threads=3 (measured)
Actual CPU at num_thread=6 ~5.7 cores saturated ~4.3–4.6 cores, no saturation
  • The default changed: unspecified now means 3 threads — it never grabs all cores in the first place.
  • Even explicit num_thread=6 doesn't saturate — decode is memory- bound and the new engine's threads simply wait instead of spinning.
  • Headroom exists now: 1.5–2 spare cores at all times, so the renderer/OS/other apps can preempt. UI starvation doesn't happen.
  • Prefill behaves the same way — UI response stayed 90–180ms during prompt processing.

The "all cores make it stutter" experience was a real property of the
old engine's design, and the new engine fixed it internally.

So is a thread cap useless? Not quite

Two reasons to keep the setting:

  1. Guarantee, not tuning. Engine internals change per version — the unspecified default moved 6→3 in one release. num_thread is an explicit ceiling that stays put regardless of version, model, or whatever else is running on the machine.
  2. API reach. The option only exists on Ollama's native /api/chat.

Bonus finding: the "thinking" switch

While measuring, I noticed the model (qwen3.5:4b, a thinking model)
produced a long reasoning trace — input analysis, greeting candidates,
style polish — before answering a plain "hello." On a modest PC that
trace alone costs tens of seconds to minutes. Ollama's native API has
a think flag; think:false suppresses it (verified: zero
pre-content chunks in our runs). For practical local-LLM speed it's as
effective a switch as the thread cap.

Where this becomes a tooling difference

Here's the part that turned this into a product feature: these
options only exist on Ollama's native /api/chat
.

What you want Ollama /api/chat OpenAI-compatible /v1
Cap CPU cores options.num_thread No way to pass it
Thinking on/off think No way to pass it
Output length / seed options.* Partially

Most VS Code-extension agents don't expose "CPU core count for local
LLM" as a setting — with an OpenAI-compatible client there's simply no
channel to send it through. To run a local LLM without killing the
host PC, you need a design that reaches the native API's options.

I shipped this in Teaspoon IDE v1.2.0 — a free, standalone
Electron AI coding IDE I maintain — as "CPU Threads" and "Thinking"
settings, plus the IDE's Ollama path moved to /api/chat. It's BYOK
(bring your own Gemini key) or fully offline via Ollama, and
source-available under FSL-1.1-MIT.
Product page: https://cuculhart.com/teaspoon.en.html

Lessons

  • tok/s explains only half the experience — same generation speed, but starved rendering changed perceived speed by a measured 90x.
  • "Use everything" isn't always optimal — and "how much to use" depends on engine internals that change between versions.
  • Re-verify on the latest environment — re-testing on the new release showed the phenomenon was engine-dependent, which made the article stronger, not weaker.
  • UI responsiveness is measurable — split wire time and DOM time, and "feels slow" becomes a number.
  • Local-LLM features are bounded by API reach — a design limited to the compatibility API's common denominator can't implement num_thread or think.

Methodology: harness e2e/bench-threads.mjs, raw data in
e2e/bench-threads-results.json — reproduce with
BENCH_THREADS=6,4 BENCH_RUNS=3 BENCH_MODEL=qwen3.5:4b node e2e/bench-threads.mjs.
n=2–3 on a single Ryzen 5 4500U, so treat the exact numbers as
one-machine data; the measurement method is the reusable part.

Top comments (0)