On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow
Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should
This part focuses on thread scheduling and CPU frequency scaling.
1. Where part 1 left off
In part 1 I benchmarked Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) and measured 0.5 tok/s — 60-120x off the theoretical limits for that SoC.
That investigation found one real bug: llama.cpp cross-compilation passes no -march by default, so quantized matmul falls back to a slow path. Fixing it cut prefill from 114 s to 45 s.
But throughput was still 0.5 tok/s. The biggest hole was filled; the problem remained.
2. A reader points in a new direction
After publishing on DEV, a reader offered two traps that look like a slow model but aren't:
Scheduling: a llama.cpp server whose child processes were parked on efficiency cores under load — a 3-second transcription became 37 seconds. Same binary, same flags; the only difference was who launched it.
Storage: running off a USB stick, the whole model is re-read at every start — 2.6 GB on a 34 MB/s drive means over a minute before the first token.
And the line that stung:
The number that looks like "the model is slow" is almost never the model.
Both point the same way: maybe the problem isn't what is being computed, but where and how fast.
3. Know your silicon first: 1+3+4
for c in /sys/devices/system/cpu/cpu[0-9]*; do
echo "$(basename $c) $(cat $c/cpufreq/cpuinfo_max_freq)"
done
| Cores | Max freq | Type |
|---|---|---|
| cpu0–3 | 1.79 GHz | A55 little |
| cpu4–6 | 2.42 GHz | A76 mid |
| cpu7 | 2.84 GHz | A76 big |
Classic 1+3+4. If inference threads land on cpu0–3, you lose 3-4x immediately.
4. Test 1 — which core are the threads actually on?
Method: read /proc, no tooling required
Field 39 of /proc/<pid>/task/<tid>/stat is the CPU the thread is currently running on.
pid=$(pidof com.example.localai)
for t in /proc/$pid/task/*; do
tid=$(basename $t)
comm=$(cat $t/comm)
cpu=$(awk '{print $39}' $t/stat)
ticks=$(awk '{print $14+$15}' $t/stat)
echo "$tid $comm cpu=$cpu ticks=$ticks"
done
Sample repeatedly — a single snapshot tells you nothing.
Result: all on big cores. Scheduling is not the problem.
main thread 9556: cpu5 -> cpu5 -> cpu5 -> cpu3 -> cpu5 -> cpu4
19524 openmp_worker cpu6
19525 openmp_worker cpu7 <- big core
19526 openmp_worker cpu5
All four inference threads stayed on cpu4–7. None touched the little cores.
Also confirmed OpenMP is genuinely working:
19524 openmp_worker 5607 ticks
19525 openmp_worker 5606 ticks
19526 openmp_worker 5602 ticks
9556 DefaultDispatch 12132 ticks (main)
n_threads=4 → 3 workers + 1 main thread. Exactly as configured.
Verdict: "threads parked on efficiency cores" does not happen on this device. Ruled out.
5. Test 2 — is the CPU being throttled?
The symptom: frequency swinging wildly
cpu4: 710 ~ 1920 MHz (max 2419) <- as low as 29%
cpu5: 1612 ~ 2419 MHz (max 2419)
cpu6: 710 ~ 2419 MHz (max 2419)
cpu7: 1612 ~ 2016 MHz (max 2841) <- never above 71%
First rule out thermals
adb shell dumpsys battery | grep temperature
# temperature: 339 -> 33.9 C
33.9 °C. Not thermal.
The culprit is schedutil, Android's default governor. It scales frequency based on utilization estimates — and LLM inference is bursty: small work units, back to back. The estimate lags, the frequency hunts, and you get exactly this oscillation.
The decisive experiment: pin everything
# root required; SELinux Enforcing blocks a direct write, so chmod first
for c in 0 1 2 3 4 5 6 7; do
g=/sys/devices/system/cpu/cpu$c/cpufreq/scaling_governor
chmod 666 "$g"
echo performance > "$g"
done
Confirmed pinned: cpu0-3 @1785, cpu4-6 @2419, cpu7 @2841 MHz.
Result: only 20% faster
| Metric | schedutil | performance | Delta |
|---|---|---|---|
| Generation | 0.5 tok/s | 0.6 tok/s | +20% |
| First token | 1863 ms | 1619 ms | −13% |
20%. Throttling is real, but it is not the dominant factor.
6. Updated elimination table
| Hypothesis | Verdict |
|---|---|
Missing ARM optimization (-march) |
✅ Real, fixed — 2.5x on prefill |
| Threads scheduled onto little cores | ❌ Ruled out |
| OpenMP not actually threading | ❌ Ruled out |
| mmap pages evicted | ❌ Ruled out (LLAMA_LOAD_MODE_MLOCK changed nothing) |
| CPU frequency throttling | ⚠️ Real, but only ~20% |
7. What's left
CPU at full clock (2419/2841 MHz)
threads on big cores (cpu4-7)
dotprod instructions present (verified by disassembly)
|
v
Utilization of peak compute capacity: ~1%
The arithmetic:
per token: 1.5B params x 2 = 3 GFLOPs
measured: 1.67 s per token
effective: 1.8 GFLOPS
peak: ~182 GFLOPS (4x A76 at full clock)
utilization: 1%
And memory:
per token reads all weights = 1.06 GB
1.06 GB / 1.67 s = 635 MB/s
available bandwidth: 17-34 GB/s
utilization: 2-4%
Neither the CPU nor the memory bus is anywhere near saturated. It's just slow.
The remaining explanation points inside ggml's quantized kernels on this platform — memory access patterns, cache behaviour, or OpenMP barrier overhead on small work units.
I can't fix that from configuration. It stays on the list as the next thing to dig into.
8. Why this post is about method, not answers
The investigation produced no final answer — but the process is reusable.
Zero-dependency diagnostics
awk '{print $39}' /proc/<pid>/task/<tid>/stat # current CPU per thread
awk '{print $14+$15}' /proc/<pid>/task/<tid>/stat # cumulative CPU time per thread
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq # live frequency
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor # governor
Works on any Android device. No extra tooling.
Negative results are results
Ruling out A, B and C is progress. It shrinks the search space from "anything" to "inside ggml's kernels".
If I only published the successful fix, you'd think this was a clean path. Real engineering is mostly wrong hypotheses and the value is eliminating them quickly.
Compute the ceiling first
Most people say "my model is slow" without quantifying it. Work out the memory-bandwidth ceiling and the compute ceiling, then measure. Knowing the gap is 60-120x tells you the problem isn't "needs tuning" — it's "something isn't running at all".
9. Restore your environment
| Item | Restored to |
|---|---|
| CPU governor | schedutil |
| SELinux | Enforcing |
| stay-awake flag | false |
| screen timeout | 60000 ms |
| temp scripts | deleted |
Don't leave side effects on a test device — especially SELinux.
10. Next
Next up: ggml quantized kernel efficiency on ARM.
- Which code path does Q4_K_M take?
- What's the cache miss rate?
- How much does the OpenMP barrier cost per generated token?
If you've dug into this, I'd love to hear it.
Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should
Part 2: this post
Code: https://github.com/Pingredsai/local-ai-android
Top comments (0)