DEV Community

Pingredsai
Pingredsai

Posted on

On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow

On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow

Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should
This part focuses on thread scheduling and CPU frequency scaling.


1. Where part 1 left off

In part 1 I benchmarked Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) and measured 0.5 tok/s — 60-120x off the theoretical limits for that SoC.

That investigation found one real bug: llama.cpp cross-compilation passes no -march by default, so quantized matmul falls back to a slow path. Fixing it cut prefill from 114 s to 45 s.

But throughput was still 0.5 tok/s. The biggest hole was filled; the problem remained.


2. A reader points in a new direction

After publishing on DEV, a reader offered two traps that look like a slow model but aren't:

Scheduling: a llama.cpp server whose child processes were parked on efficiency cores under load — a 3-second transcription became 37 seconds. Same binary, same flags; the only difference was who launched it.

Storage: running off a USB stick, the whole model is re-read at every start — 2.6 GB on a 34 MB/s drive means over a minute before the first token.

And the line that stung:

The number that looks like "the model is slow" is almost never the model.

Both point the same way: maybe the problem isn't what is being computed, but where and how fast.


3. Know your silicon first: 1+3+4

for c in /sys/devices/system/cpu/cpu[0-9]*; do
  echo "$(basename $c) $(cat $c/cpufreq/cpuinfo_max_freq)"
done
Enter fullscreen mode Exit fullscreen mode
Cores Max freq Type
cpu0–3 1.79 GHz A55 little
cpu4–6 2.42 GHz A76 mid
cpu7 2.84 GHz A76 big

Classic 1+3+4. If inference threads land on cpu0–3, you lose 3-4x immediately.


4. Test 1 — which core are the threads actually on?

Method: read /proc, no tooling required

Field 39 of /proc/<pid>/task/<tid>/stat is the CPU the thread is currently running on.

pid=$(pidof com.example.localai)

for t in /proc/$pid/task/*; do
  tid=$(basename $t)
  comm=$(cat $t/comm)
  cpu=$(awk '{print $39}' $t/stat)
  ticks=$(awk '{print $14+$15}' $t/stat)
  echo "$tid $comm cpu=$cpu ticks=$ticks"
done
Enter fullscreen mode Exit fullscreen mode

Sample repeatedly — a single snapshot tells you nothing.

Result: all on big cores. Scheduling is not the problem.

main thread 9556:  cpu5 -> cpu5 -> cpu5 -> cpu3 -> cpu5 -> cpu4
19524 openmp_worker  cpu6
19525 openmp_worker  cpu7   <- big core
19526 openmp_worker  cpu5
Enter fullscreen mode Exit fullscreen mode

All four inference threads stayed on cpu4–7. None touched the little cores.

Also confirmed OpenMP is genuinely working:

19524 openmp_worker   5607 ticks
19525 openmp_worker   5606 ticks
19526 openmp_worker   5602 ticks
9556  DefaultDispatch 12132 ticks (main)
Enter fullscreen mode Exit fullscreen mode

n_threads=4 → 3 workers + 1 main thread. Exactly as configured.

Verdict: "threads parked on efficiency cores" does not happen on this device. Ruled out.


5. Test 2 — is the CPU being throttled?

The symptom: frequency swinging wildly

cpu4:  710 ~ 1920 MHz   (max 2419)   <- as low as 29%
cpu5: 1612 ~ 2419 MHz   (max 2419)
cpu6:  710 ~ 2419 MHz   (max 2419)
cpu7: 1612 ~ 2016 MHz   (max 2841)   <- never above 71%
Enter fullscreen mode Exit fullscreen mode

First rule out thermals

adb shell dumpsys battery | grep temperature
# temperature: 339  ->  33.9 C
Enter fullscreen mode Exit fullscreen mode

33.9 °C. Not thermal.

The culprit is schedutil, Android's default governor. It scales frequency based on utilization estimates — and LLM inference is bursty: small work units, back to back. The estimate lags, the frequency hunts, and you get exactly this oscillation.

The decisive experiment: pin everything

# root required; SELinux Enforcing blocks a direct write, so chmod first
for c in 0 1 2 3 4 5 6 7; do
  g=/sys/devices/system/cpu/cpu$c/cpufreq/scaling_governor
  chmod 666 "$g"
  echo performance > "$g"
done
Enter fullscreen mode Exit fullscreen mode

Confirmed pinned: cpu0-3 @1785, cpu4-6 @2419, cpu7 @2841 MHz.

Result: only 20% faster

Metric schedutil performance Delta
Generation 0.5 tok/s 0.6 tok/s +20%
First token 1863 ms 1619 ms −13%

20%. Throttling is real, but it is not the dominant factor.


6. Updated elimination table

Hypothesis Verdict
Missing ARM optimization (-march) ✅ Real, fixed — 2.5x on prefill
Threads scheduled onto little cores ❌ Ruled out
OpenMP not actually threading ❌ Ruled out
mmap pages evicted ❌ Ruled out (LLAMA_LOAD_MODE_MLOCK changed nothing)
CPU frequency throttling ⚠️ Real, but only ~20%

7. What's left

CPU at full clock (2419/2841 MHz)
threads on big cores (cpu4-7)
dotprod instructions present (verified by disassembly)
        |
        v
Utilization of peak compute capacity: ~1%
Enter fullscreen mode Exit fullscreen mode

The arithmetic:

per token: 1.5B params x 2 = 3 GFLOPs
measured: 1.67 s per token
effective: 1.8 GFLOPS
peak: ~182 GFLOPS (4x A76 at full clock)
utilization: 1%
Enter fullscreen mode Exit fullscreen mode

And memory:

per token reads all weights = 1.06 GB
1.06 GB / 1.67 s = 635 MB/s
available bandwidth: 17-34 GB/s
utilization: 2-4%
Enter fullscreen mode Exit fullscreen mode

Neither the CPU nor the memory bus is anywhere near saturated. It's just slow.

The remaining explanation points inside ggml's quantized kernels on this platform — memory access patterns, cache behaviour, or OpenMP barrier overhead on small work units.

I can't fix that from configuration. It stays on the list as the next thing to dig into.


8. Why this post is about method, not answers

The investigation produced no final answer — but the process is reusable.

Zero-dependency diagnostics

awk '{print $39}' /proc/<pid>/task/<tid>/stat      # current CPU per thread
awk '{print $14+$15}' /proc/<pid>/task/<tid>/stat  # cumulative CPU time per thread
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq   # live frequency
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor   # governor
Enter fullscreen mode Exit fullscreen mode

Works on any Android device. No extra tooling.

Negative results are results

Ruling out A, B and C is progress. It shrinks the search space from "anything" to "inside ggml's kernels".

If I only published the successful fix, you'd think this was a clean path. Real engineering is mostly wrong hypotheses and the value is eliminating them quickly.

Compute the ceiling first

Most people say "my model is slow" without quantifying it. Work out the memory-bandwidth ceiling and the compute ceiling, then measure. Knowing the gap is 60-120x tells you the problem isn't "needs tuning" — it's "something isn't running at all".


9. Restore your environment

Item Restored to
CPU governor schedutil
SELinux Enforcing
stay-awake flag false
screen timeout 60000 ms
temp scripts deleted

Don't leave side effects on a test device — especially SELinux.


10. Next

Next up: ggml quantized kernel efficiency on ARM.

  • Which code path does Q4_K_M take?
  • What's the cache miss rate?
  • How much does the OpenMP barrier cost per generated token?

If you've dug into this, I'd love to hear it.


Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should
Part 2: this post
Code: https://github.com/Pingredsai/local-ai-android

Top comments (0)