Why Your Phone Runs LLMs 80x Slower Than It Should (A Debugging Log)
Tags: on-device inference / llama.cpp / Android / performance
Status: draft
I ran Qwen2.5-1.5B-Instruct Q4_K_M fully offline on a Google Pixel 4 (Snapdragon 855, 2019).
| Metric | Measured |
|---|---|
| Generation | 0.5 tok/s |
| First token | 1.9 s |
| Peak RSS | 1.3-2.0 GB |
| CPU | 402% (4 threads saturated) |
Theoretical limits say this device should do 30-60 tok/s.
It is 60-120x slower. And the CPU is already maxed out.
Here is the full debugging log — including three mistakes I made.
1. Do the math first
Memory bandwidth bound:
model size 1.06 GB
SD855 bandwidth ~34 GB/s
one token reads all weights once
-> theoretical ceiling ~32 tok/s
Compute bound:
1.5B params x 2 FLOPs = 3 GFLOPs per token
SD855 CPU peak ~182 GFLOPS
-> theoretical ceiling ~60 tok/s
Both ceilings land in the 30-60 tok/s range. Measured: 0.5.
Knowing the gap (60-120x) tells you what to suspect.
2. Suspect 1: thread count? (ruled out)
adb shell top -b -n 1 -o PID,%CPU,%MEM,ARGS
PID %CPU %MEM ARGS
17834 402% 24.3% com.example.localai
402% — four threads fully saturated. Not a threading problem. Ruled out.
3. Suspect 2: missing ARM optimization? (HIT)
This is the valuable one.
llama.cpp does its heavy lifting in ggml. From ggml/src/ggml-cpu/CMakeLists.txt:
if (GGML_NATIVE)
# probe host CPU, add -mcpu=native
...
else()
if (GGML_CPU_ARM_ARCH)
list(APPEND ARCH_FLAGS -march=${GGML_CPU_ARM_ARCH}) # <-- key
elseif(GGML_CPU_ALL_VARIANTS)
...
# neither set -> NO -march at all
endif()
When cross-compiling for Android:
-
GGML_NATIVEisOFF(can't probe the target) -
GGML_CPU_ARM_ARCHdefaults to empty - Result: no
-marchis passed at all
Compiler falls back to the baseline ARMv8-A — no dotprod, no fp16.
Quantized matrix multiplication (99% of LLM compute) degrades to a slow path.
Fix:
arguments += listOf(
"-DUSE_LLAMA_CPP=ON",
"-DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16"
)
Result: prefill went from 114 s to 45 s — 2.5x faster.
Better. But not enough.
Verify the optimization actually landed
Don't trust the flag. Disassemble:
llvm-objdump -d libggml-cpu.so | grep sdot
18f558: 4e829420 sdot v0.4s, v1.16b, v2.16b
sdot is there. dotprod is active. Ruled out.
Still 0.5 tok/s.
4. Suspect 3: mmap pages getting evicted? (ruled out)
Android reclaims mmap'ed pages under memory pressure. Then every weight read falls back to flash:
flash read ~1.5 GB/s
1.06 GB / 1.5 GB/s ~ 700 ms/token
Right order of magnitude for the 2 s/token we saw.
So I forced the model to stay resident:
mparams.load_mode = LLAMA_LOAD_MODE_MLOCK;
Result: still 0.5 tok/s. Ruled out.
5. Three mistakes I made (worth avoiding)
Mistake 1: assuming llama_batch_get_one pins position to 0
I claimed: "llama_batch_get_one() sets pos to 0, so every token recomputes the whole sequence."
Then I hand-rolled llama_batch_init() with explicit positions.
What actually happened:
- Reading the header:
llama_batch_get_one'sposisnullptr— llama.cpp assigns positions automatically. Not 0. - My hand-rolled version hung the prefill (43 tokens, never returned).
- Reverting to
llama_batch_get_onefixed it.
Lesson: for perf issues, suspect compiler flags and memory access first — not a detail you "remember".
Mistake 2: using a removed API
llama_kv_cache_clear(g_ctx); // removed in current llama.cpp
Now:
llama_memory_clear(llama_get_memory(g_ctx), true);
Mistake 3: guessing struct field names
mparams.use_mmap = false; // fields no longer exist
mparams.use_mlock = true;
Merged into one enum:
mparams.load_mode = LLAMA_LOAD_MODE_MLOCK;
Common lesson: llama.cpp's API moves fast. Always read include/llama.h. Never write from memory.
6. Where it landed
After ruling out three suspects:
3 GFLOPs per token
2 s per token measured
-> ~0.8% of peak CPU utilization
And it is not a bandwidth wall (~500 MB/s measured vs ~34 GB/s available).
So compute efficiency itself is the problem — beyond "wrong config". Could be Android CPU scheduling, OpenMP synchronization overhead, or platform-specific behavior in llama.cpp.
That is not a one-day fix. And it should not block shipping.
7. What I did about it
- Ship it and state the number honestly — README says "Pixel 4: 0.5 tok/s"
- Explain the limitation — old SoC + large model
- Ask the community for more data — Snapdragon 8 Gen 2/3 results welcome
- Move performance work to v2 instead of blocking v1
8. Transferable lessons
| Lesson | Why |
|---|---|
| Check compiler flags first |
-march / dotprod issues are the #1 cause of slow on-device inference |
| CPU saturated ≠ efficient | 402% usage can mean you're running the slowest implementation |
| Compute theoretical limits | The gap tells you what to suspect |
| Verify with disassembly | Flags set ≠ instructions generated |
| Read the header, don't recall it | llama.cpp API churns fast |
| Set a stop-loss for the project | Perf work must not block shipping forever |
Appendix: environment
| Item | Value |
|---|---|
| Device | Google Pixel 4 (Snapdragon 855, 6GB, Android 13) |
| Model | Qwen2.5-1.5B-Instruct-Q4_K_M.gguf (1065 MB) |
| llama.cpp | commit 8a1a9b5 (2026-10-09) |
| NDK | 30.0.16248370 |
| CMake | 4.1.2 |
| Build flags |
-DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16, -DUSE_LLAMA_CPP=ON
|
| Runtime |
n_threads=4, n_ctx=2048, load_mode=MLOCK
|
Source code
Full implementation (Kotlin + JNI + llama.cpp + Jetpack Compose), MIT licensed:
https://github.com/Pingredsai/local-ai-android
Running the same model on newer hardware? Send me your numbers — that is exactly what this project needs.
Top comments (1)
Great debugging log, and the "do the math first" step is the part most people skip. Two related traps from the desktop side that look identical to a slow model and aren't:
Scheduling, not compute. A llama.cpp server I started from a tool that runs at
nice 5got its child processes parked on the efficiency cores under load, and a 3-second transcription took 37 seconds. Same binary, same flags; the only difference was who launched it.Storage, not compute. Running off a USB stick, the whole model is read at every start, so on a 34 MB/s drive a 2.6 GB model took over a minute before the first token, and seconds on a USB 3 drive. Tokens per second was fine both times; the stopwatch said otherwise.
Your
-marchfinding is the same shape: the number that looks like "the model is slow" is almost never the model.