DEV Community

Pingredsai
Pingredsai

Posted on

Why Your Phone Runs LLMs 80x Slower Than It Should (And What I Found)

Why Your Phone Runs LLMs 80x Slower Than It Should (A Debugging Log)

Tags: on-device inference / llama.cpp / Android / performance
Status: draft


I ran Qwen2.5-1.5B-Instruct Q4_K_M fully offline on a Google Pixel 4 (Snapdragon 855, 2019).

Metric Measured
Generation 0.5 tok/s
First token 1.9 s
Peak RSS 1.3-2.0 GB
CPU 402% (4 threads saturated)

Theoretical limits say this device should do 30-60 tok/s.
It is 60-120x slower. And the CPU is already maxed out.

Here is the full debugging log — including three mistakes I made.


1. Do the math first

Memory bandwidth bound:

model size      1.06 GB
SD855 bandwidth ~34 GB/s
one token reads all weights once
-> theoretical ceiling ~32 tok/s
Enter fullscreen mode Exit fullscreen mode

Compute bound:

1.5B params x 2 FLOPs = 3 GFLOPs per token
SD855 CPU peak ~182 GFLOPS
-> theoretical ceiling ~60 tok/s
Enter fullscreen mode Exit fullscreen mode

Both ceilings land in the 30-60 tok/s range. Measured: 0.5.

Knowing the gap (60-120x) tells you what to suspect.


2. Suspect 1: thread count? (ruled out)

adb shell top -b -n 1 -o PID,%CPU,%MEM,ARGS
Enter fullscreen mode Exit fullscreen mode
PID    %CPU   %MEM   ARGS
17834  402%   24.3%  com.example.localai
Enter fullscreen mode Exit fullscreen mode

402% — four threads fully saturated. Not a threading problem. Ruled out.


3. Suspect 2: missing ARM optimization? (HIT)

This is the valuable one.

llama.cpp does its heavy lifting in ggml. From ggml/src/ggml-cpu/CMakeLists.txt:

if (GGML_NATIVE)
    # probe host CPU, add -mcpu=native
    ...
else()
    if (GGML_CPU_ARM_ARCH)
        list(APPEND ARCH_FLAGS -march=${GGML_CPU_ARM_ARCH})   # <-- key
    elseif(GGML_CPU_ALL_VARIANTS)
        ...
    # neither set -> NO -march at all
endif()
Enter fullscreen mode Exit fullscreen mode

When cross-compiling for Android:

  • GGML_NATIVE is OFF (can't probe the target)
  • GGML_CPU_ARM_ARCH defaults to empty
  • Result: no -march is passed at all

Compiler falls back to the baseline ARMv8-A — no dotprod, no fp16.

Quantized matrix multiplication (99% of LLM compute) degrades to a slow path.

Fix:

arguments += listOf(
    "-DUSE_LLAMA_CPP=ON",
    "-DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16"
)
Enter fullscreen mode Exit fullscreen mode

Result: prefill went from 114 s to 45 s — 2.5x faster.

Better. But not enough.

Verify the optimization actually landed

Don't trust the flag. Disassemble:

llvm-objdump -d libggml-cpu.so | grep sdot
Enter fullscreen mode Exit fullscreen mode
18f558: 4e829420    sdot   v0.4s, v1.16b, v2.16b
Enter fullscreen mode Exit fullscreen mode

sdot is there. dotprod is active. Ruled out.

Still 0.5 tok/s.


4. Suspect 3: mmap pages getting evicted? (ruled out)

Android reclaims mmap'ed pages under memory pressure. Then every weight read falls back to flash:

flash read ~1.5 GB/s
1.06 GB / 1.5 GB/s ~ 700 ms/token
Enter fullscreen mode Exit fullscreen mode

Right order of magnitude for the 2 s/token we saw.

So I forced the model to stay resident:

mparams.load_mode = LLAMA_LOAD_MODE_MLOCK;
Enter fullscreen mode Exit fullscreen mode

Result: still 0.5 tok/s. Ruled out.


5. Three mistakes I made (worth avoiding)

Mistake 1: assuming llama_batch_get_one pins position to 0

I claimed: "llama_batch_get_one() sets pos to 0, so every token recomputes the whole sequence."

Then I hand-rolled llama_batch_init() with explicit positions.

What actually happened:

  1. Reading the header: llama_batch_get_one's pos is nullptr — llama.cpp assigns positions automatically. Not 0.
  2. My hand-rolled version hung the prefill (43 tokens, never returned).
  3. Reverting to llama_batch_get_one fixed it.

Lesson: for perf issues, suspect compiler flags and memory access first — not a detail you "remember".

Mistake 2: using a removed API

llama_kv_cache_clear(g_ctx);   // removed in current llama.cpp
Enter fullscreen mode Exit fullscreen mode

Now:

llama_memory_clear(llama_get_memory(g_ctx), true);
Enter fullscreen mode Exit fullscreen mode

Mistake 3: guessing struct field names

mparams.use_mmap  = false;   // fields no longer exist
mparams.use_mlock = true;
Enter fullscreen mode Exit fullscreen mode

Merged into one enum:

mparams.load_mode = LLAMA_LOAD_MODE_MLOCK;
Enter fullscreen mode Exit fullscreen mode

Common lesson: llama.cpp's API moves fast. Always read include/llama.h. Never write from memory.


6. Where it landed

After ruling out three suspects:

3 GFLOPs per token
2 s per token measured
-> ~0.8% of peak CPU utilization
Enter fullscreen mode Exit fullscreen mode

And it is not a bandwidth wall (~500 MB/s measured vs ~34 GB/s available).

So compute efficiency itself is the problem — beyond "wrong config". Could be Android CPU scheduling, OpenMP synchronization overhead, or platform-specific behavior in llama.cpp.

That is not a one-day fix. And it should not block shipping.


7. What I did about it

  1. Ship it and state the number honestly — README says "Pixel 4: 0.5 tok/s"
  2. Explain the limitation — old SoC + large model
  3. Ask the community for more data — Snapdragon 8 Gen 2/3 results welcome
  4. Move performance work to v2 instead of blocking v1

8. Transferable lessons

Lesson Why
Check compiler flags first -march / dotprod issues are the #1 cause of slow on-device inference
CPU saturated ≠ efficient 402% usage can mean you're running the slowest implementation
Compute theoretical limits The gap tells you what to suspect
Verify with disassembly Flags set ≠ instructions generated
Read the header, don't recall it llama.cpp API churns fast
Set a stop-loss for the project Perf work must not block shipping forever

Appendix: environment

Item Value
Device Google Pixel 4 (Snapdragon 855, 6GB, Android 13)
Model Qwen2.5-1.5B-Instruct-Q4_K_M.gguf (1065 MB)
llama.cpp commit 8a1a9b5 (2026-10-09)
NDK 30.0.16248370
CMake 4.1.2
Build flags -DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16, -DUSE_LLAMA_CPP=ON
Runtime n_threads=4, n_ctx=2048, load_mode=MLOCK

Source code

Full implementation (Kotlin + JNI + llama.cpp + Jetpack Compose), MIT licensed:

https://github.com/Pingredsai/local-ai-android


Running the same model on newer hardware? Send me your numbers — that is exactly what this project needs.

Top comments (1)

Collapse
 
alfred_odong_322108a5cc3d profile image
Alfred Odong •

Great debugging log, and the "do the math first" step is the part most people skip. Two related traps from the desktop side that look identical to a slow model and aren't:

  1. Scheduling, not compute. A llama.cpp server I started from a tool that runs at nice 5 got its child processes parked on the efficiency cores under load, and a 3-second transcription took 37 seconds. Same binary, same flags; the only difference was who launched it.

  2. Storage, not compute. Running off a USB stick, the whole model is read at every start, so on a 34 MB/s drive a 2.6 GB model took over a minute before the first token, and seconds on a USB 3 drive. Tokens per second was fine both times; the stopwatch said otherwise.

Your -march finding is the same shape: the number that looks like "the model is slow" is almost never the model.