<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Pingredsai</title>
    <description>The latest articles on DEV Community by Pingredsai (@pingredsai).</description>
    <link>https://dev.to/pingredsai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4173320%2Fbc0be47e-2369-458c-8a6a-164ba4b40430.png</url>
      <title>DEV Community: Pingredsai</title>
      <link>https://dev.to/pingredsai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pingredsai"/>
    <language>en</language>
    <item>
      <title>Why Your Phone Runs LLMs 80x Slower Than It Should (And What I Found)</title>
      <dc:creator>Pingredsai</dc:creator>
      <pubDate>Fri, 09 Oct 2026 11:45:01 +0000</pubDate>
      <link>https://dev.to/pingredsai/why-your-phone-runs-llms-80x-slower-than-it-should-and-what-i-found-2645</link>
      <guid>https://dev.to/pingredsai/why-your-phone-runs-llms-80x-slower-than-it-should-and-what-i-found-2645</guid>
      <description>&lt;h1&gt;
  
  
  Why Your Phone Runs LLMs 80x Slower Than It Should (A Debugging Log)
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;Tags: on-device inference / llama.cpp / Android / performance&lt;br&gt;
Status: draft&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;I ran &lt;strong&gt;Qwen2.5-1.5B-Instruct Q4_K_M&lt;/strong&gt; fully offline on a &lt;strong&gt;Google Pixel 4&lt;/strong&gt; (Snapdragon 855, 2019).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Measured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Generation&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.5 tok/s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First token&lt;/td&gt;
&lt;td&gt;1.9 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Peak RSS&lt;/td&gt;
&lt;td&gt;1.3-2.0 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;402% (4 threads saturated)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Theoretical limits say this device should do &lt;strong&gt;30-60 tok/s&lt;/strong&gt;.&lt;br&gt;
&lt;strong&gt;It is 60-120x slower.&lt;/strong&gt; And the CPU is already maxed out.&lt;/p&gt;

&lt;p&gt;Here is the full debugging log — including three mistakes I made.&lt;/p&gt;


&lt;h2&gt;
  
  
  1. Do the math first
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Memory bandwidth bound:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model size      1.06 GB
SD855 bandwidth ~34 GB/s
one token reads all weights once
-&amp;gt; theoretical ceiling ~32 tok/s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Compute bound:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1.5B params x 2 FLOPs = 3 GFLOPs per token
SD855 CPU peak ~182 GFLOPS
-&amp;gt; theoretical ceiling ~60 tok/s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both ceilings land in the 30-60 tok/s range. &lt;strong&gt;Measured: 0.5.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Knowing the gap (60-120x) tells you what to suspect.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Suspect 1: thread count? (ruled out)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;adb shell top &lt;span class="nt"&gt;-b&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 1 &lt;span class="nt"&gt;-o&lt;/span&gt; PID,%CPU,%MEM,ARGS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PID    %CPU   %MEM   ARGS
17834  402%   24.3%  com.example.localai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;402%&lt;/strong&gt; — four threads fully saturated. Not a threading problem. Ruled out.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Suspect 2: missing ARM optimization? (HIT)
&lt;/h2&gt;

&lt;p&gt;This is the valuable one.&lt;/p&gt;

&lt;p&gt;llama.cpp does its heavy lifting in &lt;code&gt;ggml&lt;/code&gt;. From &lt;code&gt;ggml/src/ggml-cpu/CMakeLists.txt&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cmake"&gt;&lt;code&gt;&lt;span class="nb"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;GGML_NATIVE&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# probe host CPU, add -mcpu=native&lt;/span&gt;
    ...
&lt;span class="nb"&gt;else&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nb"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;GGML_CPU_ARM_ARCH&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;APPEND ARCH_FLAGS -march=&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GGML_CPU_ARM_ARCH&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# &amp;lt;-- key&lt;/span&gt;
    &lt;span class="nb"&gt;elseif&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;GGML_CPU_ALL_VARIANTS&lt;span class="p"&gt;)&lt;/span&gt;
        ...
    &lt;span class="c1"&gt;# neither set -&amp;gt; NO -march at all&lt;/span&gt;
&lt;span class="nb"&gt;endif&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When cross-compiling for Android:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;GGML_NATIVE&lt;/code&gt; is &lt;code&gt;OFF&lt;/code&gt; (can't probe the target)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GGML_CPU_ARM_ARCH&lt;/code&gt; defaults to &lt;strong&gt;empty&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Result: &lt;strong&gt;no &lt;code&gt;-march&lt;/code&gt; is passed at all&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Compiler falls back to the &lt;strong&gt;baseline ARMv8-A&lt;/strong&gt; — no dotprod, no fp16.&lt;/p&gt;

&lt;p&gt;Quantized matrix multiplication (99% of LLM compute) degrades to a slow path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="n"&gt;arguments&lt;/span&gt; &lt;span class="p"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;listOf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"-DUSE_LLAMA_CPP=ON"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"-DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Result: prefill went from 114 s to 45 s — 2.5x faster.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Better. But not enough.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verify the optimization actually landed
&lt;/h3&gt;

&lt;p&gt;Don't trust the flag. Disassemble:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llvm-objdump &lt;span class="nt"&gt;-d&lt;/span&gt; libggml-cpu.so | &lt;span class="nb"&gt;grep &lt;/span&gt;sdot
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight armasm"&gt;&lt;code&gt;&lt;span class="nl"&gt;18f558&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="err"&gt;4&lt;/span&gt;&lt;span class="nb"&gt;e829420&lt;/span&gt;    &lt;span class="nv"&gt;sdot&lt;/span&gt;   &lt;span class="nv"&gt;v0&lt;/span&gt;&lt;span class="mf"&gt;.4&lt;/span&gt;&lt;span class="nv"&gt;s&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;v1&lt;/span&gt;&lt;span class="mf"&gt;.16&lt;/span&gt;&lt;span class="nv"&gt;b&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;v2&lt;/span&gt;&lt;span class="mf"&gt;.16&lt;/span&gt;&lt;span class="nv"&gt;b&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;code&gt;sdot&lt;/code&gt; is there.&lt;/strong&gt; dotprod is active. Ruled out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Still 0.5 tok/s.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Suspect 3: mmap pages getting evicted? (ruled out)
&lt;/h2&gt;

&lt;p&gt;Android reclaims mmap'ed pages under memory pressure. Then every weight read falls back to flash:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flash read ~1.5 GB/s
1.06 GB / 1.5 GB/s ~ 700 ms/token
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Right order of magnitude for the 2 s/token we saw.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So I forced the model to stay resident:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;mparams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;load_mode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;LLAMA_LOAD_MODE_MLOCK&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Result: still 0.5 tok/s. Ruled out.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Three mistakes I made (worth avoiding)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Mistake 1: assuming &lt;code&gt;llama_batch_get_one&lt;/code&gt; pins position to 0
&lt;/h3&gt;

&lt;p&gt;I claimed: &lt;em&gt;"&lt;code&gt;llama_batch_get_one()&lt;/code&gt; sets pos to 0, so every token recomputes the whole sequence."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Then I hand-rolled &lt;code&gt;llama_batch_init()&lt;/code&gt; with explicit positions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What actually happened:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reading the header: &lt;code&gt;llama_batch_get_one&lt;/code&gt;'s &lt;code&gt;pos&lt;/code&gt; is &lt;strong&gt;&lt;code&gt;nullptr&lt;/code&gt;&lt;/strong&gt; — llama.cpp assigns positions automatically. &lt;strong&gt;Not 0.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;My hand-rolled version &lt;strong&gt;hung the prefill&lt;/strong&gt; (43 tokens, never returned).&lt;/li&gt;
&lt;li&gt;Reverting to &lt;code&gt;llama_batch_get_one&lt;/code&gt; fixed it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Lesson: for perf issues, suspect compiler flags and memory access first — not a detail you "remember".&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Mistake 2: using a removed API
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;llama_kv_cache_clear&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;g_ctx&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;   &lt;span class="c1"&gt;// removed in current llama.cpp&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;llama_memory_clear&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;llama_get_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;g_ctx&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nb"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Mistake 3: guessing struct field names
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;mparams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;use_mmap&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// fields no longer exist&lt;/span&gt;
&lt;span class="n"&gt;mparams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;use_mlock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Merged into one enum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;mparams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;load_mode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;LLAMA_LOAD_MODE_MLOCK&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Common lesson: llama.cpp's API moves fast. Always read &lt;code&gt;include/llama.h&lt;/code&gt;. Never write from memory.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Where it landed
&lt;/h2&gt;

&lt;p&gt;After ruling out three suspects:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;3 GFLOPs per token
2 s per token measured
-&amp;gt; ~0.8% of peak CPU utilization
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And it is not a bandwidth wall (~500 MB/s measured vs ~34 GB/s available).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So compute efficiency itself is the problem&lt;/strong&gt; — beyond "wrong config". Could be Android CPU scheduling, OpenMP synchronization overhead, or platform-specific behavior in llama.cpp.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is not a one-day fix. And it should not block shipping.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  7. What I did about it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ship it and state the number honestly&lt;/strong&gt; — README says "Pixel 4: 0.5 tok/s"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explain the limitation&lt;/strong&gt; — old SoC + large model&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask the community for more data&lt;/strong&gt; — Snapdragon 8 Gen 2/3 results welcome&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move performance work to v2&lt;/strong&gt; instead of blocking v1&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  8. Transferable lessons
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Lesson&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Check compiler flags first&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;-march&lt;/code&gt; / dotprod issues are the #1 cause of slow on-device inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU saturated ≠ efficient&lt;/td&gt;
&lt;td&gt;402% usage can mean you're running the slowest implementation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compute theoretical limits&lt;/td&gt;
&lt;td&gt;The gap tells you what to suspect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Verify with disassembly&lt;/td&gt;
&lt;td&gt;Flags set ≠ instructions generated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read the header, don't recall it&lt;/td&gt;
&lt;td&gt;llama.cpp API churns fast&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Set a stop-loss for the project&lt;/td&gt;
&lt;td&gt;Perf work must not block shipping forever&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Appendix: environment
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Device&lt;/td&gt;
&lt;td&gt;Google Pixel 4 (Snapdragon 855, 6GB, Android 13)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Qwen2.5-1.5B-Instruct-Q4_K_M.gguf (1065 MB)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llama.cpp&lt;/td&gt;
&lt;td&gt;commit 8a1a9b5 (2026-10-09)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NDK&lt;/td&gt;
&lt;td&gt;30.0.16248370&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CMake&lt;/td&gt;
&lt;td&gt;4.1.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Build flags&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;-DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16&lt;/code&gt;, &lt;code&gt;-DUSE_LLAMA_CPP=ON&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;n_threads=4&lt;/code&gt;, &lt;code&gt;n_ctx=2048&lt;/code&gt;, &lt;code&gt;load_mode=MLOCK&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Source code
&lt;/h2&gt;

&lt;p&gt;Full implementation (Kotlin + JNI + llama.cpp + Jetpack Compose), MIT licensed:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/Pingredsai/local-ai-android" rel="noopener noreferrer"&gt;https://github.com/Pingredsai/local-ai-android&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Running the same model on newer hardware? Send me your numbers — that is exactly what this project needs.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>android</category>
      <category>llm</category>
      <category>performance</category>
      <category>kotlin</category>
    </item>
  </channel>
</rss>
