<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: GINIGEN AI</title>
    <description>The latest articles on DEV Community by GINIGEN AI (@ginigen_ai_010d9180cbbdb1).</description>
    <link>https://dev.to/ginigen_ai_010d9180cbbdb1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4163333%2F8df509f4-22c1-4190-9b28-35e38a01d52a.png</url>
      <title>DEV Community: GINIGEN AI</title>
      <link>https://dev.to/ginigen_ai_010d9180cbbdb1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ginigen_ai_010d9180cbbdb1"/>
    <language>en</language>
    <item>
      <title>4-bit GGUF Quality for MoE Models: Why Only 3B of 180B Params Fire, and How to Prove Parity</title>
      <dc:creator>GINIGEN AI</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:10:26 +0000</pubDate>
      <link>https://dev.to/ginigen_ai_010d9180cbbdb1/4-bit-gguf-quality-for-moe-models-why-only-3b-of-180b-params-fire-and-how-to-prove-parity-16a9</link>
      <guid>https://dev.to/ginigen_ai_010d9180cbbdb1/4-bit-gguf-quality-for-moe-models-why-only-3b-of-180b-params-fire-and-how-to-prove-parity-16a9</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;Mixture-of-Experts (MoE) models look enormous on disk, but only a small slice of the weights does work on any single token. A 180B-parameter MoE can activate roughly 3B parameters per forward pass. That sparsity is exactly why 4-bit GGUF quantization behaves so differently on MoE than on a dense model, and why a mixed-precision scheme such as the UD-Q4_K_XL style graft can hold accuracy that a flat 4-bit cast would lose.&lt;/p&gt;

&lt;p&gt;The honest part of the story is not the compression ratio. It is verification. A quantized file that loads and produces fluent text is not proof of anything. You prove parity by scoring the quantized model on a held-out benchmark and comparing it to the full-precision baseline, number against number. In the example below, a 4-bit mixed graft reaches MMLU-Pro 87.65, equal to the full-precision baseline inside noise. This post shows the reasoning and the exact &lt;code&gt;llama.cpp&lt;/code&gt; commands to reproduce that check yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does only a fraction of an MoE model activate per token?
&lt;/h2&gt;

&lt;p&gt;A dense transformer runs every weight for every token. A Mixture-of-Experts transformer replaces the big feed-forward block in each layer with many smaller expert blocks plus a router. For each token the router picks a few experts (often two) and ignores the rest. The attention layers and the router still run every time, but the overwhelming majority of the feed-forward parameters sit idle on any given token.&lt;/p&gt;

&lt;p&gt;The practical consequence: the "180B" headline counts total parameters, while the compute and the active memory traffic track the much smaller active set, on the order of 3B. Two numbers describe an MoE and they answer different questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Total parameters set the disk and RAM footprint you must store.&lt;/li&gt;
&lt;li&gt;Active parameters per token set the arithmetic you actually perform.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Quantization touches storage, so it operates on the total. Quality, however, is decided token by token on the active path. This gap is the whole reason MoE quantization is its own topic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does 4-bit hurt MoE differently than a dense model?
&lt;/h2&gt;

&lt;p&gt;Two forces pull in opposite directions.&lt;/p&gt;

&lt;p&gt;First, MoE is more forgiving in aggregate. Averaged over a corpus, errors introduced into rarely used experts barely move the output, because those experts rarely fire. A dense model has no such luxury: every weight is on the hot path for every token, so quantization error accumulates everywhere at once.&lt;/p&gt;

&lt;p&gt;Second, MoE is more fragile in specific places. The router is tiny but decisive. If quantization noise flips a routing decision, the token is suddenly served by a different expert, and the error is not a small perturbation but a discrete change of computation. Shared or "always-on" experts, the attention projections, and the embedding and output layers are also on the hot path for every token, so error there behaves like dense-model error.&lt;/p&gt;

&lt;p&gt;So the right mental model is not "MoE tolerates 4-bit." It is "MoE has a small hot path that must stay precise and a large cold path that can be squeezed hard." A flat 4-bit cast ignores that structure and spends the same few bits on the router as on a seldom-touched expert. That is where quality quietly leaks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does a UD-Q4_K_XL style mixed-precision graft preserve?
&lt;/h2&gt;

&lt;p&gt;Unsloth's "UD" (Unsloth Dynamic) GGUF variants, and similar mixed-precision schemes, treat the model as a graft rather than a uniform block. The idea is simple to state: keep the sensitive, always-active tensors at higher precision and push the rarely active expert tensors down to 4-bit. The &lt;code&gt;_K_XL&lt;/code&gt; suffix signals a K-quant layout with an extra-large share of high-precision tensors retained.&lt;/p&gt;

&lt;p&gt;In practice a graft like this tends to protect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The token embedding and the final output projection, which every token reads and writes.&lt;/li&gt;
&lt;li&gt;Attention query, key, value, and output projections, which run on every token in every layer.&lt;/li&gt;
&lt;li&gt;The router or gate weights, so routing decisions stay stable and tokens keep landing on the intended experts.&lt;/li&gt;
&lt;li&gt;Often the first and last transformer blocks, which empirically carry more than their share of the quality.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else, principally the large pool of expert feed-forward weights, goes to 4-bit. Because that pool is both the biggest contributor to file size and the least active per token, you capture most of the compression while leaving the hot path near full precision. The result is a file a fraction of the original size whose per-token behavior barely changed, which is the only kind of compression worth shipping to an edge device.&lt;/p&gt;

&lt;p&gt;The important caveat: which tensors you keep high is a model-specific choice, and vendors publish ready-made UD builds precisely so you do not have to guess. The reasoning above is the general principle, not a recipe to hand-tune.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you verify quality parity instead of trusting the quant?
&lt;/h2&gt;

&lt;p&gt;Here is the rule that matters: a quant is guilty until it passes a benchmark. Perplexity on a tiny sample and a few good-looking chat replies are not evidence. You need a discriminating, held-out benchmark scored identically on both the full-precision model and the quantized model, then you compare the scores.&lt;/p&gt;

&lt;p&gt;MMLU-Pro is a reasonable choice because it is broad, hard enough to separate models, and standard enough that the harness is well tested. The workflow is: build the quant, score the baseline, score the quant, and accept only if the gap is within benchmark noise.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: produce the quantized GGUF
&lt;/h3&gt;

&lt;p&gt;If you are starting from a full-precision GGUF and want to roll your own 4-bit file (rather than download a vendor UD build), &lt;code&gt;llama.cpp&lt;/code&gt; does the cast:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Convert a Hugging Face checkpoint to a full-precision GGUF first (if needed)&lt;/span&gt;
python llama.cpp/convert_hf_to_gguf.py ./model-full &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--outfile&lt;/span&gt; model-f16.gguf &lt;span class="nt"&gt;--outtype&lt;/span&gt; f16

&lt;span class="c"&gt;# Quantize to a 4-bit K-quant. Q4_K_M is the common balanced target.&lt;/span&gt;
./llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M

&lt;span class="c"&gt;# Inspect which tensors landed at which precision&lt;/span&gt;
./llama-gguf model-Q4_K_M.gguf | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-iE&lt;/span&gt; &lt;span class="s2"&gt;"type|ffn|attn|token_embd"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For MoE specifically, prefer a published mixed-precision build (UD-Q4_K_XL or equivalent) when one exists, because the per-tensor precision map is the hard part and it has already been tuned for that architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: score both models on the same benchmark
&lt;/h3&gt;

&lt;p&gt;Run the identical harness against the baseline and the quant. Keep every knob fixed: same prompt template, same number of shots, same decoding settings, same subset of questions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Baseline: full-precision model&lt;/span&gt;
llama-perplexity &lt;span class="nt"&gt;--multiple-choice&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-m&lt;/span&gt; model-f16.gguf &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-bf&lt;/span&gt; mmlu-pro.bin &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--ctx-size&lt;/span&gt; 4096 2&amp;gt;&amp;amp;1 | &lt;span class="nb"&gt;tee &lt;/span&gt;baseline.log

&lt;span class="c"&gt;# Candidate: 4-bit mixed graft&lt;/span&gt;
llama-perplexity &lt;span class="nt"&gt;--multiple-choice&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-m&lt;/span&gt; model-UD-Q4_K_XL.gguf &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-bf&lt;/span&gt; mmlu-pro.bin &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--ctx-size&lt;/span&gt; 4096 2&amp;gt;&amp;amp;1 | &lt;span class="nb"&gt;tee &lt;/span&gt;quant.log

&lt;span class="c"&gt;# Compare the final accuracy lines&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"final"&lt;/span&gt; baseline.log quant.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 3: accept or reject on the number
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;baseline (f16)        MMLU-Pro: 87.6x
UD-Q4_K_XL (4-bit)    MMLU-Pro: 87.65
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Equal inside noise is a pass. If instead the quant drops several points, do not ship it: either the graft protected the wrong tensors, the router drifted, or the target precision was too aggressive for this model. The benchmark told you something the demo never would.&lt;/p&gt;

&lt;p&gt;A few guardrails so the comparison stays honest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fix the random seed and the question subset so both runs see the same items.&lt;/li&gt;
&lt;li&gt;Report the margin, not just "it passed." A positive gap is not automatically a win if it is smaller than run-to-run variance.&lt;/li&gt;
&lt;li&gt;Run at least two passes per model. If the two baseline runs disagree by more than the baseline-to-quant gap, your measurement is too noisy to conclude anything.&lt;/li&gt;
&lt;li&gt;Beware caches and warm-up effects that help one run and not the other.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When is 4-bit MoE the right call, and when is it not?
&lt;/h2&gt;

&lt;p&gt;It shines when the total parameter count is what blocks you (you cannot fit the file in RAM) while the active path is small enough that the device can actually run it. That is the classic edge case: a big MoE that no consumer machine can hold at full precision becomes loadable at 4-bit because the cold experts shrink dramatically.&lt;/p&gt;

&lt;p&gt;It is the wrong call when the model is dense, when you need the last fraction of a point on a sensitive task, or when you have not budgeted time to verify. The compression is free; the confidence is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does a smaller GGUF always mean lower quality?&lt;/strong&gt;&lt;br&gt;
No. File size tracks total parameters, but MoE quality is decided on the small active path. A mixed-precision graft can shrink the file a lot while keeping the active path near full precision, so a much smaller file can score the same on a held-out benchmark.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between Q4_K_M and a UD-Q4_K_XL style build?&lt;/strong&gt;&lt;br&gt;
Q4_K_M applies a mostly uniform 4-bit K-quant. A UD-Q4_K_XL style build keeps more of the sensitive, always-active tensors (embeddings, attention, router) at higher precision and pushes only the rarely active experts to 4-bit, which protects quality on MoE at a similar overall size.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not just read the perplexity number?&lt;/strong&gt;&lt;br&gt;
Perplexity on a small sample is a weak proxy and can look fine while task accuracy slips. A discriminating multiple-choice benchmark like MMLU-Pro, scored identically on both models, is far harder to fool and tells you whether real task behavior survived.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How close is close enough to call it parity?&lt;/strong&gt;&lt;br&gt;
Within the benchmark's own run-to-run noise. Score each model at least twice. If the baseline-to-quant gap is smaller than the spread between two baseline runs, treat it as equal. Report the actual margin rather than a pass or fail label.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I just download a vendor quant instead of making my own?&lt;/strong&gt;&lt;br&gt;
Usually yes, and for MoE it is often the better choice, because the per-tensor precision map is model specific and tuning it is the hard part. Still run the verification step yourself: trust the build, but confirm the number on your own harness.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this apply to dense models too?&lt;/strong&gt;&lt;br&gt;
The verification workflow applies to everything. The sparsity argument does not: dense models run every weight on every token, so 4-bit error accumulates across the whole network and there is no large cold path to squeeze cheaply.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Running a 180B Model on a Laptop With No GPU: How 4-bit GGUF Keeps Full Accuracy&lt;/li&gt;
&lt;li&gt;Offline AI on a Phone: Thread Pinning, Lazy Loading, and a Safety Gate That Knows When to Stay Quiet&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>quantization</category>
      <category>edgeai</category>
    </item>
    <item>
      <title>Building an AI That Refuses to Say "Safe": Reliability Engineering for an Offline Disaster App</title>
      <dc:creator>GINIGEN AI</dc:creator>
      <pubDate>Tue, 06 Oct 2026 13:12:57 +0000</pubDate>
      <link>https://dev.to/ginigen_ai_010d9180cbbdb1/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline-disaster-app-4eel</link>
      <guid>https://dev.to/ginigen_ai_010d9180cbbdb1/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline-disaster-app-4eel</guid>
      <description>&lt;p&gt;Most AI reliability writing assumes a network. You call a bigger model to double-check, you log to a server, you push a hotfix when something goes wrong. A disaster app breaks every one of those assumptions. By the time someone opens it, the cell tower may be rubble and the phone may be in airplane mode for the next two days. There is no bigger model to call, no server to log to, and no hotfix that will arrive in time.&lt;/p&gt;

&lt;p&gt;This post is about the engineering we did for HeliGO, an offline disaster and distress app built around an on-device 4B model. The runtime side (thread pinning and lazy loading on a phone) was covered earlier in this series. Here I want to focus on the harder problem: how do you make a small model's answers trustworthy when a wrong answer can get someone killed and nobody is online to catch it?&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;HeliGO runs entirely on-device in airplane mode: maps, contour lines, altitude, compass, rescue coordinates, and an AI chat, with no network call at any point.&lt;/li&gt;
&lt;li&gt;The model is Edge-4B-TELL, built on Gemma-4 E4B, small enough to run on a phone chip.&lt;/li&gt;
&lt;li&gt;Safety is not a prompt. It is a 3-stage gate around the model: block fatal outputs, substitute official guidance text, or refuse.&lt;/li&gt;
&lt;li&gt;The gate is asymmetric on purpose. It never says a thing is safe to eat, but it always warns when something is poisonous.&lt;/li&gt;
&lt;li&gt;Confidence comes from reading the model's hidden state (the TELL technique), not from asking the model how sure it is, because asking does not work.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why "just prompt it to be careful" fails offline
&lt;/h2&gt;

&lt;p&gt;The obvious approach is a system prompt: "You are a disaster assistant. Never give dangerous advice. Say you are unsure when you are unsure." On a 4B model with no fallback, this is wishful thinking for two reasons.&lt;/p&gt;

&lt;p&gt;First, a prompt is a suggestion, not a constraint. A small model under distribution shift (a panicked, ungrammatical question about a mushroom in the dark) will still occasionally produce a confident, wrong, fatal answer. There is no second model downstream to catch it.&lt;/p&gt;

&lt;p&gt;Second, and this is the part that surprised us most, the model's own stated confidence is not just useless, it is inverted. When we asked the model to attach a confidence number to its answers, wrong answers tended to come with &lt;em&gt;higher&lt;/em&gt; stated confidence than right ones. Asking a model "are you sure?" and trusting the reply is worse than a coin flip in exactly the cases that matter.&lt;/p&gt;

&lt;p&gt;So the reliability layer cannot live inside the model's text output. It has to wrap the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 3-stage safety gate
&lt;/h2&gt;

&lt;p&gt;Every answer the model produces passes through three stages before it reaches the screen. Think of it as a decision pipeline, not a filter.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model draft
   |
   v
[1] fatal-output block   -- is this a category that can kill if wrong?
   |  (if yes and risky) -&amp;gt; suppress, escalate to stage 2
   v
[2] official-text substitute -- is there a matching MOIS guideline?
   |  (if yes) -&amp;gt; answer with the official text, not the model's words
   v
[3] refuse               -- still not confident enough?
   |  (if yes) -&amp;gt; say "I am not sure", do not guess
   v
answer shown to user
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Stage 1, fatal-output block.&lt;/strong&gt; Some answer categories are life-or-death: "yes, you can eat this", "yes, this water is safe", "yes, go that way". For these, the model's confidence is irrelevant. If the content falls in a fatal category and carries any risk signal, it is suppressed regardless of how sure the model sounds. This is the direct consequence of the inverted-confidence finding: in the categories where being wrong is fatal, we do not let the model's certainty vote at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage 2, official-text substitution.&lt;/strong&gt; We packaged 75 official response guidelines from Korea's Ministry of the Interior and Safety (MOIS) on the device. When a question maps to one of them (earthquake, flood, wildfire, first aid, and so on), the app answers with the official text verbatim instead of the model's paraphrase. The model is still useful here: it does the matching and retrieval. But the words the user reads are the authoritative source, not a 4B paraphrase that might drift.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage 3, refuse.&lt;/strong&gt; If an answer is neither safely substitutable nor confidently verifiable, the app declines. "I am not sure" is a valid and often correct output for a disaster tool. A refusal costs the user a little time. A confident wrong answer can cost a life. The gate is built so refusal is the default failure mode, never a guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  The asymmetry principle
&lt;/h2&gt;

&lt;p&gt;The gate is deliberately not symmetric, and this is the design decision I am most sure about.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The app never says "this is safe." It always says "this is poisonous" when it has reason to.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The reasoning is a straightforward expected-cost argument. Consider the two ways the app can be wrong about something edible:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;App says&lt;/th&gt;
&lt;th&gt;If app is wrong&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Edible plant&lt;/td&gt;
&lt;td&gt;"I cannot confirm this is safe"&lt;/td&gt;
&lt;td&gt;User stays hungry&lt;/td&gt;
&lt;td&gt;Low, recoverable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Poisonous plant&lt;/td&gt;
&lt;td&gt;"This may be poisonous"&lt;/td&gt;
&lt;td&gt;User avoids safe food&lt;/td&gt;
&lt;td&gt;Low, recoverable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Poisonous plant&lt;/td&gt;
&lt;td&gt;"This is safe to eat"&lt;/td&gt;
&lt;td&gt;User is poisoned&lt;/td&gt;
&lt;td&gt;Catastrophic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three of these outcomes are survivable. Only one is not. So the app is tuned to make the unsurvivable mistake nearly impossible, at the cost of being overly cautious in the harmless direction. The costs of silence are not equal on the two sides, so the gate is not equal either. We ship warnings for 12 venomous and poisonous species in the same spirit: a false alarm is cheap, a missed warning is not. Any edge safety system where errors have asymmetric cost should bias toward the cheap error.&lt;/p&gt;

&lt;h2&gt;
  
  
  TELL: reading the model instead of asking it
&lt;/h2&gt;

&lt;p&gt;Stages 1 and 3 both need a trustworthy confidence signal, and we already established that the model's spoken confidence is inverted. The fix is to stop asking.&lt;/p&gt;

&lt;p&gt;TELL reads the model's internal state while it produces an answer and derives a reliability estimate from that, rather than from any number the model writes in its text. The intuition: the computation a transformer performs on its way to an answer carries signal about whether that answer is well-grounded, and that signal is more honest than the self-report the model emits at the end. Internal state does not have an incentive to sound confident; the output text, apparently, does.&lt;/p&gt;

&lt;p&gt;I am not going to publish the thresholds or the exact internal signals we read, for the obvious reason that they are what makes the gate hard to game. The point for this article is architectural: the confidence check lives outside the token stream. We treat the model as an instrument to be read, not a witness to be interviewed.&lt;/p&gt;

&lt;p&gt;The TELL technique was released alongside the model. In HeliGO it currently backs the gate's suppress-or-refuse decisions, and a future update will surface low-confidence filtering more directly to the user.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually ships on the device
&lt;/h2&gt;

&lt;p&gt;For an offline tool, "on-device" has to be literal. Everything the gate and the map need is packaged into the app so it works with the radio off:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;4,998 terrain tiles with elevation data, and contour lines that redraw automatically by zoom range&lt;/li&gt;
&lt;li&gt;16,568 mountain summit elevations&lt;/li&gt;
&lt;li&gt;1,255 shelter, medical, water, and fire facilities, plus nearest-high-ground routing for floods&lt;/li&gt;
&lt;li&gt;75 MOIS official response guidelines (the stage-2 substitution corpus)&lt;/li&gt;
&lt;li&gt;12 venomous and poisonous species warnings&lt;/li&gt;
&lt;li&gt;The Edge-4B-TELL model itself, for offline chat, photo identification, and voice input&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this phones home. Location history in particular never leaves the device, which is both a privacy property and a reliability property: there is no network dependency to fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  H2: How do you make an offline AI trustworthy without a bigger model to check it?
&lt;/h2&gt;

&lt;p&gt;You move trust out of the model and into a deterministic wrapper. The model drafts and retrieves; a gate you control decides what reaches the user. Fatal categories are blocked on content, not on the model's confidence. Authoritative answers come from packaged official text. When neither applies, you refuse. The model never gets the final word on anything that can kill.&lt;/p&gt;

&lt;h2&gt;
  
  
  H2: Why not just use a larger model on-device?
&lt;/h2&gt;

&lt;p&gt;Because the phone is the budget. A 4B model is what runs in airplane mode on a consumer chip with room left for maps and rendering. The interesting claim of this project is not that a big model is safe, it is that a small model plus the right wrapper can be safe enough to ship for life-or-death use. The reliability came from architecture, not from parameter count.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does HeliGO send anything to a server?&lt;/strong&gt;&lt;br&gt;
No. Maps, contours, altitude, compass, rescue coordinates, and the AI chat all run on-device in airplane mode. Location history does not leave the phone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the model?&lt;/strong&gt;&lt;br&gt;
Edge-4B-TELL, built on Google's Gemma-4 E4B, published on Hugging Face. It handles offline chat, photo identification, and voice input on the phone's own chip.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does it never say something is safe to eat?&lt;/strong&gt;&lt;br&gt;
Because the cost of being wrong is asymmetric. A missed safe meal is hunger; a wrong "safe" is poisoning. The gate biases toward the recoverable error and only ever issues the poison warning side confidently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I trust the AI's confidence when it answers?&lt;/strong&gt;&lt;br&gt;
The app does not trust the model's stated confidence, and neither should you in general. Measured on this model, wrong answers tended to be stated more confidently than right ones. HeliGO derives confidence from the model's internal state (TELL) instead, outside the text it writes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this a replacement for calling emergency services?&lt;/strong&gt;&lt;br&gt;
No. HeliGO does not replace rescue services. If you have any signal, call emergency services first. It is a tool for when you have no connection.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways for edge builders
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Put the safety logic outside the model. A prompt is not a constraint; a deterministic gate is.&lt;/li&gt;
&lt;li&gt;Do not trust self-reported confidence. Measure whether it is calibrated before you rely on it. It may be inverted.&lt;/li&gt;
&lt;li&gt;Make your error handling asymmetric when the costs are asymmetric. Bias toward the cheap mistake.&lt;/li&gt;
&lt;li&gt;Package everything. For a true offline tool, every datum the app needs must live on the device.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Further reading (same account):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Offline AI on a Phone: Thread Pinning, Lazy Loading, and a Safety Gate That Knows When to Stay Quiet&lt;/li&gt;
&lt;li&gt;Running a 180B Model on a Laptop With No GPU: How 4-bit GGUF Keeps Full Accuracy&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>edgeai</category>
      <category>llm</category>
      <category>android</category>
    </item>
    <item>
      <title>Running a 180B Model on a Laptop With No GPU: How 4-bit GGUF Keeps Full Accuracy</title>
      <dc:creator>GINIGEN AI</dc:creator>
      <pubDate>Tue, 06 Oct 2026 00:42:56 +0000</pubDate>
      <link>https://dev.to/ginigen_ai_010d9180cbbdb1/running-a-180b-model-on-a-laptop-with-no-gpu-how-4-bit-gguf-keeps-full-accuracy-31c9</link>
      <guid>https://dev.to/ginigen_ai_010d9180cbbdb1/running-a-180b-model-on-a-laptop-with-no-gpu-how-4-bit-gguf-keeps-full-accuracy-31c9</guid>
      <description>&lt;p&gt;A frontier-class model used to mean a rack of data-center GPUs. That assumption is what this post takes apart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;POCKET-Darwin-180B-GGUF&lt;/strong&gt; is a 4-bit build of Darwin-180B-RSI, a 180-billion-parameter model, packaged so it runs &lt;strong&gt;without a GPU&lt;/strong&gt;. It ships as GGUF on Hugging Face and on ModelScope. The headline numbers, all measured in October 2026:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Size:&lt;/strong&gt; 111 GB (4-bit GGUF, 4 files), down from 360 GB (BF16, 131 files)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CPU only:&lt;/strong&gt; a single server CPU (16 threads) generates &lt;strong&gt;18.4 to 21.0 tokens/s&lt;/strong&gt;, peak memory 78.8 GB&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Laptop:&lt;/strong&gt; an RTX 5060 Laptop (8 GB VRAM) with 32 GB RAM runs it at &lt;strong&gt;4.17 tokens/s&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mini PC:&lt;/strong&gt; a 128 GB-RAM mini PC holds the whole model in memory, no GPU&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accuracy:&lt;/strong&gt; MMLU-Pro, 2,000 questions, paired comparison, &lt;strong&gt;87.65% original vs 87.65% quantized&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;A 180B mixture-of-experts model activates only about 3B parameters per token. Combine that sparsity with a 4-bit quantization that keeps the small fraction of the weights that actually changed during training at high precision, and you get a model that fits in laptop-class memory and still answers like the full model. You run it with llama.cpp in one command. This post explains why it works and how to reproduce it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does a 180B model fit on a laptop?
&lt;/h2&gt;

&lt;p&gt;Two things make this possible, and neither is magic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, the architecture is sparse.&lt;/strong&gt; Darwin-180B is a mixture-of-experts (MoE) model: 512 expert sub-networks, of which only 10 are selected per token. Of the 180B total parameters, roughly &lt;strong&gt;3B are active for any given token&lt;/strong&gt;. The dense parts (attention, routing, embeddings) run every step; the expert weights are read on demand. llama.cpp can memory-map the file and let the operating system page expert tensors in and out, so you do not need all 111 GB resident at once. On a 128 GB mini PC the whole thing sits in RAM; on a 32 GB laptop the OS pages against the GGUF file on disk, which is why the laptop is slower but still runs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, the quantization is selective.&lt;/strong&gt; The build starts from a community-verified 4-bit quantization of the base weights. On top of that, only the portion that the self-improvement training actually modified (about &lt;strong&gt;3% of total volume&lt;/strong&gt;) is kept at higher precision. You spend your bit budget where the signal is, not uniformly. That is the difference between "4-bit and lossy" and "4-bit and lossless on the benchmark."&lt;/p&gt;

&lt;h2&gt;
  
  
  Does 4-bit quantization hurt accuracy here?
&lt;/h2&gt;

&lt;p&gt;On the measured benchmark, no. The honest way to check quantization damage is a &lt;strong&gt;paired comparison&lt;/strong&gt;: same questions, same decoding, original versus quantized, item by item. On MMLU-Pro (2,000 questions) both the original and the 4-bit build scored &lt;strong&gt;87.65%&lt;/strong&gt;. Identical.&lt;/p&gt;

&lt;p&gt;That result is specific and worth reading carefully. It does not mean 4-bit is free in general. It means that for this model, on this benchmark, with a selective scheme that protects the trained delta, the drop is below the measurement's resolution. Quantization error concentrates in the weights that carry the most task-relevant signal, so protecting that 3% is what preserves the score.&lt;/p&gt;

&lt;p&gt;There is a second, subtler result. On SuperGPQA (1,000 graduate-level science questions, never used in training), the quantized POCKET build was compared against the same 4-bit quantization of its base model. POCKET scored &lt;strong&gt;61.55% vs 59.10%&lt;/strong&gt;, a 2.45-point gain that held as statistically significant, while using about &lt;strong&gt;13% fewer tokens per question&lt;/strong&gt;. The self-improvement training taught the model to commit to an answer instead of re-deriving it, and quantization preserved that behavior. A third-party developer reproduced this at the file level in a public thread.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you run POCKET-Darwin-180B-GGUF?
&lt;/h2&gt;

&lt;p&gt;You need a recent &lt;strong&gt;llama.cpp&lt;/strong&gt; build (b11048 or later), the four GGUF files (111 GB total), and enough memory to either hold or page the model. No cloud account, no GPU driver stack required for the CPU path.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Build llama.cpp (b11048 or later)&lt;/span&gt;
git clone https://github.com/ggml-org/llama.cpp
&lt;span class="nb"&gt;cd &lt;/span&gt;llama.cpp
cmake &lt;span class="nt"&gt;-B&lt;/span&gt; build
cmake &lt;span class="nt"&gt;--build&lt;/span&gt; build &lt;span class="nt"&gt;--config&lt;/span&gt; Release &lt;span class="nt"&gt;-j&lt;/span&gt;

&lt;span class="c"&gt;# 2. Download the GGUF (4 files, ~111 GB) from Hugging Face&lt;/span&gt;
&lt;span class="c"&gt;#    FINAL-Bench/POCKET-Darwin-180B-GGUF&lt;/span&gt;
&lt;span class="c"&gt;#    (any standard huggingface download works; place all parts&lt;/span&gt;
&lt;span class="c"&gt;#     in one directory so the loader finds the shards)&lt;/span&gt;

&lt;span class="c"&gt;# 3. Run, CPU only, 16 threads&lt;/span&gt;
./build/bin/llama-cli &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-m&lt;/span&gt; ./POCKET-Darwin-180B-Q4.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-t&lt;/span&gt; 16 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt; 8192 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"Explain mixture-of-experts routing in two sentences."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few practical notes from the measurements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Threads matter.&lt;/strong&gt; The server CPU result used 16 threads. Match &lt;code&gt;-t&lt;/code&gt; to your physical cores. Oversubscribing threads past your core count usually slows generation, not speeds it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory decides your path.&lt;/strong&gt; With 128 GB of RAM the model stays resident and you get the top of the range. With 32 GB the OS pages expert tensors from disk, so throughput drops to a few tokens per second. A fast NVMe SSD helps the paging path noticeably.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context length costs memory too.&lt;/strong&gt; The 78.8 GB figure is peak at a working context. Raising &lt;code&gt;-c&lt;/code&gt; raises the KV-cache footprint on top of the weights.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPU offload is optional.&lt;/strong&gt; On the 8 GB laptop GPU, only a fraction of layers fit in VRAM, so most of the work still lands on CPU and RAM. The 4.17 tok/s number reflects that mixed path.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When is a CPU-only 180B model actually the right call?
&lt;/h2&gt;

&lt;p&gt;Throughput is the honest tradeoff. 21 tok/s on a server CPU is fine for batch jobs, drafting, extraction, and agent steps that tolerate latency. 4 tok/s on a laptop is a developer convenience, not a chat product. If you need interactive speed at scale, a GPU still wins.&lt;/p&gt;

&lt;p&gt;Where the CPU path wins is &lt;strong&gt;where the data cannot leave the building&lt;/strong&gt;. An air-gapped server, a mini PC on a factory floor, a defense or finance or public-sector box with no outbound connection: there a model that runs entirely on local CPU and RAM, with weights you downloaded once, is not a compromise. It is the only option that clears the policy. Edge AI is often less about milliseconds and more about where the bytes are allowed to sit.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can I run this in LM Studio or Ollama?&lt;/strong&gt;&lt;br&gt;
This post only verifies the &lt;strong&gt;llama.cpp&lt;/strong&gt; path (b11048 or later), which is the engine GGUF targets. Treat other runners as untested here until you confirm they load this specific build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need the full 111 GB in RAM?&lt;/strong&gt;&lt;br&gt;
No. With 128 GB you can keep it resident for best speed. With 32 GB the OS memory-maps the GGUF and pages expert tensors on demand, trading throughput for a smaller footprint. You do need the full 111 GB on disk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is the laptop so much slower than the server CPU?&lt;/strong&gt;&lt;br&gt;
Memory. The 32 GB laptop cannot hold the model, so it pages from disk, and only a few layers fit on the 8 GB GPU. The 16-thread server CPU with ample RAM keeps far more of the working set hot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is 4-bit always this lossless?&lt;/strong&gt;&lt;br&gt;
No, and do not generalize it. The 87.65% match is a measured paired result for this model and benchmark, enabled by keeping the trained delta (about 3% of the weights) at higher precision. Always verify your own quantization with a paired comparison rather than assuming parity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What makes the active parameter count so low?&lt;/strong&gt;&lt;br&gt;
The MoE router picks 10 of 512 experts per token, so roughly 3B of the 180B parameters do work on any single token. That sparsity is what makes CPU inference and on-demand expert loading practical.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where do I get it?&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF&lt;/code&gt; and &lt;code&gt;modelscope.cn/models/FINAL-Bench/POCKET-Darwin-180B-GGUF&lt;/code&gt;. The original is &lt;code&gt;huggingface.co/FINAL-Bench/Darwin-180B-RSI&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Earlier in this series: &lt;em&gt;Offline AI on a Phone: Thread Pinning, Lazy Loading, and a Safety Gate That Knows When to Stay Quiet&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Measurements in this post are from October 2026 device testing. Quantization parity was checked with paired, same-question comparisons; treat quantization as model-specific and verify your own.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>edgeai</category>
      <category>quantization</category>
    </item>
    <item>
      <title>Offline AI on a Phone: Thread Pinning, Lazy Loading, and a Safety Gate That Knows When to Stay Quiet</title>
      <dc:creator>GINIGEN AI</dc:creator>
      <pubDate>Mon, 05 Oct 2026 08:56:19 +0000</pubDate>
      <link>https://dev.to/ginigen_ai_010d9180cbbdb1/offline-ai-on-a-phone-thread-pinning-lazy-loading-and-a-safety-gate-that-knows-when-to-stay-quiet-5a9k</link>
      <guid>https://dev.to/ginigen_ai_010d9180cbbdb1/offline-ai-on-a-phone-thread-pinning-lazy-loading-and-a-safety-gate-that-knows-when-to-stay-quiet-5a9k</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;HeliGO&lt;/strong&gt; is Ginigen's offline disaster and survival app. In airplane mode it still gives you maps, contour lines, altitude, a compass, rescue coordinates, and an AI chat, all running from data and a model stored on the phone.&lt;/li&gt;
&lt;li&gt;The model is &lt;strong&gt;Edge-4B-TELL&lt;/strong&gt;, a 4-bit GGUF build based on Google's Gemma-4 E4B, published on Hugging Face (1,526 downloads and 35 likes at the time of writing).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thread count is not a tuning detail on Android.&lt;/strong&gt; With the screen off, our process was allowed only 4 cores. Whisper with 6 threads took &lt;strong&gt;49.4 s&lt;/strong&gt;; with 4 threads it took &lt;strong&gt;2.1 s&lt;/strong&gt; on the same phone (Galaxy S25). Read &lt;code&gt;Cpus_allowed_list&lt;/code&gt; and size your thread pool from it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lazy-load rare features.&lt;/strong&gt; Making the 0.92 GB vision projector always resident pushed model load time from about &lt;strong&gt;10 s to about 1 minute&lt;/strong&gt;, and slowed plain text chat too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not trust a model's stated confidence.&lt;/strong&gt; Asked to self-report confidence, the model was more confident on wrong answers (AUROC 0.441, worse than a coin flip). HeliGO uses a three-stage safety gate, and our TELL signal reads the model's internal state instead of its opinion.&lt;/li&gt;
&lt;li&gt;Bonus from our sister company VIDRAFT: &lt;strong&gt;POCKET-Darwin-180B-GGUF&lt;/strong&gt; shows the same "quantize without losing quality" idea at 180B scale: 111.3 GB, MMLU-Pro 87.65% vs 87.65% for the BF16 original.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why build an AI app that never touches the network?
&lt;/h2&gt;

&lt;p&gt;In a disaster, the network is usually the first thing to go. Cell towers fail in an earthquake. A hiker who loses the trail in the mountains loses signal at the same time. Yet the information you need right then (where to evacuate, how to treat an injury, which plant or mushroom not to eat) mostly lives on the internet.&lt;/p&gt;

&lt;p&gt;HeliGO flips that assumption. Everything the app needs ships inside the phone. If you turn on airplane mode, nothing degrades. That is not a fallback mode; it is the only mode we design for.&lt;/p&gt;

&lt;p&gt;Here is what is stored on the device:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Records&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terrain tiles with elevation (auto-switching contour lines)&lt;/td&gt;
&lt;td&gt;4,998&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Summit elevations&lt;/td&gt;
&lt;td&gt;16,568&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shelters, medical, water supply, and fire facilities&lt;/td&gt;
&lt;td&gt;1,255&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Korea Ministry of the Interior and Safety (MOIS) public action guidelines&lt;/td&gt;
&lt;td&gt;75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Venomous species warnings&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On top of that: nearest high ground for flood situations, sunrise and sunset, moon phase, and breadcrumb backtracking along the path you walked.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26bkg%3Dwhite%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522Summit%2520elevations%2522%252C%2522Terrain%2520tiles%2522%252C%2522Facilities%2522%252C%2522MOIS%2520guidelines%2522%252C%2522Venomous%2520species%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Records%2520stored%2520on%2520device%2522%252C%2522data%2522%253A%255B16568%252C4998%252C1255%252C75%252C12%255D%252C%2522backgroundColor%2522%253A%2522%25232563eb%2522%257D%255D%257D%252C%2522options%2522%253A%257B%2522scales%2522%253A%257B%2522y%2522%253A%257B%2522type%2522%253A%2522logarithmic%2522%257D%257D%252C%2522plugins%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522HeliGO%2520offline%2520data%2520inventory%2520%2528log%2520scale%2529%2522%257D%257D%257D%257D" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26bkg%3Dwhite%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522Summit%2520elevations%2522%252C%2522Terrain%2520tiles%2522%252C%2522Facilities%2522%252C%2522MOIS%2520guidelines%2522%252C%2522Venomous%2520species%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Records%2520stored%2520on%2520device%2522%252C%2522data%2522%253A%255B16568%252C4998%252C1255%252C75%252C12%255D%252C%2522backgroundColor%2522%253A%2522%25232563eb%2522%257D%255D%257D%252C%2522options%2522%253A%257B%2522scales%2522%253A%257B%2522y%2522%253A%257B%2522type%2522%253A%2522logarithmic%2522%257D%257D%252C%2522plugins%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522HeliGO%2520offline%2520data%2520inventory%2520%2528log%2520scale%2529%2522%257D%257D%257D%257D" alt="HeliGO offline data inventory" width="1600" height="840"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The AI is the newest layer, and it is also the one that needs the most care. Maps do not hallucinate. Language models do.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you run an LLM fully offline on a phone?
&lt;/h2&gt;

&lt;p&gt;The short answer: a quantized GGUF model, llama.cpp compiled for ARM64, and a lot of discipline about memory and threads.&lt;/p&gt;

&lt;p&gt;Our stack for the on-device model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; Edge-4B-TELL. The inference weights are a verbatim mirror of Google's &lt;code&gt;gemma-4-E4B-it-qat-q4_0-gguf&lt;/code&gt;. We host our own copy so that a disaster app does not break if an upstream path moves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Files:&lt;/strong&gt; a 4.8 GB text-generation GGUF and a 0.92 GB vision and audio projector (mmproj).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; llama.cpp, running as a native process under the app.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speech:&lt;/strong&gt; whisper.cpp for voice input.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Quantization-aware training (QAT) at 4 bits is what makes a 4B-class multimodal model fit a phone at all. A GGUF file packs the weights and the metadata llama.cpp needs, so the app can load it straight from local storage with no conversion step.&lt;/p&gt;

&lt;p&gt;A generic way to try the same pattern on a development device (not our app's code, just the standard llama.cpp workflow):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Build llama.cpp for Android with the NDK (on your workstation)&lt;/span&gt;
cmake &lt;span class="nt"&gt;-B&lt;/span&gt; build-android &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DCMAKE_TOOLCHAIN_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$ANDROID_NDK&lt;/span&gt;/build/cmake/android.toolchain.cmake &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DANDROID_ABI&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;arm64-v8a &lt;span class="nt"&gt;-DANDROID_PLATFORM&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;android-28 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DGGML_OPENMP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;OFF
cmake &lt;span class="nt"&gt;--build&lt;/span&gt; build-android &lt;span class="nt"&gt;-j&lt;/span&gt; &lt;span class="nt"&gt;--target&lt;/span&gt; llama-server llama-cli

&lt;span class="c"&gt;# Push binary and model to the device&lt;/span&gt;
adb push build-android/bin/llama-server /data/local/tmp/
adb push gemma-4-E4B_q4_0-it.gguf /data/local/tmp/

&lt;span class="c"&gt;# Run fully offline; bind to localhost only&lt;/span&gt;
adb shell &lt;span class="s2"&gt;"cd /data/local/tmp &amp;amp;&amp;amp; ./llama-server &lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
  -m gemma-4-E4B_q4_0-it.gguf -c 4096 -t 4 &lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
  --host 127.0.0.1 --port 8080"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That flag &lt;code&gt;-t 4&lt;/code&gt; looks innocent. It is the single most important number in this article.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does thread count matter for on-device inference on Android?
&lt;/h2&gt;

&lt;p&gt;Here is the incident. Our voice pipeline ran whisper with 4 threads. The Galaxy S25 has more cores than that, so we raised it to 6, expecting a speedup.&lt;/p&gt;

&lt;p&gt;Translation latency jumped to about 56 seconds.&lt;/p&gt;

&lt;p&gt;The cause: when the phone screen is off, Android moves a foreground-service app and its native child processes into a restricted cpuset. On the S25 that set was &lt;strong&gt;4 cores&lt;/strong&gt;. ggml's thread pool uses spin-waiting, and when you run more worker threads than you have cores, threads that are spinning steal time from threads that are doing real work. The pool does not degrade gracefully. It collapses.&lt;/p&gt;

&lt;p&gt;Measured on the same phone, same 4 allowed cores, screen off:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;whisper threads&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-t 6&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49.4 s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-t 4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.1 s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is roughly a 23x difference from one integer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26bkg%3Dwhite%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522-t%25206%2520%2528more%2520threads%2520than%2520allowed%2520cores%2529%2522%252C%2522-t%25204%2520%2528matches%2520Cpus_allowed_list%2529%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Seconds%252C%2520lower%2520is%2520better%2522%252C%2522data%2522%253A%255B49.4%252C2.1%255D%252C%2522backgroundColor%2522%253A%255B%2522%2523dc2626%2522%252C%2522%252316a34a%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522plugins%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522whisper%2520on%2520Galaxy%2520S25%2520with%2520screen%2520off%2520%25284%2520cores%2520allowed%2529%2522%257D%257D%257D%257D" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26bkg%3Dwhite%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522-t%25206%2520%2528more%2520threads%2520than%2520allowed%2520cores%2529%2522%252C%2522-t%25204%2520%2528matches%2520Cpus_allowed_list%2529%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Seconds%252C%2520lower%2520is%2520better%2522%252C%2522data%2522%253A%255B49.4%252C2.1%255D%252C%2522backgroundColor%2522%253A%255B%2522%2523dc2626%2522%252C%2522%252316a34a%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522plugins%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522whisper%2520on%2520Galaxy%2520S25%2520with%2520screen%2520off%2520%25284%2520cores%2520allowed%2529%2522%257D%257D%257D%257D" alt="Thread count vs latency on Galaxy S25" width="1600" height="840"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Why our benchmarks missed it
&lt;/h3&gt;

&lt;p&gt;We benchmarked through &lt;code&gt;adb shell run-as &amp;lt;package&amp;gt;&lt;/code&gt;. A process started that way inherits the &lt;strong&gt;shell's&lt;/strong&gt; cpuset, which allows all cores. So the benchmark said 6 threads was fine. The real app, with the screen off, saw only 4. &lt;strong&gt;adb benchmarks hide this class of bug.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The fix: read the allowed cores at runtime
&lt;/h3&gt;

&lt;p&gt;Do not hardcode thread counts, and do not use the total core count. Ask the kernel what this process is actually allowed to use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Inside the app's own process context&lt;/span&gt;
&lt;span class="nb"&gt;grep &lt;/span&gt;Cpus_allowed_list /proc/self/status
&lt;span class="c"&gt;# Cpus_allowed_list:  0-1,4-5   (example: 4 cores)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A small helper that turns that into a thread count:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;allowedCoreCount&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nc"&gt;Int&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;line&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;java&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;File&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/proc/self/status"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;readLines&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;firstOrNull&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;it&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Cpus_allowed_list"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;?:&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;
    &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;spec&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;substringAfter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;":"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;","&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sumOf&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
        &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;r&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"-"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;it&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toInt&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="p"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;threads&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;minOf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;allowedCoreCount&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;// pass "-t $threads" to whisper / llama.cpp&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two practical rules we now follow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Compute threads from &lt;code&gt;Cpus_allowed_list&lt;/code&gt;, and cap them.&lt;/strong&gt; Re-check when the app moves between foreground and background if your workload can run in both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Benchmark under the real restriction.&lt;/strong&gt; Pin the process to the same cores the app will get (for example &lt;code&gt;taskset 33&lt;/code&gt; for cores 0, 1, 4, 5) and test with the screen off, not only from an adb shell.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Reproduce the restricted cpuset during a benchmark (mask 0x33 = cores 0,1,4,5)&lt;/span&gt;
adb shell &lt;span class="s2"&gt;"taskset 33 /data/local/tmp/whisper-cli -m ggml-base.bin -f test.wav -t 4"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why should rare features be lazy-loaded on mobile?
&lt;/h2&gt;

&lt;p&gt;The second lesson came from a fix that made things worse.&lt;/p&gt;

&lt;p&gt;The vision projector (the mmproj file, 0.92 GB) is what lets the model look at a photo. We had a bug in the photo feature, and while fixing it we made the projector &lt;strong&gt;always load&lt;/strong&gt; with the model. The photo feature worked again. But model load time went from about &lt;strong&gt;10 seconds to about 1 minute&lt;/strong&gt;, and ordinary text questions slowed down too.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26bkg%3Dwhite%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522Projector%2520lazy-loaded%2522%252C%2522Projector%2520always%2520loaded%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Model%2520load%2520time%2520%2528seconds%2529%2522%252C%2522data%2522%253A%255B10%252C60%255D%252C%2522backgroundColor%2522%253A%255B%2522%252316a34a%2522%252C%2522%2523dc2626%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522plugins%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522HeliGO%2520model%2520load%2520time%253A%2520always-on%25200.92%2520GB%2520vision%2520projector%2522%257D%257D%257D%257D" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26bkg%3Dwhite%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522Projector%2520lazy-loaded%2522%252C%2522Projector%2520always%2520loaded%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Model%2520load%2520time%2520%2528seconds%2529%2522%252C%2522data%2522%253A%255B10%252C60%255D%252C%2522backgroundColor%2522%253A%255B%2522%252316a34a%2522%252C%2522%2523dc2626%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522plugins%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522HeliGO%2520model%2520load%2520time%253A%2520always-on%25200.92%2520GB%2520vision%2520projector%2522%257D%257D%257D%257D" alt="Model load time: lazy vs always-loaded projector" width="1600" height="840"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What went wrong was not the code. It was the decision. We looked only at the broken feature and never measured the main path. Users chat with text all the time. They send a photo occasionally. "Keep it loaded so it is always ready" sounds safe, but it charges every user, every session, for a feature most sessions never touch.&lt;/p&gt;

&lt;p&gt;The rule we adopted:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Before adding anything to global initialization, ask &lt;strong&gt;how often it is used.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Rare features load &lt;strong&gt;when first used&lt;/strong&gt; (lazy loading), and can be released afterwards under memory pressure.&lt;/li&gt;
&lt;li&gt;Measure the &lt;strong&gt;main path&lt;/strong&gt; (cold start to first token for a text question) &lt;strong&gt;before and after&lt;/strong&gt; every change. If you did not measure it, you did not fix it.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Lazy pattern with llama.cpp: start text-only by default&lt;/span&gt;
./llama-server &lt;span class="nt"&gt;-m&lt;/span&gt; model.gguf &lt;span class="nt"&gt;-t&lt;/span&gt; &lt;span class="nv"&gt;$THREADS&lt;/span&gt;            &lt;span class="c"&gt;# fast cold start&lt;/span&gt;

&lt;span class="c"&gt;# Only when the user opens the camera feature, start (or switch to) a&lt;/span&gt;
&lt;span class="c"&gt;# vision-capable instance with the projector:&lt;/span&gt;
./llama-server &lt;span class="nt"&gt;-m&lt;/span&gt; model.gguf &lt;span class="nt"&gt;--mmproj&lt;/span&gt; mmproj.gguf &lt;span class="nt"&gt;-t&lt;/span&gt; &lt;span class="nv"&gt;$THREADS&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a phone, memory and load time are the product. Every megabyte you keep resident is a megabyte the OS can use as a reason to kill you in the background.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you make an on-device AI refuse safely?
&lt;/h2&gt;

&lt;p&gt;In a disaster app, a wrong answer can hurt someone. So HeliGO never passes the model's output straight to the user. Every answer goes through a three-stage safety gate:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Block life-threatening answers, regardless of confidence.&lt;/strong&gt; An answer like "this is safe to eat" is stopped even if the model sounds certain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer from official text when it applies.&lt;/strong&gt; If the question matches one of the 75 MOIS public action guidelines, the app answers with the guideline text instead of the model's paraphrase.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Abstain when uncertain.&lt;/strong&gt; If neither of the above settles it and the answer is not reliable, the app says it does not know.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  The asymmetry principle
&lt;/h3&gt;

&lt;p&gt;The gate is deliberately lopsided: &lt;strong&gt;it never says "safe", and it always says "poisonous".&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The two possible errors do not cost the same. If the app wrongly withholds "safe", you skip a meal and stay hungry. If it wrongly says "safe" about something venomous or toxic, someone can die. When silence and speech carry such different costs, the design should not treat them symmetrically. This is less a machine learning decision than a product ethics decision, and we think every safety-critical on-device AI needs one written down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you trust an LLM when it says how confident it is?
&lt;/h2&gt;

&lt;p&gt;No, and our measurements were sharper than we expected.&lt;/p&gt;

&lt;p&gt;We prompted the model to state its confidence with each answer on 665 Korean disaster-procedure questions. The average stated confidence was &lt;strong&gt;0.863&lt;/strong&gt;. When we used that number to separate correct from incorrect answers, the AUROC was &lt;strong&gt;0.441&lt;/strong&gt;. Anything below 0.5 means the signal points the wrong way: the model tended to sound &lt;strong&gt;more&lt;/strong&gt; certain on the answers it got &lt;strong&gt;wrong&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So "Are you sure?" is not a weak safety check. It is a misleading one.&lt;/p&gt;

&lt;p&gt;That is why we built &lt;strong&gt;TELL&lt;/strong&gt;. Instead of asking the model for its opinion, TELL reads the model's internal state during generation and estimates whether the answer is likely to be correct. The readout ships alongside the weights in the Edge-4B-TELL repository, and the public numbers on the model card (held-out cross-validation on the same 665 questions) are:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;AUROC&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Self-reported confidence&lt;/td&gt;
&lt;td&gt;0.441&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Surface features (length, digits, formatting)&lt;/td&gt;
&lt;td&gt;0.736&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TELL (internal-state readout)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.759&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;We publish the surface-feature baseline on purpose. TELL beats it by a real but modest margin, and we would rather show that honestly than claim a bigger win. Inside the app, the comparison that matters is TELL versus having no calibration signal at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does 4-bit quantization destroy model quality?
&lt;/h2&gt;

&lt;p&gt;Not if you are careful, and the evidence goes well beyond 4B.&lt;/p&gt;

&lt;p&gt;Our sister company &lt;strong&gt;VIDRAFT&lt;/strong&gt; recently released &lt;strong&gt;POCKET-Darwin-180B-GGUF&lt;/strong&gt;, a 4-bit GGUF build of the 180B mixture-of-experts model Darwin-180B-RSI. Public facts from its model card:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Size:&lt;/strong&gt; 111.3 GB in 4 files, down from 360 GB for the BF16 original.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Base layout:&lt;/strong&gt; built on the widely used UD-Q4_K_XL GGUF layout.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality:&lt;/strong&gt; MMLU-Pro, 2,000 paired questions: &lt;strong&gt;87.65% vs 87.65%&lt;/strong&gt; for the BF16 original.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shorter answers:&lt;/strong&gt; on the same 2,000 questions, answers averaged 3,694 tokens vs 4,322 for the same-format parent build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hardware:&lt;/strong&gt; runs with llama.cpp (b11048 or newer), including a gaming laptop with an 8 GB GPU and 32 GB RAM (4.17 tokens/s) and a CPU-only server (18 to 21 tokens/s).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sparsity:&lt;/strong&gt; about 3B active parameters per token, with experts read on demand from SSD through memory mapping.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The phone and the 180B laptop build share one idea: &lt;strong&gt;edge AI is a systems problem, not just a model problem.&lt;/strong&gt; Quantization makes the model fit. Threads, memory mapping, and load paths decide whether it is actually usable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is the practical checklist for shipping on-device LLMs?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Quantize with a format your runtime loads natively (GGUF for llama.cpp), and verify quality against the full-precision model on paired questions.&lt;/li&gt;
&lt;li&gt;Size thread pools from &lt;code&gt;/proc/self/status&lt;/code&gt; &lt;code&gt;Cpus_allowed_list&lt;/code&gt;, not from the core count.&lt;/li&gt;
&lt;li&gt;Benchmark with the screen off, under the real cpuset, not only through &lt;code&gt;adb shell run-as&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Lazy-load rare modalities (vision projectors, extra adapters). Measure cold start to first token before and after every change.&lt;/li&gt;
&lt;li&gt;Never ship raw model output in a safety-critical domain. Put a gate in front: block, defer to official text, or abstain.&lt;/li&gt;
&lt;li&gt;Do not use self-reported confidence as a safety signal. Validate any confidence signal against a simple surface baseline.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does HeliGO really work in airplane mode?&lt;/strong&gt;&lt;br&gt;
Yes. Maps, contours, altitude, compass, rescue coordinates, facility lookups, and AI chat all run from data and a model stored on the phone. HeliGO does not replace emergency services; it is a tool for when you cannot reach them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which model runs inside HeliGO?&lt;/strong&gt;&lt;br&gt;
Edge-4B-TELL on Hugging Face. The inference weights are an unmodified mirror of Google's Gemma-4 E4B QAT Q4_0 GGUF; Ginigen's contribution is the TELL calibration readout shipped with it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many threads should I give llama.cpp or whisper.cpp on Android?&lt;/strong&gt;&lt;br&gt;
No more than the number of cores in your process's &lt;code&gt;Cpus_allowed_list&lt;/code&gt;. On a Galaxy S25 with the screen off that was 4, and going to 6 made whisper about 23x slower.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not just ask the model if it is sure?&lt;/strong&gt;&lt;br&gt;
Because on our disaster-procedure questions the model's stated confidence ranked answers worse than chance (AUROC 0.441). Wrong answers came with higher stated confidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can POCKET-Darwin-180B-GGUF run without a GPU?&lt;/strong&gt;&lt;br&gt;
According to its model card, yes: 18 to 21 tokens/s on one server CPU socket with llama.cpp b11048 or newer, and it also runs on a laptop with an 8 GB GPU and 32 GB RAM.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does the safety gate never say "safe"?&lt;/strong&gt;&lt;br&gt;
Because the costs are asymmetric. Wrongly withholding "safe" costs a meal; wrongly saying "safe" can cost a life.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Ginigen: &lt;a href="https://www.ginigen.net" rel="noopener noreferrer"&gt;https://www.ginigen.net&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Edge-4B-TELL model: &lt;a href="https://huggingface.co/ginigen-ai/Edge-4B-TELL" rel="noopener noreferrer"&gt;https://huggingface.co/ginigen-ai/Edge-4B-TELL&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;POCKET-Darwin-180B-GGUF (VIDRAFT): &lt;a href="https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF" rel="noopener noreferrer"&gt;https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>android</category>
      <category>llm</category>
      <category>edgeai</category>
    </item>
  </channel>
</rss>
