There is a quiet assumption in the on-device AI ecosystem: that the models shipping from Hugging Face and litert-community are good enough. Drop them into your app, write a good system prompt, and you are done.
I believed that too, until I benchmarked it.
This post is about what I learned when a strong prompt and an under-resourced fine-tune went head to head, why task-aligned fine-tuning still matters for serious on-device apps, and why the toolchain makes it hard to ship fine-tuned models to the hardware where they would matter most.
General models are general. Your app is not.
Gemma 4 E2B is a strong general-purpose model. Instruction-tuned, INT4-quantized, 2.59 GB - it runs comfortably on a phone GPU at about 25 tokens per second. For chatbot-style apps, it works out of the box.
But Redacto is not a chatbot. It is a zero-trust PII/PHI redaction engine with 7 specialized document categories: medical records, police reports, financial statements, journalism source notes, field service logs, legal filings, and a general fallback. Each category has its own detection rules. A patient's daughter's name is PHI in a medical record. A suspect's description is explicitly not PII in a police report - it needs to be preserved. These distinctions are domain-specific, and a general model does not know them by default.
Prompt engineering gets you surprisingly far. With carefully crafted system prompts, explicit preserve lists, and the structured [CATEGORY_N] output format, the standard Gemma 4 E2B model scores 80.5% overall on our 85-entry benchmark suite. That is real and usable.
But there are domains where prompt engineering hits a ceiling.
The numbers, and what they do and do not prove
We ran a head-to-head benchmark: standard Gemma 4 E2B versus a QLoRA fine-tuned variant, both on GPU, same device (Galaxy S25 Ultra), same 85-entry dataset across all 5 modes (the benchmark covers 5 of the 7 document categories; legal filings and the general fallback were not part of this run).
The overall scores tell one story - the fine-tuned model scored 70.3% versus the standard model's 80.5%, largely because of an output format mismatch (the fine-tuned model was trained on generic [REDACTED] labels, not Redacto's [CATEGORY_N] format). That is a training data problem, not a fine-tuning problem. In other words, a strong prompt on the stock model beat a quick, under-resourced fine-tune overall.
A caveat before the table: every number here comes from a single directional benchmarking session in May 2026 on one Galaxy S25 Ultra. No raw per-run logs were saved and the device is no longer available, so treat these as directional signals, not a rigorous multi-run study.
The per-mode entity recall tells a different story:
| Mode | Standard | Fine-tuned | Delta |
|---|---|---|---|
| TACTICAL | 63.7% | 76.8% | +13.1% |
| FIELD_SERVICE | 82.1% | 95.3% | +13.2% |
| FINANCIAL | 83.8% | 85.5% | +1.7% |
| JOURNALISM | 71.1% | 61.3% | -9.8% |
| HIPAA | 95.7% | 39.9% | -55.8% |
Look at TACTICAL. The standard model catches 63.7% of entities - victims, witnesses, minors - that should be redacted. The fine-tuned model catches 76.8%. That is a 13-point improvement.
For a police report, that 13% gap is the difference between protecting a victim's identity and exposing it. An app that misses 1 in 3 names is not a privacy tool. It is a liability.
FIELD_SERVICE shows the same pattern: +13.2%, driven by the fine-tuned model's ability to recognize contextual security information - "the key is under the mat," WiFi passwords embedded in natural language, gate codes in conversational text. These are not standard PII patterns. A general model does not look for them. A fine-tuned model does.
The HIPAA regression (-55.8%) and JOURNALISM regression (-9.8%) are instructive too. They show that fine-tuning on generic PII data actively hurts domains where the standard model's specialized prompt engineering is already strong. The standard model's HIPAA system prompt was specifically engineered for relational PHI - "the patient's daughter Lisa" - which the fine-tuned model's generic training data does not cover. The lesson is not "don't fine-tune." The lesson is "fine-tune on your app's actual task data, with your app's actual output format." If we retrained with Redacto's system prompts and [CATEGORY_N] labels, the gains in TACTICAL and FIELD_SERVICE would likely hold while the regressions would disappear. Both quantity and alignment of the training data matter here, and this run was short on both.
Consistency is the hidden killer
An 80% entity recall score sounds decent in a benchmark table. In production, it means 1 in 5 sensitive identifiers leaks through. Every time.
Users do not experience averages. They experience individual documents. A medical provider using Redacto to redact a discharge summary before faxing it to an insurance company does not care that the model works 80% of the time. They care that this document, right now, is clean.
This is why task-aligned fine-tuning matters for apps where accuracy matters. The jump from 63.7% to 76.8% on TACTICAL is not just a benchmark improvement - it is the difference between an app that is trustworthy and one that is not. And with proper task-specific training data, those numbers should climb further.
The toolchain works - up to a point
Here is the fine-tuning pipeline, step by step, with what actually happens at each stage:
Step 1: QLoRA on Colab. Works. We fine-tuned Gemma 4 E2B using QLoRA (LoRA rank 8, alpha 16, dropout 0.05) on 3,000 examples from the ai4privacy/pii-masking-400k dataset. Trainable parameters: 2,850,816 - just 0.06% of the model. Training time on an NVIDIA RTX PRO 6000: 217 seconds. Final loss: 4.3064. The tooling here is mature. peft, transformers, bitsandbytes - it all works.
It is worth being honest about what this fine-tune was: 3,000 of the roughly 400,000 samples in ai4privacy/pii-masking-400k, a single epoch, only a few minutes of actual training (about 217 seconds, though wrestling the toolchain end to end took the better part of a day), and trained on generic [REDACTED] labels rather than Redacto's [CATEGORY_N] output format. That is an under-resourced and partly misaligned run, not a serious training effort. So the results above are best read as "what a strong prompt did against a quick, under-resourced fine-tune," not as the ceiling of what fine-tuning can do. With the full dataset and task-aligned labels, the overall picture could flip.
Step 2: Export to .litertlm. Works, with a trap. The litert_torch.generative.export_hf tool converts the merged weights to a quantized .litertlm bundle. But there is an undocumented incompatibility: the Gemma 4 chat template from Hugging Face uses map.get() in Jinja, which LiteRT-LM's on-device template parser does not support. The error - Failed to apply template: unknown method: map has no method named get - will stop every developer who tries this pipeline. The fix: swap in Gemma 3's chat template before export. Not obvious. Not documented. But it works.
Step 3: GPU inference. This is where the story gets muddy. In our notes the fine-tuned .litertlm loaded on the Adreno 830 GPU and ran inference at 9.0 tokens per second (slower than the standard model's roughly 25 tok/s, which we attributed to the fine-tuned export being 4.7 GB versus the standard bundle's 2.59 GB - a quantization granularity difference, not a fundamental limitation). The fine-tuned export is 4.7 GB and GPU-only; it is too large and the wrong shape to compile for the NPU.
Step 4: NPU inference. Blocked.
The wall
The standard Gemma 4 E2B model runs on the Hexagon V79 NPU at about 42 tokens per second - roughly 1.7x faster than GPU. Because NPU decode is memory-bandwidth-bound, that rate stays roughly flat rather than scaling with much else, which is exactly what you would expect. Time to first token drops from 366ms to 92ms. For a redaction app processing sensitive documents, that speed difference is the gap between a noticeable pause and an instant response.
But the standard model was compiled for the NPU by Qualcomm's hardware team. The .litertlm file at litert-community contains a QNN-prepared bundle (QNN being Qualcomm AI Engine Direct) with custom DISPATCH_OP operations compiled specifically for the Hexagon V79 DSP on the Snapdragon 8 Elite. This compilation step - converting quantized weights into a QNN graph that the NPU can execute - requires the AIMET (AI Model Efficiency Toolkit) and an ahead-of-time (AoT) QNN toolchain.
That toolchain is not part of the public LiteRT-LM SDK.
Community projects have since demonstrated ahead-of-time QNN compilation for LLMs through their own toolchains, which suggests the gap is specific to the LiteRT-LM export path rather than QNN itself. For developers working inside the LiteRT-LM pipeline, though, the wall is the same.
This is not an engineering shortcut we missed. This is a hardware-team integration boundary. The tools required to compile a model into a QNN-prepared NPU bundle are internal to Qualcomm's AI stack. Developers can download the Qualcomm AI Engine Direct SDK and access QNN APIs for classical ML models, but the specific pipeline for compiling LLM-class models - with their attention mechanisms, KV caches, and multi-head architectures - into Hexagon-optimized graphs is not publicly available for independent use on arbitrary fine-tuned weights.
The net result, at least in our attempt: a fine-tuned Gemma 4 E2B that showed a double-digit TACTICAL entity-recall gain in a directional benchmark, but that we could not compile for the NPU. The gap between "I improved my model" and "my users benefit from it on the fastest hardware" is hard to bridge today without vendor support.
What this means for every developer
This is not a Redacto-specific problem. Every developer who needs domain-specific accuracy is likely to hit this wall in the same sequence:
- Ship the standard model. It works.
- Discover domain-specific accuracy gaps. They exist.
- Fine-tune. It can help, if the data is aligned.
- Try to deploy the fine-tuned model to NPU. Blocked.
The result: fine-tuning becomes a GPU-only optimization. The penalty is not just speed, either. Our fine-tuned model drew roughly 3x the current of the standard model on GPU (301 mA vs 101 mA) - a system-wide reading, so directional rather than precise, but the direction matters for battery life too. And the entire value proposition of the NPU - purpose-built silicon that runs inference about 1.7x faster than GPU (and several times faster than CPU), and by widely reported figures at lower power - stays tied to the standard models that were pre-compiled for it.
This is less a villain than two reasonable design choices colliding. Ahead-of-time NPU compilation lives inside the vendor's AI stack, which was not built with third-party fine-tunes in mind; and the open export path targets generic INT4, not a QNN graph. The practical effect is still awkward: the developers who invest in accuracy by fine-tuning end up on the slower runtime, while those who ship the stock model get the faster one. That is an outcome to fix, not a conspiracy to blame.
What the ecosystem needs
The fix is conceptually simple: a public, documented pipeline from fine-tuned Hugging Face weights to NPU-compiled .litertlm.
In concrete terms:
A public QNN AoT compilation path for LLMs. Today, Qualcomm's QNN SDK supports classical model compilation. Extending this to LLM architectures - or providing a hosted compilation service - would let developers compile their fine-tuned weights for Hexagon without needing internal toolchain access.
AIMET integration with the LiteRT-LM export pipeline. AIMET handles quantization-aware training and post-training quantization for Qualcomm targets. If litert_torch.generative.export_hf could target QNN as a backend (not just generic INT4), the export step would produce NPU-ready bundles directly.
A compilation service, if open-sourcing the toolchain is not feasible. Upload merged weights, receive a QNN-compiled .litertlm. This is how some cloud TPU compilation workflows operate - the developer never touches the compiler directly, but they get hardware-optimized artifacts back.
Documentation of the compilation requirements. Even knowing what the pipeline needs - AIMET version, QNN SDK version, target SoC specification, supported quantization schemes, memory layout constraints - would let the community start building bridges.
The bigger picture
On-device AI is at an inflection point. The models are good enough. The hardware is fast enough. The runtimes are stable enough. The missing piece is the development loop: train, evaluate, compile, deploy, measure, iterate.
Right now, that loop is broken at the compile step for anyone who fine-tunes. And task-aligned fine-tuning is often what closes the last domain-specific accuracy gap. Which means the loop is broken for exactly the developers who are building the most demanding applications.
The standard Gemma 4 E2B on NPU at about 42 tokens per second with 92ms time-to-first-token is genuinely impressive hardware. Getting fine-tuned models onto that hardware - with the kind of domain-specific gains we saw in the benchmark - is the difference between on-device AI as a demo and on-device AI as a product.
The models are ready. The silicon is ready. The remaining gap is a public path through the toolchain.
Related in this series of "Edge AI from the Trenches"
- Prompt Engineering Beat My Fine-Tuned Model. Here Is Why. - the decision framework for when fine-tuning is the right lever to pull
- My Fine-Tuned Gemma 4 Loaded Fine, Then Broke on the First Message - the undocumented export failure that blocks fine-tuned models before they even reach the NPU wall
- I Opened a .litertlm File. Here Is What Is Actually in There. - why the compiled bundle is the end of the pipeline, and what gets sealed at export time
Jaydeep Shah is a developer with roots in embedded systems, Android platform internals, and silicon-level AI optimization. He now explores on-device AI inference - bringing models from the cloud to phones and edge hardware. Along with his team Edge Artists, he builds applications using LiteRT-LM and Gemma models on mobile hardware, and writes about what works, what breaks, and what he learns along the way. This post is part of the Edge AI from the Trenches series.
Last updated: Oct 2026
22nd of 23 posts in the "Edge AI from the Trenches" series
Top comments (0)