DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring Android's Jetpack Compose to a Quantized On-Device LLM for Real-Time Code Completion

---
title: "On-Device Code Completion in Jetpack Compose: Hitting the 200ms Latency Budget"
published: true
description: "Wire CodeGemma 2B INT4 to a Compose editor with StateFlow token streaming, debounced triggers, and a sub-200ms first-token latency target on Pixel 9."
tags: kotlin, android, mobile, architecture
canonical_url: https://mvpfactory.co/blog/on-device-llm-compose-latency-budget
---

## What We Are Building

By the end of this tutorial you will have a working on-device code completion pipeline in Jetpack Compose: a quantized CodeGemma 2B INT4 model feeding tokens into a `StateFlow`, consumed by a Compose editor overlay, gated behind a debounced trigger that prevents inference storms. The hard target is **p95 first-token latency under 200ms** on Pixel 9-class hardware. That is the line between a feature that feels fluid and one users ignore.

## Prerequisites

- Android project targeting API 31+
- Pixel 9 or Snapdragon 8 Gen 3 / Tensor G4 device for profiling
- TFLite runtime with GPU delegate dependency
- Familiarity with Kotlin coroutines and `StateFlow`
- CodeGemma 2B INT4 exported via TFLite (roughly 1.1 GB on disk)

---

## Step 1 — Choose Your Model (and Why INT4 Is Non-Negotiable)

Let me show you a pattern I use in every project: profile the model *before* writing any app code.

| Model | Size (INT4) | p50 First Token (Pixel 9) | Code MBPP Pass@1 |
|---|---|---|---|
| CodeGemma 2B INT4 | ~1.1 GB | ~120ms | Competitive for 2B class |
| Generic 3B INT8 | ~3.0 GB | ~280ms | Lower code-specific accuracy |
| 7B INT4 | ~4.2 GB | >500ms | Out-of-budget for real-time |

INT8 doubles your memory footprint and pushes you past the latency ceiling. INT4 is the only configuration that fits flagship VRAM headroom while leaving room for the host app.

## Step 2 — Build the Streaming Pipeline with StateFlow

Token streaming from a TFLite inference loop maps naturally onto a `StateFlow`. Each decoded token emits a partial buffer update; your Compose editor observes it with `collectAsState()`.

Enter fullscreen mode Exit fullscreen mode


kotlin
// ViewModel
private val _completionTokens = MutableStateFlow("")
val completionTokens: StateFlow = _completionTokens.asStateFlow()

fun streamCompletion(prompt: String) {
viewModelScope.launch(Dispatchers.Default) {
_completionTokens.value = ""
inferenceEngine.streamTokens(prompt).collect { token ->
_completionTokens.update { it + token }
}
}
}


Wire this to your editor overlay in Compose:

Enter fullscreen mode Exit fullscreen mode


kotlin
val completion by viewModel.completionTokens.collectAsState()
// Render completion as a ghost-text overlay anchored to cursor position


`StateFlow` conflation is the underrated win here. If your UI frame drops, you process the latest accumulated buffer — never a stale intermediate state.

## Step 3 — Gate Inference With Debounce

Here is the gotcha that will save you hours: firing inference on every keystroke is catastrophic on-device. At 15–20% CPU per inference run, an unbounded trigger will crater your frame rate within seconds.

The production trigger model:

1. Debounce cursor idle — 150–250ms after the last keystroke
2. Gate on syntactic signal — `.`, `(`, space after a keyword, or newline
3. Cancel in-flight inference on any new keystroke via `Job.cancel()`

Enter fullscreen mode Exit fullscreen mode


kotlin
editorState
.onEach { cancelCurrentInference() }
.debounce(180)
.filter { isTriggerContext(it.cursorContext) }
.collectLatest { state ->
viewModel.streamCompletion(state.buildPrompt())
}


`collectLatest` handles cancellation automatically — any new emission cancels the previous coroutine, which propagates into the inference loop.

## Step 4 — Pick the Right Delegate

The docs do not mention this, but naive "always prefer GPU" breaks on mid-range devices with silent fallback latency spikes. Use this decision tree:

Enter fullscreen mode Exit fullscreen mode


plaintext
Is device flagship-tier (Snapdragon 8 Gen 3 / Tensor G4+)?
├─ YES → Try GPU Delegate → if init < 2s, proceed
│ └─ FAIL → Fall back to NNAPI with INT8 cast
└─ NO → Try NNAPI → benchmark on first run
└─ p50 > 350ms → fall back to CPU (disable feature)


Instrument delegate init time at startup, cache the result in `SharedPreferences`, and skip fallback detection on subsequent launches. Cold delegate initialization is a one-time penalty — paying it on every inference is an architecture bug.

## Step 5 — Decouple Editor State From TextFieldValue

Maintain a separate `EditorState` data class tracking cursor offset, visible line range, and last accepted completion. This decouples inference triggers from raw `TextFieldValue` and gives you deterministic test cases for trigger logic without a running model.

---

## Gotchas

- **Skipping early profiling.** Set the 200ms first-token budget before writing inference code. If the model misses it in a synthetic benchmark, no Compose optimization will recover it.
- **Not caching delegate selection.** Re-running delegate detection on every cold start adds 1–2 seconds of invisible latency users will blame on something else.
- **Missing `collectLatest`.** Using plain `collect` instead of `collectLatest` means stale inference jobs accumulate. The cancellation semantics are built in — use them.

---

## Conclusion

Here is the minimal setup to get this working: CodeGemma 2B INT4 for the model, `StateFlow` for the streaming backbone, 180ms debounce with `collectLatest` for storm prevention, and cached delegate selection for startup perf. Get all three right and on-device completion is genuinely competitive with cloud-hosted alternatives — and it works on a plane.

**Further reading:** [TFLite GPU delegate docs](https://www.tensorflow.org/lite/performance/gpu) · [CodeGemma model card](https://ai.google.dev/gemma/docs/codegemma) · [Kotlin StateFlow reference](https://kotlinlang.org/api/kotlinx.coroutines/kotlinx-coroutines-core/kotlinx.coroutines.flow/-state-flow/)
Enter fullscreen mode Exit fullscreen mode

Top comments (0)