DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring Android's MediaPipe LLM Inference API to a Streaming Compose UI

---
title: "Streaming On-Device LLMs to Compose with MediaPipe and Kotlin Flow"
published: true
description: "Wire MediaPipe's LLM Inference API to Jetpack Compose using Kotlin Flow — covering SharedFlow buffer sizing, JNI-to-coroutine dispatch, and context window compression for Pixel 8-class devices."
tags: kotlin, android, mobile, architecture
canonical_url: https://mvpfactory.co/blog/mediapipe-llm-compose-streaming
---
Enter fullscreen mode Exit fullscreen mode

What We Are Building

Let me show you a pattern I use in every on-device AI feature: a production-grade pipeline that takes raw token callbacks from MediaPipe's LLM Inference API and streams them into a reactive Compose UI without dropped tokens, UI jank, or silent prompt truncation.

MediaPipe gives you on-device Gemma and Phi-3 inference on Android. The hard part is not the inference — it is the plumbing between the native engine and your UI layer.


Prerequisites

  • Android project targeting API 26+
  • MediaPipe Tasks dependency added to your build
  • Basic familiarity with Kotlin coroutines and StateFlow

Step 1: Understand the Token's Journey

Every token MediaPipe generates travels across four boundaries:

  1. JNI callback → Kotlin lambda
  2. Kotlin lambda → SharedFlow emission
  3. SharedFlow → StateFlow accumulation in ViewModel
  4. Compose recomposition on the main thread

Each boundary is a potential data loss or jank point. Here is the gotcha that will save you hours: the callback originates from a JNI thread — not the main thread, not a coroutine dispatcher, but a raw native thread with no coroutine context.


Step 2: Fix the JNI Dispatcher Boundary

Most teams get this wrong on the first pass.

// ❌ Dangerous — emitting from unknown native thread
inferenceSession.generateAsync(prompt) { partialResult, _ ->
    _tokenFlow.tryEmit(partialResult) // silently drops tokens under load
}

// ✅ Correct — dispatch through a dedicated coroutine scope
inferenceSession.generateAsync(prompt) { partialResult, _ ->
    emissionScope.launch(Dispatchers.Default) {
        _tokenFlow.emit(partialResult) // suspends if buffer full
    }
}
Enter fullscreen mode Exit fullscreen mode

The distinction between tryEmit and emit is load-bearing. tryEmit returns false and drops the token if the buffer is full — with no exception, no log, no signal. On a Pixel 8 running Gemma 2B at roughly 15–20 tokens per second, you will see drops under any meaningful UI load if you rely on tryEmit with a default buffer.


Step 3: Size Your SharedFlow Buffer Correctly

Here is the minimal setup to get this working on capable hardware:

private val _tokenFlow = MutableSharedFlow<String>(
    replay = 0,
    extraBufferCapacity = 64,
    onBufferOverflow = BufferOverflow.SUSPEND
)
Enter fullscreen mode Exit fullscreen mode

The docs do not mention this, but buffer requirements vary significantly by device class:

Device Class Approx Tokens/sec Recommended Buffer Overflow Policy
Pixel 8 / Snapdragon 8 Gen 2 15–25 64 SUSPEND
Mid-range (Dimensity 700) 5–12 32 SUSPEND
Low-end (< 4 GB RAM) 2–6 16 DROP_OLDEST

Use SUSPEND on capable hardware so backpressure signals the emission scope to slow down. On constrained devices where inference is already the bottleneck, DROP_OLDEST prevents unbounded coroutine queue growth at the cost of occasional visual glitching — which is less harmful than an OOM.


Step 4: Accumulate to StateFlow in the ViewModel

Expose a StateFlow<String> to the UI layer, never the raw SharedFlow. Raw SharedFlow collection in Compose can produce redundant recompositions when multiple collectors exist.

val responseState: StateFlow<String> = _tokenFlow
    .runningFold("") { acc, token -> acc + token }
    .stateIn(viewModelScope, SharingStarted.Eagerly, "")
Enter fullscreen mode Exit fullscreen mode

One collectAsState call in your composable gives you stable, lifecycle-aware streaming output with no boilerplate.


Step 5: Handle the Context Window Ceiling

MediaPipe's LLM Inference API enforces a hard context window configured at model initialization — typically 1024 to 4096 tokens depending on model variant and available RAM. Exceeding this limit does not throw an exception. It silently truncates the prompt from the beginning.

Build prompt compression before you need it. Here is a rolling window strategy that takes an afternoon to implement:

fun compressHistory(history: List<Message>, maxTokens: Int): List<Message> {
    var tokenCount = estimateTokens(systemPrompt)
    return history.reversed()
        .takeWhile { msg ->
            tokenCount += estimateTokens(msg.content)
            tokenCount <= maxTokens * 0.85 // 15% safety margin
        }
        .reversed()
}
Enter fullscreen mode Exit fullscreen mode

The 15% safety margin is not cosmetic — character-based token estimation is approximate, and hitting the hard limit mid-generation produces corrupted partial output that is worse than a clean truncation.


Gotchas

  • Silent token drops with tryEmit — there is nothing to grep for. You will only notice when the UI output looks incomplete under load.
  • Wrong dispatcher on JNI callback — always dispatch through Dispatchers.Default before emitting. Do not assume you are on any known thread.
  • Exposing raw SharedFlow to Compose — multiple collectors trigger redundant recompositions. Always terminate at a StateFlow via runningFold.
  • Context window truncation — it is silent and it corrupts output. Retrofit a rolling window before you ship multi-turn conversations, not after user complaints.

Conclusion

On-device LLM inference on Android is production-ready. The gap between a working prototype and a shipping feature is deliberate engineering at three specific layers: the JNI dispatcher boundary, SharedFlow buffer configuration, and context window management. Each of these fails silently in ways that will not surface during early development.

Wire these correctly once and the pattern holds across every on-device AI feature you ship after it.

Top comments (0)