DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring Android's CameraX to a Quantized On-Device OCR Model for Real-Time Document Intelligence

---
title: "Real-Time On-Device OCR with CameraX: Staying Under 18ms"
published: true
description: "Wire Android CameraX to a quantized CRAFT+CRNN OCR pipeline with NNAPI delegation and region batching — staying under 18ms on mid-range devices."
tags: [kotlin, android, mobile, architecture]
canonical_url: https://blog.mvp-factory.dev/camerax-ocr-under-18ms
---
Enter fullscreen mode Exit fullscreen mode

What We Are Building

By the end of this tutorial you will have a real-time on-device OCR pipeline that stays under 18ms on mid-range Android hardware. We are wiring CameraX to a two-stage quantized model: CRAFT for text region detection and CRNN for character recognition. The naive implementation blows your latency budget in the first 50ms. I will show you the three decisions that keep it from doing that.


Prerequisites

  • Android Studio Hedgehog or later
  • TensorFlow Lite with NNAPI and GPU delegate dependencies
  • A mid-range test device (Snapdragon 680-class is the benchmark target)
  • Familiarity with Kotlin coroutines and CameraX ImageAnalysis

Step 1: Pick the Right Delegate Per Model, Not Per App

Here is the pattern I use in every project: benchmark delegate selection independently for each model stage.

CRAFT is convolutional and spatial — it maps well to GPU parallelism. CRNN is recurrent-heavy — it benefits from NNAPI on devices where the DSP vendor driver is mature.

Model Preferred delegate Reason
CRAFT (detection) GPU delegate Dense convolutions, spatial pooling
CRNN (recognition) NNAPI (DSP path) Sequential ops, lower memory bandwidth
CRNN fallback CPU (XNNPACK) NNAPI driver instability on some OEMs

Do not assume. Build a runtime delegate probe that runs a warm-up inference pass and selects based on measured latency:

val options = Interpreter.Options().apply {
    val nnApiDelegate = NnApiDelegate()
    addDelegate(nnApiDelegate)
    setNumThreads(2)
}
// Probe: run 3 warm-up passes, measure median
Enter fullscreen mode Exit fullscreen mode

NNAPI performance varies significantly across OEM driver implementations. Measure on your target device tier before committing.


Step 2: Batch Your Region Proposals

This is the gotcha that will save you hours. If you run CRNN inference once per detected text region, you pay interpreter initialization and memory transfer overhead on every single call. On a dense document that means 20–40 serial inference calls per frame.

The fix is to batch all region proposals from CRAFT into a single CRNN inference pass. Pad or resize all candidate crops to a fixed input height of 32px, stack them into a batch tensor, and run one forward pass:

val batchTensor = Array(regions.size) { i ->
    preprocessCrop(regions[i], targetHeight = 32)
}
crnnInterpreter.runForMultipleInputsOutputs(
    arrayOf(batchTensor), outputMap
)
Enter fullscreen mode Exit fullscreen mode

On a Snapdragon 680 device with batches of 10–20 regions, batching reduces recognition latency by roughly 60–70% versus serial calls. This is not a micro-optimization. It is the difference between a usable demo and something you would actually ship.


Step 3: Wire the CameraX Frame Pipeline

CameraX ImageAnalysis delivers frames at the camera's native rate — typically 30fps. Use STRATEGY_KEEP_ONLY_LATEST and gate with an AtomicBoolean:

val isProcessing = AtomicBoolean(false)

val imageAnalysis = ImageAnalysis.Builder()
    .setBackpressureStrategy(ImageAnalysis.STRATEGY_KEEP_ONLY_LATEST)
    .setOutputImageFormat(ImageAnalysis.OUTPUT_IMAGE_FORMAT_YUV_420_888)
    .build()
    .also { analysis ->
        analysis.setAnalyzer(executor) { imageProxy ->
            if (!isProcessing.compareAndSet(false, true)) {
                imageProxy.close()
                return@setAnalyzer
            }
            scope.launch(Dispatchers.Default) {
                try {
                    runOcrPipeline(imageProxy)
                } finally {
                    isProcessing.set(false)
                    imageProxy.close()
                }
            }
        }
    }
Enter fullscreen mode Exit fullscreen mode

The docs do not mention this, but STRATEGY_KEEP_ONLY_LATEST alone is not enough. Without the compareAndSet gate you can still dispatch a second coroutine before the first completes if the executor has available threads. The backpressure strategy handles the queue. The atomic flag handles in-flight concurrency. You need both.

Use YUV_420_888 and convert only the Y-plane for CRAFT input. Skipping the chroma planes cuts preprocessing cost by roughly 40% on the same benchmark device.


Latency Budget Breakdown

Stage Target
YUV → grayscale crop ~1ms
CRAFT detection (GPU) ~8ms
Region proposal batching ~1ms
CRNN recognition (NNAPI) ~6ms
Result post-processing ~1ms
Total ~17ms

Gotchas

Per-region serial inference kills you silently. The latency looks acceptable in unit tests against single crops. It only becomes catastrophic when a real camera frame lands with 15 detected regions.

NNAPI driver quality is OEM-specific. A device that benchmarks beautifully on your desk may perform worse on a different OEM's firmware. Always include the CPU/XNNPACK fallback path for CRNN.

Frame drop is not failure. STRATEGY_KEEP_ONLY_LATEST is doing its job when it drops frames. Fighting it with a larger queue only builds up backpressure that delays your results without improving accuracy.


Conclusion

Hitting 18ms end-to-end on mid-range hardware is achievable, but only if every layer pulls in the same direction. Benchmark delegate selection per model stage, not globally. Batch your CRAFT region proposals into a single CRNN pass. Trust STRATEGY_KEEP_ONLY_LATEST and guard in-flight concurrency with an AtomicBoolean. Get all three right and your pipeline stays inside the budget.

Here is the minimal setup to get this working — the rest is tuning for your specific document domain.

Top comments (0)