DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring Android's CameraX to a Quantized Depth-Estimation Model for Real-Time Spatial Audio

---
title: "CameraX + Quantized Depth AI: The Frame Pipeline That Stays Under 22ms"
published: true
description: "Wire Android CameraX to a quantized MiDaS-small INT8 model and binaural audio rendering. Here is the exact frame pipeline that stays under 22ms on mid-range Android devices."
tags: [android, kotlin, mobile, architecture]
canonical_url: https://blog.mvp-factory.dev/camerax-depth-ai-spatial-audio-22ms
---
Enter fullscreen mode Exit fullscreen mode

What We Are Building

Today I am going to show you a pattern I use when building real-time spatial awareness apps on Android. We are wiring CameraX YUV frame extraction to a quantized MiDaS-small INT8 model running on the GPU delegate, then mapping the depth output to binaural audio parameters via AAudio — all under a 22ms end-to-end budget on mid-range hardware.

Miss that budget consistently and you get drift between visual and audio cues. On mid-range devices that is not a UX preference — it is a functional failure.

Prerequisites:

  • Android project targeting API 26+
  • TensorFlow Lite with GPU delegate dependency
  • CameraX 1.3.x or later
  • A quantized MiDaS-small INT8 .tflite model (~4MB)

The 22ms Budget — Where Every Millisecond Goes

Before touching code, let me show you the budget breakdown. Teams that skip this benchmark in isolation, ship to a real device, and wonder why their audio drifts.

Stage Target Overage Risk
CameraX YUV capture + callback 3ms Low
YUV → RGB + resize to 256×256 4ms Medium (CPU path)
MiDaS-small INT8 inference (GPU) 8ms High on older GPUs
Depth map → stereo position params 2ms Low
AAudio parameter update 2ms Low
Total 19ms 3ms headroom

Design for headroom, not the happy path. Thermal throttling on mid-range devices can push GPU inference from 8ms to 14ms under sustained load.


Step 1 — CameraX Frame Extraction

Here is the minimal setup to get this working. Use ImageAnalysis with STRATEGY_KEEP_ONLY_LATEST. This is non-negotiable.

val analysis = ImageAnalysis.Builder()
    .setTargetResolution(Size(640, 480))
    .setBackpressureStrategy(ImageAnalysis.STRATEGY_KEEP_ONLY_LATEST)
    .setOutputImageFormat(ImageAnalysis.OUTPUT_IMAGE_FORMAT_YUV_420_888)
    .build()

analysis.setAnalyzer(inferenceExecutor) { imageProxy ->
    processFrame(imageProxy) // owns close()
}
Enter fullscreen mode Exit fullscreen mode

YUV_420_888 avoids an extra GPU copy compared to RGBA. For the resize step, do it on the GPU — CameraX's built-in effect pipeline or a Vulkan compute shader. A CPU bicubic resize at 640×480 → 256×256 will consistently blow your 4ms budget.


Step 2 — MiDaS-small INT8 via GPU Delegate

INT8 quantization cuts MiDaS-small from ~13MB to ~4MB and yields roughly 1.8× throughput on Adreno and Mali GPUs.

val gpuDelegate = GpuDelegate(
    GpuDelegate.Options().apply {
        setPrecisionLossAllowed(true) // FP16 fallback on unsupported ops
        setQuantizedModelsAllowed(true)
    }
)

val options = Interpreter.Options().addDelegate(gpuDelegate)
val interpreter = Interpreter(loadModelFile(), options)
Enter fullscreen mode Exit fullscreen mode

The docs do not mention this clearly, but setPrecisionLossAllowed(true) enables FP16 fallback for ops the GPU delegate cannot handle in INT8. The depth artifacts are imperceptible to the downstream audio stage.


Step 3 — Depth Map to Spatial Audio Parameters

fun depthMapToAzimuth(depthMap: FloatArray, width: Int, height: Int): Float {
    val threshold = depthMap.max()!! * 0.9f
    var weightedX = 0f; var totalWeight = 0f
    depthMap.forEachIndexed { i, v ->
        if (v >= threshold) { weightedX += (i % width) * v; totalWeight += v }
    }
    val normalizedX = (weightedX / totalWeight) / width
    return (normalizedX - 0.5f) * 180f // -90 to +90 degrees
}
Enter fullscreen mode Exit fullscreen mode

Feed azimuth and a depth-derived distance parameter into AAudio with an HRTF convolution effect. AAudio's lower-latency path gives you 6–10ms vs. 12–20ms for audio parameter propagation compared to OpenSL ES.


Gotchas

Here is the gotcha that will save you hours: MiDaS output is inverse relative depth, not metric distance. Closer objects have higher values, but the scale is relative. Most teams expect metric and build broken mapping logic. For spatial audio positioning you only need relative values — so lean into it.

Three more before you ship:

  1. Drop frames deliberately. STRATEGY_KEEP_ONLY_LATEST is your primary latency defense. An analyzer that processes every frame causes ANRs on mid-range devices within minutes of sustained use.

  2. Benchmark on thermal-stressed hardware. Run your inference loop for 10 minutes before measuring. A model that hits 8ms cold will often hit 14ms hot. Design around the steady-state number.

  3. Prefer GPU delegate over NNAPI. NNAPI performance variance across OEM drivers makes it a poor default for production builds. GPU delegate with setPrecisionLossAllowed(true) consistently outperforms NNAPI on mid-range Snapdragon and MediaTek when you account for driver variance.


Wrapping Up

The pipeline is: CameraX YUV → INT8 MiDaS via GPU delegate → depth centroid mapping → AAudio binaural update. With 3ms headroom baked in, this holds under 22ms on mid-range hardware even under thermal load.

Further reading:

Top comments (0)