<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: SoftwareDevs mvpfactory.io</title>
    <description>The latest articles on DEV Community by SoftwareDevs mvpfactory.io (@software_mvp-factory).</description>
    <link>https://dev.to/software_mvp-factory</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3790305%2F141f30ba-972f-4b17-9b03-c77343f2747d.png</url>
      <title>DEV Community: SoftwareDevs mvpfactory.io</title>
      <link>https://dev.to/software_mvp-factory</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/software_mvp-factory"/>
    <language>en</language>
    <item>
      <title>Wiring Android's CameraX to a Quantized On-Device OCR Model for Real-Time Document Intelligence</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Thu, 08 Oct 2026 08:39:23 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-androids-camerax-to-a-quantized-on-device-ocr-model-for-real-time-document-intelligence-1fl6</link>
      <guid>https://dev.to/software_mvp-factory/wiring-androids-camerax-to-a-quantized-on-device-ocr-model-for-real-time-document-intelligence-1fl6</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Real-Time&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;On-Device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;OCR&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CameraX:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Staying&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Under&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;18ms"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wire&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Android&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CameraX&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;quantized&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CRAFT+CRNN&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;OCR&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pipeline&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;NNAPI&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;delegation&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;region&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;batching&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;—&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;staying&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;under&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;18ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;mid-range&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;devices."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;kotlin&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;android&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;mobile&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;architecture&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://blog.mvp-factory.dev/camerax-ocr-under-18ms&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What We Are Building
&lt;/h2&gt;

&lt;p&gt;By the end of this tutorial you will have a real-time on-device OCR pipeline that stays under 18ms on mid-range Android hardware. We are wiring CameraX to a two-stage quantized model: CRAFT for text region detection and CRNN for character recognition. The naive implementation blows your latency budget in the first 50ms. I will show you the three decisions that keep it from doing that.&lt;/p&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Android Studio Hedgehog or later&lt;/li&gt;
&lt;li&gt;TensorFlow Lite with NNAPI and GPU delegate dependencies&lt;/li&gt;
&lt;li&gt;A mid-range test device (Snapdragon 680-class is the benchmark target)&lt;/li&gt;
&lt;li&gt;Familiarity with Kotlin coroutines and CameraX &lt;code&gt;ImageAnalysis&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Step 1: Pick the Right Delegate Per Model, Not Per App
&lt;/h2&gt;

&lt;p&gt;Here is the pattern I use in every project: benchmark delegate selection independently for each model stage.&lt;/p&gt;

&lt;p&gt;CRAFT is convolutional and spatial — it maps well to GPU parallelism. CRNN is recurrent-heavy — it benefits from NNAPI on devices where the DSP vendor driver is mature.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Preferred delegate&lt;/th&gt;
&lt;th&gt;Reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CRAFT (detection)&lt;/td&gt;
&lt;td&gt;GPU delegate&lt;/td&gt;
&lt;td&gt;Dense convolutions, spatial pooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CRNN (recognition)&lt;/td&gt;
&lt;td&gt;NNAPI (DSP path)&lt;/td&gt;
&lt;td&gt;Sequential ops, lower memory bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CRNN fallback&lt;/td&gt;
&lt;td&gt;CPU (XNNPACK)&lt;/td&gt;
&lt;td&gt;NNAPI driver instability on some OEMs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Do not assume. Build a runtime delegate probe that runs a warm-up inference pass and selects based on measured latency:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;options&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Interpreter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Options&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;apply&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;nnApiDelegate&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;NnApiDelegate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;addDelegate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nnApiDelegate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;setNumThreads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;// Probe: run 3 warm-up passes, measure median&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NNAPI performance varies significantly across OEM driver implementations. Measure on your target device tier before committing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 2: Batch Your Region Proposals
&lt;/h2&gt;

&lt;p&gt;This is the gotcha that will save you hours. If you run CRNN inference once per detected text region, you pay interpreter initialization and memory transfer overhead on every single call. On a dense document that means 20–40 serial inference calls per frame.&lt;/p&gt;

&lt;p&gt;The fix is to batch all region proposals from CRAFT into a single CRNN inference pass. Pad or resize all candidate crops to a fixed input height of 32px, stack them into a batch tensor, and run one forward pass:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;batchTensor&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;regions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
    &lt;span class="nf"&gt;preprocessCrop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;regions&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;targetHeight&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;crnnInterpreter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;runForMultipleInputsOutputs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nf"&gt;arrayOf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batchTensor&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;outputMap&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a Snapdragon 680 device with batches of 10–20 regions, batching reduces recognition latency by roughly 60–70% versus serial calls. This is not a micro-optimization. It is the difference between a usable demo and something you would actually ship.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3: Wire the CameraX Frame Pipeline
&lt;/h2&gt;

&lt;p&gt;CameraX &lt;code&gt;ImageAnalysis&lt;/code&gt; delivers frames at the camera's native rate — typically 30fps. Use &lt;code&gt;STRATEGY_KEEP_ONLY_LATEST&lt;/code&gt; and gate with an &lt;code&gt;AtomicBoolean&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;isProcessing&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;AtomicBoolean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;imageAnalysis&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ImageAnalysis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Builder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setBackpressureStrategy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ImageAnalysis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;STRATEGY_KEEP_ONLY_LATEST&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setOutputImageFormat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ImageAnalysis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;OUTPUT_IMAGE_FORMAT_YUV_420_888&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;build&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;also&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;analysis&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
        &lt;span class="n"&gt;analysis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setAnalyzer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;imageProxy&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(!&lt;/span&gt;&lt;span class="n"&gt;isProcessing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compareAndSet&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="n"&gt;imageProxy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="nd"&gt;@setAnalyzer&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="n"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;launch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Dispatchers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Default&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="nf"&gt;runOcrPipeline&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;imageProxy&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="n"&gt;isProcessing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                    &lt;span class="n"&gt;imageProxy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The docs do not mention this, but &lt;code&gt;STRATEGY_KEEP_ONLY_LATEST&lt;/code&gt; alone is not enough. Without the &lt;code&gt;compareAndSet&lt;/code&gt; gate you can still dispatch a second coroutine before the first completes if the executor has available threads. The backpressure strategy handles the queue. The atomic flag handles in-flight concurrency. You need both.&lt;/p&gt;

&lt;p&gt;Use &lt;code&gt;YUV_420_888&lt;/code&gt; and convert only the Y-plane for CRAFT input. Skipping the chroma planes cuts preprocessing cost by roughly 40% on the same benchmark device.&lt;/p&gt;




&lt;h2&gt;
  
  
  Latency Budget Breakdown
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;YUV → grayscale crop&lt;/td&gt;
&lt;td&gt;~1ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CRAFT detection (GPU)&lt;/td&gt;
&lt;td&gt;~8ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Region proposal batching&lt;/td&gt;
&lt;td&gt;~1ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CRNN recognition (NNAPI)&lt;/td&gt;
&lt;td&gt;~6ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result post-processing&lt;/td&gt;
&lt;td&gt;~1ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~17ms&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Per-region serial inference kills you silently.&lt;/strong&gt; The latency looks acceptable in unit tests against single crops. It only becomes catastrophic when a real camera frame lands with 15 detected regions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NNAPI driver quality is OEM-specific.&lt;/strong&gt; A device that benchmarks beautifully on your desk may perform worse on a different OEM's firmware. Always include the CPU/XNNPACK fallback path for CRNN.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frame drop is not failure.&lt;/strong&gt; &lt;code&gt;STRATEGY_KEEP_ONLY_LATEST&lt;/code&gt; is doing its job when it drops frames. Fighting it with a larger queue only builds up backpressure that delays your results without improving accuracy.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Hitting 18ms end-to-end on mid-range hardware is achievable, but only if every layer pulls in the same direction. Benchmark delegate selection per model stage, not globally. Batch your CRAFT region proposals into a single CRNN pass. Trust &lt;code&gt;STRATEGY_KEEP_ONLY_LATEST&lt;/code&gt; and guard in-flight concurrency with an &lt;code&gt;AtomicBoolean&lt;/code&gt;. Get all three right and your pipeline stays inside the budget.&lt;/p&gt;

&lt;p&gt;Here is the minimal setup to get this working — the rest is tuning for your specific document domain.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>PostgreSQL Connection Pooling at Scale</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Wed, 07 Oct 2026 14:55:27 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/postgresql-connection-pooling-at-scale-19bb</link>
      <guid>https://dev.to/software_mvp-factory/postgresql-connection-pooling-at-scale-19bb</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PostgreSQL&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Connection&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Pooling&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;at&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Scale:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;PgBouncer&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Transaction&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Mode&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Deep&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Dive"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Master&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;PgBouncer&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;transaction&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pooling,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fix&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;prepared&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;statement&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;errors,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;right-size&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;your&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pool&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;proven&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;formula&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;stop&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;connection&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;exhaustion&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;high-concurrency&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;backends."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgresql, architecture, devops, api&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/pgbouncer-transaction-pooling-deep-dive&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gu"&gt;## What We Will Cover&lt;/span&gt;

By the end of this workshop you will understand why PostgreSQL's default &lt;span class="sb"&gt;`max_connections = 100`&lt;/span&gt; will kill a medium-traffic API before it ever touches CPU limits, how PgBouncer's transaction mode works and why it breaks prepared statements by design, and how to size your pool correctly the first time using a formula that holds up in production.

Let me show you a pattern I use in every backend project that expects real traffic.
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Prerequisites&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; A running PostgreSQL instance (13+)
&lt;span class="p"&gt;-&lt;/span&gt; PgBouncer installed (&lt;span class="sb"&gt;`apt install pgbouncer`&lt;/span&gt; or equivalent)
&lt;span class="p"&gt;-&lt;/span&gt; Basic familiarity with connection pooling concepts
&lt;span class="p"&gt;-&lt;/span&gt; Your application driver ready: pgjdbc (Kotlin/Java), asyncpg (Python), or psycopg2/3
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## The Problem That Paged Me at 2am&lt;/span&gt;

Each PostgreSQL connection consumes roughly 5–10 MB of RAM. At 500 open connections you are burning up to 5 GB in overhead before a single query runs. A traffic spike hits, Postgres reaches its connection ceiling, new connections queue, timeouts cascade. You know the rest.

The fix is transaction-mode pooling in PgBouncer. Here is the minimal setup to get this working.
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Step 1 — Pick the Right Pooling Mode&lt;/span&gt;

PgBouncer offers three modes. For modern APIs and mobile backends, transaction mode is what you want.

| Mode | Prepared Statements | Use Case |
|---|---|---|
| Session | ✅ Supported | Legacy apps, long sessions |
| Transaction | ❌ Broken by design | High-concurrency APIs |
| Statement | ❌ Broken | Rarely recommended |

Transaction mode lets dozens of application threads share a much smaller pool of actual Postgres connections. The tradeoff is prepared statements — and this is where most teams burn time.
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Step 2 — Fix Prepared Statements at the Driver Level&lt;/span&gt;

PostgreSQL scopes prepared statements to a session. In transaction mode, your logical session maps to &lt;span class="ge"&gt;*different*&lt;/span&gt; physical backend connections across transactions. When your ORM runs &lt;span class="sb"&gt;`EXECUTE stmt_abc`&lt;/span&gt;, that statement was prepared on connection #3, but PgBouncer handed you connection #7. Postgres returns &lt;span class="sb"&gt;`ERROR: prepared statement "stmt_abc" does not exist`&lt;/span&gt;.

The docs do not make this obvious, but the fix is always at the driver level — prevent the extended query protocol from issuing &lt;span class="sb"&gt;`Parse`&lt;/span&gt;/&lt;span class="sb"&gt;`Bind`&lt;/span&gt; commands that create named server-side prepared statements.

&lt;span class="gs"&gt;**Kotlin/Java with HikariCP + pgjdbc:**&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
val config = HikariConfig().apply {&lt;br&gt;
    jdbcUrl = "jdbc:postgresql://pgbouncer:5432/mydb?preferQueryMode=simple"&lt;br&gt;
    maximumPoolSize = 10&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
**Python with asyncpg:**
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
conn = await asyncpg.connect(&lt;br&gt;
    "postgresql://user:pass@pgbouncer:5432/db",&lt;br&gt;
    prepared_statement_cache_size=0&lt;br&gt;
)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
**psycopg2 (libpq-based):**
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;/p&gt;
&lt;h1&gt;
  
  
  For full safety, use psycopg3 with prepared=False per statement
&lt;/h1&gt;

&lt;p&gt;conn = psycopg2.connect(dsn, options="-c standard_conforming_strings=on")&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
Do this before any migration, not after the errors start appearing.

---

## Step 3 — Size Your Pool With the Formula That Actually Works

This formula from the HikariCP team has held up across production deployments:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
pool_size = (core_count * 2) + effective_spindle_count&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
For a 4-core server with SSDs (`effective_spindle_count ≈ 1`):

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
pool_size = (4 * 2) + 1 = 9&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
This feels too small. Every team pushes back. Here is the gotcha that will save you hours of misguided tuning: PostgreSQL is I/O bound on reads and CPU bound on complex queries. Beyond a threshold, you are not adding parallelism — you are adding context-switching overhead and lock contention. A pool of 9–10 on a 4-core machine outperforms a pool of 100 under sustained load.

Your `pgbouncer.ini`:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
ini&lt;br&gt;
[pgbouncer]&lt;br&gt;
pool_mode = transaction&lt;br&gt;
max_client_conn = 1000&lt;br&gt;
default_pool_size = 9&lt;br&gt;
min_pool_size = 2&lt;br&gt;
reserve_pool_size = 2&lt;br&gt;
reserve_pool_timeout = 3&lt;br&gt;
server_idle_timeout = 600&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
---

## Step 4 — Choose the Right Pooler for Your Scale

| Feature | PgBouncer | pgpool-II | Odyssey |
|---|---|---|---|
| Transaction mode | ✅ Excellent | ✅ Supported | ✅ Excellent |
| Read replica routing | ❌ No | ✅ Built-in | ⚠️ Limited |
| Protocol overhead | Very low | Moderate | Low |
| Operational complexity | Low | High | Medium |

Use PgBouncer for lightweight, high-performance pooling with minimal operational surface area. Use pgpool-II when you need read/write splitting and can absorb the configuration work. Odyssey from Yandex is worth evaluating at very high connection counts — its multi-threaded model handles scale that PgBouncer's single-threaded architecture cannot match above ~50k clients.

---

## Gotchas

**`server_reset_query` in transaction mode.** In session mode, PgBouncer runs `DISCARD ALL` between clients. In transaction mode this is skipped — but a custom `server_reset_query` will still fire and break your session state assumptions.

**`max_client_conn` lower than your thread count.** If your application has 200 threads but `max_client_conn = 100`, the remaining 100 threads block, timeout, and retry. That retry storm amplifies load exactly when you can least afford it.

**No client-side `connect_timeout`.** Without it, blocked connections queue indefinitely. Set `connect_timeout=3000ms` at the pool level, not just at the Postgres level.

---

## Conclusion

Most connection storms are self-inflicted through misconfiguration, not load. Audit your ORM's prepared statement behavior first, apply the pool size formula, then deploy PgBouncer with monitoring on `SHOW POOLS` and `SHOW STATS`. Pool saturation should trigger alerts before clients ever see errors.

**Resources:**
- [PgBouncer documentation](https://www.pgbouncer.org/config.html)
- [HikariCP pool sizing analysis](https://github.com/brettwooldridge/HikariCP/wiki/About-Pool-Sizing)
- [pgjdbc `preferQueryMode` reference](https://jdbc.postgresql.org/documentation/use/)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring PostgreSQL LISTEN/NOTIFY to a Kotlin Multiplatform Mobile Client</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Wed, 07 Oct 2026 08:51:19 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-postgresql-listennotify-to-a-kotlin-multiplatform-mobile-client-1mfg</link>
      <guid>https://dev.to/software_mvp-factory/wiring-postgresql-listennotify-to-a-kotlin-multiplatform-mobile-client-1mfg</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wiring&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;PostgreSQL&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;LISTEN/NOTIFY&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;KMP&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Mobile&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Client:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Real-Time&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Push&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Without&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;WebSockets"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Map&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pg_notify&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;channels&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;KMP&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;SharedFlows&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;real-time&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;mobile&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;push.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Zero&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;broker&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;overhead,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;sub-100ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;latency,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;reconnect&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;strategy&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;that&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;survives&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;flaky&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;networks."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kotlin, android, mobile, postgresql&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/postgres-listen-notify-kmp-sharedflow&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gu"&gt;## What We Will Build&lt;/span&gt;

By the end of this tutorial, you will have a working real-time push pipeline from a PostgreSQL trigger through a Kotlin Multiplatform &lt;span class="sb"&gt;`SharedFlow`&lt;/span&gt; to your Android and iOS clients — with no WebSockets, no polling, and no message broker added to your stack.

Let me show you a pattern I use in every project that needs real-time updates before the infrastructure bill gets out of hand.

&lt;span class="gu"&gt;## Prerequisites&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; PostgreSQL 14+ (though this works back to 6.4)
&lt;span class="p"&gt;-&lt;/span&gt; A KMP project targeting Android and iOS
&lt;span class="p"&gt;-&lt;/span&gt; JDBC datasource on the JVM backend (HikariCP or similar)
&lt;span class="p"&gt;-&lt;/span&gt; Familiarity with Kotlin coroutines and &lt;span class="sb"&gt;`Flow`&lt;/span&gt;
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Step 1: Set Up the Database Trigger&lt;/span&gt;

Here is the minimal setup to get this working on the Postgres side. This trigger fires on every insert into an &lt;span class="sb"&gt;`events`&lt;/span&gt; table and broadcasts the new row as JSON:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
sql&lt;br&gt;
CREATE OR REPLACE FUNCTION notify_event_insert()&lt;br&gt;
RETURNS trigger AS $$&lt;br&gt;
BEGIN&lt;br&gt;
  PERFORM pg_notify(&lt;br&gt;
    'event_channel',&lt;br&gt;
    row_to_json(NEW)::text&lt;br&gt;
  );&lt;br&gt;
  RETURN NEW;&lt;br&gt;
END;&lt;br&gt;
$$ LANGUAGE plpgsql;&lt;/p&gt;

&lt;p&gt;CREATE TRIGGER on_event_insert&lt;br&gt;
AFTER INSERT ON events&lt;br&gt;
FOR EACH ROW EXECUTE FUNCTION notify_event_insert();&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
Any client subscribed to `event_channel` receives the payload the moment the transaction commits. No polling interval, no separate write path.

---

## Step 2: Map the Channel to a KMP SharedFlow

A `SharedFlow` fits naturally here — it is hot, it multicasts, and it survives subscriber churn without dropping the upstream connection. In `commonMain`:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
class PgNotifyChannel(private val channelName: String) {&lt;br&gt;
    private val _events = MutableSharedFlow(&lt;br&gt;
        replay = 0,&lt;br&gt;
        extraBufferCapacity = 64,&lt;br&gt;
        onBufferOverflow = BufferOverflow.DROP_OLDEST&lt;br&gt;
    )&lt;br&gt;
    val events: SharedFlow = _events.asSharedFlow()&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;suspend fun emit(payload: String) = _events.emit(payload)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;p&gt;}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
`DROP_OLDEST` is a deliberate choice here. A slow subscriber should never back-pressure the notification pipeline. Size `extraBufferCapacity` to your expected burst volume, not the average.

---

## Step 3: Hold the JDBC LISTEN Connection

On the JVM backend, a dedicated coroutine owns the connection and forwards payloads into the flow:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
fun listenOnChannel(scope: CoroutineScope, channel: PgNotifyChannel) {&lt;br&gt;
    scope.launch(Dispatchers.IO + SupervisorJob()) {&lt;br&gt;
        val conn = dataSource.connection&lt;br&gt;
        conn.createStatement().execute("LISTEN ${channel.channelName}")&lt;br&gt;
        val pgConn = conn.unwrap(PGConnection::class.java)&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    while (isActive) {
        val notifications = pgConn.getNotifications(1000) ?: continue
        notifications.forEach { channel.emit(it.parameter) }
    }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;p&gt;}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
`SupervisorJob()` ensures a failure here does not cascade to the parent scope. For a long-lived connection managed alongside other application work, this matters.

---

## Step 4: Reconnect Without Thundering-Herd

The docs do not mention this, but a persistent JDBC connection in `LISTEN` mode will die silently under network interruptions, Postgres restarts, or load balancer idle timeouts. Here is the gotcha that will save you hours: a blind retry loop causes a thundering-herd problem on outage recovery. Use exponential backoff with full jitter instead:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
suspend fun withReconnect(block: suspend () -&amp;gt; Unit) {&lt;br&gt;
    var attempt = 0&lt;br&gt;
    while (true) {&lt;br&gt;
        try {&lt;br&gt;
            block()&lt;br&gt;
            attempt = 0&lt;br&gt;
        } catch (e: Exception) {&lt;br&gt;
            val delay = minOf(30_000L, (500L shl attempt)) + Random.nextLong(500)&lt;br&gt;
            delay(delay)&lt;br&gt;
            attempt++&lt;br&gt;
        }&lt;br&gt;
    }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
On Android, gate reconnection attempts on `NetworkCallback`. On iOS, use `NWPathMonitor`. Expose both through `expect`/`actual` declarations and skip the backoff loop entirely while the device is offline — there is no point cycling through attempts when the path is not satisfied.

---

## Gotchas

**Buffer overflow is silent by default.** `DROP_OLDEST` prevents back-pressure but means a slow consumer loses events. Instrument your `SharedFlow` buffer fill level under load before going to production.

**Connection count is the real ceiling.** LISTEN/NOTIFY is latency-competitive with a WebSocket layer (~50–150ms) and beats polling (0–5s) while adding zero infrastructure. Redis Pub/Sub wins on raw latency (~5–20ms), but that tradeoff only matters once listener counts stress Postgres connection limits — typically in the tens of thousands. Benchmark your actual listener count before deploying a broker.

**`unwrap` is driver-specific.** `conn.unwrap(PGConnection::class.java)` works with the `pgjdbc` driver. If you are using R2DBC or another driver, the notification polling API differs.

---

## Conclusion

Most teams reach for a message broker before they have tested whether Postgres already does the job. `LISTEN/NOTIFY` has been in Postgres since version 6.4, costs nothing to run, and delivers latency that mobile users will not notice. Taking breaks to actually think this through helps too — I keep HealthyDesk running during long architecture sessions for exactly that reason.

The three things worth getting right: benchmark `pg_notify` before deploying a broker, supervise the `LISTEN` connection with `SupervisorJob()` and gate reconnection on device connectivity, and use `SharedFlow` with `DROP_OLDEST` sized to burst volume — not average load.

**Resources:**
- [PostgreSQL NOTIFY docs](https://www.postgresql.org/docs/current/sql-notify.html)
- [KMP SharedFlow docs](https://kotlinlang.org/api/kotlinx.coroutines/kotlinx-coroutines-core/kotlinx.coroutines.flow/-shared-flow/)
- [pgjdbc PGConnection API](https://jdbc.postgresql.org/documentation/publicapi/org/postgresql/PGConnection.html)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Kotlin Coroutines Structured Concurrency Under the Hood</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Tue, 06 Oct 2026 14:28:59 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/kotlin-coroutines-structured-concurrency-under-the-hood-4c1p</link>
      <guid>https://dev.to/software_mvp-factory/kotlin-coroutines-structured-concurrency-under-the-hood-4c1p</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Kotlin&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Coroutines:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Job&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Trees,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Cancellation&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Propagation,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;SupervisorJob&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Boundary"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Learn&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;how&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Kotlin&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;coroutines&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Job&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;hierarchy&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;works&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;at&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;runtime,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;how&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;cancellation&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;propagates&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;up&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;down&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tree,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;why&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;SupervisorJob&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;boundary&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;that&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;prevents&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;cascade&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;failures&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;production&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;apps."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kotlin, android, architecture, mobile&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/kotlin-coroutines-job-trees-cancellation-supervisorjob&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What We Are Building
&lt;/h2&gt;

&lt;p&gt;By the end of this tutorial you will understand exactly how the &lt;code&gt;Job&lt;/code&gt; tree works at runtime, why a single unhandled exception can silently kill your entire &lt;code&gt;ViewModel&lt;/code&gt;, and how &lt;code&gt;SupervisorJob&lt;/code&gt; breaks the propagation contract to prevent cascade failures. We will also cover the three specific failure modes — swallowed exceptions, leaked coroutines, and frozen scopes — that show up in production Android and KMP apps.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Kotlin coroutines basics (you have used &lt;code&gt;launch&lt;/code&gt; and &lt;code&gt;async&lt;/code&gt; before)&lt;/li&gt;
&lt;li&gt;Android &lt;code&gt;ViewModel&lt;/code&gt; and &lt;code&gt;viewModelScope&lt;/code&gt; familiarity&lt;/li&gt;
&lt;li&gt;A Kotlin project with &lt;code&gt;kotlinx-coroutines-core&lt;/code&gt; on the classpath&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Step 1 — The Job Tree Is a Runtime Graph, Not a Naming Convention
&lt;/h2&gt;

&lt;p&gt;Most teams get this wrong: structured concurrency is a live, mutable tree in memory, not just a code convention. Every &lt;code&gt;launch&lt;/code&gt; or &lt;code&gt;async&lt;/code&gt; call creates a &lt;code&gt;Job&lt;/code&gt; instance that registers itself as a child of the scope's current &lt;code&gt;Job&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;scope&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;CoroutineScope&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;SupervisorJob&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Dispatchers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Default&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;parent&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;launch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;       &lt;span class="c1"&gt;// Job A&lt;/span&gt;
    &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;child1&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;launch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;..&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;// Job B — child of A&lt;/span&gt;
    &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;child2&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;launch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;..&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;   &lt;span class="c1"&gt;// Job C — child of A&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cancelling Job A pushes &lt;code&gt;CancellationException&lt;/code&gt; down to B and C. That is the downward contract. The upward contract is the dangerous one.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 2 — Cancellation Travels Both Directions by Default
&lt;/h2&gt;

&lt;p&gt;A non-cancellation exception thrown by any child propagates &lt;strong&gt;up&lt;/strong&gt; to the parent. The parent treats it as its own failure, cancels itself, then cancels all remaining children.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;scope&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;CoroutineScope&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Job&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Dispatchers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Default&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;launch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;launch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nc"&gt;RuntimeException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"child failed"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="c1"&gt;// Kills the entire scope&lt;/span&gt;
    &lt;span class="nf"&gt;launch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nf"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;                          &lt;span class="c1"&gt;// Also cancelled&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a medium-complexity Android &lt;code&gt;ViewModel&lt;/code&gt; with 4–6 active coroutines, one unhandled exception leaves the scope permanently cancelled — silent, no crash, no log unless you instrument it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3 — Use SupervisorJob to Break the Upward Contract
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;SupervisorJob&lt;/code&gt; overrides upward propagation. A child failure does not reach the parent or its siblings.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope type&lt;/th&gt;
&lt;th&gt;Propagates up?&lt;/th&gt;
&lt;th&gt;Siblings cancelled?&lt;/th&gt;
&lt;th&gt;Parent cancelled?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Job()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;SupervisorJob()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;supervisorScope { }&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;No (local block)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;viewModelScope&lt;/code&gt; in AndroidX is already backed by &lt;code&gt;SupervisorJob()&lt;/code&gt;. That is why a failed data-fetch coroutine does not kill an ongoing animation coroutine in the same &lt;code&gt;ViewModel&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;MyViewModel&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;ViewModel&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;loadData&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;viewModelScope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;launch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// Failure here does not cancel other viewModelScope children&lt;/span&gt;
        &lt;span class="n"&gt;repository&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Gotchas — Three Production Failure Modes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Swallowed exceptions under SupervisorJob&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;SupervisorJob&lt;/code&gt; prevents cascade failures but also silences them. Without a &lt;code&gt;CoroutineExceptionHandler&lt;/code&gt;, the exception goes to the thread's uncaught handler — invisible to most Android crash reporters.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;scope&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;CoroutineScope&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nc"&gt;SupervisorJob&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Dispatchers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Default&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;CoroutineExceptionHandler&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
        &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Coroutine failure"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Always pair &lt;code&gt;SupervisorJob&lt;/code&gt; with a handler. Logging is not optional — it is part of the contract.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Leaked coroutines from frozen scopes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A cancelled &lt;code&gt;Job()&lt;/code&gt; scope silently rejects new &lt;code&gt;launch&lt;/code&gt; calls — they complete immediately without executing. This surfaces as features that "just stopped working" after the first error, with no exception thrown and no obvious log.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. &lt;code&gt;async&lt;/code&gt; + &lt;code&gt;SupervisorJob&lt;/code&gt; = deferred time bomb&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The docs do not emphasise this enough: &lt;code&gt;async&lt;/code&gt; under a &lt;code&gt;SupervisorJob&lt;/code&gt; stores the exception in the &lt;code&gt;Deferred&lt;/code&gt; rather than throwing it. The exception only materialises at &lt;code&gt;.await()&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;deferred&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;supervisorScope&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;async&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nc"&gt;IOException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"disk full"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;// Exception stored — NOT thrown yet&lt;/span&gt;
&lt;span class="n"&gt;deferred&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;await&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;// Throws HERE, possibly far from the original call site&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In production KMP apps, this is the single most common source of ghost failures in shared data-layer modules. Audit every &lt;code&gt;async&lt;/code&gt; call under supervisor scopes and enforce &lt;code&gt;await()&lt;/code&gt; at the call site.&lt;/p&gt;




&lt;h2&gt;
  
  
  Choosing the Right Boundary
&lt;/h2&gt;

&lt;p&gt;Here is the pattern I use in every project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application root scope     → SupervisorJob (isolate feature-level failures)
ViewModel scope            → SupervisorJob (already provided by viewModelScope)
Request/transaction scope  → Job (one failure should abort the whole operation)
supervisorScope { }        → Inline supervisor for async fan-out with mixed failure tolerance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use &lt;code&gt;Job()&lt;/code&gt; at transaction boundaries where atomicity matters. Use &lt;code&gt;SupervisorJob()&lt;/code&gt; at feature boundaries where independent operations should not cascade.&lt;/p&gt;




&lt;h2&gt;
  
  
  Wrapping Up
&lt;/h2&gt;

&lt;p&gt;The Job tree is always there. The only question is whether you designed it or inherited it by accident. Pair every &lt;code&gt;SupervisorJob&lt;/code&gt; with a &lt;code&gt;CoroutineExceptionHandler&lt;/code&gt;, audit every unawaited &lt;code&gt;Deferred&lt;/code&gt;, and use &lt;code&gt;Job()&lt;/code&gt; anywhere a partial failure should abort the whole operation.&lt;/p&gt;

&lt;p&gt;Further reading: &lt;a href="https://kotlinlang.org/docs/cancellation-and-timeouts.html" rel="noopener noreferrer"&gt;Kotlin coroutines guide — Cancellation and timeouts&lt;/a&gt; and the &lt;a href="https://kotlinlang.org/docs/coroutines-basics.html#structured-concurrency" rel="noopener noreferrer"&gt;structured concurrency article&lt;/a&gt; in the official docs.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring Android's Jetpack Compose to a Quantized On-Device LLM for Real-Time Code Completion</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Tue, 06 Oct 2026 08:19:26 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-androids-jetpack-compose-to-a-quantized-on-device-llm-for-real-time-code-completion-4gog</link>
      <guid>https://dev.to/software_mvp-factory/wiring-androids-jetpack-compose-to-a-quantized-on-device-llm-for-real-time-code-completion-4gog</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;On-Device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Code&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Completion&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Jetpack&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Compose:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Hitting&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;200ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Latency&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Budget"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wire&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CodeGemma&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;2B&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;INT4&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Compose&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;editor&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;StateFlow&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;token&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;streaming,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;debounced&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;triggers,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;sub-200ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;first-token&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;latency&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;target&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Pixel&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;9."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kotlin, android, mobile, architecture&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/on-device-llm-compose-latency-budget&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gu"&gt;## What We Are Building&lt;/span&gt;

By the end of this tutorial you will have a working on-device code completion pipeline in Jetpack Compose: a quantized CodeGemma 2B INT4 model feeding tokens into a &lt;span class="sb"&gt;`StateFlow`&lt;/span&gt;, consumed by a Compose editor overlay, gated behind a debounced trigger that prevents inference storms. The hard target is &lt;span class="gs"&gt;**p95 first-token latency under 200ms**&lt;/span&gt; on Pixel 9-class hardware. That is the line between a feature that feels fluid and one users ignore.

&lt;span class="gu"&gt;## Prerequisites&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Android project targeting API 31+
&lt;span class="p"&gt;-&lt;/span&gt; Pixel 9 or Snapdragon 8 Gen 3 / Tensor G4 device for profiling
&lt;span class="p"&gt;-&lt;/span&gt; TFLite runtime with GPU delegate dependency
&lt;span class="p"&gt;-&lt;/span&gt; Familiarity with Kotlin coroutines and &lt;span class="sb"&gt;`StateFlow`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; CodeGemma 2B INT4 exported via TFLite (roughly 1.1 GB on disk)
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Step 1 — Choose Your Model (and Why INT4 Is Non-Negotiable)&lt;/span&gt;

Let me show you a pattern I use in every project: profile the model &lt;span class="ge"&gt;*before*&lt;/span&gt; writing any app code.

| Model | Size (INT4) | p50 First Token (Pixel 9) | Code MBPP Pass@1 |
|---|---|---|---|
| CodeGemma 2B INT4 | ~1.1 GB | ~120ms | Competitive for 2B class |
| Generic 3B INT8 | ~3.0 GB | ~280ms | Lower code-specific accuracy |
| 7B INT4 | ~4.2 GB | &amp;gt;500ms | Out-of-budget for real-time |

INT8 doubles your memory footprint and pushes you past the latency ceiling. INT4 is the only configuration that fits flagship VRAM headroom while leaving room for the host app.

&lt;span class="gu"&gt;## Step 2 — Build the Streaming Pipeline with StateFlow&lt;/span&gt;

Token streaming from a TFLite inference loop maps naturally onto a &lt;span class="sb"&gt;`StateFlow`&lt;/span&gt;. Each decoded token emits a partial buffer update; your Compose editor observes it with &lt;span class="sb"&gt;`collectAsState()`&lt;/span&gt;.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
// ViewModel&lt;br&gt;
private val _completionTokens = MutableStateFlow("")&lt;br&gt;
val completionTokens: StateFlow = _completionTokens.asStateFlow()&lt;/p&gt;

&lt;p&gt;fun streamCompletion(prompt: String) {&lt;br&gt;
    viewModelScope.launch(Dispatchers.Default) {&lt;br&gt;
        _completionTokens.value = ""&lt;br&gt;
        inferenceEngine.streamTokens(prompt).collect { token -&amp;gt;&lt;br&gt;
            _completionTokens.update { it + token }&lt;br&gt;
        }&lt;br&gt;
    }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
Wire this to your editor overlay in Compose:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
val completion by viewModel.completionTokens.collectAsState()&lt;br&gt;
// Render &lt;code&gt;completion&lt;/code&gt; as a ghost-text overlay anchored to cursor position&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
`StateFlow` conflation is the underrated win here. If your UI frame drops, you process the latest accumulated buffer — never a stale intermediate state.

## Step 3 — Gate Inference With Debounce

Here is the gotcha that will save you hours: firing inference on every keystroke is catastrophic on-device. At 15–20% CPU per inference run, an unbounded trigger will crater your frame rate within seconds.

The production trigger model:

1. Debounce cursor idle — 150–250ms after the last keystroke
2. Gate on syntactic signal — `.`, `(`, space after a keyword, or newline
3. Cancel in-flight inference on any new keystroke via `Job.cancel()`

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
editorState&lt;br&gt;
    .onEach { cancelCurrentInference() }&lt;br&gt;
    .debounce(180)&lt;br&gt;
    .filter { isTriggerContext(it.cursorContext) }&lt;br&gt;
    .collectLatest { state -&amp;gt;&lt;br&gt;
        viewModel.streamCompletion(state.buildPrompt())&lt;br&gt;
    }&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
`collectLatest` handles cancellation automatically — any new emission cancels the previous coroutine, which propagates into the inference loop.

## Step 4 — Pick the Right Delegate

The docs do not mention this, but naive "always prefer GPU" breaks on mid-range devices with silent fallback latency spikes. Use this decision tree:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
plaintext&lt;br&gt;
Is device flagship-tier (Snapdragon 8 Gen 3 / Tensor G4+)?&lt;br&gt;
├─ YES → Try GPU Delegate → if init &amp;lt; 2s, proceed&lt;br&gt;
│         └─ FAIL → Fall back to NNAPI with INT8 cast&lt;br&gt;
└─ NO  → Try NNAPI → benchmark on first run&lt;br&gt;
          └─ p50 &amp;gt; 350ms → fall back to CPU (disable feature)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
Instrument delegate init time at startup, cache the result in `SharedPreferences`, and skip fallback detection on subsequent launches. Cold delegate initialization is a one-time penalty — paying it on every inference is an architecture bug.

## Step 5 — Decouple Editor State From TextFieldValue

Maintain a separate `EditorState` data class tracking cursor offset, visible line range, and last accepted completion. This decouples inference triggers from raw `TextFieldValue` and gives you deterministic test cases for trigger logic without a running model.

---

## Gotchas

- **Skipping early profiling.** Set the 200ms first-token budget before writing inference code. If the model misses it in a synthetic benchmark, no Compose optimization will recover it.
- **Not caching delegate selection.** Re-running delegate detection on every cold start adds 1–2 seconds of invisible latency users will blame on something else.
- **Missing `collectLatest`.** Using plain `collect` instead of `collectLatest` means stale inference jobs accumulate. The cancellation semantics are built in — use them.

---

## Conclusion

Here is the minimal setup to get this working: CodeGemma 2B INT4 for the model, `StateFlow` for the streaming backbone, 180ms debounce with `collectLatest` for storm prevention, and cached delegate selection for startup perf. Get all three right and on-device completion is genuinely competitive with cloud-hosted alternatives — and it works on a plane.

**Further reading:** [TFLite GPU delegate docs](https://www.tensorflow.org/lite/performance/gpu) · [CodeGemma model card](https://ai.google.dev/gemma/docs/codegemma) · [Kotlin StateFlow reference](https://kotlinlang.org/api/kotlinx.coroutines/kotlinx-coroutines-core/kotlinx.coroutines.flow/-state-flow/)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring iOS Core ML to a Quantized On-Device Speech Synthesis Model for Real-Time TTS</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Mon, 05 Oct 2026 13:26:03 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-ios-core-ml-to-a-quantized-on-device-speech-synthesis-model-for-real-time-tts-6fa</link>
      <guid>https://dev.to/software_mvp-factory/wiring-ios-core-ml-to-a-quantized-on-device-speech-synthesis-model-for-real-time-tts-6fa</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;On-Device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;TTS&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;iPhone:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Core&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ML,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Neural&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Engine&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Scheduling,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Sub-200ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Latency&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Ceiling"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Run&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;quantized&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;TTS&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;iPhone&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;via&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Core&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ML.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Learn&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;phoneme&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;buffer&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;design,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ANE&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;vs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;GPU&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tradeoffs,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;how&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;hit&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;sub-200ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;first-audio&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;latency&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on-device."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ios, swift, mobile, architecture&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/quantized-tts-ios-core-ml-latency-ceiling&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gu"&gt;## What We Are Building&lt;/span&gt;

By the end of this walkthrough, you will have a working chunked phoneme synthesis pipeline that feeds a quantized VITS or Kokoro-class TTS model through Core ML — split deliberately across the Neural Engine and GPU — and delivers first audio in 120–180ms on A15 and newer. Not pseudocode. Not theory. A real architecture you can drop into a production iOS app.

&lt;span class="gu"&gt;## Prerequisites&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Xcode 15+, deployment target iOS 16+
&lt;span class="p"&gt;-&lt;/span&gt; A distilled VITS or Kokoro model converted to &lt;span class="sb"&gt;`.mlpackage`&lt;/span&gt; (encoder + vocoder split as two separate assets)
&lt;span class="p"&gt;-&lt;/span&gt; Basic familiarity with &lt;span class="sb"&gt;`AVAudioEngine`&lt;/span&gt; and &lt;span class="sb"&gt;`MLModel`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; A device with a Neural Engine (iPhone XS or later — the simulator will not reflect real latency)
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Why On-Device TTS Right Now&lt;/span&gt;

Cloud TTS is getting complicated. OpenAI announced in 2025 it would test sponsored content inside ChatGPT, and that trajectory is unlikely to reverse. The on-device case was already compelling: no API cost, no latency jitter from network round-trips, no audio leaving the device.

The numbers are concrete. A typical cloud TTS round-trip runs &lt;span class="gs"&gt;**300–600ms**&lt;/span&gt; on a good connection. Core ML on a Neural Engine-capable iPhone hits &lt;span class="gs"&gt;**120–180ms to first audio**&lt;/span&gt; for a quantized model — if you architect the pipeline correctly.
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## The Phoneme-to-Mel Pipeline&lt;/span&gt;

Most distilled TTS architectures share a common spine:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;text → G2P → duration predictor → mel spectrogram → vocoder&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;
&lt;span class="kt"&gt;For&lt;/span&gt; &lt;span class="kt"&gt;Core&lt;/span&gt; &lt;span class="kt"&gt;ML&lt;/span&gt; &lt;span class="n"&gt;deployment&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;split&lt;/span&gt; &lt;span class="n"&gt;this&lt;/span&gt; &lt;span class="n"&gt;into&lt;/span&gt; &lt;span class="n"&gt;two&lt;/span&gt; &lt;span class="n"&gt;inference&lt;/span&gt; &lt;span class="nv"&gt;passes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="kt"&gt;Encoder&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="kt"&gt;Duration&lt;/span&gt; &lt;span class="kt"&gt;Predictor&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="err"&gt;—&lt;/span&gt; &lt;span class="n"&gt;runs&lt;/span&gt; &lt;span class="n"&gt;once&lt;/span&gt; &lt;span class="n"&gt;per&lt;/span&gt; &lt;span class="n"&gt;utterance&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;produces&lt;/span&gt; &lt;span class="n"&gt;aligned&lt;/span&gt; &lt;span class="n"&gt;mel&lt;/span&gt; &lt;span class="n"&gt;frames&lt;/span&gt;
&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="kt"&gt;HiFi&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="kt"&gt;GAN&lt;/span&gt; &lt;span class="n"&gt;or&lt;/span&gt; &lt;span class="kt"&gt;MB&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="kt"&gt;MelGAN&lt;/span&gt; &lt;span class="kt"&gt;Vocoder&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="err"&gt;—&lt;/span&gt; &lt;span class="n"&gt;converts&lt;/span&gt; &lt;span class="n"&gt;mel&lt;/span&gt; &lt;span class="n"&gt;frames&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="mf"&gt;22.05&lt;/span&gt;&lt;span class="n"&gt;kHz&lt;/span&gt; &lt;span class="kt"&gt;PCM&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;streamed&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;

&lt;span class="kt"&gt;Here&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;minimal&lt;/span&gt; &lt;span class="n"&gt;setup&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt; &lt;span class="n"&gt;this&lt;/span&gt; &lt;span class="n"&gt;working&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt; &lt;span class="kt"&gt;The&lt;/span&gt; &lt;span class="n"&gt;trick&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;sub&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;never&lt;/span&gt; &lt;span class="n"&gt;wait&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;full&lt;/span&gt; &lt;span class="n"&gt;utterance&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt; &lt;span class="kt"&gt;Fire&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;vocoder&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;first&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="err"&gt;–&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt; &lt;span class="n"&gt;mel&lt;/span&gt; &lt;span class="n"&gt;frames&lt;/span&gt; &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;encoder&lt;/span&gt; &lt;span class="n"&gt;continues&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="nv"&gt;rest&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
swift&lt;br&gt;
func synthesizeChunked(phonemes: [Int32], chunkSize: Int = 64) async throws {&lt;br&gt;
    var offset = 0&lt;br&gt;
    while offset &amp;lt; phonemes.count {&lt;br&gt;
        let slice = Array(phonemes[offset..&amp;lt;min(offset + chunkSize, phonemes.count)])&lt;br&gt;
        let melFrames = try await encoderModel.predict(phonemes: slice)&lt;br&gt;
        let audio = try await vocoderModel.predict(mel: melFrames)&lt;br&gt;
        audioEngine.scheduleBuffer(audio)&lt;br&gt;
        offset += chunkSize&lt;br&gt;
    }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
First audio hits the speaker before synthesis completes. That is where the latency budget is won.

---

## ANE vs GPU: Split the Model, Don't Trust `.all`

Let me show you a pattern I use in every project. Core ML's `MLComputeUnits` gives you three paths. Here is what I measured on iPhone 13 Pro (A15, VITS-small at 22kHz, 5-word utterances):

| Compute Unit | First-Audio Latency | Power Draw | Best For |
|---|---|---|---|
| `.cpuAndNeuralEngine` | 130–180ms | Low | Attention-heavy encoder |
| `.cpuAndGPU` | 200–280ms | High | Upsampling vocoder |
| `.all` (auto) | 140–200ms | Medium | Baseline only |

The ANE excels at the encoder's attention layers. The vocoder, with its upsampling convolutions, consistently runs faster on GPU. So split the model:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
swift&lt;br&gt;
let encoderConfig = MLModelConfiguration()&lt;br&gt;
encoderConfig.computeUnits = .cpuAndNeuralEngine&lt;/p&gt;

&lt;p&gt;let vocoderConfig = MLModelConfiguration()&lt;br&gt;
vocoderConfig.computeUnits = .cpuAndGPU&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
The combined pipeline lands at **150–175ms** on A15 and newer. Do not trust `.all` — measure and override.

---

## Gotchas

**INT8 quantization will break your prosody.** This is the gotcha that will save you hours. Quantizing the duration predictor to INT8 degrades speech quality significantly — irregular pauses, flattened intonation, clipped phoneme boundaries. The quality cliff is sharp, not gradual.

| Layer Group | Safe Quantization | Notes |
|---|---|---|
| Text encoder | INT8 | Minimal perceptible impact |
| Duration predictor | **FP16 only** | INT8 breaks prosody |
| Mel decoder | INT8 | Acceptable with calibration |
| Vocoder upsampling | **FP16 only** | Audible artifacts at INT8 |

Mixed-precision lands your model at **35–55MB** — well within the 80MB ceiling I treat as the on-device viability threshold for non-game apps.

**Double-buffer your audio or you will get gaps.** Use `AVAudioPlayerNode.scheduleBuffer(_:completionHandler:)` with one chunk playing and one synthesizing. The completion handler triggers the next dispatch. Keep the audio thread hot.

**Thermal throttling is a real constraint, not an edge case.** Sustained ANE load will throttle over extended sessions — especially relevant for accessibility tooling or hands-free workflows. Design synthesis as burst-plus-pause, not a continuous stream. (Speaking of sustained screen work: apps like [HealthyDesk](https://play.google.com/store/apps/details?id=com.healthydesk) exist precisely because continuous focused sessions have physiological costs worth designing around.)

---

## Conclusion

Three things to ship with:

1. **Split compute units.** Encoder on ANE, vocoder on GPU. Measure independently on your target device — `.all` is a starting point, not a final answer.
2. **Protect the duration predictor.** Keep it at FP16 regardless of model size pressure. The perceptual cost of INT8 here far outweighs the storage savings.
3. **Stream mel chunks, not complete utterances.** 50–80 frame slices are the architectural difference between 150ms and 400ms first-audio latency.

**Relevant docs:** [Core ML Performance documentation](https://developer.apple.com/documentation/coreml) covers ANE scheduling behavior and per-chip thresholds — the authoritative source for anything that changes between silicon generations.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring Android's Vulkan Backend to a Quantized On-Device Diffusion Model for Real-Time Texture Synthesis</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:52:41 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-androids-vulkan-backend-to-a-quantized-on-device-diffusion-model-for-real-time-texture-53a9</link>
      <guid>https://dev.to/software_mvp-factory/wiring-androids-vulkan-backend-to-a-quantized-on-device-diffusion-model-for-real-time-texture-53a9</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Vulkan&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Compute&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Diffusion&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Under&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;33ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Snapdragon&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;8&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Gen&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;3"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bypass&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;NNAPI&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;scheduling&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;overhead&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;drive&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;quantized&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;INT8&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;diffusion&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;through&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Vulkan&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;compute&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;shaders,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;staying&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;under&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;33ms&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;per&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;frame&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Snapdragon&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;8&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Gen&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;3."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;android, mobile, architecture, performance&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/vulkan-compute-diffusion-snapdragon-33ms&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What You Will Build
&lt;/h2&gt;

&lt;p&gt;By the end of this guide you will have a tile-based Vulkan compute pipeline that drives a quantized INT8 stable diffusion model directly on Adreno hardware — no NNAPI dispatcher in the hot path, sub-33ms inference per frame, and the memory barrier placement that keeps your 99th percentile latency from blowing your render budget.&lt;/p&gt;

&lt;p&gt;Let me show you a pattern I use in every project that needs deterministic GPU scheduling on Android.&lt;/p&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Android device running Snapdragon 8 Gen 3 (Adreno 750)&lt;/li&gt;
&lt;li&gt;NDK r25+ with Vulkan headers&lt;/li&gt;
&lt;li&gt;A quantized INT8 diffusion model — specifically a stripped-down UNet operating on 64×128 or 128×128 tiles&lt;/li&gt;
&lt;li&gt;Familiarity with &lt;code&gt;VkCommandBuffer&lt;/code&gt; submission and basic Vulkan synchronization&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The Problem with NNAPI on a Tight Frame Budget
&lt;/h2&gt;

&lt;p&gt;NNAPI is easy to treat as a black box. Plug it in, let it handle hardware negotiation, then spend an afternoon debugging inconsistent latency under load with no obvious culprit.&lt;/p&gt;

&lt;p&gt;That inconsistency is baked into the design. NNAPI's internal dispatcher negotiates between DSP, NPU, and GPU backends at runtime. On Snapdragon 8 Gen 3, that negotiation introduces overhead that compounds badly in a render loop. Mean latency of ~28ms sounds workable until you see the 99th percentile spike to ~52ms. That is not a rounding error — that is a visible stutter.&lt;/p&gt;

&lt;p&gt;The fix is to bypass NNAPI entirely and drive your quantized model through Vulkan compute pipelines you own.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Tile Pipeline Architecture
&lt;/h2&gt;

&lt;p&gt;Here is the minimal setup to get this working. The execution path for each frame looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ CPU: Tile Scheduler ]
        |
        v  VkCommandBuffer submission
[ Compute Shader: INT8 UNet Forward Pass ]
        |
        v  VkImageMemoryBarrier (COMPUTE_SHADER -&amp;gt; TRANSFER)
[ Transfer: Tile Blit to Texture Atlas ]
        |
        v  VkImageMemoryBarrier (TRANSFER -&amp;gt; FRAGMENT_SHADER)
[ Fragment Shader: Final Composite ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Barrier placement between compute and transfer is the single biggest lever you have on stall time.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 1 — Scope Your Pipeline Barriers Precisely
&lt;/h2&gt;

&lt;p&gt;Here is the gotcha that will save you hours: a barrier using &lt;code&gt;VK_PIPELINE_STAGE_ALL_COMMANDS_BIT&lt;/code&gt; as its source stage serializes the entire GPU pipeline unnecessarily.&lt;/p&gt;

&lt;p&gt;For the inference-to-blit transition, scope it tightly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;VkImageMemoryBarrier&lt;/span&gt; &lt;span class="n"&gt;barrier&lt;/span&gt;&lt;span class="p"&gt;{};&lt;/span&gt;
&lt;span class="n"&gt;barrier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;srcAccessMask&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;VK_ACCESS_SHADER_WRITE_BIT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;barrier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dstAccessMask&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;VK_ACCESS_TRANSFER_READ_BIT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;barrier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;oldLayout&lt;/span&gt;      &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;VK_IMAGE_LAYOUT_GENERAL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;barrier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;newLayout&lt;/span&gt;      &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;VK_IMAGE_LAYOUT_TRANSFER_SRC_OPTIMAL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;vkCmdPipelineBarrier&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;VK_PIPELINE_STAGE_COMPUTE_SHADER_BIT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;// srcStageMask&lt;/span&gt;
    &lt;span class="n"&gt;VK_PIPELINE_STAGE_TRANSFER_BIT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;         &lt;span class="c1"&gt;// dstStageMask&lt;/span&gt;
    &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;nullptr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;nullptr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;barrier&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The GPU's fixed-function hardware — texture units, rasterizer — keeps running while the barrier resolves only what needs to resolve. This single change can recover several milliseconds at the 99th percentile.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 2 — Pre-Allocate Descriptor Sets at Pipeline Init
&lt;/h2&gt;

&lt;p&gt;Dynamic descriptor set allocation mid-frame is a silent killer. Every &lt;code&gt;vkAllocateDescriptorSets&lt;/code&gt; call against an undersized pool can trigger a driver-side reallocation. The docs do not always make clear how badly this compounds over a tile grid.&lt;/p&gt;

&lt;p&gt;Pre-allocate at pipeline initialization time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;VkDescriptorPoolCreateInfo&lt;/span&gt; &lt;span class="n"&gt;poolInfo&lt;/span&gt;&lt;span class="p"&gt;{};&lt;/span&gt;
&lt;span class="n"&gt;poolInfo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;maxSets&lt;/span&gt;       &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;MAX_TILES_IN_FLIGHT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;FRAMES_IN_FLIGHT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;poolInfo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;poolSizeCount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;static_cast&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;uint32_t&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;poolSizes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="n"&gt;poolInfo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pPoolSizes&lt;/span&gt;    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;poolSizes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a fixed tile grid, skip &lt;code&gt;VK_DESCRIPTOR_POOL_CREATE_FREE_DESCRIPTOR_SET_BIT&lt;/code&gt; — the simpler reset path is faster. Only add the flag if you genuinely need per-tile recycling.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3 — Double-Buffer Tiles to Hide Blit Latency
&lt;/h2&gt;

&lt;p&gt;With two tiles in flight and double-buffered descriptor sets, you can overlap the transfer of tile N with compute for tile N+1. This effectively hides blit latency behind inference latency.&lt;/p&gt;

&lt;p&gt;On Adreno 750, a 128×128 INT8 tile inference pass through a quantized UNet fits inside 18–22ms in isolation. The remaining budget breaks down like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Command buffer recording and submission: ~2ms&lt;/li&gt;
&lt;li&gt;Barrier resolution and blit: ~3ms&lt;/li&gt;
&lt;li&gt;Final composite fragment pass: ~4ms&lt;/li&gt;
&lt;li&gt;Headroom for driver variance: ~4ms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Total: comfortably under 33ms.&lt;/p&gt;




&lt;h2&gt;
  
  
  Gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Coarse barrier stages.&lt;/strong&gt; Replacing &lt;code&gt;VK_PIPELINE_STAGE_ALL_COMMANDS_BIT&lt;/code&gt; with narrowest-applicable stage flags is not optional for a real-time pipeline. It is the first thing to audit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Descriptor pool sizing.&lt;/strong&gt; Size your pool to &lt;code&gt;tile_count × frames_in_flight&lt;/code&gt; at startup. Any mid-frame reallocation invalidates your latency budget immediately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;99th percentile, not mean.&lt;/strong&gt; NNAPI averages ~28ms but spikes to ~52ms at p99. Vulkan direct averages ~18ms and holds ~26ms at p99. A pipeline that stutters on every tenth frame is not acceptable regardless of its mean.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quantization path matters.&lt;/strong&gt; The barrier and descriptor set patterns above apply equally to TFLite GPU delegate and ONNX Runtime QNN backends — only the dispatch layer differs.&lt;/p&gt;




&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;NNAPI (GPU backend)&lt;/th&gt;
&lt;th&gt;Vulkan Direct&lt;/th&gt;
&lt;th&gt;Delta&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mean tile inference&lt;/td&gt;
&lt;td&gt;~28ms&lt;/td&gt;
&lt;td&gt;~18ms&lt;/td&gt;
&lt;td&gt;−10ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;99th percentile&lt;/td&gt;
&lt;td&gt;~52ms&lt;/td&gt;
&lt;td&gt;~26ms&lt;/td&gt;
&lt;td&gt;−26ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frame budget headroom&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;~7ms&lt;/td&gt;
&lt;td&gt;Available&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;If you own the target GPU and your model is quantized, Vulkan compute gives you deterministic scheduling that NNAPI cannot match. Scope your pipeline barriers to the narrowest applicable stages, pre-allocate descriptor sets at init time, and double-buffer your tile grid to hide transfer latency. Those three changes are what close the gap from 45ms to under 33ms on Snapdragon 8 Gen 3.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring Core ML's Neural Engine to a Quantized On-Device Embedding Model for Real-Time Semantic Search</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:03:48 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-core-mls-neural-engine-to-a-quantized-on-device-embedding-model-for-real-time-semantic-3gif</link>
      <guid>https://dev.to/software_mvp-factory/wiring-core-mls-neural-engine-to-a-quantized-on-device-embedding-model-for-real-time-semantic-3gif</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;On-Device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Semantic&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Search&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Core&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ML&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ANE:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Quantized&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Embeddings,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;HNSW&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Indexing,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Memory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Ceiling&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;That&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Actually&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Matters"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Deploy&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;quantized&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MiniLM-L6&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;INT8&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;iPhone&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;via&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Core&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ML,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;force&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ANE&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;scheduling,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;profile&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;token&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;throughput&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Instruments,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;keep&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;your&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;HNSW&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;index&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;under&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on-device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;memory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ceiling."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;ios&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;swift&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;mobile&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;architecture&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/on-device-semantic-search-core-ml-ane-quantized-embeddings&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What We Are Building
&lt;/h2&gt;

&lt;p&gt;By the end of this tutorial, you will have a fully on-device semantic search pipeline: a quantized MiniLM-L6 INT8 sentence-transformer converted to Core ML, wired explicitly to the Apple Neural Engine, and backed by an HNSW index — all running inside your iPhone app with no cloud round-trip. On iPhone 15 Pro, that gets you ~18ms per query at recall@10 of ~0.91. Let me show you the pattern I use in every project that handles sensitive user data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Xcode 15+, iOS 17 deployment target&lt;/li&gt;
&lt;li&gt;Python environment with &lt;code&gt;coremltools&lt;/code&gt; 7.x installed&lt;/li&gt;
&lt;li&gt;Basic familiarity with Core ML model loading in Swift&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;hnswlib&lt;/code&gt; or &lt;code&gt;usearch&lt;/code&gt; for the vector index layer&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Step 1 — Convert and Quantize the Model
&lt;/h2&gt;

&lt;p&gt;Start with &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt;. Convert it to Core ML with INT8 weight quantization using &lt;code&gt;coremltools&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;coremltools&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;ct&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ct&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;traced_model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;ct&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;TensorType&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input_ids&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;))],&lt;/span&gt;
    &lt;span class="n"&gt;compute_precision&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ct&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;precision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;FLOAT16&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;minimum_deployment_target&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ct&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;iOS17&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;spec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ct&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;optimize&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;coreml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;linear_quantize_weights&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ct&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;optimize&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;coreml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;OptimizationConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;global_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ct&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;optimize&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;coreml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;OpLinearQuantizerConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;linear_symmetric&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;int8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gets you a ~22MB model artifact. That is not your memory problem — we will come back to what actually is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2 — Force ANE Scheduling in Swift
&lt;/h2&gt;

&lt;p&gt;Here is the gotcha that will save you hours: do not use &lt;code&gt;.all&lt;/code&gt; for compute units.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;MLModelConfiguration&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;computeUnits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cpuAndNeuralEngine&lt;/span&gt;  &lt;span class="c1"&gt;// NOT .all — avoids GPU fallback&lt;/span&gt;

&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="kt"&gt;MLModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;contentsOf&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;modelURL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;configuration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using &lt;code&gt;.all&lt;/code&gt; risks silent GPU fallback during thermal throttling. &lt;code&gt;.cpuAndNeuralEngine&lt;/code&gt; is more predictable on A17 Pro and M-series chips. Thermal events will still affect ANE availability at extremes, but you lose the random variance that &lt;code&gt;.all&lt;/code&gt; introduces.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3 — Profile with Instruments' Core ML Template
&lt;/h2&gt;

&lt;p&gt;Do not reach for Time Profiler here. Use the &lt;strong&gt;Core ML&lt;/strong&gt; template in Xcode Instruments — it exposes per-layer compute unit attribution and lets you verify ANE utilization rather than guessing.&lt;/p&gt;

&lt;p&gt;Here is the minimal setup to get this working on iPhone 15 Pro with a 64-token sequence:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quantization&lt;/th&gt;
&lt;th&gt;Avg Latency&lt;/th&gt;
&lt;th&gt;ANE Utilization&lt;/th&gt;
&lt;th&gt;Recall@10&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;FP32 (baseline)&lt;/td&gt;
&lt;td&gt;47ms&lt;/td&gt;
&lt;td&gt;~20%&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FP16&lt;/td&gt;
&lt;td&gt;28ms&lt;/td&gt;
&lt;td&gt;~55%&lt;/td&gt;
&lt;td&gt;0.92&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT8&lt;/td&gt;
&lt;td&gt;18ms&lt;/td&gt;
&lt;td&gt;~82%&lt;/td&gt;
&lt;td&gt;0.91&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT4&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;~78%&lt;/td&gt;
&lt;td&gt;0.83&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;INT8 is the pragmatic sweet spot. The docs do not mention this, but INT4's recall degradation becomes more pronounced on short queries — anecdotally, queries under 8 tokens appear more sensitive to quantization error relative to embedding variance. Measure against your actual query distribution before committing to INT4.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4 — Budget Your HNSW Index, Not Your Model
&lt;/h2&gt;

&lt;p&gt;This is where most teams get surprised. The embedding model is not your memory ceiling. Your HNSW index is.&lt;/p&gt;

&lt;p&gt;At 384 float32 dimensions per vector with M=16, ef=200:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Corpus Size&lt;/th&gt;
&lt;th&gt;HNSW Memory&lt;/th&gt;
&lt;th&gt;Fits on iPhone?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10K docs&lt;/td&gt;
&lt;td&gt;~23MB&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50K docs&lt;/td&gt;
&lt;td&gt;~115MB&lt;/td&gt;
&lt;td&gt;Marginal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100K docs&lt;/td&gt;
&lt;td&gt;~230MB&lt;/td&gt;
&lt;td&gt;No — jettison risk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The practical ceiling for a foreground app is around &lt;strong&gt;150MB total&lt;/strong&gt; for the index before iOS memory pressure events start terminating background processes. On older devices that ceiling is tighter — always profile on your minimum-supported hardware.&lt;/p&gt;

&lt;p&gt;At 50K documents you are marginal. Consider product quantization (PQ) on the stored vectors — distinct from the weight quantization you already applied — to compress by 4–8x. Integrate &lt;code&gt;usearch&lt;/code&gt;, which ships a native Swift API with INT8 vector storage, or bridge &lt;code&gt;hnswlib&lt;/code&gt; via Swift/C++.&lt;/p&gt;

&lt;p&gt;I work on projects where this architecture runs offline entirely — health and productivity apps like &lt;a href="https://play.google.com/store/apps/details?id=com.healthydesk" rel="noopener noreferrer"&gt;HealthyDesk&lt;/a&gt; are exactly the kind of context where eliminating cloud round-trips matters both for latency and for user trust. No network dependency, no privacy surface.&lt;/p&gt;




&lt;h2&gt;
  
  
  Gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Never leave compute units as &lt;code&gt;.all&lt;/code&gt; in production.&lt;/strong&gt;&lt;br&gt;
Silent GPU fallback during a thermal event will destroy your latency SLA and you will not catch it in the simulator. Always specify &lt;code&gt;.cpuAndNeuralEngine&lt;/code&gt; explicitly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Measure recall@10 before shipping INT4.&lt;/strong&gt;&lt;br&gt;
The 7ms gain is real. So is the ~8-point recall drop. Benchmark on your corpus with your actual query distribution — not synthetic embeddings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The memory budget is for the index, not the model.&lt;/strong&gt;&lt;br&gt;
If your corpus exceeds ~40K documents, design for PQ compression or index sharding from the start. Retrofitting memory architecture post-launch is expensive.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The ANE is fast enough for production semantic search entirely on-device. The discipline is in measurement, not assumption: measure ANE utilization, measure recall at your actual quantization level, and measure index memory against your minimum-supported hardware before you ship. Get those three right and you have a genuinely fast, private, network-independent search experience.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://coremltools.readme.io" rel="noopener noreferrer"&gt;Core ML Tools documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/unum-cloud/usearch" rel="noopener noreferrer"&gt;usearch Swift API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/coreml" rel="noopener noreferrer"&gt;Apple — Optimizing your Core ML usage&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>PostgreSQL Advisory Locks for Distributed Job Scheduling: Preventing Double-Execution Without a Queue</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Thu, 01 Oct 2026 07:27:50 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/postgresql-advisory-locks-for-distributed-job-scheduling-preventing-double-execution-without-a-3i4d</link>
      <guid>https://dev.to/software_mvp-factory/postgresql-advisory-locks-for-distributed-job-scheduling-preventing-double-execution-without-a-3i4d</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PostgreSQL&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Advisory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Locks:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Distributed&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Job&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Safety&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Without&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Queue"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pg_try_advisory_lock&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;prevents&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;double-execution&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;using&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;only&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;your&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;existing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;database.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Learn&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;session&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;vs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;transaction&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;scopes,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;lock&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;key&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;hashing,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;deadlock&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;patterns&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;10+&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;workers."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgresql, architecture, cloud, api&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://blog.mvpfactory.co/postgresql-advisory-locks-distributed-job-safety&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What We Are Building
&lt;/h2&gt;

&lt;p&gt;Let me show you a pattern I use in every project that needs distributed job coordination: PostgreSQL advisory locks. By the end of this tutorial, you will prevent double-execution across multiple workers — no Redis, no Sidekiq, no message queue infrastructure. Just your existing database doing more work.&lt;/p&gt;

&lt;p&gt;We will cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The difference between session and transaction-scoped locks (and why you almost always want the latter)&lt;/li&gt;
&lt;li&gt;A safe lock key hashing strategy that survives 50,000+ jobs per day&lt;/li&gt;
&lt;li&gt;The deadlock pattern that surfaces past 10 concurrent workers, and the canonical ordering fix&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;PostgreSQL 9.5+&lt;/li&gt;
&lt;li&gt;A jobs table with integer or UUID primary keys&lt;/li&gt;
&lt;li&gt;Basic understanding of SQL transactions&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Step 1 — Understand the Lock Scope Choices
&lt;/h2&gt;

&lt;p&gt;Here is the gotcha that will save you hours. PostgreSQL gives you two flavors of advisory locks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Function&lt;/th&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Released by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pg_advisory_lock(key)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Session&lt;/td&gt;
&lt;td&gt;Explicit unlock or connection close&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pg_advisory_xact_lock(key)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Transaction&lt;/td&gt;
&lt;td&gt;Automatic on &lt;code&gt;COMMIT&lt;/code&gt; / &lt;code&gt;ROLLBACK&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pg_try_advisory_lock(key)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Session&lt;/td&gt;
&lt;td&gt;Explicit unlock or connection close&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pg_try_advisory_xact_lock(key)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Transaction&lt;/td&gt;
&lt;td&gt;Automatic on &lt;code&gt;COMMIT&lt;/code&gt; / &lt;code&gt;ROLLBACK&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Transaction-scoped locks self-clean on failure, require no explicit unlock logic, and compose naturally with your existing transaction management. Default to these.&lt;/p&gt;

&lt;p&gt;Session-scoped locks are treacherous in pooled environments. If your pool (PgBouncer in transaction mode, for instance) recycles a crashed worker's connection, the lock vanishes before you intended. If it holds the connection open, the lock persists indefinitely. Neither outcome is what you want.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 2 — The Safe Acquisition Pattern
&lt;/h2&gt;

&lt;p&gt;Here is the minimal setup to get this working:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Safe pattern: transaction-scoped, non-blocking&lt;/span&gt;
&lt;span class="k"&gt;BEGIN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;pg_try_advisory_xact_lock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hashtext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'job:invoice_sync:'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;acquired&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="n"&gt;acquired&lt;/span&gt; &lt;span class="k"&gt;THEN&lt;/span&gt;
  &lt;span class="k"&gt;ROLLBACK&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="c1"&gt;-- Another worker owns this job — skip it&lt;/span&gt;
  &lt;span class="k"&gt;RETURN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- Do the work inside the transaction&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;jobs&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'processing'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;started_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;NOW&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;COMMIT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pg_try_advisory_xact_lock&lt;/code&gt; is non-blocking. It returns &lt;code&gt;true&lt;/code&gt; if it acquired the lock, &lt;code&gt;false&lt;/code&gt; if another worker already holds it. Your worker skips cleanly — no waiting, no queuing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3 — Key Hashing That Does Not Bite You in Production
&lt;/h2&gt;

&lt;p&gt;The docs do not mention this, but &lt;code&gt;hashtext&lt;/code&gt; returns a 32-bit integer. With 10 concurrent workers running 50,000 jobs per day, the birthday paradox gives you roughly a 1-in-400,000 collision probability per acquisition. That adds up.&lt;/p&gt;

&lt;p&gt;Use the two-part bigint variant instead, with a namespaced job type in the upper 32 bits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Bit-shift namespace into upper 32 bits&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;pg_try_advisory_xact_lock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job_type_id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;bigint&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;bigint&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="s1"&gt;'FFFFFFFF'&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;bigint&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your job type is a string rather than an integer, derive a stable namespace from it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Two-part key: string namespace + row ID&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;pg_try_advisory_xact_lock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'x'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="n"&gt;substr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'invoice_sync'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;))::&lt;/span&gt;&lt;span class="nb"&gt;bit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;)::&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two-part bigint approach cuts collision probability by roughly a factor of 4 billion compared to a single &lt;code&gt;hashtext&lt;/code&gt; call.&lt;/p&gt;




&lt;h2&gt;
  
  
  Gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Deadlocks past 10 workers.&lt;/strong&gt; When workers acquire multiple locks per job — say, locking both a job record and a dependent resource — you create deadlock conditions. PostgreSQL detects these and raises an exception, but the detection cycle defaults to 1 second (&lt;code&gt;deadlock_timeout&lt;/code&gt;). At high concurrency, this is expensive.&lt;/p&gt;

&lt;p&gt;Fix it with canonical lock ordering: always acquire locks in the same sorted order across all workers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Workers must always lock lower ID first&lt;/span&gt;
&lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="n"&gt;lock_key&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;key&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;job_locks&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;job_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;key&lt;/span&gt; &lt;span class="k"&gt;ASC&lt;/span&gt; &lt;span class="n"&gt;LOOP&lt;/span&gt;
  &lt;span class="n"&gt;PERFORM&lt;/span&gt; &lt;span class="n"&gt;pg_try_advisory_xact_lock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lock_key&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="n"&gt;LOOP&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Connection pool starvation.&lt;/strong&gt; If your pool has fewer connections than workers, lock acquisition queues behind connection acquisition. A worker holding a lock while waiting for a second connection will block others indefinitely. Size your pool to &lt;code&gt;workers × max_locks_per_job + headroom&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reaching for the blocking variant.&lt;/strong&gt; Most teams instinctively use &lt;code&gt;pg_advisory_lock&lt;/code&gt; (blocking) instead of &lt;code&gt;pg_try_advisory_lock&lt;/code&gt; (non-blocking). The blocking variant under load causes worker pile-ups. Always prefer the non-blocking &lt;code&gt;try&lt;/code&gt; variant and let workers skip and retry on their own schedule.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Advisory locks eliminate an entire category of queue-related operational complexity — no additional services, no at-least-once delivery concerns, no message visibility timeouts to tune.&lt;/p&gt;

&lt;p&gt;Three rules before you ship this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Default to &lt;code&gt;pg_try_advisory_xact_lock&lt;/code&gt; over the session variant&lt;/li&gt;
&lt;li&gt;Use two-part bigint keys with a namespaced job type in the upper 32 bits&lt;/li&gt;
&lt;li&gt;Enforce canonical lock ordering across all workers acquiring multiple locks&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Get connection pool sizing right too — &lt;code&gt;workers × locks_per_job + headroom&lt;/code&gt; — or starvation will surface as symptoms that look nothing like the actual problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Further reading:&lt;/strong&gt; &lt;a href="https://www.postgresql.org/docs/current/explicit-locking.html#ADVISORY-LOCKS" rel="noopener noreferrer"&gt;PostgreSQL Advisory Locks documentation&lt;/a&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring Android's MediaPipe LLM Inference API to a Quantized On-Device Reranker for RAG</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Wed, 30 Sep 2026 14:28:24 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-androids-mediapipe-llm-inference-api-to-a-quantized-on-device-reranker-for-rag-2bjg</link>
      <guid>https://dev.to/software_mvp-factory/wiring-androids-mediapipe-llm-inference-api-to-a-quantized-on-device-reranker-for-rag-2bjg</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;On-Device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;RAG&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Android:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;FAISS&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Retrieval&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;+&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MediaPipe&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Cross-Encoder&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Reranker"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Build&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;full&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on-device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;RAG&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pipeline&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Android&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;using&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;FAISS-lite&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;int8&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;embeddings&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MediaPipe&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;LLM&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;cross-encoder&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;reranker&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;—&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;real&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;memory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;budgets&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;latency&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Pixel&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;9."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kotlin, android, architecture, mobile&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://blog.mvp-factory.dev/on-device-rag-android-mediapipe-faiss-reranker&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gu"&gt;## What We Are Building&lt;/span&gt;

By the end of this tutorial, you will have a complete retrieval-augmented generation pipeline running entirely on-device on Android — no network calls, no server dependency. We will wire together a FAISS-lite ANN index using int8 bi-encoder embeddings for fast retrieval, then rerank the top candidates with a quantized cross-encoder loaded through MediaPipe's LLM Inference API. All of it fits in ~2.1 GB RAM and completes retrieval in ~132 ms on a Pixel 9 (Snapdragon 8 Gen 3).

Let me show you a pattern I use in every project: optimize retrieval before you optimize the model. Most teams do this backwards.
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Prerequisites&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Android device with Snapdragon 8 Gen 3 or equivalent (Pixel 9 used here)
&lt;span class="p"&gt;-&lt;/span&gt; Android Studio Hedgehog or later
&lt;span class="p"&gt;-&lt;/span&gt; MediaPipe Tasks Android library (&lt;span class="sb"&gt;`com.google.mediapipe:tasks-genai`&lt;/span&gt;)
&lt;span class="p"&gt;-&lt;/span&gt; TensorFlow Lite runtime
&lt;span class="p"&gt;-&lt;/span&gt; FAISS-lite Android binding
&lt;span class="p"&gt;-&lt;/span&gt; A pre-exported int8 bi-encoder TFLite model (e.g., quantized MiniLM, 384 dimensions)
&lt;span class="p"&gt;-&lt;/span&gt; A quantized INT4 cross-encoder (~180 MB on-disk)
&lt;span class="p"&gt;
---
&lt;/span&gt;
&lt;span class="gu"&gt;## Step 1: Understand the Two-Stage Pipeline&lt;/span&gt;

Here is the minimal architecture that makes this work:

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Query&lt;br&gt;
  │&lt;br&gt;
  ▼&lt;br&gt;
[int8 Bi-Encoder]  ──►  FAISS-lite ANN  ──►  Top-K Candidates (K=10)&lt;br&gt;
                                                      │&lt;br&gt;
                                                      ▼&lt;br&gt;
                                           [Quantized Cross-Encoder]&lt;br&gt;
                                           MediaPipe LLM Inference API&lt;br&gt;
                                                      │&lt;br&gt;
                                                      ▼&lt;br&gt;
                                             Re-ranked Top-3 Chunks&lt;br&gt;
                                                      │&lt;br&gt;
                                                      ▼&lt;br&gt;
                                             On-Device LLM Context&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;
&lt;span class="nc"&gt;Stage&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="n"&gt;fast&lt;/span&gt; &lt;span class="n"&gt;but&lt;/span&gt; &lt;span class="n"&gt;imprecise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="nc"&gt;Stage&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="n"&gt;what&lt;/span&gt; &lt;span class="n"&gt;separates&lt;/span&gt; &lt;span class="n"&gt;actually&lt;/span&gt; &lt;span class="n"&gt;relevant&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="n"&gt;from&lt;/span&gt; &lt;span class="n"&gt;plausible-looking&lt;/span&gt; &lt;span class="n"&gt;noise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="nc"&gt;You&lt;/span&gt; &lt;span class="n"&gt;need&lt;/span&gt; &lt;span class="n"&gt;both&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;span class="p"&gt;---&lt;/span&gt;

&lt;span class="err"&gt;##&lt;/span&gt; &lt;span class="nc"&gt;Step&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Stage&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="err"&gt;—&lt;/span&gt; &lt;span class="nc"&gt;Bi-Encoder&lt;/span&gt; &lt;span class="nc"&gt;Retrieval&lt;/span&gt; &lt;span class="n"&gt;with&lt;/span&gt; &lt;span class="nc"&gt;FAISS-lite&lt;/span&gt;

&lt;span class="nc"&gt;Pre-compute&lt;/span&gt; &lt;span class="n"&gt;embeddings&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;your&lt;/span&gt; &lt;span class="n"&gt;corpus&lt;/span&gt; &lt;span class="n"&gt;offline&lt;/span&gt; &lt;span class="n"&gt;and&lt;/span&gt; &lt;span class="n"&gt;store&lt;/span&gt; &lt;span class="n"&gt;them&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;flat&lt;/span&gt; &lt;span class="nc"&gt;INT8&lt;/span&gt; &lt;span class="nc"&gt;FAISS&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="nc"&gt;At&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;embed&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="n"&gt;with&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;same&lt;/span&gt; &lt;span class="n"&gt;bi-encoder&lt;/span&gt; &lt;span class="n"&gt;and&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt; &lt;span class="nc"&gt;ANN&lt;/span&gt; &lt;span class="n"&gt;search&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
val embeddingInterpreter = Interpreter(loadModelFile("bi_encoder_int8.tflite"))&lt;br&gt;
val queryEmbedding = FloatArray(384)&lt;br&gt;
embeddingInterpreter.run(tokenize(query), queryEmbedding)&lt;/p&gt;

&lt;p&gt;val index = FaissIndex.load("corpus.index") // flat int8, ~12 MB for 50k chunks&lt;br&gt;
val topK = index.search(queryEmbedding, k = 10)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
At ~18 ms per query on Pixel 9, this stage is effectively free. The 50k-chunk index costs ~12 MB on-disk and ~48 MB in RAM.

---

## Step 3: Stage 2 — Cross-Encoder Reranking via MediaPipe

Load your INT4 quantized cross-encoder through MediaPipe LLM Inference API. Score each (query, chunk) pair and sort descending. Take the top 3.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
kotlin&lt;br&gt;
val llmInference = LlmInference.createFromOptions(&lt;br&gt;
    context,&lt;br&gt;
    LlmInference.LlmInferenceOptions.builder()&lt;br&gt;
        .setModelPath("/data/local/tmp/reranker_int4.bin")&lt;br&gt;
        .setMaxTokens(512)&lt;br&gt;
        .build()&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;val scores = topK.map { chunk -&amp;gt;&lt;br&gt;
    val prompt = "Relevance score 0-10 for:\nQuery: $query\nPassage: $chunk\nScore:"&lt;br&gt;
    llmInference.generateResponse(prompt).trim().toFloatOrNull() ?: 0f&lt;br&gt;
}&lt;br&gt;
val reranked = topK.zip(scores).sortedByDescending { it.second }.take(3)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
Ten forward passes through the cross-encoder costs ~110 ms total on Snapdragon 8 Gen 3 with NPU delegation. In practice this is invisible behind main LLM generation time.

---

## Step 4: Budget Your Memory and Context Window

Here is the full cost breakdown on Pixel 9:

| Component | On-Disk | RAM | Latency |
|---|---|---|---|
| int8 Bi-Encoder (TFLite) | ~22 MB | ~85 MB | ~18 ms |
| FAISS-lite Index (50k chunks) | ~12 MB | ~48 MB | ~4 ms |
| Cross-Encoder Reranker (INT4) | ~180 MB | ~420 MB | ~110 ms |
| Main LLM (INT4, 1B param) | ~800 MB | ~1,400 MB | varies |
| **Total** | **~1.01 GB** | **~1.95 GB** | **~132 ms retrieval** |

Pixel 9 ships with 12 GB RAM. You have comfortable headroom.

For your context window with a 2,048-token LLM:

- System prompt: ~150 tokens
- 3 chunks × ~350 tokens each: ~1,050 tokens
- Query + response buffer: ~500 tokens
- **Remaining for generation: ~350 tokens**

The docs do not mention this, but if your model supports 4,096 tokens — increasingly common in 1B–3B quantized models — expanding to top-5 chunks and leaving ~800 tokens for generation moves quality more than switching to a larger model entirely.

---

## Gotchas

**Retrieval quality gates everything downstream.** A fast LLM producing confident nonsense because the context window is stuffed with irrelevant chunks is unfixable at the generation stage. Budget your context window *before* you budget your model size.

**Do not skip the reranker to save 110 ms.** The bi-encoder is optimized for recall, not precision. The cross-encoder is what makes the top-3 chunks actually top-3. Running a server-side reranker defeats the point of on-device inference — MediaPipe handles this cleanly without drama.

**INT4 quantization requires NPU delegation to hit these numbers.** Without explicit NPU delegation in your `LlmInferenceOptions`, you fall back to CPU and latency balloons. Snapdragon 8 Gen 3's Hexagon NPU delivers ~45 TOPS — use it.

**Three to five high-quality chunks beat ten mediocre ones.** The temptation is to widen K and stuff the context. Resist it. Rerank tightly and pass fewer, better chunks.

---

## Conclusion

You now have a working two-stage on-device RAG pipeline: FAISS-lite ANN retrieval at ~18 ms, MediaPipe cross-encoder reranking at ~110 ms, and a full memory footprint of ~1.95 GB — well within Pixel 9's budget. The Snapdragon 8 Gen 3 NPU makes INT4 cross-encoder inference practical without a server in the loop.

**Further reading:**
- [MediaPipe LLM Inference API docs](https://developers.google.com/mediapipe/solutions/genai/llm_inference/android)
- [TFLite quantization guide](https://www.tensorflow.org/lite/performance/post_training_quantization)
- [FAISS documentation](https://faiss.ai/)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring iOS CoreML to a Quantized On-Device Diffusion Model for Real-Time Image Editing</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:39:09 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-ios-coreml-to-a-quantized-on-device-diffusion-model-for-real-time-image-editing-39l8</link>
      <guid>https://dev.to/software_mvp-factory/wiring-ios-coreml-to-a-quantized-on-device-diffusion-model-for-real-time-image-editing-39l8</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wiring&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CoreML&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Quantized&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Diffusion&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Model&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Real-Time&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;iOS&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Editing"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Ship&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;quantized&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Stable&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Diffusion&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;iOS&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CoreML&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ML&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Program&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;format,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;cross-attention&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;KV-cache&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;reuse,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;per-chip&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;memory-tier&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fallback&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;A16,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;A17,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;M-series."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ios, swift, mobile, architecture&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/coreml-quantized-diffusion-ios&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h1&gt;
  
  
  Wiring CoreML to a Quantized Diffusion Model for Real-Time iOS Editing
&lt;/h1&gt;

&lt;p&gt;Let me show you a pattern I use in every on-device ML project: treat memory pressure as a first-class constraint from day one, not after your first TestFlight crash.&lt;/p&gt;

&lt;p&gt;Shipping a real-time diffusion-based image editor on iOS is achievable. The UI canvas runs at 60fps because inference executes asynchronously off the main thread — actual per-step latency runs from 280ms to 800ms depending on chip tier and quantization depth. The gap between smooth and jittery comes down to three decisions: quantization depth, attention KV-cache reuse, and a hard per-chip memory ceiling that triggers quality fallback before the OS kills your process.&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Are Building
&lt;/h2&gt;

&lt;p&gt;A CoreML inference pipeline for a quantized Stable Diffusion model that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Splits the model into four separately compiled &lt;code&gt;.mlpackage&lt;/code&gt; segments&lt;/li&gt;
&lt;li&gt;Reuses cross-attention K/V tensors across denoising steps&lt;/li&gt;
&lt;li&gt;Detects chip tier at runtime and configures compute units accordingly&lt;/li&gt;
&lt;li&gt;Falls back to lower resolution gracefully instead of silently degrading to CPU&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Xcode 15+, iOS 17+ deployment target&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;coremltools&lt;/code&gt; 7.x installed in your Python environment&lt;/li&gt;
&lt;li&gt;A converted SD 1.5 model (text encoder, VAE encoder, U-Net, VAE decoder as separate &lt;code&gt;.mlpackage&lt;/code&gt; files)&lt;/li&gt;
&lt;li&gt;Basic familiarity with &lt;code&gt;MLModel&lt;/code&gt; and &lt;code&gt;MLModelConfiguration&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Step 1 — Pick Your Quantization Floor
&lt;/h2&gt;

&lt;p&gt;A full FP32 SD 1.5 U-Net is unusable on-device. Here is the precision table that matters:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Precision&lt;/th&gt;
&lt;th&gt;U-Net Size&lt;/th&gt;
&lt;th&gt;ANE Eligible&lt;/th&gt;
&lt;th&gt;Step Latency (A17)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;FP16&lt;/td&gt;
&lt;td&gt;~2.5 GB&lt;/td&gt;
&lt;td&gt;Partial&lt;/td&gt;
&lt;td&gt;~800 ms/step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT8 weights&lt;/td&gt;
&lt;td&gt;~1.3 GB&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;~420 ms/step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT4 weights&lt;/td&gt;
&lt;td&gt;~700 MB&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;~280 ms/step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT4 + attention FP16&lt;/td&gt;
&lt;td&gt;~750 MB&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;~295 ms/step&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last row is the production choice. Keep attention projections at FP16 — palettizing them to INT4 compounds error across 20 denoising steps in ways that are visually obvious. Palettize everything else with &lt;code&gt;coremltools.optimize.coreml.palettize_weights&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 2 — Cache Cross-Attention K/V Across Denoising Steps
&lt;/h2&gt;

&lt;p&gt;This is the highest-leverage optimization most iOS ML engineers skip. In a guided diffusion edit, text conditioning does not change between steps. That means cross-attention K and V projections are identical on every step — recomputing them is pure waste.&lt;/p&gt;

&lt;p&gt;Here is the minimal setup to get this working:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;cachedKV&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;MLMultiArray&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[:]&lt;/span&gt;

&lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;denoisingStep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;latent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;MLMultiArray&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;step&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throws&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="kt"&gt;MLMultiArray&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;inputDict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"latent_input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;latent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="s"&gt;"timestep"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;MLMultiArray&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt;
        &lt;span class="s"&gt;"use_cached_kv"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;MLMultiArray&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;cachedKV&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;inputDict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;provider&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="kt"&gt;MLDictionaryFeatureProvider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;dictionary&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;inputDict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="n"&gt;unet&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;prediction&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;cachedKV&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;extractKV&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;guard&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;featureValue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;for&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"latent_output"&lt;/span&gt;&lt;span class="p"&gt;)?&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;multiArrayValue&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="kt"&gt;InferenceError&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;missingOutput&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This alone cuts cross-attention compute by 35–45% on a 20-step schedule with no quality cost.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3 — Set Compute Units Per Chip Tier
&lt;/h2&gt;

&lt;p&gt;The docs do not mention this clearly, but &lt;code&gt;MLModelConfiguration.computeUnits&lt;/code&gt; should reflect the chip tier detected at runtime. iOS will not crash your app when you breach the Neural Engine's working-set limit — it silently delegates layers to CPU, which is 4–8x slower.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Chip&lt;/th&gt;
&lt;th&gt;ANE Budget&lt;/th&gt;
&lt;th&gt;Safe Model Budget&lt;/th&gt;
&lt;th&gt;Fallback Trigger&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A16 Bionic&lt;/td&gt;
&lt;td&gt;~1.0 GB&lt;/td&gt;
&lt;td&gt;~700 MB&lt;/td&gt;
&lt;td&gt;CPU delegation above ~1.1 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A17 Pro&lt;/td&gt;
&lt;td&gt;~1.4 GB&lt;/td&gt;
&lt;td&gt;~1.0 GB&lt;/td&gt;
&lt;td&gt;CPU delegation above ~1.5 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;M2 / M4 (iPad)&lt;/td&gt;
&lt;td&gt;~3.5 GB&lt;/td&gt;
&lt;td&gt;~2.5 GB&lt;/td&gt;
&lt;td&gt;Rarely triggered&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;resolvedComputeUnits&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="kt"&gt;MLComputeUnits&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;chip&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;ChipTierDetector&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;current&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;// wrapper around sysctlbyname("hw.optional.*")&lt;/span&gt;
    &lt;span class="k"&gt;switch&lt;/span&gt; &lt;span class="n"&gt;chip&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;a16&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;a17&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cpuAndNeuralEngine&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;m2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;m4&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;all&lt;/span&gt; &lt;span class="c1"&gt;// GPU path enabled for non-ANE ops&lt;/span&gt;
    &lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cpuAndNeuralEngine&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before each inference pass, call &lt;code&gt;os_proc_available_memory()&lt;/code&gt;. If headroom drops below your model's activation footprint, drop to 384×384 instead of 512×512 rather than letting the runtime decide for you.&lt;/p&gt;




&lt;h2&gt;
  
  
  Gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Silent CPU delegation is your real enemy.&lt;/strong&gt; There is no error thrown — just latency doubling and frames dropping. Instrument memory headroom before every pass.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benchmarking on M2 iPad and shipping to A16 iPhones.&lt;/strong&gt; The Neural Engine tier gap is brutal — not just in raw TOPS but in on-chip SRAM for intermediate activations. Always test on your lowest supported chip.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Palettizing attention projections.&lt;/strong&gt; INT4 attention layers compound error visually across 20+ steps. Exempt them explicitly in your &lt;code&gt;palettize_weights&lt;/code&gt; config and keep them at FP16.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Three decisions determine whether your diffusion app ships or gets shelved: INT4 weight palettization with FP16 attention, K/V cache reuse from step zero, and per-chip &lt;code&gt;computeUnits&lt;/code&gt; configuration backed by runtime memory monitoring. Each one independently improves the pipeline; together they are what separates a 280ms interactive editor from an 800ms thermally-throttled demo.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://coremltools.readme.io/docs/optimizing-models" rel="noopener noreferrer"&gt;CoreML Tools &lt;code&gt;palettize_weights&lt;/code&gt; docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/coreml/mlprogram" rel="noopener noreferrer"&gt;Apple ML Program format guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/os/3191911-os_proc_available_memory" rel="noopener noreferrer"&gt;&lt;code&gt;os_proc_available_memory&lt;/code&gt; reference&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Wiring Android's MediaPipe LLM Inference API to a Streaming Compose UI</title>
      <dc:creator>SoftwareDevs mvpfactory.io</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:47:33 +0000</pubDate>
      <link>https://dev.to/software_mvp-factory/wiring-androids-mediapipe-llm-inference-api-to-a-streaming-compose-ui-255k</link>
      <guid>https://dev.to/software_mvp-factory/wiring-androids-mediapipe-llm-inference-api-to-a-streaming-compose-ui-255k</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Streaming&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;On-Device&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;LLMs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Compose&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MediaPipe&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Kotlin&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Flow"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Wire&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;MediaPipe's&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;LLM&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Inference&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;API&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Jetpack&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Compose&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;using&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Kotlin&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Flow&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;—&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;covering&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;SharedFlow&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;buffer&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;sizing,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;JNI-to-coroutine&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;dispatch,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;context&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;window&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;compression&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Pixel&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;8-class&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;devices."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kotlin, android, mobile, architecture&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://mvpfactory.co/blog/mediapipe-llm-compose-streaming&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What We Are Building
&lt;/h2&gt;

&lt;p&gt;Let me show you a pattern I use in every on-device AI feature: a production-grade pipeline that takes raw token callbacks from MediaPipe's LLM Inference API and streams them into a reactive Compose UI without dropped tokens, UI jank, or silent prompt truncation.&lt;/p&gt;

&lt;p&gt;MediaPipe gives you on-device Gemma and Phi-3 inference on Android. The hard part is not the inference — it is the plumbing between the native engine and your UI layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Android project targeting API 26+&lt;/li&gt;
&lt;li&gt;MediaPipe Tasks dependency added to your build&lt;/li&gt;
&lt;li&gt;Basic familiarity with Kotlin coroutines and &lt;code&gt;StateFlow&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Step 1: Understand the Token's Journey
&lt;/h2&gt;

&lt;p&gt;Every token MediaPipe generates travels across four boundaries:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;JNI callback → Kotlin lambda&lt;/li&gt;
&lt;li&gt;Kotlin lambda → SharedFlow emission&lt;/li&gt;
&lt;li&gt;SharedFlow → StateFlow accumulation in ViewModel&lt;/li&gt;
&lt;li&gt;Compose recomposition on the main thread&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each boundary is a potential data loss or jank point. Here is the gotcha that will save you hours: &lt;strong&gt;the callback originates from a JNI thread&lt;/strong&gt; — not the main thread, not a coroutine dispatcher, but a raw native thread with no coroutine context.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 2: Fix the JNI Dispatcher Boundary
&lt;/h2&gt;

&lt;p&gt;Most teams get this wrong on the first pass.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="c1"&gt;// ❌ Dangerous — emitting from unknown native thread&lt;/span&gt;
&lt;span class="n"&gt;inferenceSession&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generateAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;partialResult&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
    &lt;span class="n"&gt;_tokenFlow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;tryEmit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;partialResult&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;// silently drops tokens under load&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// ✅ Correct — dispatch through a dedicated coroutine scope&lt;/span&gt;
&lt;span class="n"&gt;inferenceSession&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generateAsync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;partialResult&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
    &lt;span class="n"&gt;emissionScope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;launch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Dispatchers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Default&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;_tokenFlow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;emit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;partialResult&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;// suspends if buffer full&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The distinction between &lt;code&gt;tryEmit&lt;/code&gt; and &lt;code&gt;emit&lt;/code&gt; is load-bearing. &lt;code&gt;tryEmit&lt;/code&gt; returns &lt;code&gt;false&lt;/code&gt; and drops the token if the buffer is full — with no exception, no log, no signal. On a Pixel 8 running Gemma 2B at roughly 15–20 tokens per second, you will see drops under any meaningful UI load if you rely on &lt;code&gt;tryEmit&lt;/code&gt; with a default buffer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3: Size Your SharedFlow Buffer Correctly
&lt;/h2&gt;

&lt;p&gt;Here is the minimal setup to get this working on capable hardware:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;_tokenFlow&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MutableSharedFlow&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;(&lt;/span&gt;
    &lt;span class="n"&gt;replay&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;extraBufferCapacity&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;onBufferOverflow&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BufferOverflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;SUSPEND&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The docs do not mention this, but buffer requirements vary significantly by device class:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Device Class&lt;/th&gt;
&lt;th&gt;Approx Tokens/sec&lt;/th&gt;
&lt;th&gt;Recommended Buffer&lt;/th&gt;
&lt;th&gt;Overflow Policy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pixel 8 / Snapdragon 8 Gen 2&lt;/td&gt;
&lt;td&gt;15–25&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;td&gt;&lt;code&gt;SUSPEND&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mid-range (Dimensity 700)&lt;/td&gt;
&lt;td&gt;5–12&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;td&gt;&lt;code&gt;SUSPEND&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Low-end (&amp;lt; 4 GB RAM)&lt;/td&gt;
&lt;td&gt;2–6&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;&lt;code&gt;DROP_OLDEST&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Use &lt;code&gt;SUSPEND&lt;/code&gt; on capable hardware so backpressure signals the emission scope to slow down. On constrained devices where inference is already the bottleneck, &lt;code&gt;DROP_OLDEST&lt;/code&gt; prevents unbounded coroutine queue growth at the cost of occasional visual glitching — which is less harmful than an OOM.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 4: Accumulate to StateFlow in the ViewModel
&lt;/h2&gt;

&lt;p&gt;Expose a &lt;code&gt;StateFlow&amp;lt;String&amp;gt;&lt;/code&gt; to the UI layer, never the raw SharedFlow. Raw SharedFlow collection in Compose can produce redundant recompositions when multiple collectors exist.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="kd"&gt;val&lt;/span&gt; &lt;span class="py"&gt;responseState&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;StateFlow&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_tokenFlow&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;runningFold&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;acc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;acc&lt;/span&gt; &lt;span class="p"&gt;+&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stateIn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;viewModelScope&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;SharingStarted&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Eagerly&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One &lt;code&gt;collectAsState&lt;/code&gt; call in your composable gives you stable, lifecycle-aware streaming output with no boilerplate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 5: Handle the Context Window Ceiling
&lt;/h2&gt;

&lt;p&gt;MediaPipe's LLM Inference API enforces a hard context window configured at model initialization — typically 1024 to 4096 tokens depending on model variant and available RAM. Exceeding this limit does not throw an exception. It silently truncates the prompt from the beginning.&lt;/p&gt;

&lt;p&gt;Build prompt compression before you need it. Here is a rolling window strategy that takes an afternoon to implement:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight kotlin"&gt;&lt;code&gt;&lt;span class="k"&gt;fun&lt;/span&gt; &lt;span class="nf"&gt;compressHistory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;List&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;,&lt;/span&gt; &lt;span class="n"&gt;maxTokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Int&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nc"&gt;List&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;var&lt;/span&gt; &lt;span class="py"&gt;tokenCount&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;estimateTokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;systemPrompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reversed&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;takeWhile&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="p"&gt;-&amp;gt;&lt;/span&gt;
            &lt;span class="n"&gt;tokenCount&lt;/span&gt; &lt;span class="p"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;estimateTokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;tokenCount&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;maxTokens&lt;/span&gt; &lt;span class="p"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.85&lt;/span&gt; &lt;span class="c1"&gt;// 15% safety margin&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reversed&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 15% safety margin is not cosmetic — character-based token estimation is approximate, and hitting the hard limit mid-generation produces corrupted partial output that is worse than a clean truncation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Gotchas
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Silent token drops with &lt;code&gt;tryEmit&lt;/code&gt;&lt;/strong&gt; — there is nothing to grep for. You will only notice when the UI output looks incomplete under load.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wrong dispatcher on JNI callback&lt;/strong&gt; — always dispatch through &lt;code&gt;Dispatchers.Default&lt;/code&gt; before emitting. Do not assume you are on any known thread.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exposing raw SharedFlow to Compose&lt;/strong&gt; — multiple collectors trigger redundant recompositions. Always terminate at a &lt;code&gt;StateFlow&lt;/code&gt; via &lt;code&gt;runningFold&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context window truncation&lt;/strong&gt; — it is silent and it corrupts output. Retrofit a rolling window before you ship multi-turn conversations, not after user complaints.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;On-device LLM inference on Android is production-ready. The gap between a working prototype and a shipping feature is deliberate engineering at three specific layers: the JNI dispatcher boundary, SharedFlow buffer configuration, and context window management. Each of these fails silently in ways that will not surface during early development.&lt;/p&gt;

&lt;p&gt;Wire these correctly once and the pattern holds across every on-device AI feature you ship after it.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
