DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring iOS CoreML to a Quantized On-Device Diffusion Model for Real-Time Image Editing

---
title: "Wiring CoreML to a Quantized Diffusion Model for Real-Time iOS Editing"
published: true
description: "Ship quantized Stable Diffusion on iOS with CoreML ML Program format, cross-attention KV-cache reuse, and per-chip memory-tier fallback for A16, A17, and M-series."
tags: ios, swift, mobile, architecture
canonical_url: https://mvpfactory.co/blog/coreml-quantized-diffusion-ios
---
Enter fullscreen mode Exit fullscreen mode

Wiring CoreML to a Quantized Diffusion Model for Real-Time iOS Editing

Let me show you a pattern I use in every on-device ML project: treat memory pressure as a first-class constraint from day one, not after your first TestFlight crash.

Shipping a real-time diffusion-based image editor on iOS is achievable. The UI canvas runs at 60fps because inference executes asynchronously off the main thread — actual per-step latency runs from 280ms to 800ms depending on chip tier and quantization depth. The gap between smooth and jittery comes down to three decisions: quantization depth, attention KV-cache reuse, and a hard per-chip memory ceiling that triggers quality fallback before the OS kills your process.


What We Are Building

A CoreML inference pipeline for a quantized Stable Diffusion model that:

  • Splits the model into four separately compiled .mlpackage segments
  • Reuses cross-attention K/V tensors across denoising steps
  • Detects chip tier at runtime and configures compute units accordingly
  • Falls back to lower resolution gracefully instead of silently degrading to CPU

Prerequisites

  • Xcode 15+, iOS 17+ deployment target
  • coremltools 7.x installed in your Python environment
  • A converted SD 1.5 model (text encoder, VAE encoder, U-Net, VAE decoder as separate .mlpackage files)
  • Basic familiarity with MLModel and MLModelConfiguration

Step 1 — Pick Your Quantization Floor

A full FP32 SD 1.5 U-Net is unusable on-device. Here is the precision table that matters:

Precision U-Net Size ANE Eligible Step Latency (A17)
FP16 ~2.5 GB Partial ~800 ms/step
INT8 weights ~1.3 GB Yes ~420 ms/step
INT4 weights ~700 MB Yes ~280 ms/step
INT4 + attention FP16 ~750 MB Yes ~295 ms/step

The last row is the production choice. Keep attention projections at FP16 — palettizing them to INT4 compounds error across 20 denoising steps in ways that are visually obvious. Palettize everything else with coremltools.optimize.coreml.palettize_weights.


Step 2 — Cache Cross-Attention K/V Across Denoising Steps

This is the highest-leverage optimization most iOS ML engineers skip. In a guided diffusion edit, text conditioning does not change between steps. That means cross-attention K and V projections are identical on every step — recomputing them is pure waste.

Here is the minimal setup to get this working:

var cachedKV: [String: MLMultiArray] = [:]

func denoisingStep(latent: MLMultiArray, step: Int) throws -> MLMultiArray {
    var inputDict: [String: Any] = [
        "latent_input": latent,
        "timestep": MLMultiArray([step]),
        "use_cached_kv": MLMultiArray([step > 0 ? 1 : 0])
    ]

    if step > 0 {
        for (key, value) in cachedKV { inputDict[key] = value }
    }

    let provider = try MLDictionaryFeatureProvider(dictionary: inputDict)
    let output = try unet.prediction(from: provider)

    if step == 0 {
        cachedKV = extractKV(from: output)
    }

    guard let result = output.featureValue(for: "latent_output")?.multiArrayValue else {
        throw InferenceError.missingOutput
    }
    return result
}
Enter fullscreen mode Exit fullscreen mode

This alone cuts cross-attention compute by 35–45% on a 20-step schedule with no quality cost.


Step 3 — Set Compute Units Per Chip Tier

The docs do not mention this clearly, but MLModelConfiguration.computeUnits should reflect the chip tier detected at runtime. iOS will not crash your app when you breach the Neural Engine's working-set limit — it silently delegates layers to CPU, which is 4–8x slower.

Chip ANE Budget Safe Model Budget Fallback Trigger
A16 Bionic ~1.0 GB ~700 MB CPU delegation above ~1.1 GB
A17 Pro ~1.4 GB ~1.0 GB CPU delegation above ~1.5 GB
M2 / M4 (iPad) ~3.5 GB ~2.5 GB Rarely triggered
func resolvedComputeUnits() -> MLComputeUnits {
    let chip = ChipTierDetector.current() // wrapper around sysctlbyname("hw.optional.*")
    switch chip {
    case .a16, .a17:
        return .cpuAndNeuralEngine
    case .m2, .m4:
        return .all // GPU path enabled for non-ANE ops
    default:
        return .cpuAndNeuralEngine
    }
}
Enter fullscreen mode Exit fullscreen mode

Before each inference pass, call os_proc_available_memory(). If headroom drops below your model's activation footprint, drop to 384×384 instead of 512×512 rather than letting the runtime decide for you.


Gotchas

Silent CPU delegation is your real enemy. There is no error thrown — just latency doubling and frames dropping. Instrument memory headroom before every pass.

Benchmarking on M2 iPad and shipping to A16 iPhones. The Neural Engine tier gap is brutal — not just in raw TOPS but in on-chip SRAM for intermediate activations. Always test on your lowest supported chip.

Palettizing attention projections. INT4 attention layers compound error visually across 20+ steps. Exempt them explicitly in your palettize_weights config and keep them at FP16.


Conclusion

Three decisions determine whether your diffusion app ships or gets shelved: INT4 weight palettization with FP16 attention, K/V cache reuse from step zero, and per-chip computeUnits configuration backed by runtime memory monitoring. Each one independently improves the pipeline; together they are what separates a 280ms interactive editor from an 800ms thermally-throttled demo.

Resources:

Top comments (0)