DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring iOS Core ML to a Quantized On-Device Speech Synthesis Model for Real-Time TTS

---
title: "On-Device TTS on iPhone: Core ML, Neural Engine Scheduling, and the Sub-200ms Latency Ceiling"
published: true
description: "Run a quantized TTS model on iPhone via Core ML. Learn phoneme buffer design, ANE vs GPU tradeoffs, and how to hit sub-200ms first-audio latency on-device."
tags: ios, swift, mobile, architecture
canonical_url: https://mvpfactory.co/blog/quantized-tts-ios-core-ml-latency-ceiling
---

## What We Are Building

By the end of this walkthrough, you will have a working chunked phoneme synthesis pipeline that feeds a quantized VITS or Kokoro-class TTS model through Core ML — split deliberately across the Neural Engine and GPU — and delivers first audio in 120–180ms on A15 and newer. Not pseudocode. Not theory. A real architecture you can drop into a production iOS app.

## Prerequisites

- Xcode 15+, deployment target iOS 16+
- A distilled VITS or Kokoro model converted to `.mlpackage` (encoder + vocoder split as two separate assets)
- Basic familiarity with `AVAudioEngine` and `MLModel`
- A device with a Neural Engine (iPhone XS or later — the simulator will not reflect real latency)

---

## Why On-Device TTS Right Now

Cloud TTS is getting complicated. OpenAI announced in 2025 it would test sponsored content inside ChatGPT, and that trajectory is unlikely to reverse. The on-device case was already compelling: no API cost, no latency jitter from network round-trips, no audio leaving the device.

The numbers are concrete. A typical cloud TTS round-trip runs **300–600ms** on a good connection. Core ML on a Neural Engine-capable iPhone hits **120–180ms to first audio** for a quantized model — if you architect the pipeline correctly.

---

## The Phoneme-to-Mel Pipeline

Most distilled TTS architectures share a common spine:

Enter fullscreen mode Exit fullscreen mode

text → G2P → duration predictor → mel spectrogram → vocoder


For Core ML deployment, split this into two inference passes:

1. **Encoder + Duration Predictor** — runs once per utterance chunk, produces aligned mel frames
2. **HiFi-GAN or MB-MelGAN Vocoder** — converts mel frames to 22.05kHz PCM, streamed in chunks

Here is the minimal setup to get this working. The trick for sub-200ms is to never wait for the full utterance. Fire the vocoder on the first 50–80 mel frames while the encoder continues on the rest:

Enter fullscreen mode Exit fullscreen mode


swift
func synthesizeChunked(phonemes: [Int32], chunkSize: Int = 64) async throws {
var offset = 0
while offset < phonemes.count {
let slice = Array(phonemes[offset..<min(offset + chunkSize, phonemes.count)])
let melFrames = try await encoderModel.predict(phonemes: slice)
let audio = try await vocoderModel.predict(mel: melFrames)
audioEngine.scheduleBuffer(audio)
offset += chunkSize
}
}


First audio hits the speaker before synthesis completes. That is where the latency budget is won.

---

## ANE vs GPU: Split the Model, Don't Trust `.all`

Let me show you a pattern I use in every project. Core ML's `MLComputeUnits` gives you three paths. Here is what I measured on iPhone 13 Pro (A15, VITS-small at 22kHz, 5-word utterances):

| Compute Unit | First-Audio Latency | Power Draw | Best For |
|---|---|---|---|
| `.cpuAndNeuralEngine` | 130–180ms | Low | Attention-heavy encoder |
| `.cpuAndGPU` | 200–280ms | High | Upsampling vocoder |
| `.all` (auto) | 140–200ms | Medium | Baseline only |

The ANE excels at the encoder's attention layers. The vocoder, with its upsampling convolutions, consistently runs faster on GPU. So split the model:

Enter fullscreen mode Exit fullscreen mode


swift
let encoderConfig = MLModelConfiguration()
encoderConfig.computeUnits = .cpuAndNeuralEngine

let vocoderConfig = MLModelConfiguration()
vocoderConfig.computeUnits = .cpuAndGPU


The combined pipeline lands at **150–175ms** on A15 and newer. Do not trust `.all` — measure and override.

---

## Gotchas

**INT8 quantization will break your prosody.** This is the gotcha that will save you hours. Quantizing the duration predictor to INT8 degrades speech quality significantly — irregular pauses, flattened intonation, clipped phoneme boundaries. The quality cliff is sharp, not gradual.

| Layer Group | Safe Quantization | Notes |
|---|---|---|
| Text encoder | INT8 | Minimal perceptible impact |
| Duration predictor | **FP16 only** | INT8 breaks prosody |
| Mel decoder | INT8 | Acceptable with calibration |
| Vocoder upsampling | **FP16 only** | Audible artifacts at INT8 |

Mixed-precision lands your model at **35–55MB** — well within the 80MB ceiling I treat as the on-device viability threshold for non-game apps.

**Double-buffer your audio or you will get gaps.** Use `AVAudioPlayerNode.scheduleBuffer(_:completionHandler:)` with one chunk playing and one synthesizing. The completion handler triggers the next dispatch. Keep the audio thread hot.

**Thermal throttling is a real constraint, not an edge case.** Sustained ANE load will throttle over extended sessions — especially relevant for accessibility tooling or hands-free workflows. Design synthesis as burst-plus-pause, not a continuous stream. (Speaking of sustained screen work: apps like [HealthyDesk](https://play.google.com/store/apps/details?id=com.healthydesk) exist precisely because continuous focused sessions have physiological costs worth designing around.)

---

## Conclusion

Three things to ship with:

1. **Split compute units.** Encoder on ANE, vocoder on GPU. Measure independently on your target device — `.all` is a starting point, not a final answer.
2. **Protect the duration predictor.** Keep it at FP16 regardless of model size pressure. The perceptual cost of INT8 here far outweighs the storage savings.
3. **Stream mel chunks, not complete utterances.** 50–80 frame slices are the architectural difference between 150ms and 400ms first-audio latency.

**Relevant docs:** [Core ML Performance documentation](https://developer.apple.com/documentation/coreml) covers ANE scheduling behavior and per-chip thresholds — the authoritative source for anything that changes between silicon generations.
Enter fullscreen mode Exit fullscreen mode

Top comments (0)