I'm building a Mac app where a parent clones their own voice on their own machine, and the app turns reading homework into audio with word-by-word highlighting — aimed at kids with dyslexia who can follow along while listening. The hard requirement came from the use case: a child's voice samples and reading materials never leave the laptop. No cloud calls, no subscription, no telemetry. Everything runs locally on Apple Silicon.
That meant finding a TTS engine that can (a) clone a voice from ~6 seconds of reference audio, (b) run at or near real time on a base M1 Pro, (c) be commercially usable, and (d) be small enough to ship inside a desktop app.
It took three bake-offs, five engines, and a pile of traps. Here's what I learned — including one lesson that flipped the whole result.
First, the license minefield
Before any benchmark: most of the famous zero-shot voice cloning models are legally unusable for a commercial product.
- F5-TTS — CC-BY-NC. Non-commercial. Out.
- XTTS-v2 (Coqui) — CPML. Out.
- self-hosted fish-speech — research license. Out.
That leaves the Apache/MIT family: Kokoro-82M, ZipVoice, Chatterbox (MIT), and Qwen3-TTS. All benchmarks below ran on a MacBook Pro M1 Pro / 32GB, cloning the same reference voice (a LibriVox narrator, 60s of dense speech) reading ~40 seconds of Peter Rabbit.
Round 1: The obvious candidates
Chatterbox: one-shot generation of the full 40s passage. The output sounded fine. The problem was speed: 808 seconds of compute for 40 seconds of audio (RTF 20). On MPS it's an order of magnitude too slow for a consumer app. Out.
ZipVoice (123M, from the k2-fsa folks): genuinely fast — per-sentence wall clock of 10–16s including ~7s of model loading per call, which works out to roughly RTF 0.6–1.5 once you keep the model resident. But it had sharp edges:
- The prompt text must exactly match the prompt audio. A 28s reference clip paired with only the first sentence's transcript produced 44 seconds of that one sentence repeated forever. A 6s clip paired with a 28s full transcript produced 0.4–2s fragments. Exact 6s/6s — perfect.
- Batch mode on MPS misbehaves: the same inputs that work per-sentence went into repetition loops in batch mode. Suspected padding bug. Production pipeline calls it one sentence at a time.
- The
zipvoicepackage on PyPI is an empty shell — the real code ispip install git+https://github.com/k2-fsa/ZipVoice, plus a handful of dependencies that aren't declared.
At this point ZipVoice was the front-runner: near-real-time, Apache-2.0, 468MB.
Round 2: Qwen3-TTS-0.6B on PyTorch — "local is impossible"
Qwen3-TTS-0.6B-Base is Apache-2.0 and the community reported RTF 0.86 for cloning. But every one of those numbers came from CUDA. On the M1 Pro, with the official qwen-tts package:
-
MPS + bf16: crashes in the embedding layer (
Placeholder storage has not been allocated), on both torch 2.9 and 2.13. - CPU, fp32: seven minutes, zero output.
- MPS + fp32 with device_map: runs — at RTF ≈ 9. A ten-word sentence took ~6 minutes. Fifteen times slower than ZipVoice.
There's a second trap with Qwen's in-context-learning (ICL) cloning: it's extremely sensitive to reference audio quality. A trimmed reference clip fed with the full-length transcript generated garbage that ran to the token limit — mostly silence and noise.
I wrote it off: local Qwen3-TTS on M-series is unusable today. I planned to fall back to a cloud API for the premium tier and keep watching for MPS support upstream.
Round 3: Same weights, different runtime
The round-2 verdict was only true for PyTorch. The community had already ported the weights to MLX: mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16 (2.4GB), running through mlx-audio.
Same model, same machine:
| Metric | PyTorch (MPS) | MLX |
|---|---|---|
| Average RTF | ~9 | 0.88 |
| Model load | minutes / crashes | 1–2 seconds |
| Content correctness | — | 7/7 sentences correct on whisper round-trip |
Faster than real time. Suddenly the "unusable" model was the fastest cloning engine of everything I'd tested, with the best quality-to-size ratio.
The biggest lesson of the whole project: the model card is not the performance story — the framework port is. If a model looks too slow on your platform, check whether someone has re-implemented it for your runtime before you write it off.
The gotcha that cost the most time
Qwen's ICL cloning demands that the reference audio be sentence-aligned at the start, with a verbatim transcript. Feed it a sloppy reference and it doesn't degrade gracefully — it fills the output to the token ceiling with noise. Interesting contrast: ZipVoice eats dirty references without blinking; Qwen does not.
So the product pipeline is now mandatory for any engine:
record → whisper transcript → trim at sentence boundaries → feed engine
A few more hard-won environment notes, in case you're walking the same road:
-
Hugging Face downloads through a proxy can hang forever on the new Xet channel — zero bytes, zero errors.
export HF_HUB_DISABLE_XET=1forces the classic HTTP path and fixes it. - Reference length is your biggest speed lever. A 28s reference re-encodes per sentence (RTF 7–94 in the worst case); 6s drops it to 1–5. Guide users to record 6–10 clean seconds.
- In
mlx_audio.generate_audio,output_pathis a directory andfile_prefixis the filename.
4-bit quantization won on every axis
Before shipping, I compared the 4-bit quantized weights against bf16, expecting a quality cliff:
- Speaker similarity (resemblyzer): 0.928 (4-bit) vs 0.923 (bf16)
- Content fidelity (whisper round-trip): 0.992 vs 0.992 — parity
- Speed: RTF 0.63 vs 1.11 — 4-bit is nearly 2× faster
- Size: 1.71GB vs 2.4GB
Cross-similarity between the two arms was 0.937 — effectively the same voice. An F0 analysis confirmed quantization didn't flatten prosody. The download pack ships 4-bit.
The supporting cast
- Kokoro-82M (Apache-2.0, 0.39GB as MLX bf16) for the 54 preset voices — story-time character voices, script mode, etc. English is its A-tier language; we only use it as a complement.
- whisper.cpp (MIT, static build) for word-level alignment — the thing that makes karaoke-style highlighting possible. 58.8s of audio → complete word timestamps in 2.87s. One subtlety: the timeline comes from whisper, but the word shapes come from the source text; a SequenceMatcher maps one to the other (166/166 words covered on the reference passage), so whisper mishearings don't corrupt the highlighting.
- Shell: Tauri 2 + React/TypeScript + Python sidecar, with the engine abstracted behind a provider interface — because MLX is Apple-only, and the Windows path (ZipVoice on CPU, or a cloud fallback) is phase two.
Real-world throughput on the M1 Pro: ~1–2.5 minutes per 1,000 characters (idle to loaded), a full ~10-minute chapter processes in the background in 10–25 minutes, and power draw is a rounding error (~0.2Wh per chapter by estimate — the battery doesn't notice).
What this adds up to
A parent records 6 seconds of reading, and a few minutes later a self-contained HTML player exists with their voice reading the chapter, words highlighting as they're spoken. No cloud, no API keys, no monthly fee — which is only possible because of how good the Apache/MIT TTS ecosystem has gotten, and because of the MLX port of Qwen3-TTS.
If you're evaluating local TTS on Apple Silicon, the short version: Kokoro for preset voices, Qwen3-TTS-0.6B (4-bit, via mlx-audio) for cloning, whisper.cpp for alignment — and budget a day for the reference-audio pipeline, because that's where the quality actually lives.
I'm building this as Elmren Voice ($39.99 one-time, macOS 13+). There's a live demo player exported from the real product if you want to hear the word-highlighting in action.
Thanks to the maintainers of mlx-audio, Qwen3-TTS, Kokoro, ZipVoice, and whisper.cpp — this product only exists because your licenses and ports do.
Top comments (0)