I've been using cloud text-to-speech for voiceover work for a while, and two things kept bothering me.
First, every regeneration costs money — even when I'm only fixing one word in a 500-word script. Second, for NDA client work, sending unpublished scripts to a third-party API isn't an option at all.
So I built SoundScript, a desktop app that runs both cloud and local TTS engines behind one interface. This is what I learned.
What it does
Four engines, picked per task:
| Engine | Runs | Cost | Requires |
|---|---|---|---|
| ElevenLabs | Cloud | Pay per character | API key |
| Anhad | Local | Free | ~340 MB model |
| Gemini | Cloud | Free tier | API key |
| Nuqta | Local | Free | 400 MB – 2.5 GB model |
Anhad is built on Kokoro-82M, an Apache-2.0 82M-parameter TTS model running via ONNX Runtime. Nuqta is Qwen served through llama.cpp.
The caching layer
The core idea is simple: hash the inputs, use the hash as the filename, check disk before hitting the API.
def cache_key(text, voice_id, stability, similarity, speed):
payload = f"{text}|{voice_id}|{stability}|{similarity}|{speed}"
return hashlib.sha256(payload.encode()).hexdigest()
def generate(text, voice_id, stability, similarity, speed):
key = cache_key(text, voice_id, stability, similarity, speed)
path = cache_dir / f"{key}.mp3"
if path.exists():
return path # instant, zero cost
audio = backend.synthesize(text, voice_id, stability, similarity, speed)
path.write_bytes(audio)
return path
What this changes is the workflow. Generating the same line twice becomes free. You stop being precious about iterating — you can A/B two readings and keep the one you like, or tweak a comma without thinking about it.
The cache sits above both cloud and local engines, so it works identically either way. That took almost no extra code, which was the nicest surprise of the project.
The second leak: burning credits guessing sliders
Stability, similarity, style — voice direction is a skill. If you don't already know what "sadness with high stability" sounds like, you find out by generating. Which costs money.
So I added an assistant that reads the script and suggests settings:
PROMPT = """You are a voice director. Given a script excerpt,
recommend: (1) the dominant emotion, (2) stability 0-1,
(3) similarity 0-1. Output JSON."""
def suggest(script_excerpt):
return llm.complete(PROMPT, script_excerpt)
The suggestions aren't perfect — voice direction is subjective, and the model doesn't hear the voice you've chosen. But "start from a suggestion, adjust once" is much cheaper than "start from defaults, iterate six times."
Going offline
Everything above assumes internet, an API key, and a willingness to send your script to a third party. For a lot of people — for me — that assumption breaks. I work from a cabin with unreliable internet half the year, and I do client work under NDA where cloud APIs aren't an option.
So I added local engines alongside the cloud ones. The engine abstraction ended up looking like this:
ENGINES = {
"tts": {
"elevenlabs": CloudTTS(api_key=...),
"anhad": LocalTTS(model_path=...),
},
"llm": {
"gemini": CloudLLM(api_key=...),
"nuqta": LocalLLM(gguf_path=...),
},
}
The cache sits above both, so it doesn't care which one you're using. Same workflow, same hash, same instant replay.
Anhad ships 15 voices — US and UK English — and runs entirely on CPU. Nuqta comes in multiple sizes so you can trade quality for speed depending on your machine.
Packaging was the hard part
The stack:
- Python for everything
- PySide6 (Qt for Python) for the GUI — native feel on Windows and Linux without maintaining two codebases
- onnxruntime for Kokoro inference
- llama-cpp-python for Qwen inference
- lameenc for MP3 encoding, since Qt doesn't ship one by default for licensing reasons
Getting PyInstaller to bundle ONNX Runtime and llama.cpp into a single Windows executable — without dragging in a 2 GB CUDA dependency — took several attempts.
If you're doing this: build CPU-only first. Ship GPU as a separate distribution channel, not a build flag. The binary stays small, the install stays fast, and the majority of users don't notice the difference.
What I'd do differently
Ship the cache first, UI second. I spent the first month on the GUI. The caching layer was the actual product. Everything else is presentation.
Design for offline from day one. Retrofitting local engines meant rewriting the engine abstraction twice. If "no network" had been a first-class assumption from the start, the interface would have looked different and the code would be cleaner.
Cut scope ruthlessly. Four engines is four rabbit holes. The last two almost killed the project. A version with one cloud engine and one local engine would have shipped in half the time.
Where it landed
SoundScript is a paid desktop app — $29 one-time, no subscription, no telemetry, no account. Cloud engines use your own ElevenLabs and Gemini API keys. Local engines run fully offline after a one-time model download. Windows 10/11 and modern Linux.
bitprogrammer.gumroad.com/l/soundscript
If you're building something similar and want to talk about ONNX packaging, embedding llama.cpp, or the caching layer, I'm in the comments.
Top comments (0)