Voice features look like a thin wrapper around a model until you ship one. The model is the easy part; the latency budget, the turn-taking and the failure modes are the work. Here is the shape that holds up in production.
TTS: stream, do not wait
A text-to-speech call that returns a whole file is fine for a podcast, and wrong for a conversation. Streaming synthesis starts playing while the rest of the sentence is still being generated, which is the difference between a responsive assistant and a pause that feels broken.
Practical rules:
- Split long text into sentences and synthesize them as a queue, so the first sentence plays early.
- Pick two voices at most. Switching voices mid-conversation destroys the illusion faster than any artifact.
- Handle numbers, units and abbreviations before the model sees them. "2.5 GB" read as "two point five gigabytes" is a preprocessing job, not a voice problem.
- Insert explicit pauses instead of punctuation tricks; a comma cannot express a two-second beat.
ASR: spend the latency budget deliberately
Transcription has three modes, and most apps pick the wrong one:
- Batch, when the user is done talking and you need the final text.
- Streaming partial results, when you want the transcript to appear while speaking.
- Realtime bidirectional, when the app must hear and speak at once.
Partial results are for the interface. Do not run your business logic on them; they change. Run it on the final segment, and design the UI so a corrected transcript does not invalidate what the user already saw.
Punctuation, casing and speaker labels all cost extra latency. Turn them on only where the output is read by a person.
Turn-taking is a product decision
When is the user finished? A fixed silence threshold fails on slow speakers and on people who pause mid-sentence. A model-based turn detector fails on background noise. The workable compromise is an adaptive threshold plus an explicit "push to talk" escape hatch for noisy rooms.
If the assistant speaks, support interruption. Barge-in - stopping playback the moment the user starts talking - is one of the few voice features users immediately notice.
Cache and cap
Cache synthesis for fixed strings: greetings, error messages, onboarding prompts. Cap the length of a single generation so one pasted document cannot stall the session. Log audio durations per request; that number, not token count, is what your bill scales with.
Test with bad audio
Record a test set with three people, one fan, one open window and one phone call. Play it through the pipeline before every release. Clean studio audio passes everything and proves nothing.
Tooling
If you want to compare models before committing to one, a studio that exposes several routes in one place helps - speech synthesis, transcription, real-time voice and music generation under one roof, such as StepAudio 3. Use it to measure latency and quality on your own test set, then wire the winner into your app.
The short version
Stream the audio, keep one or two voices, run logic on final transcripts only, make turn-taking adaptive, and test on recordings you would apologise for.
Top comments (0)