Simba 3.2 and Simba 3.0 just dropped to around 100ms time-to-first-audio on Coval's benchmark. Same voices, same quality, same price. It's already the cheapest model in Artificial Analysis' top ten, and the cheapest in model on Vapi's Humanness Index with a score of 97/100.
Sign-up today to try our models and hear for yourself.
FAQ
Is $10 per million characters the streaming price too?
Yes. Simba 3.2 and Simba 3.0 are one per-character rate whichever endpoint you call, and the efficiency has not changed it.
Does the faster path change how Simba sounds?
No. The release changed where and how inference is served, not the weights or the voices. We've also snuck in watermarking to meet EU regs on generative content.
Simba 3.2 or Simba 3.0?
Simba 3.2 for English, because it has the most expressive delivery and now the lowest time-to-first-audio. Simba 3.0 if you need German, Spanish, French, Italian or Brazilian Portuguese, or self-serve zero-shot cloned voices. Our legacy Simba 1.6 supports more languages, and those languages will be finding their way into Simba 3.0 before mid-November hard switch-off of Simba 1.6. Simba 3.2 and Simba 3.0 both sit at around 100ms now, so pick on language and voice support.
How does Coval measure the 100ms?
Time-to-first-audio, from the request being sent to the first audible sample.

Top comments (1)
If you're interested in trying out the SpeechifyAI API, sign up today and ping me for an extra $50 in credit (another 5M TTS characters) - first come first serve.