DEV Community

Cover image for Simba 3.2 Frontier TTS API achieves 100ms for $10/1M Characters

Simba 3.2 Frontier TTS API achieves 100ms for $10/1M Characters

Simba 3.2 and Simba 3.0 just dropped to around 100ms time-to-first-audio on Coval's benchmark. Same voices, same quality, same price. It's already the cheapest model in Artificial Analysis' top ten, and the cheapest in model on Vapi's Humanness Index with a score of 97/100.

Coval time-to-first-audio over 24 hours. Speechify Simba 3.2 and Simba 3.0 step down to ~0.10s. Cartesia Sonic 3.6 and ElevenLabs Eleven v3 Conversational stay at ~0.35s.

Sign-up today to try our models and hear for yourself.

FAQ

Is $10 per million characters the streaming price too?

Yes. Simba 3.2 and Simba 3.0 are one per-character rate whichever endpoint you call, and the efficiency has not changed it.

Does the faster path change how Simba sounds?

No. The release changed where and how inference is served, not the weights or the voices. We've also snuck in watermarking to meet EU regs on generative content.

Simba 3.2 or Simba 3.0?

Simba 3.2 for English, because it has the most expressive delivery and now the lowest time-to-first-audio. Simba 3.0 if you need German, Spanish, French, Italian or Brazilian Portuguese, or self-serve zero-shot cloned voices. Our legacy Simba 1.6 supports more languages, and those languages will be finding their way into Simba 3.0 before mid-November hard switch-off of Simba 1.6. Simba 3.2 and Simba 3.0 both sit at around 100ms now, so pick on language and voice support.

How does Coval measure the 100ms?

Time-to-first-audio, from the request being sent to the first audible sample.

Top comments (1)

Collapse
 
lukeocodes profile image
@lukeocodes 🕹👨‍💻 SpeechifyAI

If you're interested in trying out the SpeechifyAI API, sign up today and ping me for an extra $50 in credit (another 5M TTS characters) - first come first serve.