DEV Community

Mikhail Savchenko
Mikhail Savchenko

Posted on Originally published at inite.ai

Google Ships Two New TTS Models Built for Production Voice Agents

Google DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026, expanding its Gemini Audio family with what it calls its most expressive voice generation models to date.

Flash TTS is built for creative direction: developers can generate entirely new character voices from natural language prompts, direct delivery line by line (pacing, dialect, acting cues), and script two-speaker scenes with vocal bursts like <laughs> or <sigh> for realistic turn-taking. Flash-Lite TTS is positioned for high-volume, cost-efficient use cases such as dubbing and voice agents, with the same fine-grained tone and pacing controls.

Both models support over 100 languages and dialects, and Google says the library scales from 30 original voices to more than 2,000 production-ready voices, including regional varieties like Mexican Spanish, Quebec French, and Scots English. Voice replication lets a user recreate a consistent vocal profile from a 30-second sample, gated by consent verification, SynthID watermarking, and C2PA credentials. Voice remixing — fine-tuning an existing library voice with prompts — is coming soon. Voice replication is not available in Illinois, Texas, the EEA, UK, Switzerland, or India.

On benchmarks, Google reports Flash TTS took the top spot on Hume AI's Voice Design Benchmark (71.4) and led accent modeling (60.8), while Flash and Flash-Lite placed first and second respectively on Hume AI's Overall Quality Index. In blind evaluations on Voice Arena, both models led among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

Rollout starts today for developers through the Gemini API and Google AI Studio for both models. Enterprise access via API in Gemini Enterprise is listed as coming soon. Flash TTS also reaches general users through Gemini Notebook, while Flash-Lite TTS reaches general users through Google Vids. Developer platforms including Agora, LiveKit, Pipecat, and Vercel are enabling deployment of these models, and Google names Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang as partners integrating the TTS models for dubbing, localization, and conversational voice agents.

Top comments (0)