DEV Community

M Shahzad Qamar
M Shahzad Qamar

Posted on

Google Launches Gemini 3.8 Live: Advanced Real-Time Voice Models for Developers and Agents

Google announced Gemini 3.8 Live and the higher-capability Gemini 3.8 Live Extended Thinking on September 15, 2026. These native audio models are built for continuous, low-latency conversation: they can process visual input in near real time, execute function and API calls asynchronously in the background, switch among more than 90 languages mid-dialogue, and keep speaking while performing multi-step reasoning. Extended Thinking currently sits at the top of independent speech-to-speech quality and agentic task benchmarks, while the base Live model prioritizes scale and cost efficiency. Developers can access both models today through the Gemini API and Google AI Studio; they also power Search Live and are rolling into Gemini Live, Gmail, Docs, and Keep for eligible users. Pricing is set at roughly half a cent per minute of audio input and under two cents per minute of output, making sustained voice agents more economical than many competing offerings. The release moves voice interfaces from scripted assistants toward true agentic partners that can book, search, draft, and act without breaking conversational flow. For teams building customer-facing voice experiences, internal support bots, or mobile apps that need hands-free intelligence, the combination of quality, tool use, and pricing is a meaningful step forward. If you are prototyping real-time voice features in an iOS or Flutter app, or automating multi-step workflows that still require human-like dialogue, a short architecture review can map the new models to your existing stack and data boundaries.

Top comments (0)