Google DeepMind announced Gemini 3.8 Live and a companion mode, Gemini 3.8 Live Extended Thinking, extending its real-time voice and video API with the reasoning capability previously reserved for non-live, text-based Gemini models.
Gemini Live has powered low-latency, spoken conversational agents — the kind used in voice assistants, in-app support, and customer-facing phone bots. Until now, that low latency came at a cost: the model responded quickly but with limited ability to reason through multi-step or conditional requests mid-conversation. Extended Thinking closes that gap by letting the live model insert a reasoning step before answering, so it can work through sequential logic — checking one condition before acting on another — without dropping out of the live conversation into a separate, slower reasoning call.
Google DeepMind's post frames this as narrowing the difference between live, spoken interaction and the deeper reasoning available in standard Gemini reasoning models. Specific pricing, latency benchmarks, and availability timelines were not detailed in the announcement and remain unconfirmed at time of writing.
The practical shift is for anyone building or buying voice-based automation. Live voice agents have generally been positioned for simple, bounded tasks — appointment booking, basic FAQ, order lookups — precisely because they couldn't reliably reason through anything more complex without breaking the real-time flow. A reasoning-capable live mode suggests vendors and in-house teams can push voice agents further into support triage, multi-step sales qualification, or operations tasks that involve checking one thing against another before responding.
The caveat is that reasoning takes time, and voice interactions are latency-sensitive in a way text chat is not. A customer on a phone line notices a two-second pause more than someone typing in a chat window. Companies piloting this should benchmark Extended Thinking against their own call flows for both accuracy and perceived response delay before assuming it's a drop-in upgrade to an existing voice deployment.
Top comments (0)