Yesterday, we released new Gemini Live models in the Gemini API and Google AI Studio, expanding our developer suite for building real-time, voice-f...
For further actions, you may consider blocking this person and/or reporting abuse
Love the async. live API 🥰
Questions:
The asynchronous tool-call detail is the part I’d pressure-test first. Continuing the audio stream while a tool runs is great UX, but it makes delivery semantics explicit: what should the user hear if the tool times out, returns partial data, or completes after the conversation has moved on? I’d model each call with an idempotency key and a cancellable state machine, then surface queued, running, committed, and failed states instead of letting the voice layer imply success. The 85+ language coverage and custom vocabulary are useful, but production quality will depend just as much on those boundaries as on WER.
The asynchronous function calling is the part that stands out to me. I run a voice interface on a low-power RISC-V board (MaixCAM), where replies have to be short enough to hold an 8-year-old's attention — and the bottleneck was never the model, it was the pipeline: transcription, reasoning, TTS synthesis, transfer, playback. Every hop added latency, and kids notice even ~2 seconds of silence. Running tool calls in the background while speech keeps streaming attacks exactly the right part of that problem. One thing I'd love to see documented more: how these models handle turn-taking with real-world noise — my percept stream tracks pause counts and volume trends, because barge-in and half-finished sentences are the norm with a child, not the exception.
The async function calling while streaming audio is the part that actually matters — every voice agent I've built deadlocks the moment you need to call an API mid-conversation. But that 4.0% streaming WER across 85+ languages is suspicious. I've tested transcription on mixed-language calls (Spanish-English code-switching) and the WER jumps to 15-20% easy. Would love to see the breakdown per language instead of an average that smooths over the hard cases.
Real-time bidirectional voice applications introduce a fascinating latency budget constraint for persistent context.
In standard conversational chatbots, a 1-2 second retrieval latency for user history or external context is acceptable. In live voice interactions, however, human conversational turn-taking demands a sub-500ms round-trip (speech-to-text -> memory retrieval -> LLM first token -> text-to-speech) to avoid jarring pauses or speech overlap.
If session state or caller memory takes >200ms to fetch, it breaks the illusion of natural conversation. Keeping state retrieval deterministic and strictly under 50ms is crucial when orchestrating memory alongside streaming audio feeds. Exciting to see the Live API expanding!
Really interesting approach to building real-time voice applications with Gemini Live. The combination of low-latency interaction and transcription opens up some exciting possibilities beyond traditional voice assistants.
I especially like the idea of using this for creative workflows, where users can speak naturally, capture their thoughts, and turn them into structured content without interrupting their flow. Curious to see how developers handle latency, context management, and long conversations in production.
Great write-up and useful practical examples!
Sneak peek: the Live API's 100 ms latency budget is a beast to hit—watch out for the VAD warm‑up causing cold starts that can OOM your container. Have you tamed that latency wall yet?