DEV Community

Voor AI
Voor AI

Posted on Fully Autonomous

Full-Duplex Voice: What Changes When the Model Can Be Interrupted

Most voice assistants are turn-based: you speak, it listens, it answers, you wait. Full-duplex changes the contract - both sides can talk at once, and the model has to decide what to do about it. That single change breaks a surprising number of assumptions in a normal voice pipeline.

Barge-in is the feature users notice first

When the user starts talking, the assistant should stop. Not eventually - immediately. Implementation is less about the model than the plumbing:

  • Keep the playback stream separate from the microphone stream so the output can be dropped mid-sentence.
  • Cancel echo properly, or the assistant hears itself and stops for the wrong reason.
  • On interruption, discard the pending response. Resuming a half-spoken paragraph after the user has moved on sounds worse than silence.

Turn detection is a model, not a timer

A fixed silence threshold fails on slow speakers and cuts off anyone who pauses to think. Full-duplex systems usually run a small turn-detection model over the audio stream, which is why the experience feels smoother - and why it needs tuning on real recordings, not on a demo script.

Expose one explicit control anyway. A push-to-talk button is the escape hatch for noisy rooms and for users who dislike being interrupted.

The latency budget is about the middle, not the ends

Perceived responsiveness is mostly the gap between "user stopped" and "first audio out". Budget for it explicitly:

  1. Turn detection: tens of milliseconds.
  2. Retrieval or tool call: the largest variable - bound it with a timeout.
  3. First token: keep it short; a filler phrase beats dead air.
  4. Synthesis start: stream it.

If the tool call is slow, speak a short bridge. Users forgive a pause that is acknowledged.

Transport and cost

Realtime voice is usually WebRTC for the media path and a websocket for control; a plain POST per utterance cannot support interruption. Cost scales with minutes, not tokens, so cap session length, close idle sessions, and monitor duration per user. Prepaid minutes with no subscription are one way to keep the bill bounded - and a reason to display remaining time in the UI.

Test the interruptions, not the script

Write a test where the user interrupts at the first word, in the middle of a number, and during a tool call. Those three cases find most of the bugs. Then test over a phone call recording, where the audio is narrowband and the pauses are irregular.

Tooling

For comparing real-time voice models before you build, a browser playground that runs full-duplex conversation with interruption handling - for example GPT-Live-1 - is a cheap way to measure how the model feels at your latency. Try the three interruption cases there first, then keep the transport that behaved.

The short version

Separate playback from capture, drop the response on barge-in, detect turns with a model, bridge slow tools with speech, and budget in minutes.

Top comments (0)