DEV Community

Mirage Cloud IA
Mirage Cloud IA

Posted on

How AI Voice Agents Actually Work

AI phone agents went from obviously robotic to occasionally indistinguishable from a person in about two years. The reason is not one breakthrough. It is that three separate components got good at the same time, and the engineering problem of connecting them got solved well enough.

Here is what is actually happening during a call, and why some systems feel natural and others do not.
The pipeline
A voice agent is three models in a loop, plus telephony.

  1. Speech to text. Audio arrives from the phone network and is transcribed. Modern systems do this in a streaming fashion, transcribing continuously rather than waiting for the caller to finish, which matters enormously for perceived speed.

  2. The language model. The transcript, plus the conversation history and whatever instructions and business knowledge the agent has been given, goes to a language model, which decides what to say and sometimes what to do.

  3. Text to speech. The response is synthesised as audio and sent back down the line. Also streamed, so speech starts before the full response has been generated.

Around all of that sits telephony: SIP trunking, a phone number, call routing, and transfer to a human.
Latency is the whole game
The single number that determines whether a call feels natural is the gap between the caller finishing and the agent starting to speak.

Human conversation runs on gaps of roughly 200 to 300 milliseconds. Under about 800 milliseconds feels acceptable. Past about a second and a half, people start talking again because they assume they were not heard, which breaks the conversation entirely.

That budget has to cover audio transport, transcription, the language model producing at least its first tokens, speech synthesis starting, and audio transport back. Each component has to be fast, and the architecture has to overlap them rather than run them in sequence.

This is why streaming matters at every stage. A system that waits for the complete transcript, then waits for the complete response, then synthesises the complete audio, will feel sluggish even if every individual component is fast.

It is also why model choice involves a trade-off. A larger model gives better answers and takes longer. Many production systems use a smaller, faster model for conversation and escalate to a larger one only when the request warrants it.
Turn-taking is harder than it looks
Knowing when the caller has finished speaking is a genuinely difficult problem, and it is where most systems reveal themselves.

Simple approaches use voice activity detection: if there is silence for some threshold, assume the turn is over. Set the threshold short and the agent interrupts people who paused to think. Set it long and every exchange feels laggy.

Better systems use semantic endpointing, judging from the content whether an utterance sounds complete. "My name is Sarah and my number is" is clearly unfinished even followed by a two-second pause. "That's all, thanks" is clearly finished immediately.

Barge-in is the related capability: letting the caller interrupt the agent mid-sentence and having the agent stop talking and listen. Real callers interrupt constantly. A system that talks over an interruption feels immediately robotic, and it is one of the clearest tests when evaluating a product.
Voice synthesis
The output quality difference between a good and a poor voice agent is largely a text-to-speech question.

What separates the current generation from older systems is prosody: the rhythm, emphasis and intonation of speech. Older synthesis produced correct words with flat delivery. Current models handle emphasis, natural pauses and rising intonation on questions.

The remaining tells are usually specific: proper nouns, street names, unusual surnames, and numbers read in the wrong grouping. This is why testing a voice agent on your actual local addresses matters, particularly outside English. A system trained predominantly on English will mangle French street names and struggle with regional pronunciation, and you will only find out by trying it.
Knowledge and actions
A voice agent that only converses is a limited product. The useful ones can look things up and do things.

Two mechanisms:

Retrieval. The agent is given access to reference material, business hours, services, pricing, policies, and pulls the relevant part into context when answering. This is what keeps it accurate and on-topic rather than improvising.

Tool calling. The model can invoke functions: check a calendar, create a ticket, look up an order, send a message, transfer the call. This is what turns a conversation into an outcome.

The second one is where most of the practical value sits, and it is worth asking about specifically when evaluating a product, because a demo of natural conversation says nothing about whether the system can actually book anything.
Configuration is a product decision
Early voice agent products expected you to write prompts. That works for developers and fails for the businesses that most need this, since a plumber does not want to learn prompt engineering.

The better approach asks structured questions about the business, activity, services, hours, tone, escalation rules, and builds the configuration from the answers. Mirage Cloud's receptionist agent works this way, with a guided setup and a browser test before a phone number is attached, which is a sensible sequence: you find the problems yourself rather than through a customer complaint.
What to test before buying
Interrupt it mid-sentence. Does it stop and listen, or talk over you?

Pause mid-sentence. Does it wait, or jump in?

Give it a local address and an unusual surname. Does it handle them?

Ask something outside its scope. Does it say so and transfer, or invent an answer? This is the most important test, because in production you will not know when it is wrong.

Ask it to actually do something. Book, look up, transfer. Conversation quality and task completion are different capabilities.
The compliance point
Since 2 August 2026, EU transparency rules require that people are told when they are interacting with an AI rather than a person, unless it would be obvious. On a phone call it is not obvious, which makes this a required line in the greeting rather than an optional courtesy.

Worth checking that any product you evaluate handles this by default rather than leaving it to you to remember.

Top comments (1)

Collapse
 
octyn profile image
OCTYN •

the browser test before attaching a real number is a good call. i'd make the test include an interruption right as it says a price or appointment time, then ask it to repeat the final details. sounding natural is nice, but a wrong booking said smoothly is still a wrong booking.