DEV Community

Cover image for πŸ€–πŸ“žHow AI Calling Agents Actually Work: STT, LLM, TTS & the 1-Second Rule Nobody Talks About
lokesh singh tanwar
lokesh singh tanwar

Posted on

πŸ€–πŸ“žHow AI Calling Agents Actually Work: STT, LLM, TTS & the 1-Second Rule Nobody Talks About

Ever wondered how an AI can call you, understand what you say, respond naturally, and even book an appointment : all in a few seconds?

Let's break down what's actually happening behind that phone call. No PhD required. No complicated mathematics. Just the technology, explained in plain English.


A Phone Call That Made Me Think

Last week, my phone rang.

A polite voice said:

"Hi Lokesh, your car's free service is due next week. Should I book Saturday at 11 AM?"

I said yes.

It confirmed the appointment.

The entire conversation took less than a minute.

Then I realized something:

I had never spoken to a human.

There was no "Press 1 for service."

No robotic pause.

I could interrupt it.

I could change the appointment.

And somehow, it understood what I meant.

That's when the interesting question came to mind:

How does an AI actually have a phone conversation with a human?

Let's open the black box.


The 30-Second Explanation

At the simplest level, an AI calling agent repeatedly performs three jobs:

1. Listen β†’ Speech-to-Text (STT)

2. Think β†’ Large Language Model (LLM)

3. Speak β†’ Text-to-Speech (TTS)

Think of it like this:

              You speak
                  β”‚
                  β–Ό
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β”‚   STT    β”‚
            β”‚  Ears πŸ‘‚ β”‚
            β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β”‚   LLM    β”‚
            β”‚ Brain 🧠 β”‚
            β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β”‚   TTS    β”‚
            β”‚ Mouth πŸ”Š β”‚
            β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
             You hear
Enter fullscreen mode Exit fullscreen mode

That's the basic architecture.

But here's the difficult part:

Making STT + LLM + TTS work is relatively easy. Making all three work fast enough to feel human is the real engineering challenge.

And that's where things get interesting.


Part 1: STT β€” How Does an AI Hear You? πŸ‘‚

Speech-to-Text (STT), also called Automatic Speech Recognition (ASR), converts your voice into text.

For example:

Your voice
    ↓
"Can you book my service for Saturday?"
    ↓
Text
Enter fullscreen mode Exit fullscreen mode

But what's happening in between?

Step 1: Your voice becomes numbers

Your voice is basically air vibrating.

A microphone captures those vibrations and converts them into digital audio.

A traditional phone call typically uses 8 kHz audio, meaning thousands of audio samples are captured every second.


Step 2: Audio is broken into tiny pieces

The audio isn't processed as one giant file.

It's divided into very small windows, often around 20–25 milliseconds.

This allows the system to process speech continuously.


Step 3: The system analyzes the sound

The audio can be represented as a spectrogram β€” essentially a visual representation of which frequencies are present at different moments.

Think of it as a heat map of your voice.

A neural network can then analyze these patterns and predict what words you're saying.


Step 4: Context makes the transcription smarter

Consider these two phrases:

"Book my service."

"Book my surface."
Enter fullscreen mode Exit fullscreen mode

They can sound surprisingly similar.

The model uses language context to determine which interpretation makes sense.

The final result might become:

"Book my service on Saturday."
Enter fullscreen mode Exit fullscreen mode

Batch STT vs Streaming STT

This is one of the most important differences in voice AI.

Type How it works Best for
Batch STT Waits for the complete audio Videos, meeting recordings
Streaming STT Transcribes while you speak Live conversations

A phone agent needs streaming STT.

Otherwise, the system would wait for you to finish speaking completely before it even started understanding you.

Instead, it can receive partial results:

"Book"

"Book my"

"Book my ser..."

"Book my service on"

"Book my service on Saturday"
              ↑
            Final
Enter fullscreen mode Exit fullscreen mode

That small difference has a huge impact on how responsive the agent feels.


But How Does It Know When You're Done Talking?

Humans don't speak perfectly.

We pause.

We say:

"Umm... actually..."

We stop for a second to think.

So the AI needs to determine:

Did the user finish speaking, or are they just pausing?

Two important components help with this.

VAD β€” Voice Activity Detection

VAD answers:

"Is someone speaking right now?"

It distinguishes speech from silence and, depending on the system, background noise.

Endpointing

Endpointing decides:

"Okay, the user has finished their turn. We can respond now."

This is surprisingly important.

Set the timeout too short:

The AI interrupts you.

Set it too long:

You sit there listening to awkward silence.

Good voice AI needs to get this balance right.


Part 2: The LLM β€” The Brain 🧠

Once your voice becomes text, the transcript goes to an LLM β€” Large Language Model.

But a production voice agent isn't simply:

User β†’ ChatGPT β†’ Response
Enter fullscreen mode Exit fullscreen mode

There's much more happening.

A voice agent usually has three important things:

1. System instructions

These define the agent's behavior.

For example:

You are Priya, a service advisor.

Speak politely.
Keep responses under two sentences.

Your goal is to confirm the customer's
service appointment.

Never invent pricing information.
Enter fullscreen mode Exit fullscreen mode

Short responses matter.

Nobody wants to listen to an AI reading a six-paragraph essay over the phone.


2. Conversation memory

The agent needs context.

If you said:

"Saturday works."

And two turns later you say:

"Actually, make that Sunday."

The system needs to understand what "that" refers to.

That's why conversation history matters.


3. Tools

This is where a chatbot becomes an agent.

An LLM can call external functions to perform real-world actions.

For example:

User:
"Can you book it for Saturday at 11?"

        ↓

LLM decides:
"I should call book_appointment()"

        ↓

Tool:
book_appointment(
    date="Saturday",
    time="11:00"
)

        ↓

Database / Calendar:
Success

        ↓

AI:
"Done! You're booked for Saturday at 11 AM."
Enter fullscreen mode Exit fullscreen mode

The agent can potentially:

  • Check a calendar
  • Look up customer information
  • Query a CRM
  • Create an appointment
  • Send an SMS
  • Check service availability
  • Transfer the call to a human

And that's the important distinction:

A voice bot talks. An AI agent can talk AND take action.


Part 3: TTS β€” How Does an AI Talk? πŸ”Š

Now the system has generated a response.

But text isn't useful to someone on a phone call.

We need to convert:

"Your service is booked for Saturday."
Enter fullscreen mode Exit fullscreen mode

into actual audio.

That's the job of Text-to-Speech (TTS).


Old TTS vs Modern TTS

Remember those old GPS voices?

"Turn... left... in... two... hundred... meters."

They often sounded robotic because the system relied heavily on pre-recorded or concatenated speech segments.

Modern neural TTS works very differently.

The model generates speech dynamically.

It needs to understand things such as:

β‚Ή2,500
Enter fullscreen mode Exit fullscreen mode

should sound like:

"two thousand five hundred rupees"
Enter fullscreen mode Exit fullscreen mode

And:

RJ14 CD 4521
Enter fullscreen mode Exit fullscreen mode

may need to be spoken as a vehicle registration number.

The system also needs to determine:

  • Rhythm
  • Pitch
  • Stress
  • Pauses
  • Pronunciation
  • Intonation

That's called prosody.

And prosody is one reason a good AI voice can sound surprisingly natural.


Streaming TTS: Another Speed Trick

A slow system might do this:

Generate entire response
        ↓
Generate entire audio
        ↓
Start speaking
Enter fullscreen mode Exit fullscreen mode

That's going to feel slow.

A better system does this:

Generate sentence 1
        ↓
Send sentence 1 to TTS
        ↓
Start speaking

Meanwhile...

Generate sentence 2
        ↓
Send sentence 2 to TTS
Enter fullscreen mode Exit fullscreen mode

Everything happens concurrently.

And that brings us to the biggest challenge in voice AI.


The 1-Second Rule ⏱️

Here's something many beginners don't realize:

Humans are extremely sensitive to conversational delays.

If someone takes several seconds to respond after you finish speaking, the interaction starts feeling unnatural.

It doesn't matter how intelligent the answer is.

If the AI says:

"Umm... let me think..."

for three seconds before responding, you immediately know you're talking to a machine.

A production voice system tries to minimize the time between:

User stops talking
        ↓
Agent starts talking
Enter fullscreen mode Exit fullscreen mode

A simplified latency budget might look like:

Endpointing ............. 200–500 ms
STT ..................... 100–300 ms
LLM first tokens ........ 200–500 ms
TTS first audio ......... 100–200 ms
Network / telephony ..... 50–150 ms
Enter fullscreen mode Exit fullscreen mode

These numbers vary significantly depending on the architecture, provider, model, and network.

And if you simply add everything sequentially, the latency quickly becomes too high.

So what's the solution?


Everything Streams. Everything Overlaps.

Think of it like a relay race.

The next runner doesn't wait for the previous runner to finish the entire race.

The handoff happens as early as possible.

A voice agent works similarly:

Time ───────────────────────────────────────▢

STT    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
             β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
LLM             β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
TTS                 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
Audio                   β–Άβ–Άβ–Άβ–Άβ–Άβ–Άβ–Άβ–Άβ–Άβ–Άβ–Άβ–Άβ–Ά
Enter fullscreen mode Exit fullscreen mode

The pipeline overlaps:

  • STT processes speech while you're talking.
  • LLM starts generating as soon as the turn is complete.
  • TTS receives the first available text.
  • Audio starts playing before the complete response exists.

That overlap is one of the biggest secrets behind natural voice AI.


Production Tricks for Lower Latency

There are several other optimizations teams use.

Use infrastructure close to users

If most users are in India, routing every audio packet to a distant server can add unnecessary latency.

Use smaller models when possible

A simple appointment confirmation doesn't necessarily need the largest available LLM.

Cache common responses

Phrases like:

"Sure, give me a moment."

can potentially be pre-generated or optimized.

Use natural filler responses

Instead of silence while a tool executes:

"Let me quickly check that for you."

A short filler can make the interaction feel much more natural.


Barge-In β€” The Feature That Makes AI Feel Human

Here's another major difference between a basic voice bot and a good one.

Humans interrupt each other.

All the time.

Imagine:

AI:

"Your appointment is scheduled for Saturday at 11 AM, and the estimated cost isβ€”"

You:

"Wait, can we make it Sunday?"

A bad voice bot:

continues talking.

A good voice agent:

stops immediately.

This is called barge-in or interruption handling.


What happens during a barge-in?

While the AI is speaking:

  1. VAD continues listening.
  2. The system detects that the user has started speaking.
  3. TTS playback is stopped.
  4. The current LLM generation may be cancelled.
  5. The system continues the conversation from the user's new input.

That last part is especially important.

Suppose the AI was saying:

"The total cost will be β‚Ή2,500 and your appointment..."

But you interrupted before hearing the price.

The conversation state shouldn't incorrectly assume that you heard:

"β‚Ή2,500"

That's a subtle implementation detail β€” but a very important one.


So How Does AI Connect to a Real Phone? πŸ“±

So far we've only talked about the AI.

But there's another layer:

Telephony.

The architecture typically looks something like this:

πŸ“± Customer
     β”‚
     β”‚ Phone Network
     β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Telephony Layer    β”‚
β”‚ SIP / PSTN          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚
          β”‚ Live Audio
          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚       Voice Agent Server       β”‚
β”‚                                β”‚
β”‚   VAD β†’ STT β†’ LLM β†’ TTS       β”‚
β”‚             β”‚                  β”‚
β”‚             β–Ό                  β”‚
β”‚        Tools / CRM / DB        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                β”‚
                β–Ό
          πŸ“± Customer
Enter fullscreen mode Exit fullscreen mode

PSTN

The traditional phone network used by mobile phones.

SIP

A protocol used to establish and manage voice communication over IP networks.

Telephony provider

The provider acts as the bridge between the phone network and your application.

This is how your AI application can actually participate in a real phone conversation.


The Core Voice Agent Loop β€” In Code

Here's a simplified Python-style version of what the core logic can look like:

async def handle_call(call):

    history = [system_prompt]

    # 1. LISTEN
    async for user_text in stt.stream(call.audio_in):

        if not user_text.is_final:
            continue

        history.append({
            "role": "user",
            "content": user_text.text
        })

        # 2. THINK
        reply_stream = llm.stream(
            history,
            tools=[
                book_appointment,
                transfer_to_human
            ]
        )

        # 3. SPEAK
        spoken = ""

        async for sentence in split_into_sentences(reply_stream):

            # Barge-in
            if call.user_started_talking():
                tts.stop()
                break

            await call.audio_out.play(
                tts.stream(sentence)
            )

            spoken += sentence

        # Save only what was actually spoken
        history.append({
            "role": "assistant",
            "content": spoken
        })
Enter fullscreen mode Exit fullscreen mode

There are three important ideas here:

1. Everything is asynchronous.

2. Responses are streamed instead of waiting for the entire answer.

3. Barge-in is handled explicitly.

These small implementation details can make a huge difference in the final user experience.


What About Speech-to-Speech Models?

There's another approach becoming increasingly popular.

Instead of:

Voice
 ↓
STT
 ↓
Text
 ↓
LLM
 ↓
Text
 ↓
TTS
 ↓
Voice
Enter fullscreen mode Exit fullscreen mode

you can use a speech-to-speech architecture:

Voice
   β”‚
   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ One AI Model β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
     Voice
Enter fullscreen mode Exit fullscreen mode

Potential advantages

  • Lower latency
  • More natural conversation
  • Better ability to reason about vocal tone
  • Potentially more human-like interactions

But there's a trade-off

The traditional pipeline gives developers more control.

You can independently change:

  • STT
  • LLM
  • TTS
  • Prompts
  • Tools
  • Logging
  • Safety rules

For production systems, that flexibility can be extremely valuable.

So there isn't one universal "best" architecture.

It depends on the problem you're solving.


The Hard Problems Nobody Shows in Demos

AI voice demos usually happen in perfect environments.

Real phone calls are not perfect.

They're messy.

1. Background noise

Traffic.

TV.

Kids.

Construction.

A noisy workshop.

All of this can affect STT and VAD.


2. Mixed languages

This is especially interesting in India.

Someone might say:

"Mera service kab due hai, next week?"

That's a mixture of Hindi and English in one sentence.

Your STT, LLM, and TTS all need to handle that naturally.


3. Numbers, Names & Dates

Consider:

15 vs 50
Enter fullscreen mode Exit fullscreen mode

Or:

RJ14AB1234
Enter fullscreen mode Exit fullscreen mode

Or:

β‚Ή25,000
Enter fullscreen mode Exit fullscreen mode

A single incorrect digit can cause a real-world problem.

That's why reliable agents often repeat important information back to the customer.

For example:

"Just to confirm, your appointment is Saturday at 11 AM. Correct?"


4. Knowing When to Transfer to a Human

Not every problem should be handled by AI.

An angry customer.

A complicated complaint.

A sensitive request.

A situation outside the agent's authority.

A good system should know when to say:

"I'll connect you with a human representative."

Knowing when not to continue is just as important as knowing what to say.


5. Hallucinations

LLMs can confidently generate incorrect information.

On a phone call, there's no "edit message" button.

That's why production voice agents should rely on:

  • Strong system instructions
  • Verified data
  • Tool calls
  • Validation
  • Guardrails
  • Human escalation

The agent shouldn't guess when it can check.


Where Are AI Calling Agents Being Used?

We're already seeing this architecture across many industries:

πŸš— Automotive

  • Service reminders
  • Insurance renewals
  • Test-drive follow-ups
  • Customer feedback
  • Appointment booking

πŸ₯ Healthcare

  • Appointment reminders
  • Scheduling
  • Basic patient communication

🏦 Banking

  • Payment reminders
  • Basic account queries
  • Customer support

πŸ›’ E-commerce

  • Delivery confirmation
  • Return pickup coordination
  • Customer notifications

🎧 Support

  • Frequently asked questions
  • First-level support
  • Call routing
  • Human escalation

The common pattern?

High-volume + repetitive + clearly defined tasks = a strong use case for voice AI.


Want to Build One? Here's a Beginner Roadmap

You don't need to train an AI model from scratch.

Start small.

Week 1 β€” Learn STT

Record your voice.

Send it to an STT API.

Look at the transcript.

Then test it with:

  • Background noise
  • Different accents
  • Different speaking speeds

Week 2 β€” Learn TTS

Take text and convert it into speech.

Experiment with:

  • Different voices
  • Numbers
  • Dates
  • Prices
  • Names

Week 3 β€” Build the Basic Loop

Build:

Microphone
    ↓
STT
    ↓
LLM
    ↓
TTS
    ↓
Speaker
Enter fullscreen mode Exit fullscreen mode

Then measure the latency of every step.


Week 4 β€” Add Streaming & Barge-In

This is where the project starts feeling like a real voice agent.

Implement:

  • Streaming STT
  • Streaming LLM
  • Streaming TTS
  • VAD
  • Endpointing
  • Barge-in

Week 5 β€” Connect a Phone Number

Finally, connect your application to a telephony provider.

Now your laptop project becomes an actual phone-based AI agent.

And trust me β€” this is where things get really interesting.


One Tip I'd Give Anyone Building Voice AI

Log latency from day one.

Don't just measure:

"The call works."

Measure:

User stopped speaking
        ↓
Endpoint detected
        ↓
STT completed
        ↓
LLM first token
        ↓
TTS first audio
        ↓
Agent started speaking
Enter fullscreen mode Exit fullscreen mode

Because if the conversation feels slow, you need to know exactly where the delay came from.

You can't optimize what you don't measure.


Quick Recap

Component Job Key Idea
VAD Detects speech Controls turn-taking
STT Voice β†’ Text Streaming matters
LLM Thinks & decides Tools make it an agent
TTS Text β†’ Voice Stream audio early
Telephony Connects phone calls Real-world audio is messy
Orchestration Connects everything Overlap everything

The architecture may look simple:

VAD
 ↓
STT
 ↓
LLM
 ↓
TTS
Enter fullscreen mode Exit fullscreen mode

But making those components work together quickly, reliably and naturally is where the real engineering happens.


Final Thought

The next time an AI calls you to remind you about a service appointment, don't just listen to the words.

Listen to the timing.

Try interrupting it.

Change your answer halfway through.

Give it an unexpected question.

You'll start noticing all the engineering happening underneath what sounds like a simple conversation.

Because behind that one sentence β€”

"Sure, I can book that for you."

β€” there may be an entire pipeline of speech recognition, language models, tool calls, streaming audio, telephony infrastructure, latency optimization and interruption handling working together in real time.

And that's what makes voice AI so fascinating.

The conversation sounds simple.

The engineering behind it isn't.


I'm Building Voice AI in the Real World

I'm a Software Developer working with AI voice agents for the automobile industry at Alpever and that's what made this technology especially interesting to me.

As a Software Developer, I get to see how all these pieces come together in a real-world product β€” not just the AI model, but the entire system behind the conversation.

The more I work with voice AI, the more I realize that building a good voice agent isn't just about choosing a powerful LLM.

It's about the entire system:

Latency + Audio + AI + Tools + Telephony + Reliability + UX

If you're also exploring voice AI, I'd love to hear what you're building.

Would you be comfortable talking to an AI on a phone call?

Yes or No β€” and why?

Drop your answer in the comments.

If this post helped you understand what's happening behind AI phone calls, save it for later and share it with someone who's getting into voice AI.

Top comments (0)