DEV Community

Gautam
Gautam

Posted on

Building Bharat Voice AI: My 10-Day Journey from a Voice Agent to a Multilingual Multi-Agent System

 # Building Bharat Voice AI: My 10-Day Journey from a Voice Agent to a Multilingual Multi-Agent System

Over the last 10 days, I took a simple idea, a voice assistant, and gradually turned it into a more complete multilingual voice-agent system.

This project is called Bharat Voice AI.

The goal was to build a voice agent that can communicate naturally with users in India, understand different languages and code-mixed speech, remember returning users, use real-world tools, make outbound calls, escalate to humans, measure its performance, and hand specialized tasks to another agent.

This journey was part of:

10 Days of Voice Agents — VoiceForBharat Edition

The project uses Murf Falcon for voice generation together with LiveKit, Gemini, Deepgram, Python, SQLite, and real external tools.


What is Bharat Voice AI?

Bharat Voice AI is a multilingual voice assistant designed around the idea that voice interfaces should feel natural for Indian users.

Instead of forcing users to interact only through a text interface, the system allows them to speak naturally.

It supports:

  • English
  • Hindi
  • Gujarati
  • Hinglish and code-mixed conversations
  • Persistent user memory
  • Real-time tools
  • Human escalation
  • Outbound voice calls
  • Call analytics
  • Specialist-agent handoff

The system is designed so that the main agent does not have to do everything itself.

When a task requires specialized knowledge, it can hand the conversation to a specialist agent.


Why Voice?

India has a huge diversity of languages, communication styles, and levels of digital literacy.

A voice interface can make technology easier to access because users do not have to type everything.

For Bharat Voice AI, I wanted the interaction to feel closer to a normal conversation.

For example:

"Veraval mein aaj weather kaisa hai?"

The system should understand the intent even though the sentence mixes Hindi and English.

It should then respond naturally in the appropriate language.


Architecture

The basic architecture looks like this:

                    Bharat Voice AI
                           |
                     LiveKit Agents
                           |
              +------------+------------+
              |                         |
           Deepgram                   Gemini
             STT                       LLM
              |                         |
              +------------+------------+
                           |
                         Tools
              +------------+------------+
              |            |            |
           Weather      Memory      Escalation
              |            |            |
          Open-Meteo     SQLite       SQLite
                           |
                    Specialist Agents
                           |
                     Murf Falcon TTS
                           |
                         User
Enter fullscreen mode Exit fullscreen mode

Day 1 - Building the Voice Agent
The first step was getting the basic voice pipeline working.
The core pipeline became:

User Speech
     |
     v
Deepgram STT
     |
     v
Gemini
     |
     v
Murf Falcon
     |
     v
User Voice
Enter fullscreen mode Exit fullscreen mode

The agent could listen to the user and respond using voice.
The main technologies were:
LiveKit Agents
Deepgram
Gemini
Murf Falcon
Day 2 - Persona and Guardrails

The next step was defining who the agent actually is.

The agent became:

Bharat Voice AI

I added clear instructions covering:

Identity
Objectives
Knowledge boundaries
Language behavior
Guardrails
Escalation
Conversation style

The agent should not pretend to know something it does not know.

It should also know when a request is outside its role.

This became especially important later when I added human escalation and specialist agents.
Multilingual Conversations

One of the important goals was supporting Indian languages.

The agent supports:

English

"What's the weather today?"

Hindi

"आज वेरावल में मौसम कैसा है?"

Gujarati

"આજે વેરાવળમાં હવામાન કેવું છે?"

I also tested code-mixed speech such as:

"Veraval mein aaj weather kaisa hai?"

The important part is not simply detecting the language.

The agent should also respond in the appropriate script.

For example, Hindi should be written and spoken naturally rather than being forced into Romanized Hindi.
Day 3 - Frontend

After the voice pipeline worked, I customized the frontend.

The frontend clearly represents the state of the agent.

The main states are:

Ready
  |
  v
Connecting
  |
  v
Listening
  |
  v
Speaking
  |
  v
Call Ended
Enter fullscreen mode Exit fullscreen mode


t
Day 4 - Persistent Memory

A voice agent that forgets everything after every call isn't very useful for returning users.

So I added persistent memory using SQLite.

The profile stores information such as:

user_id
name
language_preference
facts
last_interaction

The important design decision was that memory is handled by backend functions rather than being written into the LLM prompt.

The agent can:

Look up a caller.
Ask permission before saving information.
Save approved information.
Retrieve the profile during a later conversation.

For example:

First conversation:

"My name is Gautam."

The agent asks whether it should remember the name.

After permission is granted, the information is stored.

During a later conversation, the agent can recognize the returning user.

Day 5 - Real Tools

The agent needed to do more than generate answers from the LLM.

I added a real weather tool.

The weather flow is:

User
 |
 | "What is the weather in Veraval today?"
 v
Gemini
 |
 | tool call
 v
get_weather()
 |
 v
Weather API
 |
 v
Real weather data
 |
 v
Gemini
 |
 v
Murf Falcon
 |
 v
User
Enter fullscreen mode Exit fullscreen mode

The important rule is:

The agent must never invent current weather information.

If the weather service is unavailable, the agent should say that it could not retrieve the latest information.

It should not guess.

One Real Debugging Lesson

One of the problems I encountered during development was a Deepgram connection failure.

The voice pipeline produced an error similar to:

APIConnectionError:
failed to connect to deepgram

Instead of assuming the code was broken, I tested the network connection from Windows PowerShell.

I used:

Test-NetConnection api.deepgram.com -Port 443

Initially the connection failed.

After investigating the network/DNS path, I eventually got:

TcpTestSucceeded : True

This was a good reminder that voice-agent problems are not always LLM problems.

The failure can be caused by:

Network connectivity
DNS
Firewall
API availability
WebSocket connections
Authentication
Tool execution
Frontend state

Debugging the complete pipeline is essential.

Day 6 - Outbound Voice Calls

The next step was making the agent call a user instead of waiting for the user to open the browser.

I integrated outbound calling using LiveKit telephony/SIP.

I also tested the Linphone route.

The basic flow became:

Bharat Voice AI
      |
      v
LiveKit SIP
      |
      v
Linphone
      |
      v
    Phone
Enter fullscreen mode Exit fullscreen mode


The outbound conversation needs a different opening from a browser conversation.

The agent must immediately explain:

Who is calling
Why it is calling
How the user can end the call

I also tested the SIP connection and call lifecycle.

Day 7 - Human Escalation

A voice agent should not try to solve everything.

I added a human escalation workflow.

When a user says:

"I want to talk to a human."

the agent can start the escalation process.

But it does not automatically share information.

It first asks for permission.

For example:

"I can create a request for human assistance. Before I do, I would share your name, the issue you described, what I checked, your preferred language, and the urgency. Would you like me to create the request?"

If the user says yes, the backend creates a real escalation record.

A reference ID is generated by the backend.

The important rule is:

Never tell the user that an escalation was created unless the database operation actually succeeded.

This avoids fake success messages.

Day 8 - Call Analytics

After adding many features, I needed a way to measure what was happening.

So I built a Call Analytics Dashboard.

The dashboard tracks:


The data comes from real calls stored in SQLite.

It is not hardcoded.

A call is considered successful when the user's intended task is successfully completed.

For example:

User asks for weather
        |
Weather tool succeeds
        |
Weather information delivered
        |
     SUCCESS
Enter fullscreen mode Exit fullscreen mode

If the required tool fails or the task is incomplete, the call can be recorded as failed or incomplete.

The dashboard can therefore show actual agent performance.

Day 9 - Specialist Agent Handoff

This was one of the most interesting parts of the project.

Instead of forcing the main agent to handle every type of request, I created a specialist:

Bharat Weather Specialist

The architecture became:

                   User
                     |
                     v
              Bharat Voice AI
                     |
              Weather request?
                 /       \
               No         Yes
               |           |
               v           v
          Main Agent   Weather Specialist
                           |
                           v
                      get_weather()
                           |
                           v
                       Real Data
Enter fullscreen mode Exit fullscreen mode


python
For example:

User:

"What is the weather today in Veraval?"

The main agent says:

"For detailed weather information, I'll connect you with our weather specialist."

The specialist then takes over.

The important part is that the user should not have to explain the entire question again.

The conversation context is passed to the specialist.

Native Agent Handoff

The LiveKit agent-handoff pattern allows the main agent to return a specialist agent together with an announcement.

Conceptually:

@function_tool
async def handoff_to_weather_specialist(
    self,
    context: RunContext,
) -> tuple[Agent, str]:
    specialist = BharatWeatherSpecialist(
        chat_ctx=self.chat_ctx.copy(
            exclude_instructions=True
        )
    )

    return (
        specialist,
        "I'll connect you with our weather specialist."
    )
Enter fullscreen mode Exit fullscreen mode

The actual implementation in the repository follows the project's LiveKit version and architecture.

The important concept is that the specialist receives the previous conversation context.

A Day 9 Bug That Taught Me Something

During testing, the handoff itself worked.

However, I encountered a problem where the internal function-call information could appear as visible text instead of the actual weather response.

For example, the user could see something similar to:

The function called is get_weather
location = Veraval
forecast_days = 1
Enter fullscreen mode Exit fullscreen mode

That is not acceptable for a real voice assistant.

The user should hear:

"The latest weather in Veraval is..."

not internal tool-call information.

This led me to investigate the difference between:

Function registration
Tool execution
LLM output
LiveKit agent handoff
Frontend message rendering

It was another reminder that a production voice agent is a complete system, not simply an LLM with speech.

The Complete SystemAfter the nine days, Bharat Voice AI looks roughly like this:

                         USER
                          |
                    Voice / Browser
                          |
                          v
                    LiveKit Agents
                          |
                    Bharat Voice AI
                          |
        +-----------------+------------------+
        |                 |                  |
      Memory           Tools              Routing
        |                 |                  |
      SQLite       +------+-------+          |
                   |              |          |
                Weather      Other Tools     |
                   |                         |
              Open-Meteo                     |
                                             |
                                      Specialist Agent
                                             |
                                      Weather Specialist
                                             |
                                      get_weather()
                                             |
        +------------------------------------+
        |
   Human Escalation
        |
      SQLite
        |
    Murf Falcon
        |
        v
       USER
Enter fullscreen mode Exit fullscreen mode

Outbound calls extend the system through SIP/telephony.

Analytics records the outcome of calls and makes it visible through the dashboard.

Technology Stack

The main technologies used in the project include:

Python
LiveKit Agents
Gemini
Deepgram
Murf Falcon
SQLite
Open-Meteo
SIP / Linphone
Browser frontend
GitHub

Each component has a different responsibility.

Deepgram
Speech → Text

Gemini
Reasoning + Conversation

Tools
Real-world data and actions

SQLite
Persistent state

LiveKit
Real-time voice transport and agent orchestration

Murf Falcon
Text → Natural Voice
Enter fullscreen mode Exit fullscreen mode

Security

One of the most important lessons is to never put API credentials directly into the source code.

Use environment variables.

For example:

DEEPGRAM_API_KEY=your_key_here
MURF_API_KEY=your_key_here
LIVEKIT_API_KEY=your_key_here
LIVEKIT_API_SECRET=your_secret_here
Enter fullscreen mode Exit fullscreen mode

The real .env file should never be committed to GitHub.

Use:

.env
Enter fullscreen mode Exit fullscreen mode

locally and:

.env.example
Enter fullscreen mode Exit fullscreen mode

for documentation.

Never publish:

  • API keys
  • SIP credentials
  • passwords
  • OTPs
  • PINs
  • caller information
  • private database records

Running the Project

The complete project is available on GitHub:

https://github.com/gunmasterg9/bharat-voice-ai

A typical setup is:

git clone https://github.com/gunmasterg9/bharat-voice-ai.git

cd bharat-voice-ai

cd backend

uv sync
Enter fullscreen mode Exit fullscreen mode

Create your environment configuration:

.env
Enter fullscreen mode Exit fullscreen mode

Add the required API credentials.

Then start the backend using the project's configured startup command.

Start the frontend and open the browser interface.

Allow microphone access.

Then start a conversation.

For example:

"Hello Bharat Voice AI."

Then test:

"What is the weather today in Veraval?"

Then test multilingual interaction:

"આજે વેરાવળમાં હવામાન કેવું છે?"

Then test memory:

"My name is Gautam."

Then restart the application and verify that the saved profile can be retrieved.

Testing

I created tests covering important parts of the system.

The project includes testing for:

Agent behavior
Memory
Persistent memory
Language switching
Weather tools
Escalation
Outbound calls
Linphone
Analytics
Specialist handoff

The Day 9 development test suite included specialist handoff tests alongside the previous functionality.

The important lesson was that every new feature should be tested without breaking the previous days' work.

What I Learned

The biggest lesson from this challenge is that building a voice agent is much more than connecting an LLM to a TTS API.

A reliable voice agent needs:

Voice
+
Reasoning
+
Memory
+
Tools
+
State
+
Error Handling
+
Security
+
Human Handoff
+
Observability
Enter fullscreen mode Exit fullscreen mode

A model can generate a great response, but the surrounding system determines whether the product is actually reliable.

I also learned to debug the entire pipeline instead of assuming every problem is caused by the LLM.

What I Would Build Next

There is still a lot I would like to add.

Future improvements could include:

More specialist agents
Better interruption handling
More Indian languages
Better low-bandwidth support
Advanced call analytics
Conversation quality scoring
Better tool observability
More real-world Indian datasets
Improved outbound-call workflows
Specialist-to-specialist routing
More sophisticated RAG
Production deployment and monitoring

The long-term goal would be to turn Bharat Voice AI into a platform where different specialized voice agents can work together.

Final Thoughts

The most interesting part of the challenge wasn't building the first voice conversation.

It was everything that came afterward.

Making the agent remember.

Making it use real data.

Making it call a phone.

Making it know when to ask a human.

Measuring whether conversations actually succeeded.

And finally, teaching one agent when another agent is better suited to help.

That progression changed how I think about voice AI.

A voice agent isn't just a chatbot that speaks.

It can become a complete software system with memory, tools, workflows, specialized agents, and real-world actions.

That is what I wanted to explore with Bharat Voice AI.

Project Links
GitHub

https://github.com/gunmasterg9/bharat-voice-ai
All Videos in Linkedin
https://www.linkedin.com/in/gautam-vandar-71a75632b/

Challenge

10 Days of Voice Agents — VoiceForBharat Edition

Voice Technology

Murf Falcon

Thank You

Thank you to Murf AI for organizing the 10 Days of Voice Agents challenge and providing the opportunity to build, experiment, debug, and learn through a real voice-agent project.

Building Bharat Voice AI over these 10 days was a great experience, and I hope this project helps someone else start building their own voice agent.

Top comments (0)