# Building Bharat Voice AI: My 10-Day Journey from a Voice Agent to a Multilingual Multi-Agent System
Over the last 10 days, I took a simple idea, a voice assistant, and gradually turned it into a more complete multilingual voice-agent system.
This project is called Bharat Voice AI.
The goal was to build a voice agent that can communicate naturally with users in India, understand different languages and code-mixed speech, remember returning users, use real-world tools, make outbound calls, escalate to humans, measure its performance, and hand specialized tasks to another agent.
This journey was part of:
10 Days of Voice Agents — VoiceForBharat Edition
The project uses Murf Falcon for voice generation together with LiveKit, Gemini, Deepgram, Python, SQLite, and real external tools.
What is Bharat Voice AI?
Bharat Voice AI is a multilingual voice assistant designed around the idea that voice interfaces should feel natural for Indian users.
Instead of forcing users to interact only through a text interface, the system allows them to speak naturally.
It supports:
- English
- Hindi
- Gujarati
- Hinglish and code-mixed conversations
- Persistent user memory
- Real-time tools
- Human escalation
- Outbound voice calls
- Call analytics
- Specialist-agent handoff
The system is designed so that the main agent does not have to do everything itself.
When a task requires specialized knowledge, it can hand the conversation to a specialist agent.
Why Voice?
India has a huge diversity of languages, communication styles, and levels of digital literacy.
A voice interface can make technology easier to access because users do not have to type everything.
For Bharat Voice AI, I wanted the interaction to feel closer to a normal conversation.
For example:
"Veraval mein aaj weather kaisa hai?"
The system should understand the intent even though the sentence mixes Hindi and English.
It should then respond naturally in the appropriate language.
Architecture
The basic architecture looks like this:
Bharat Voice AI
|
LiveKit Agents
|
+------------+------------+
| |
Deepgram Gemini
STT LLM
| |
+------------+------------+
|
Tools
+------------+------------+
| | |
Weather Memory Escalation
| | |
Open-Meteo SQLite SQLite
|
Specialist Agents
|
Murf Falcon TTS
|
User
Day 1 - Building the Voice Agent
The first step was getting the basic voice pipeline working.
The core pipeline became:
User Speech
|
v
Deepgram STT
|
v
Gemini
|
v
Murf Falcon
|
v
User Voice
The agent could listen to the user and respond using voice.
The main technologies were:
LiveKit Agents
Deepgram
Gemini
Murf Falcon
Day 2 - Persona and Guardrails
The next step was defining who the agent actually is.
The agent became:
Bharat Voice AI
I added clear instructions covering:
Identity
Objectives
Knowledge boundaries
Language behavior
Guardrails
Escalation
Conversation style
The agent should not pretend to know something it does not know.
It should also know when a request is outside its role.
This became especially important later when I added human escalation and specialist agents.
Multilingual Conversations
One of the important goals was supporting Indian languages.
The agent supports:
English
"What's the weather today?"
Hindi
"आज वेरावल में मौसम कैसा है?"
Gujarati
"આજે વેરાવળમાં હવામાન કેવું છે?"
I also tested code-mixed speech such as:
"Veraval mein aaj weather kaisa hai?"
The important part is not simply detecting the language.
The agent should also respond in the appropriate script.
For example, Hindi should be written and spoken naturally rather than being forced into Romanized Hindi.
Day 3 - Frontend
After the voice pipeline worked, I customized the frontend.
The frontend clearly represents the state of the agent.
The main states are:
Ready
|
v
Connecting
|
v
Listening
|
v
Speaking
|
v
Call Ended
t
Day 4 - Persistent Memory
A voice agent that forgets everything after every call isn't very useful for returning users.
So I added persistent memory using SQLite.
The profile stores information such as:
user_id
name
language_preference
facts
last_interaction
The important design decision was that memory is handled by backend functions rather than being written into the LLM prompt.
The agent can:
Look up a caller.
Ask permission before saving information.
Save approved information.
Retrieve the profile during a later conversation.
For example:
First conversation:
"My name is Gautam."
The agent asks whether it should remember the name.
After permission is granted, the information is stored.
During a later conversation, the agent can recognize the returning user.
Day 5 - Real Tools
The agent needed to do more than generate answers from the LLM.
I added a real weather tool.
The weather flow is:
User
|
| "What is the weather in Veraval today?"
v
Gemini
|
| tool call
v
get_weather()
|
v
Weather API
|
v
Real weather data
|
v
Gemini
|
v
Murf Falcon
|
v
User
The important rule is:
The agent must never invent current weather information.
If the weather service is unavailable, the agent should say that it could not retrieve the latest information.
It should not guess.
One Real Debugging Lesson
One of the problems I encountered during development was a Deepgram connection failure.
The voice pipeline produced an error similar to:
APIConnectionError:
failed to connect to deepgram
Instead of assuming the code was broken, I tested the network connection from Windows PowerShell.
I used:
Test-NetConnection api.deepgram.com -Port 443
Initially the connection failed.
After investigating the network/DNS path, I eventually got:
TcpTestSucceeded : True
This was a good reminder that voice-agent problems are not always LLM problems.
The failure can be caused by:
Network connectivity
DNS
Firewall
API availability
WebSocket connections
Authentication
Tool execution
Frontend state
Debugging the complete pipeline is essential.
Day 6 - Outbound Voice Calls
The next step was making the agent call a user instead of waiting for the user to open the browser.
I integrated outbound calling using LiveKit telephony/SIP.
I also tested the Linphone route.
The basic flow became:
Bharat Voice AI
|
v
LiveKit SIP
|
v
Linphone
|
v
Phone

The outbound conversation needs a different opening from a browser conversation.
The agent must immediately explain:
Who is calling
Why it is calling
How the user can end the call
I also tested the SIP connection and call lifecycle.
Day 7 - Human Escalation
A voice agent should not try to solve everything.
I added a human escalation workflow.
When a user says:
"I want to talk to a human."
the agent can start the escalation process.
But it does not automatically share information.
It first asks for permission.
For example:
"I can create a request for human assistance. Before I do, I would share your name, the issue you described, what I checked, your preferred language, and the urgency. Would you like me to create the request?"
If the user says yes, the backend creates a real escalation record.
A reference ID is generated by the backend.
The important rule is:
Never tell the user that an escalation was created unless the database operation actually succeeded.
This avoids fake success messages.
Day 8 - Call Analytics
After adding many features, I needed a way to measure what was happening.
So I built a Call Analytics Dashboard.
The dashboard tracks:

The data comes from real calls stored in SQLite.
It is not hardcoded.
A call is considered successful when the user's intended task is successfully completed.
For example:
User asks for weather
|
Weather tool succeeds
|
Weather information delivered
|
SUCCESS
If the required tool fails or the task is incomplete, the call can be recorded as failed or incomplete.
The dashboard can therefore show actual agent performance.
Day 9 - Specialist Agent Handoff
This was one of the most interesting parts of the project.
Instead of forcing the main agent to handle every type of request, I created a specialist:
Bharat Weather Specialist
The architecture became:
User
|
v
Bharat Voice AI
|
Weather request?
/ \
No Yes
| |
v v
Main Agent Weather Specialist
|
v
get_weather()
|
v
Real Data
python
For example:
User:
"What is the weather today in Veraval?"
The main agent says:
"For detailed weather information, I'll connect you with our weather specialist."
The specialist then takes over.
The important part is that the user should not have to explain the entire question again.
The conversation context is passed to the specialist.
Native Agent Handoff
The LiveKit agent-handoff pattern allows the main agent to return a specialist agent together with an announcement.
Conceptually:
@function_tool
async def handoff_to_weather_specialist(
self,
context: RunContext,
) -> tuple[Agent, str]:
specialist = BharatWeatherSpecialist(
chat_ctx=self.chat_ctx.copy(
exclude_instructions=True
)
)
return (
specialist,
"I'll connect you with our weather specialist."
)
The actual implementation in the repository follows the project's LiveKit version and architecture.
The important concept is that the specialist receives the previous conversation context.
A Day 9 Bug That Taught Me Something
During testing, the handoff itself worked.
However, I encountered a problem where the internal function-call information could appear as visible text instead of the actual weather response.
For example, the user could see something similar to:
The function called is get_weather
location = Veraval
forecast_days = 1
That is not acceptable for a real voice assistant.
The user should hear:
"The latest weather in Veraval is..."
not internal tool-call information.
This led me to investigate the difference between:
Function registration
Tool execution
LLM output
LiveKit agent handoff
Frontend message rendering
It was another reminder that a production voice agent is a complete system, not simply an LLM with speech.
The Complete SystemAfter the nine days, Bharat Voice AI looks roughly like this:
USER
|
Voice / Browser
|
v
LiveKit Agents
|
Bharat Voice AI
|
+-----------------+------------------+
| | |
Memory Tools Routing
| | |
SQLite +------+-------+ |
| | |
Weather Other Tools |
| |
Open-Meteo |
|
Specialist Agent
|
Weather Specialist
|
get_weather()
|
+------------------------------------+
|
Human Escalation
|
SQLite
|
Murf Falcon
|
v
USER
Outbound calls extend the system through SIP/telephony.
Analytics records the outcome of calls and makes it visible through the dashboard.
Technology Stack
The main technologies used in the project include:
Python
LiveKit Agents
Gemini
Deepgram
Murf Falcon
SQLite
Open-Meteo
SIP / Linphone
Browser frontend
GitHub
Each component has a different responsibility.
Deepgram
Speech → Text
Gemini
Reasoning + Conversation
Tools
Real-world data and actions
SQLite
Persistent state
LiveKit
Real-time voice transport and agent orchestration
Murf Falcon
Text → Natural Voice
Security
One of the most important lessons is to never put API credentials directly into the source code.
Use environment variables.
For example:
DEEPGRAM_API_KEY=your_key_here
MURF_API_KEY=your_key_here
LIVEKIT_API_KEY=your_key_here
LIVEKIT_API_SECRET=your_secret_here
The real .env file should never be committed to GitHub.
Use:
.env
locally and:
.env.example
for documentation.
Never publish:
- API keys
- SIP credentials
- passwords
- OTPs
- PINs
- caller information
- private database records
Running the Project
The complete project is available on GitHub:
https://github.com/gunmasterg9/bharat-voice-ai
A typical setup is:
git clone https://github.com/gunmasterg9/bharat-voice-ai.git
cd bharat-voice-ai
cd backend
uv sync
Create your environment configuration:
.env
Add the required API credentials.
Then start the backend using the project's configured startup command.
Start the frontend and open the browser interface.
Allow microphone access.
Then start a conversation.
For example:
"Hello Bharat Voice AI."
Then test:
"What is the weather today in Veraval?"
Then test multilingual interaction:
"આજે વેરાવળમાં હવામાન કેવું છે?"
Then test memory:
"My name is Gautam."
Then restart the application and verify that the saved profile can be retrieved.
Testing
I created tests covering important parts of the system.
The project includes testing for:
Agent behavior
Memory
Persistent memory
Language switching
Weather tools
Escalation
Outbound calls
Linphone
Analytics
Specialist handoff
The Day 9 development test suite included specialist handoff tests alongside the previous functionality.
The important lesson was that every new feature should be tested without breaking the previous days' work.
What I Learned
The biggest lesson from this challenge is that building a voice agent is much more than connecting an LLM to a TTS API.
A reliable voice agent needs:
Voice
+
Reasoning
+
Memory
+
Tools
+
State
+
Error Handling
+
Security
+
Human Handoff
+
Observability
A model can generate a great response, but the surrounding system determines whether the product is actually reliable.
I also learned to debug the entire pipeline instead of assuming every problem is caused by the LLM.
What I Would Build Next
There is still a lot I would like to add.
Future improvements could include:
More specialist agents
Better interruption handling
More Indian languages
Better low-bandwidth support
Advanced call analytics
Conversation quality scoring
Better tool observability
More real-world Indian datasets
Improved outbound-call workflows
Specialist-to-specialist routing
More sophisticated RAG
Production deployment and monitoring
The long-term goal would be to turn Bharat Voice AI into a platform where different specialized voice agents can work together.
Final Thoughts
The most interesting part of the challenge wasn't building the first voice conversation.
It was everything that came afterward.
Making the agent remember.
Making it use real data.
Making it call a phone.
Making it know when to ask a human.
Measuring whether conversations actually succeeded.
And finally, teaching one agent when another agent is better suited to help.
That progression changed how I think about voice AI.
A voice agent isn't just a chatbot that speaks.
It can become a complete software system with memory, tools, workflows, specialized agents, and real-world actions.
That is what I wanted to explore with Bharat Voice AI.
Project Links
GitHub
https://github.com/gunmasterg9/bharat-voice-ai
All Videos in Linkedin
https://www.linkedin.com/in/gautam-vandar-71a75632b/
Challenge
10 Days of Voice Agents — VoiceForBharat Edition
Voice Technology
Murf Falcon
Thank You
Thank you to Murf AI for organizing the 10 Days of Voice Agents challenge and providing the opportunity to build, experiment, debug, and learn through a real voice-agent project.
Building Bharat Voice AI over these 10 days was a great experience, and I hope this project helps someone else start building their own voice agent.


Top comments (0)