Over the last 10 days, I participated in the 10 Days of Voice Agents — VoiceForBharat Edition. My goal was to build a production-ready, highly capable voice AI that can talk, remember past users, check inventory, file complaints, and even hand off conversations to specialized sub-agents.
Here is the story of how I built the Namma Kirana Store Assistant, the challenges I faced with multi-agent routing, and how you can build one too.
🛒 The Problem: Why Voice AI for Kirana Stores?
In India, neighborhood grocery stores run on relationships. Customers don't want to navigate complex apps or type out orders; they want to call their local shopkeeper, speak in a mix of Hindi and English (Hinglish), and say "Bhaiya, 2 kg pyaaz bhej do" (Brother, send 2 kg of onions).
However, shopkeepers are often too busy managing the physical store to answer every call.
I built the Store Assistant to solve this. It is a highly conversational Voice AI that acts as the shop's front desk. Voice is the perfect medium here because it removes the friction of typing and apps, allowing local MSMEs (Micro, Small & Medium Enterprises) to scale their customer service naturally.
✨ The Magic: What the Agent Does
To make the AI feel like a real store assistant, I focused on a few core features:
- Ultra-Fast Indian Voice (Murf Falcon): For a voice agent to feel natural in India, the voice must have the right accent and cadence. I used the blazing-fast Murf Falcon TTS API with the "Anisha" voice. It perfectly handles code-mixing (switching between Hindi and English) with near-zero latency.
- Contextual Memory: When a customer calls, the agent silently triggers a lookup_caller tool to check a local database. If it recognizes a returning user, it greets them by name and asks for feedback on their last order.
- Inventory & Order Tools: The agent doesn't hallucinate prices. It uses a check_inventory tool to fetch real-time stock and prices before confirming any order.
- Outbound Calling: The system can initiate outbound calls to returning customers to advertise new stock or notify them that a previous complaint has been resolved.
- Call Analytics Dashboard: I built a real-time web dashboard using Next.js and WebSockets that automatically logs live call transcripts, durations, and categorizes the call outcomes (e.g., "Successful Order" or "User Hang-up")
🔄 The Architecture: Multi-Agent Handoffs
The biggest breakthrough in my project was moving away from a single, monolithic agent. Putting order-taking, policy explanation, and complaint-handling into one massive system prompt caused the LLM to get confused and time out.
Instead, I implemented a Multi-Agent Handoff Architecture using LiveKit.
Here is how the audio and data move through the system:

When a user asks a complex question about returns, the Main Assistant uses a tool called transfer_to_policy. The system suspends the Main Assistant, condenses the conversation history, and instantly spins up Kavya, the Policy Specialist, who searches the knowledge base and answers the question. Once she is done, she transfers the user back to the Main Assistant to conclude the order.
🐛 The Hardest Challenge: "Identity Theft" During Handoffs
Real lessons come from bugs, and I hit a massive one during multi-agent handoffs.
When transferring a user from the Main Assistant to a Specialist, I needed to pass the conversation history so the Specialist knew what the user was asking. I wrote a helper function to copy the chat context.
Suddenly, my Policy Specialist started acting crazy. Instead of answering policy questions, she would introduce herself and then immediately try to transfer the user again, often crashing the system because she was calling tools she didn't possess.
The Root Cause
When I copied the ChatContext to pass the history, I accidentally copied the Main Agent's System Prompt into the new agent's memory.
When Kavya (the Policy Specialist) woke up, her brain contained TWO conflicting prompts:
- "You are the Main Agent. Route policy questions to the specialist." (Accidentally copied)
- "You are Kavya, the Policy Specialist."
The Fix
I dug into the LiveKit Agents SDK and found the exclude_instructions=True flag for copying contexts. By explicitly stripping the old agent's system prompt and injecting a firm system note directing the new agent to answer the pending question, the routing loop was completely fixed.
def truncate_context(chat_ctx, system_note=None):
from livekit.agents.llm import ChatContext
new_ctx = ChatContext()
# CRITICAL FIX: Strip the previous agent's system prompt so the new agent doesn't get confused!
clean_ctx = chat_ctx.copy(exclude_instructions=True)
# Separate remaining system notes and actual chat history
sys_msgs = [m for m in clean_ctx.messages() if m.role == 'system']
chat_msgs = [m for m in clean_ctx.messages() if m.role != 'system']
# ... (code to append messages to new_ctx)
if system_note:
# Inject a firm instruction so the new specialist knows exactly what to do
new_ctx.add_message(role='system', content=f'[System Note: {system_note}]')
return new_ctx
🛠️ Build Your Own Voice Agent
If you want to build something similar, it is easier than you think! A modern Voice AI requires four main components:
- Real-time Transport: LiveKit (handles the WebRTC audio streaming so latency is minimal).
- Speech-to-Text (STT): Deepgram (turns the user's voice into text instantly).
- The Brain (LLM): OpenAI / Anthropic / Gemini (processes the text, uses tools, and generates a reply).
- Text-to-Speech (TTS): Murf Falcon (turns the LLM's text reply back into natural, accented speech).
Getting Started: Build it Yourself
To build and run this exact multi-agent architecture on your own machine, follow these steps:
Clone the Starter Kit Start by cloning the repository to your local machine:
git clone https://github.com/Koushiklalala/murf-livekit-starter.git
cd murf-livekit-starter
Set up your API Keys You will need API keys from LiveKit, OpenAI, Murf, and Deepgram. Create a .env file in your backend directory. Never hardcode keys in your Python files or push them to GitHub.
LIVEKIT_URL=ws://127.0.0.1:7880
LIVEKIT_API_KEY=devkey
LIVEKIT_API_SECRET=secret
OPENAI_API_KEY=sk-...
MURF_API_KEY=murf-...
DEEPGRAM_API_KEY=dg-...
Start the LiveKit Dev Server You need a local LiveKit server to handle the WebRTC audio transport. Open your first terminal and start it:
.\livekit-dev

(Keep this terminal running. It will output your local WebSocket URL, API Key, and Secret).Start the Backend Agent Open a second terminal, activate your virtual environment, and start the main voice agent:
cd backend && uv run python src/agent.py dev

Start the Call Analytics Dashboard To see the real-time call logs and outcomes, open a third terminal and start the dashboard using uv:
cd backend && uv run python src/dashboard.py

Open localhost:8080 in your browser.

Start the Frontend Open a fourth terminal, navigate to your frontend directory, and start the UI using pnpm:
cd frontend && pnpm dev

Open localhost:3000 in your browser. Click "Connect" to talk to the Main Assistant!
Triggering the Outbound Calls
Trigger Outbound Calls (Optional) The system can make outbound calls to users! To trigger an advertising call or a ticket-resolved call, open a fifth terminal and run the specific scripts:
To trigger a marketing outbound call:
cd backend && uv run python src/outbound_call.py

To trigger a ticket resolution outbound call:
uv run python src/resolve_ticket.py <YOUR TICKIT ID>
Note : Replace with your actual Complaint ID.

🔗 Links and Resources
- GitHub Repository: Check out the full code here!
- Murf Falcon TTS:Documentation
- LiveKit Voice AI:Quickstart Guide
Building this over the last 10 days was an incredible learning experience. If you are a developer looking to get into Voice AI, I highly recommend getting your hands dirty with LiveKit and Murf!

Top comments (0)