<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gautam</title>
    <description>The latest articles on DEV Community by Gautam (@gautammax).</description>
    <link>https://dev.to/gautammax</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4078594%2F07719cb0-ffdf-41c2-8891-70a6f818408c.png</url>
      <title>DEV Community: Gautam</title>
      <link>https://dev.to/gautammax</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gautammax"/>
    <language>en</language>
    <item>
      <title>Mitra Local: A Private, Local-First AI Voice Memory &amp; Task Assistant Built for My Friend Ketan</title>
      <dc:creator>Gautam</dc:creator>
      <pubDate>Sat, 03 Oct 2026 17:03:15 +0000</pubDate>
      <link>https://dev.to/gautammax/mitra-local-a-private-local-first-ai-voice-memory-task-assistant-built-for-my-friend-ketan-3h3i</link>
      <guid>https://dev.to/gautammax/mitra-local-a-private-local-first-ai-voice-memory-task-assistant-built-for-my-friend-ketan-3h3i</guid>
      <description>&lt;p&gt;&lt;em&gt;This article is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Meet My Friend Ketan
&lt;/h3&gt;

&lt;p&gt;My childhood friend &lt;strong&gt;Ketan&lt;/strong&gt; runs a textile and garment trading firm in Ahmedabad, Gujarat. If you spend even one morning with him at his shop or warehouse, you will notice three things immediately:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;His phone rings unceasingly.&lt;/strong&gt; Between 9 AM and 7 PM, he takes over 50 phone calls from yarn spinners, dye mills, courier agencies, brokers, and retail store owners.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;He switches languages three times in a single sentence.&lt;/strong&gt; He will greet in Gujarati, negotiate quantities in Hindi, and discuss dispatch timelines or bank transactions in English.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;He is constantly making verbal commitments on the go:&lt;/strong&gt;
&amp;gt; &lt;em&gt;"હા Maheshbhai, કાલે સવારે તમને sample swatches મોકલાવી દઉં છું."&lt;/em&gt; (Yes Maheshbhai, sending you the sample swatches tomorrow morning.)
&amp;gt; &lt;em&gt;"Sureshji, aapka 45,000 ka payment RTGS kar diya hai, shaam tak credit ho jayega."&lt;/em&gt; (Sureshji, your 45,000 payment was sent via RTGS, will credit by evening.)
&amp;gt; &lt;em&gt;"Priya, please make sure the GST invoice for bill #1042 is cleared before Friday 4 PM."&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;By evening, half of those verbal promises were forgotten or scribbled hastily on scraps of paper, receipts, or scattered WhatsApp self-chats.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Off-the-Shelf Tools Failed Ketan
&lt;/h3&gt;

&lt;p&gt;Ketan tried Siri, Google Keep, Notion, and standard Todoist apps. &lt;strong&gt;None of them worked for him:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Language Inflexibility:&lt;/strong&gt; Siri and Google Assistant choked on code-mixed Gujarati-English sentences (&lt;em&gt;"Maheshbhai ne call kari de"&lt;/em&gt; became unintelligible gibberish).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High Friction:&lt;/strong&gt; Notion and Todoist demanded manual typing with date pickers—impossible while inspecting fabric rolls or driving between warehouses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy &amp;amp; Business Secrets:&lt;/strong&gt; Ketan strictly refused to upload his customer contact books, invoice figures, and verbal business memos to public corporate AI clouds.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Solution: Mitra Local
&lt;/h3&gt;

&lt;p&gt;So I promised him: &lt;em&gt;"I will build an assistant designed specifically for how you speak and how you work—running completely on your laptop, 100% private, with zero subscription fees."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That project became &lt;strong&gt;Mitra Local&lt;/strong&gt; (મિત્ર / मित्र — "Friend").&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mitra Local&lt;/strong&gt; is an open-source, local-first voice memory and task companion. Ketan simply hits a hotkey or clicks the microphone, talks naturally in whatever mix of Gujarati, Hindi, or English comes out of his mouth, and lets Mitra handle the rest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🗣️ &lt;strong&gt;Native Code-Mixed Understanding:&lt;/strong&gt; Seamlessly parses Gujarati (&lt;code&gt;ગુજરાતી&lt;/code&gt;), Hindi (&lt;code&gt;हिंदी&lt;/code&gt;), English, and hybrid mixtures without requiring manual language toggling.&lt;/li&gt;
&lt;li&gt;🎯 &lt;strong&gt;Automatic Intent &amp;amp; Action Extraction:&lt;/strong&gt; Distinguishes between actionable tasks (calls, deliveries, payments, follow-ups) and knowledge memories (client preferences, phone numbers, notes).&lt;/li&gt;
&lt;li&gt;👤 &lt;strong&gt;Entity &amp;amp; Person Graph:&lt;/strong&gt; Mentions of "Maheshbhai" or "Sureshji" automatically link to contact records with conversation histories and pending to-dos.&lt;/li&gt;
&lt;li&gt;⚡ &lt;strong&gt;Local Voice Feedback:&lt;/strong&gt; Mitra speaks back a concise confirmation in natural audio so Ketan knows his commitment was captured without staring at the screen.&lt;/li&gt;
&lt;li&gt;🔒 &lt;strong&gt;100% Offline &amp;amp; Private:&lt;/strong&gt; Powered by open-source AI (Ollama + Gemma / Llama) and local speech-to-text. Zero cloud telemetry.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/7dzVaiROmpA" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Here is how Mitra Local works in real life when Ketan speaks to it:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Real-World Voice Interaction Scenarios
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Language Input&lt;/th&gt;
&lt;th&gt;Spoken Command&lt;/th&gt;
&lt;th&gt;Extracted Action &amp;amp; Intelligence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Gujarati + Eng&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"કાલે Maheshbhai ને call કરી દે, shipment વિશે પૂછવાનું છે."&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Call Task&lt;/strong&gt; linked to contact &lt;code&gt;Maheshbhai&lt;/code&gt;, scheduled for tomorrow, categorized under shipments.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hindi + Eng&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"Sureshji ko bol dena ki 45,000 ka payment RTGS kar diya hai."&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Payment Memory &amp;amp; Verification Task&lt;/strong&gt; linked to contact &lt;code&gt;Sureshji&lt;/code&gt; with amount &lt;code&gt;₹45,000&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;English&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"Remind me to file GST tax documents by Friday 3 PM."&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;High-Priority Task&lt;/strong&gt; with absolute deadline set to Friday 15:00.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pure Gujarati&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"યાદ રાખજે કે રમેશભાઈની દુકાને નવા સેમ્પલ પહોંચાડવાના છે."&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Delivery Task&lt;/strong&gt; linked to &lt;code&gt;Rameshbhai&lt;/code&gt;'s shop record.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pure Hindi&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"याद रखना कि राजेश का जन्मदिन 15 तारीख को है."&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Personal Memory&lt;/strong&gt; tagged under &lt;code&gt;Rajesh&lt;/code&gt; in the local memory bank.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  2. The Real-Time Workflow
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Speak naturally:&lt;/strong&gt; Ketan clicks the central microphone button (or triggers the local hotkey) and speaks in Gujarati, Hindi, English, or a mix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instant Structured Extraction:&lt;/strong&gt; Within 400 milliseconds, the dashboard dynamically updates with:

&lt;ul&gt;
&lt;li&gt;Extracted task title with original code-mixed transcript preserved&lt;/li&gt;
&lt;li&gt;Categorized action type (Call, Payment, Meeting, Delivery, Note)&lt;/li&gt;
&lt;li&gt;Identified person / contact entity&lt;/li&gt;
&lt;li&gt;Normalized due date &amp;amp; priority level&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verbal Confirmation:&lt;/strong&gt; The local Web Speech API generates an instant natural voice response:
&amp;gt; &lt;em&gt;"Call task for Maheshbhai regarding shipment scheduled for tomorrow."&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relational Organization:&lt;/strong&gt; The task is instantly filed in the Task Matrix and linked to the Person's profile in the SQLite database.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  3. Ketan's Reaction
&lt;/h3&gt;

&lt;p&gt;When I handed Ketan the first build on his laptop, he laughed and immediately tested it with his fastest Kathiyawadi Gujarati:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"કાલે સવારે 11 વાગ્યે શર્માજી સાથે પેમેન્ટનું સેટલમેન્ટ કરવાનું છે, યાદ રાખજે!"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Within less than half a second, the card appeared on his screen:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Title&lt;/strong&gt;: &lt;em&gt;Payment settlement with Sharmaji&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Action&lt;/strong&gt;: &lt;em&gt;Payment / Financial&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Person&lt;/strong&gt;: &lt;em&gt;Sharmaji&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Due&lt;/strong&gt;: &lt;em&gt;Tomorrow at 11:00 AM&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audio Response&lt;/strong&gt;: &lt;em&gt;"Payment task created for Sharmaji tomorrow at 11:00 AM."&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;He looked up and said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Gautam, this is the first app that actually talks like an Indian businessman. No complex forms, no English-only barriers, and no fear of my data leaving my laptop."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;Mitra Local is 100% open source under the permissive &lt;strong&gt;MIT License&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Repository&lt;/strong&gt;: &lt;a href="https://github.com/gunmasterg9/mitra-local" rel="noopener noreferrer"&gt;https://github.com/gunmasterg9/mitra-local&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Quickstart: Running Locally in 3 Minutes
&lt;/h3&gt;

&lt;h4&gt;
  
  
  1. Clone the Repository
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/gunmasterg9/mitra-local.git
&lt;span class="nb"&gt;cd &lt;/span&gt;mitra-local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  2. Start the Backend
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;backend
python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;span class="c"&gt;# On Windows:&lt;/span&gt;
.venv&lt;span class="se"&gt;\S&lt;/span&gt;cripts&lt;span class="se"&gt;\a&lt;/span&gt;ctivate
&lt;span class="c"&gt;# On macOS/Linux:&lt;/span&gt;
&lt;span class="c"&gt;# source .venv/bin/activate&lt;/span&gt;

pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
python &lt;span class="nt"&gt;-m&lt;/span&gt; uvicorn app.main:app &lt;span class="nt"&gt;--reload&lt;/span&gt; &lt;span class="nt"&gt;--port&lt;/span&gt; 8000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Backend runs on &lt;code&gt;http://localhost:8000&lt;/code&gt; (Interactive API docs at &lt;code&gt;http://localhost:8000/docs&lt;/code&gt;).&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Start the Frontend
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ../frontend
npm &lt;span class="nb"&gt;install
&lt;/span&gt;npm run dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Frontend runs on &lt;code&gt;http://localhost:5173&lt;/code&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  4. Automated Verification Suite
&lt;/h4&gt;

&lt;p&gt;We built a comprehensive test suite covering multilingual entity extraction, relative date parsing, contact deduplication, and database transactions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;backend
python &lt;span class="nt"&gt;-m&lt;/span&gt; pytest tests/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;======================= 36 passed, 16 warnings in 8.02s =======================
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All &lt;strong&gt;36 unit and integration tests&lt;/strong&gt; pass, ensuring zero-regression reliability.&lt;/p&gt;




&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;Mitra Local is built from the ground up as a resilient, local-first system designed for zero latency and complete data sovereignty.&lt;/p&gt;

&lt;h3&gt;
  
  
  Architecture Overview
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌────────────────────────────────────────────────────────┐
│             Modern React 19 + Vite Frontend            │
│  - One-click Voice Mic (Web Audio API / MediaRecorder) │
│  - Real-time Task Matrix &amp;amp; Priority Filters            │
│  - Memory Bank &amp;amp; Person Profiles                       │
│  - Local Browser Speech Synthesis (TTS)                │
└───────────────────────────┬────────────────────────────┘
                            │ REST API
                            ▼
┌────────────────────────────────────────────────────────┐
│                FastAPI Asynchronous Engine             │
│  - Audio Transcription Pipeline (Local Whisper / API)  │
│  - Structured Multilingual Extraction Pipeline         │
│  - Heuristic Fallback Engine (Zero-Downtime Guarantee) │
└───────────────┬────────────────────────┬───────────────┘
                │                        │
                ▼                        ▼
┌──────────────────────────────┐ ┌───────────────────────┐
│     Local AI with Ollama     │ │  Local SQLite Engine  │
│  - Gemma 4 12B / Llama 3.2   │ │  - aiosqlite Async DB │
│  - Zero Cloud Telemetry      │ │  - Full-Text Search   │
│  - Structured JSON Schema    │ │  - People &amp;amp; Memories  │
└──────────────────────────────┘ └───────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Technical Stack
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontend&lt;/strong&gt;: React 19, Vite, Lucide Icons, Vanilla CSS design tokens with sleek dark mode, Web Speech Synthesis API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend&lt;/strong&gt;: FastAPI (Python 3.12), Pydantic v2 schemas, SQLAlchemy with &lt;code&gt;aiosqlite&lt;/code&gt; for asynchronous local database transactions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local AI Inference&lt;/strong&gt;: &lt;a href="https://ollama.ai" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; running open-weight models (&lt;code&gt;gemma4:12b&lt;/code&gt;, &lt;code&gt;llama3.2&lt;/code&gt;, or &lt;code&gt;mistral&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage&lt;/strong&gt;: Local SQLite database stored securely in &lt;code&gt;backend/data/mitra.db&lt;/code&gt; on Ketan's machine.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Solving the Multilingual Code-Switching Challenge
&lt;/h3&gt;

&lt;p&gt;Indian business conversations freely mix Gujarati, Hindi, and English with honorifics (&lt;em&gt;"-bhai"&lt;/em&gt;, &lt;em&gt;"-ji"&lt;/em&gt;, &lt;em&gt;"-ben"&lt;/em&gt;). Conventional prompts often fail by either translating names into English words (e.g. converting "Sureshji" into "Sir Suresh") or dropping dates.&lt;/p&gt;

&lt;p&gt;To achieve bulletproof parsing, Mitra combines two layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Structured Few-Shot Prompting with Open LLMs&lt;/strong&gt;:
We prompt local models with strict JSON output schemas that explicitly understand Indic code-mixing:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="nl"&gt;"intent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"create_task"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="nl"&gt;"task"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Call Maheshbhai regarding shipment"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="nl"&gt;"title_original"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"કાલે Maheshbhai ને call કરી દે, shipment વિશે પૂછવાનું છે."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="nl"&gt;"action_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"call"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="nl"&gt;"person_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Maheshbhai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="nl"&gt;"due_date"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-04"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="nl"&gt;"priority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
       &lt;/span&gt;&lt;span class="nl"&gt;"language"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gu-en"&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="nl"&gt;"speech_response"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Call task for Maheshbhai regarding shipment scheduled for tomorrow."&lt;/span&gt;&lt;span class="w"&gt;
   &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Heuristic Linguistic Fallback Engine&lt;/strong&gt;:
Laptops in the field don't always have GPU acceleration, and background tasks might temporarily tie up resources. Mitra includes an integrated regex- and rule-based Indic NLP engine that recognizes common action verbs in Gujarati (&lt;em&gt;"call kari de"&lt;/em&gt;, &lt;em&gt;"moklavano chhe"&lt;/em&gt;, &lt;em&gt;"poochhvanu chhe"&lt;/em&gt;), Hindi (&lt;em&gt;"bol dena"&lt;/em&gt;, &lt;em&gt;"bhej do"&lt;/em&gt;, &lt;em&gt;"yaad rakhna"&lt;/em&gt;), and English (&lt;em&gt;"remind me"&lt;/em&gt;, &lt;em&gt;"pay"&lt;/em&gt;), along with relative temporal expressions (&lt;em&gt;"kaale"&lt;/em&gt;, &lt;em&gt;"parso"&lt;/em&gt;, &lt;em&gt;"kal"&lt;/em&gt;). Even with Ollama paused, Mitra processes inputs flawlessly in under 5 milliseconds.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;When starting this project for Ketan, the easiest route would have been wiring up a quick commercial cloud API (OpenAI or Claude) with a credit card. &lt;strong&gt;But for Ketan, and millions of independent business owners like him, closed cloud AI is fundamentally the wrong answer.&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;Here is why &lt;strong&gt;open innovation and open-source AI&lt;/strong&gt; were strictly essential to solving this problem:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Absolute Privacy and Data Sovereignty
&lt;/h3&gt;

&lt;p&gt;In the textile trading world, commercial relationships and pricing are hard-won trade secrets. A typical memo contains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Supplier identities and credit balances&lt;/li&gt;
&lt;li&gt;GST numbers and pending payment amounts&lt;/li&gt;
&lt;li&gt;Private personal numbers and commitments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Sending audio snippets or text logs of these conversations to third-party public clouds exposes small business owners to data harvesting, telemetry profiling, and compliance vulnerabilities. With &lt;strong&gt;open-weight models running on Ollama and local SQLite storage&lt;/strong&gt;, Ketan has 100% mathematical certainty that not a single byte of his business data ever leaves his hardware.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Zero Recurring Overhead for Everyday People
&lt;/h3&gt;

&lt;p&gt;Proprietary AI SaaS platforms charge $20 to $50 per user per month, plus metered API token fees. For an independent merchant, recurring software subscriptions add up quickly and become a financial burden. Open-source models (like Gemma and Llama) democratize state-of-the-art intelligence: once downloaded, they run indefinitely with &lt;strong&gt;zero subscription costs, zero token meters, and zero paywalls&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Freedom to Adapt for Underrepresented Languages
&lt;/h3&gt;

&lt;p&gt;Commercial AI models are optimized primarily for formal Western corporate English. They frequently misinterpret code-mixed regional Indian dialects like Gujarati-English or Hinglish, treating common honorifics as foreign errors or hallucinations. &lt;br&gt;
Because open models and open-source code can be inspected, steered, and fine-tuned locally, we were able to build customized prompt pipelines and hybrid heuristic engines tailored specifically to Gujarati merchants. Open innovation puts the power to build AI in the hands of the communities that need it most, rather than waiting for Silicon Valley giants to prioritize regional dialects.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. True Offline Autonomy
&lt;/h3&gt;

&lt;p&gt;Wholesalers and traders frequently operate in basement warehouses, noisy textile markets, and transport hubs with spotty cellular reception. A tool that depends on roundtrips to an external cloud server will fail right when it is needed most. Open-source local inference delivers instant, reliable performance anywhere—on a flight, in a godown, or during an internet outage.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion &amp;amp; What's Next
&lt;/h2&gt;

&lt;p&gt;Building &lt;strong&gt;Mitra Local&lt;/strong&gt; for Ketan showed me that the most impactful AI tools aren't generic chatbots—they are personalized, empathetic utilities built for the specific rhythms of real people's lives.&lt;/p&gt;

&lt;p&gt;Our upcoming roadmap includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fine-tuning a quantized Whisper model on regional Gujarati and Kathiyawadi accents&lt;/li&gt;
&lt;li&gt;Local WhatsApp Desktop webhook bridge for automated reminder dispatch&lt;/li&gt;
&lt;li&gt;Offline vector search using DuckDB / ChromaDB for multi-year business archives&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Thank you to the DEV Community and Hacktoberfest for championing open innovation that makes technology truly personal and accessible.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Repository:&lt;/strong&gt; &lt;a href="https://github.com/gunmasterg9/mitra-local" rel="noopener noreferrer"&gt;https://github.com/gunmasterg9/mitra-local&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;License:&lt;/strong&gt; &lt;a href="https://github.com/gunmasterg9/mitra-local/blob/main/LICENSE" rel="noopener noreferrer"&gt;MIT License&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  devchallenge #hacktoberfest #hf26challenge
&lt;/h1&gt;

</description>
      <category>devchallenge</category>
      <category>hacktoberfest</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>Building Bharat Voice AI: My 10-Day Journey from a Voice Agent to a Multilingual Multi-Agent System</title>
      <dc:creator>Gautam</dc:creator>
      <pubDate>Sat, 15 Aug 2026 08:29:00 +0000</pubDate>
      <link>https://dev.to/gautammax/building-bharat-voice-ai-my-10-day-journey-from-a-voice-agent-to-a-multilingual-multi-agent-system-2d4b</link>
      <guid>https://dev.to/gautammax/building-bharat-voice-ai-my-10-day-journey-from-a-voice-agent-to-a-multilingual-multi-agent-system-2d4b</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0vzyg3r5f3wyiw5voj2k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0vzyg3r5f3wyiw5voj2k.png" alt=" " width="800" height="329"&gt;&lt;/a&gt;# Building Bharat Voice AI: My 10-Day Journey from a Voice Agent to a Multilingual Multi-Agent System&lt;/p&gt;

&lt;p&gt;Over the last 10 days, I took a simple idea, a voice assistant, and gradually turned it into a more complete multilingual voice-agent system.&lt;/p&gt;

&lt;p&gt;This project is called &lt;strong&gt;Bharat Voice AI&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal was to build a voice agent that can communicate naturally with users in India, understand different languages and code-mixed speech, remember returning users, use real-world tools, make outbound calls, escalate to humans, measure its performance, and hand specialized tasks to another agent.&lt;/p&gt;

&lt;p&gt;This journey was part of:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;10 Days of Voice Agents — VoiceForBharat Edition&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The project uses &lt;strong&gt;Murf Falcon&lt;/strong&gt; for voice generation together with LiveKit, Gemini, Deepgram, Python, SQLite, and real external tools.&lt;/p&gt;




&lt;h2&gt;
  
  
  What is Bharat Voice AI?
&lt;/h2&gt;

&lt;p&gt;Bharat Voice AI is a multilingual voice assistant designed around the idea that voice interfaces should feel natural for Indian users.&lt;/p&gt;

&lt;p&gt;Instead of forcing users to interact only through a text interface, the system allows them to speak naturally.&lt;/p&gt;

&lt;p&gt;It supports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;English&lt;/li&gt;
&lt;li&gt;Hindi&lt;/li&gt;
&lt;li&gt;Gujarati&lt;/li&gt;
&lt;li&gt;Hinglish and code-mixed conversations&lt;/li&gt;
&lt;li&gt;Persistent user memory&lt;/li&gt;
&lt;li&gt;Real-time tools&lt;/li&gt;
&lt;li&gt;Human escalation&lt;/li&gt;
&lt;li&gt;Outbound voice calls&lt;/li&gt;
&lt;li&gt;Call analytics&lt;/li&gt;
&lt;li&gt;Specialist-agent handoff&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The system is designed so that the main agent does not have to do everything itself.&lt;/p&gt;

&lt;p&gt;When a task requires specialized knowledge, it can hand the conversation to a specialist agent.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why Voice?
&lt;/h1&gt;

&lt;p&gt;India has a huge diversity of languages, communication styles, and levels of digital literacy.&lt;/p&gt;

&lt;p&gt;A voice interface can make technology easier to access because users do not have to type everything.&lt;/p&gt;

&lt;p&gt;For Bharat Voice AI, I wanted the interaction to feel closer to a normal conversation.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Veraval mein aaj weather kaisa hai?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The system should understand the intent even though the sentence mixes Hindi and English.&lt;/p&gt;

&lt;p&gt;It should then respond naturally in the appropriate language.&lt;/p&gt;




&lt;h1&gt;
  
  
  Architecture
&lt;/h1&gt;

&lt;p&gt;The basic architecture looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Bharat Voice AI
                           |
                     LiveKit Agents
                           |
              +------------+------------+
              |                         |
           Deepgram                   Gemini
             STT                       LLM
              |                         |
              +------------+------------+
                           |
                         Tools
              +------------+------------+
              |            |            |
           Weather      Memory      Escalation
              |            |            |
          Open-Meteo     SQLite       SQLite
                           |
                    Specialist Agents
                           |
                     Murf Falcon TTS
                           |
                         User
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvgsonwvzdnxbks8caj58.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvgsonwvzdnxbks8caj58.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 1 - Building the Voice Agent&lt;/strong&gt;&lt;br&gt;
The first step was getting the basic voice pipeline working.&lt;br&gt;
The core pipeline became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User Speech
     |
     v
Deepgram STT
     |
     v
Gemini
     |
     v
Murf Falcon
     |
     v
User Voice
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent could listen to the user and respond using voice.&lt;br&gt;
The main technologies were:&lt;br&gt;
LiveKit Agents&lt;br&gt;
Deepgram&lt;br&gt;
Gemini&lt;br&gt;
Murf Falcon&lt;br&gt;
&lt;strong&gt;Day 2 - Persona and Guardrails&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The next step was defining who the agent actually is.&lt;/p&gt;

&lt;p&gt;The agent became:&lt;/p&gt;

&lt;p&gt;Bharat Voice AI&lt;/p&gt;

&lt;p&gt;I added clear instructions covering:&lt;/p&gt;

&lt;p&gt;Identity&lt;br&gt;
Objectives&lt;br&gt;
Knowledge boundaries&lt;br&gt;
Language behavior&lt;br&gt;
Guardrails&lt;br&gt;
Escalation&lt;br&gt;
Conversation style&lt;/p&gt;

&lt;p&gt;The agent should not pretend to know something it does not know.&lt;/p&gt;

&lt;p&gt;It should also know when a request is outside its role.&lt;/p&gt;

&lt;p&gt;This became especially important later when I added human escalation and specialist agents.&lt;br&gt;
Multilingual Conversations&lt;/p&gt;

&lt;p&gt;One of the important goals was supporting Indian languages.&lt;/p&gt;

&lt;p&gt;The agent supports:&lt;/p&gt;

&lt;p&gt;English&lt;/p&gt;

&lt;p&gt;"What's the weather today?"&lt;/p&gt;

&lt;p&gt;Hindi&lt;/p&gt;

&lt;p&gt;"आज वेरावल में मौसम कैसा है?"&lt;/p&gt;

&lt;p&gt;Gujarati&lt;/p&gt;

&lt;p&gt;"આજે વેરાવળમાં હવામાન કેવું છે?"&lt;/p&gt;

&lt;p&gt;I also tested code-mixed speech such as:&lt;/p&gt;

&lt;p&gt;"Veraval mein aaj weather kaisa hai?"&lt;/p&gt;

&lt;p&gt;The important part is not simply detecting the language.&lt;/p&gt;

&lt;p&gt;The agent should also respond in the appropriate script.&lt;/p&gt;

&lt;p&gt;For example, Hindi should be written and spoken naturally rather than being forced into Romanized Hindi.&lt;br&gt;
&lt;strong&gt;Day 3 - Frontend&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;After the voice pipeline worked, I customized the frontend.&lt;/p&gt;

&lt;p&gt;The frontend clearly represents the state of the agent.&lt;/p&gt;

&lt;p&gt;The main states are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ready
  |
  v
Connecting
  |
  v
Listening
  |
  v
Speaking
  |
  v
Call Ended
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
t&lt;br&gt;
&lt;strong&gt;Day 4 - Persistent Memory&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A voice agent that forgets everything after every call isn't very useful for returning users.&lt;/p&gt;

&lt;p&gt;So I added persistent memory using SQLite.&lt;/p&gt;

&lt;p&gt;The profile stores information such as:&lt;/p&gt;

&lt;p&gt;user_id&lt;br&gt;
name&lt;br&gt;
language_preference&lt;br&gt;
facts&lt;br&gt;
last_interaction&lt;/p&gt;

&lt;p&gt;The important design decision was that memory is handled by backend functions rather than being written into the LLM prompt.&lt;/p&gt;

&lt;p&gt;The agent can:&lt;/p&gt;

&lt;p&gt;Look up a caller.&lt;br&gt;
Ask permission before saving information.&lt;br&gt;
Save approved information.&lt;br&gt;
Retrieve the profile during a later conversation.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;First conversation:&lt;/p&gt;

&lt;p&gt;"My name is Gautam."&lt;/p&gt;

&lt;p&gt;The agent asks whether it should remember the name.&lt;/p&gt;

&lt;p&gt;After permission is granted, the information is stored.&lt;/p&gt;

&lt;p&gt;During a later conversation, the agent can recognize the returning user.&lt;/p&gt;

&lt;p&gt;Day 5 - Real Tools&lt;/p&gt;

&lt;p&gt;The agent needed to do more than generate answers from the LLM.&lt;/p&gt;

&lt;p&gt;I added a real weather tool.&lt;/p&gt;

&lt;p&gt;The weather flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
 |
 | "What is the weather in Veraval today?"
 v
Gemini
 |
 | tool call
 v
get_weather()
 |
 v
Weather API
 |
 v
Real weather data
 |
 v
Gemini
 |
 v
Murf Falcon
 |
 v
User
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important rule is:&lt;/p&gt;

&lt;p&gt;The agent must never invent current weather information.&lt;/p&gt;

&lt;p&gt;If the weather service is unavailable, the agent should say that it could not retrieve the latest information.&lt;/p&gt;

&lt;p&gt;It should not guess.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One Real Debugging Lesson&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One of the problems I encountered during development was a Deepgram connection failure.&lt;/p&gt;

&lt;p&gt;The voice pipeline produced an error similar to:&lt;/p&gt;

&lt;p&gt;APIConnectionError:&lt;br&gt;
failed to connect to deepgram&lt;/p&gt;

&lt;p&gt;Instead of assuming the code was broken, I tested the network connection from Windows PowerShell.&lt;/p&gt;

&lt;p&gt;I used:&lt;/p&gt;

&lt;p&gt;Test-NetConnection api.deepgram.com -Port 443&lt;/p&gt;

&lt;p&gt;Initially the connection failed.&lt;/p&gt;

&lt;p&gt;After investigating the network/DNS path, I eventually got:&lt;/p&gt;

&lt;p&gt;TcpTestSucceeded : True&lt;/p&gt;

&lt;p&gt;This was a good reminder that voice-agent problems are not always LLM problems.&lt;/p&gt;

&lt;p&gt;The failure can be caused by:&lt;/p&gt;

&lt;p&gt;Network connectivity&lt;br&gt;
DNS&lt;br&gt;
Firewall&lt;br&gt;
API availability&lt;br&gt;
WebSocket connections&lt;br&gt;
Authentication&lt;br&gt;
Tool execution&lt;br&gt;
Frontend state&lt;/p&gt;

&lt;p&gt;Debugging the complete pipeline is essential.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 6 - Outbound Voice Calls&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The next step was making the agent call a user instead of waiting for the user to open the browser.&lt;/p&gt;

&lt;p&gt;I integrated outbound calling using LiveKit telephony/SIP.&lt;/p&gt;

&lt;p&gt;I also tested the Linphone route.&lt;/p&gt;

&lt;p&gt;The basic flow became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bharat Voice AI
      |
      v
LiveKit SIP
      |
      v
Linphone
      |
      v
    Phone
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frdmyn9j2nba14vwu3sgh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frdmyn9j2nba14vwu3sgh.png" alt=" " width="369" height="638"&gt;&lt;/a&gt;&lt;br&gt;
The outbound conversation needs a different opening from a browser conversation.&lt;/p&gt;

&lt;p&gt;The agent must immediately explain:&lt;/p&gt;

&lt;p&gt;Who is calling&lt;br&gt;
Why it is calling&lt;br&gt;
How the user can end the call&lt;/p&gt;

&lt;p&gt;I also tested the SIP connection and call lifecycle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 7 - Human Escalation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A voice agent should not try to solve everything.&lt;/p&gt;

&lt;p&gt;I added a human escalation workflow.&lt;/p&gt;

&lt;p&gt;When a user says:&lt;/p&gt;

&lt;p&gt;"I want to talk to a human."&lt;/p&gt;

&lt;p&gt;the agent can start the escalation process.&lt;/p&gt;

&lt;p&gt;But it does not automatically share information.&lt;/p&gt;

&lt;p&gt;It first asks for permission.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;"I can create a request for human assistance. Before I do, I would share your name, the issue you described, what I checked, your preferred language, and the urgency. Would you like me to create the request?"&lt;/p&gt;

&lt;p&gt;If the user says yes, the backend creates a real escalation record.&lt;/p&gt;

&lt;p&gt;A reference ID is generated by the backend.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7xsqmzmr0xfppr4jrv8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz7xsqmzmr0xfppr4jrv8.png" alt=" " width="800" height="150"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The important rule is:&lt;/p&gt;

&lt;p&gt;Never tell the user that an escalation was created unless the database operation actually succeeded.&lt;/p&gt;

&lt;p&gt;This avoids fake success messages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 8 - Call Analytics&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;After adding many features, I needed a way to measure what was happening.&lt;/p&gt;

&lt;p&gt;So I built a Call Analytics Dashboard.&lt;/p&gt;

&lt;p&gt;The dashboard tracks:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl8wk2xlfkl9suc08wjy1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl8wk2xlfkl9suc08wjy1.png" alt=" " width="800" height="375"&gt;&lt;/a&gt;&lt;br&gt;
The data comes from real calls stored in SQLite.&lt;/p&gt;

&lt;p&gt;It is not hardcoded.&lt;/p&gt;

&lt;p&gt;A call is considered successful when the user's intended task is successfully completed.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User asks for weather
        |
Weather tool succeeds
        |
Weather information delivered
        |
     SUCCESS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the required tool fails or the task is incomplete, the call can be recorded as failed or incomplete.&lt;/p&gt;

&lt;p&gt;The dashboard can therefore show actual agent performance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 9 - Specialist Agent Handoff&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This was one of the most interesting parts of the project.&lt;/p&gt;

&lt;p&gt;Instead of forcing the main agent to handle every type of request, I created a specialist:&lt;/p&gt;

&lt;p&gt;Bharat Weather Specialist&lt;/p&gt;

&lt;p&gt;The architecture became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                   User
                     |
                     v
              Bharat Voice AI
                     |
              Weather request?
                 /       \
               No         Yes
               |           |
               v           v
          Main Agent   Weather Specialist
                           |
                           v
                      get_weather()
                           |
                           v
                       Real Data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
python&lt;br&gt;
For example:&lt;/p&gt;

&lt;p&gt;User:&lt;/p&gt;

&lt;p&gt;"What is the weather today in Veraval?"&lt;/p&gt;

&lt;p&gt;The main agent says:&lt;/p&gt;

&lt;p&gt;"For detailed weather information, I'll connect you with our weather specialist."&lt;/p&gt;

&lt;p&gt;The specialist then takes over.&lt;/p&gt;

&lt;p&gt;The important part is that the user should not have to explain the entire question again.&lt;/p&gt;

&lt;p&gt;The conversation context is passed to the specialist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Native Agent Handoff&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The LiveKit agent-handoff pattern allows the main agent to return a specialist agent together with an announcement.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@function_tool&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;handoff_to_weather_specialist&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;RunContext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;specialist&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BharatWeatherSpecialist&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;chat_ctx&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat_ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;copy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;exclude_instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;specialist&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ll connect you with our weather specialist.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The actual implementation in the repository follows the project's LiveKit version and architecture.&lt;/p&gt;

&lt;p&gt;The important concept is that the specialist receives the previous conversation context.&lt;/p&gt;

&lt;p&gt;A Day 9 Bug That Taught Me Something&lt;/p&gt;

&lt;p&gt;During testing, the handoff itself worked.&lt;/p&gt;

&lt;p&gt;However, I encountered a problem where the internal function-call information could appear as visible text instead of the actual weather response.&lt;/p&gt;

&lt;p&gt;For example, the user could see something similar to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The function called is get_weather
location = Veraval
forecast_days = 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not acceptable for a real voice assistant.&lt;/p&gt;

&lt;p&gt;The user should hear:&lt;/p&gt;

&lt;p&gt;"The latest weather in Veraval is..."&lt;/p&gt;

&lt;p&gt;not internal tool-call information.&lt;/p&gt;

&lt;p&gt;This led me to investigate the difference between:&lt;/p&gt;

&lt;p&gt;Function registration&lt;br&gt;
Tool execution&lt;br&gt;
LLM output&lt;br&gt;
LiveKit agent handoff&lt;br&gt;
Frontend message rendering&lt;/p&gt;

&lt;p&gt;It was another reminder that a production voice agent is a complete system, not simply an LLM with speech.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Complete System&lt;/strong&gt;After the nine days, Bharat Voice AI looks roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                         USER
                          |
                    Voice / Browser
                          |
                          v
                    LiveKit Agents
                          |
                    Bharat Voice AI
                          |
        +-----------------+------------------+
        |                 |                  |
      Memory           Tools              Routing
        |                 |                  |
      SQLite       +------+-------+          |
                   |              |          |
                Weather      Other Tools     |
                   |                         |
              Open-Meteo                     |
                                             |
                                      Specialist Agent
                                             |
                                      Weather Specialist
                                             |
                                      get_weather()
                                             |
        +------------------------------------+
        |
   Human Escalation
        |
      SQLite
        |
    Murf Falcon
        |
        v
       USER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Outbound calls extend the system through SIP/telephony.&lt;/p&gt;

&lt;p&gt;Analytics records the outcome of calls and makes it visible through the dashboard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Technology Stack&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The main technologies used in the project include:&lt;/p&gt;

&lt;p&gt;Python&lt;br&gt;
LiveKit Agents&lt;br&gt;
Gemini&lt;br&gt;
Deepgram&lt;br&gt;
Murf Falcon&lt;br&gt;
SQLite&lt;br&gt;
Open-Meteo&lt;br&gt;
SIP / Linphone&lt;br&gt;
Browser frontend&lt;br&gt;
GitHub&lt;/p&gt;

&lt;p&gt;Each component has a different responsibility.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Deepgram
Speech → Text

Gemini
Reasoning + Conversation

Tools
Real-world data and actions

SQLite
Persistent state

LiveKit
Real-time voice transport and agent orchestration

Murf Falcon
Text → Natural Voice
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Security&lt;/p&gt;

&lt;p&gt;One of the most important lessons is to never put API credentials directly into the source code.&lt;/p&gt;

&lt;p&gt;Use environment variables.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DEEPGRAM_API_KEY=your_key_here
MURF_API_KEY=your_key_here
LIVEKIT_API_KEY=your_key_here
LIVEKIT_API_SECRET=your_secret_here
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The real .env file should never be committed to GitHub.&lt;/p&gt;

&lt;p&gt;Use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;locally and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.env.example
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;for documentation.&lt;/p&gt;

&lt;p&gt;Never publish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API keys&lt;/li&gt;
&lt;li&gt;SIP credentials&lt;/li&gt;
&lt;li&gt;passwords&lt;/li&gt;
&lt;li&gt;OTPs&lt;/li&gt;
&lt;li&gt;PINs&lt;/li&gt;
&lt;li&gt;caller information&lt;/li&gt;
&lt;li&gt;private database records&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Running the Project&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The complete project is available on GitHub:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://github.com/gunmasterg9/bharat-voice-ai

A typical setup is:

git clone https://github.com/gunmasterg9/bharat-voice-ai.git

cd bharat-voice-ai

cd backend

uv sync
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create your environment configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add the required API credentials.&lt;/p&gt;

&lt;p&gt;Then start the backend using the project's configured startup command.&lt;/p&gt;

&lt;p&gt;Start the frontend and open the browser interface.&lt;/p&gt;

&lt;p&gt;Allow microphone access.&lt;/p&gt;

&lt;p&gt;Then start a conversation.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;"Hello Bharat Voice AI."&lt;/p&gt;

&lt;p&gt;Then test:&lt;/p&gt;

&lt;p&gt;"What is the weather today in Veraval?"&lt;/p&gt;

&lt;p&gt;Then test multilingual interaction:&lt;/p&gt;

&lt;p&gt;"આજે વેરાવળમાં હવામાન કેવું છે?"&lt;/p&gt;

&lt;p&gt;Then test memory:&lt;/p&gt;

&lt;p&gt;"My name is Gautam."&lt;/p&gt;

&lt;p&gt;Then restart the application and verify that the saved profile can be retrieved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Testing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I created tests covering important parts of the system.&lt;/p&gt;

&lt;p&gt;The project includes testing for:&lt;/p&gt;

&lt;p&gt;Agent behavior&lt;br&gt;
Memory&lt;br&gt;
Persistent memory&lt;br&gt;
Language switching&lt;br&gt;
Weather tools&lt;br&gt;
Escalation&lt;br&gt;
Outbound calls&lt;br&gt;
Linphone&lt;br&gt;
Analytics&lt;br&gt;
Specialist handoff&lt;/p&gt;

&lt;p&gt;The Day 9 development test suite included specialist handoff tests alongside the previous functionality.&lt;/p&gt;

&lt;p&gt;The important lesson was that every new feature should be tested without breaking the previous days' work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I Learned&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The biggest lesson from this challenge is that building a voice agent is much more than connecting an LLM to a TTS API.&lt;/p&gt;

&lt;p&gt;A reliable voice agent needs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Voice
+
Reasoning
+
Memory
+
Tools
+
State
+
Error Handling
+
Security
+
Human Handoff
+
Observability
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A model can generate a great response, but the surrounding system determines whether the product is actually reliable.&lt;/p&gt;

&lt;p&gt;I also learned to debug the entire pipeline instead of assuming every problem is caused by the LLM.&lt;/p&gt;

&lt;p&gt;What I Would Build Next&lt;/p&gt;

&lt;p&gt;There is still a lot I would like to add.&lt;/p&gt;

&lt;p&gt;Future improvements could include:&lt;/p&gt;

&lt;p&gt;More specialist agents&lt;br&gt;
Better interruption handling&lt;br&gt;
More Indian languages&lt;br&gt;
Better low-bandwidth support&lt;br&gt;
Advanced call analytics&lt;br&gt;
Conversation quality scoring&lt;br&gt;
Better tool observability&lt;br&gt;
More real-world Indian datasets&lt;br&gt;
Improved outbound-call workflows&lt;br&gt;
Specialist-to-specialist routing&lt;br&gt;
More sophisticated RAG&lt;br&gt;
Production deployment and monitoring&lt;/p&gt;

&lt;p&gt;The long-term goal would be to turn Bharat Voice AI into a platform where different specialized voice agents can work together.&lt;/p&gt;

&lt;p&gt;Final Thoughts&lt;/p&gt;

&lt;p&gt;The most interesting part of the challenge wasn't building the first voice conversation.&lt;/p&gt;

&lt;p&gt;It was everything that came afterward.&lt;/p&gt;

&lt;p&gt;Making the agent remember.&lt;/p&gt;

&lt;p&gt;Making it use real data.&lt;/p&gt;

&lt;p&gt;Making it call a phone.&lt;/p&gt;

&lt;p&gt;Making it know when to ask a human.&lt;/p&gt;

&lt;p&gt;Measuring whether conversations actually succeeded.&lt;/p&gt;

&lt;p&gt;And finally, teaching one agent when another agent is better suited to help.&lt;/p&gt;

&lt;p&gt;That progression changed how I think about voice AI.&lt;/p&gt;

&lt;p&gt;A voice agent isn't just a chatbot that speaks.&lt;/p&gt;

&lt;p&gt;It can become a complete software system with memory, tools, workflows, specialized agents, and real-world actions.&lt;/p&gt;

&lt;p&gt;That is what I wanted to explore with Bharat Voice AI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Project Links&lt;/strong&gt;&lt;br&gt;
GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/gunmasterg9/bharat-voice-ai" rel="noopener noreferrer"&gt;https://github.com/gunmasterg9/bharat-voice-ai&lt;/a&gt;&lt;br&gt;
All Videos in Linkedin&lt;br&gt;
&lt;a href="https://www.linkedin.com/in/gautam-vandar-71a75632b/" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/gautam-vandar-71a75632b/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Challenge&lt;/p&gt;

&lt;p&gt;10 Days of Voice Agents — VoiceForBharat Edition&lt;/p&gt;

&lt;p&gt;Voice Technology&lt;/p&gt;

&lt;p&gt;Murf Falcon&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thank You&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Thank you to Murf AI for organizing the 10 Days of Voice Agents challenge and providing the opportunity to build, experiment, debug, and learn through a real voice-agent project.&lt;/p&gt;

&lt;p&gt;Building Bharat Voice AI over these 10 days was a great experience, and I hope this project helps someone else start building their own voice agent.&lt;/p&gt;

</description>
      <category>antigravity</category>
      <category>python</category>
      <category>gemini</category>
      <category>github</category>
    </item>
  </channel>
</rss>
