<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Deep Bisen</title>
    <description>The latest articles on DEV Community by Deep Bisen (@deep_bisen_06).</description>
    <link>https://dev.to/deep_bisen_06</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4078455%2Fd8209a6f-ca96-41a0-985d-5ac3e3796275.jpg</url>
      <title>DEV Community: Deep Bisen</title>
      <link>https://dev.to/deep_bisen_06</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/deep_bisen_06"/>
    <language>en</language>
    <item>
      <title>Ghar Ka Swad — A Journey Back Home 🍛</title>
      <dc:creator>Deep Bisen</dc:creator>
      <pubDate>Sun, 16 Aug 2026 11:21:15 +0000</pubDate>
      <link>https://dev.to/deep_bisen_06/ghar-ka-swad-a-journey-back-home-3jng</link>
      <guid>https://dev.to/deep_bisen_06/ghar-ka-swad-a-journey-back-home-3jng</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/frontend-2026-07-29"&gt;Frontend Challenge - Comfort Food Edition, Perfect Landing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Ghar Ka Swad 🍛
&lt;/h1&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Ghar Ka Swad&lt;/strong&gt; &lt;em&gt;(Hindi for "the taste of home")&lt;/em&gt; is an interactive journey through Indian comfort food — built around a simple idea:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Some meals don't just feed you. They bring you home.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of building another restaurant or food-delivery landing page, I wanted to create an experience that connects food with &lt;strong&gt;memory, region, family, and the feeling of home&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Visitors discover comfort foods based on mood, travel through India's regional food traditions, walk through a recipe ritual step by step, explore the ingredients behind comfort, leave a memory on a shared wall, and build their own comfort plate — before arriving at a final, quiet emotional close.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flr74lkkazj3jsum82o92.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flr74lkkazj3jsum82o92.png" alt=" " width="800" height="382"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live Demo:&lt;/strong&gt; &lt;a href="https://deepbisen-06.github.io/ghar-ka-swad/" rel="noopener noreferrer"&gt;https://deepbisen-06.github.io/ghar-ka-swad/&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/deepbisen-06/ghar-ka-swad" rel="noopener noreferrer"&gt;https://github.com/deepbisen-06/ghar-ka-swad&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  ✨ The Experience
&lt;/h2&gt;

&lt;p&gt;The site moves through a single continuous narrative rather than disconnected sections:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🍛 &lt;strong&gt;Comfort Discovery&lt;/strong&gt; — "What feels like home to you?" A short, skippable mood prompt recommends a dish and a short emotional story before revealing the full site.&lt;/li&gt;
&lt;li&gt;🥘 &lt;strong&gt;Comfort Food Collection&lt;/strong&gt; — Eight iconic dishes, each with a region, flavor profile, and a playful "comfort score" — clearly labeled as playful, not scientific.&lt;/li&gt;
&lt;li&gt;🇮🇳 &lt;strong&gt;India Journey&lt;/strong&gt; — A custom, hand-illustrated SVG map (not Google Maps) lets you travel through North, West, South, East, and Northeast India, each region revealing its signature dish and food culture.&lt;/li&gt;
&lt;li&gt;📖 &lt;strong&gt;Recipe Journey&lt;/strong&gt; — A scroll-driven walk through a Sunday Rajma Chawal ritual, step by step, with an optional distraction-free &lt;strong&gt;Cook Mode&lt;/strong&gt; for actually cooking along.&lt;/li&gt;
&lt;li&gt;🌿 &lt;strong&gt;Ingredient Constellation&lt;/strong&gt; — An SVG network connecting core Indian aromatics (turmeric, cardamom, cumin, ghee...) to the dishes they define.&lt;/li&gt;
&lt;li&gt;💭 &lt;strong&gt;Memory Wall&lt;/strong&gt; — Postcard-style cards holding real comfort-food memories, with the option to leave your own — stored locally, no backend required.&lt;/li&gt;
&lt;li&gt;🫓 &lt;strong&gt;Comfort Plate&lt;/strong&gt; — Build a plate from dishes you've discovered along the way, then share it.&lt;/li&gt;
&lt;li&gt;❤️ &lt;strong&gt;Final Journey&lt;/strong&gt; — Everything you touched during the visit reassembles into one plate, closing on: &lt;em&gt;"Maybe home was never a place. Maybe it was always the food waiting for you."&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4iw73e9iadp0f4uak3gb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4iw73e9iadp0f4uak3gb.png" alt=" " width="800" height="381"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  🎨 Design &amp;amp; Inspiration
&lt;/h2&gt;

&lt;p&gt;I wanted Ghar Ka Swad to feel like an &lt;strong&gt;interactive editorial story&lt;/strong&gt;, not a restaurant website. The visual direction combines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Warm, earthy tones (parchment, terracotta, turmeric gold, olive) used deliberately, not as a wash-of-orange&lt;/li&gt;
&lt;li&gt;Cormorant Garamond for editorial headings, Inter for restrained body text&lt;/li&gt;
&lt;li&gt;Custom SVG illustrations — the India map, ingredient constellation, and comfort-score meters are all hand-built vectors, not stock assets&lt;/li&gt;
&lt;li&gt;Subtle motion: rising steam, floating spice motes, scroll-driven reveals — used sparingly, never as decoration for its own sake&lt;/li&gt;
&lt;li&gt;Responsive layouts designed intentionally per breakpoint, not just desktop stacked down to mobile&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal throughout was for every section to feel like another page in the same story, rather than a stitched-together set of landing-page blocks.&lt;/p&gt;




&lt;h2&gt;
  
  
  ♿ Accessibility
&lt;/h2&gt;

&lt;p&gt;Accessibility wasn't a pass at the end — it shaped how each interactive feature was built:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every interactive element — including the custom SVG India map and the ingredient constellation — is fully reachable and operable by keyboard alone (&lt;code&gt;Tab&lt;/code&gt;, &lt;code&gt;Enter&lt;/code&gt;, &lt;code&gt;Space&lt;/code&gt;, arrow keys where relevant)&lt;/li&gt;
&lt;li&gt;Visible, high-contrast focus states throughout&lt;/li&gt;
&lt;li&gt;A single accessible dialog primitive handles focus trapping, &lt;code&gt;Escape&lt;/code&gt; to dismiss, and focus restoration everywhere a dialog appears, instead of ad hoc handling per feature&lt;/li&gt;
&lt;li&gt;Dynamic content changes — a Comfort Discovery recommendation appearing, a dish added to your plate — are announced via &lt;code&gt;aria-live&lt;/code&gt; regions for screen reader users&lt;/li&gt;
&lt;li&gt;Full &lt;code&gt;prefers-reduced-motion&lt;/code&gt; support: parallax, steam animation, and smooth-scroll are disabled entirely, with content remaining fully usable&lt;/li&gt;
&lt;li&gt;Color pairings verified against WCAG AA contrast minimums&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  ⚡ Performance
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuvljhuo4oiwx1wwgbub1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuvljhuo4oiwx1wwgbub1.png" alt=" " width="799" height="381"&gt;&lt;/a&gt;&lt;br&gt;
This came from keeping the dependency list deliberately small, hand-building lightweight SVGs instead of reaching for a particle/canvas library, and making sure every animation and layout was built to avoid layout shift.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Built With
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;React + TypeScript + Vite&lt;/li&gt;
&lt;li&gt;Tailwind CSS&lt;/li&gt;
&lt;li&gt;Framer Motion&lt;/li&gt;
&lt;li&gt;Lenis (smooth scrolling)&lt;/li&gt;
&lt;li&gt;Lucide React (icons)&lt;/li&gt;
&lt;li&gt;Hand-built SVG illustrations&lt;/li&gt;
&lt;li&gt;Browser &lt;code&gt;localStorage&lt;/code&gt; for memories and comfort plate persistence&lt;/li&gt;
&lt;li&gt;Native Web Share API, with a clipboard-copy fallback where unsupported&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  💭 What I Learned
&lt;/h2&gt;

&lt;p&gt;The biggest lesson from this project: a landing page doesn't need to be packed with features to feel interactive — it needs the &lt;em&gt;storytelling itself&lt;/em&gt; to be interactive.&lt;/p&gt;

&lt;p&gt;The real challenge wasn't fitting in more sections. It was making sure moving from discovering a dish, to learning where it comes from, to experiencing its recipe, to building your own plate, felt like one continuous thread rather than a checklist of widgets. A lot of the actual work went into deciding what &lt;em&gt;not&lt;/em&gt; to add — a custom cursor, autoplay audio, heavier particle effects — because each of those would have cost accessibility or focus without adding to the story.&lt;/p&gt;




&lt;h2&gt;
  
  
  ❤️ Why Ghar Ka Swad?
&lt;/h2&gt;

&lt;p&gt;For many of us, comfort food is more than a recipe.&lt;/p&gt;

&lt;p&gt;It's a Sunday lunch.&lt;br&gt;
It's something your family makes without measuring anything.&lt;br&gt;
It's the smell from the kitchen before you even walk in.&lt;br&gt;
It's a dish that somehow tastes different when you're away from home.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ghar Ka Swad is my attempt to turn that feeling into an interactive web experience.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  🔗 Links
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Live Demo:&lt;/strong&gt; &lt;a href="https://deepbisen-06.github.io/ghar-ka-swad/" rel="noopener noreferrer"&gt;https://deepbisen-06.github.io/ghar-ka-swad/&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/deepbisen-06/ghar-ka-swad" rel="noopener noreferrer"&gt;https://github.com/deepbisen-06/ghar-ka-swad&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;Built with ❤️ for the &lt;strong&gt;DEV Community Frontend Challenge: Comfort Food Edition&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Maybe home was never a place.&lt;br&gt;
Maybe it was always the food waiting for you. 🍛&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>frontendchallenge</category>
      <category>devchallenge</category>
      <category>css</category>
      <category>webdev</category>
    </item>
    <item>
      <title>🎙️ Building EduBuddy: My 10-Day Journey from a Voice Agent to a Multi-Agent AI Learning Companion</title>
      <dc:creator>Deep Bisen</dc:creator>
      <pubDate>Sat, 15 Aug 2026 05:06:29 +0000</pubDate>
      <link>https://dev.to/deep_bisen_06/building-edubuddy-my-10-day-journey-from-a-voice-agent-to-a-multi-agent-ai-learning-companion-383c</link>
      <guid>https://dev.to/deep_bisen_06/building-edubuddy-my-10-day-journey-from-a-voice-agent-to-a-multi-agent-ai-learning-companion-383c</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Ten days ago, I signed up for &lt;strong&gt;10 Days of Voice Agents — VoiceForBharat Edition&lt;/strong&gt;, a challenge by Murf AI, with a starter pipeline and a vague idea: build something that lets people learn by talking instead of typing. I didn't have a finished architecture in my head. I had a track — &lt;em&gt;Learning &amp;amp; Literacy&lt;/em&gt; — and a daily prompt to add one new capability.&lt;/p&gt;

&lt;p&gt;What came out the other end is &lt;strong&gt;EduBuddy&lt;/strong&gt;, an AI voice and text learning companion that gives learners exercises, evaluates their spoken answers, remembers who they are, calls them for practice, knows when to bring in a human, and can hand a maths question off to a specialist agent built just for that.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F47rlrrxyikl8mn20qyke.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F47rlrrxyikl8mn20qyke.png" alt=" " width="800" height="387"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'm writing this partly to document the build for myself, and partly because I think the more useful story isn't "look what I shipped" — it's how a voice agent actually gets assembled, piece by piece, and where it breaks along the way. If you're evaluating this as a project, or thinking about building something similar, this should tell you exactly what's under the hood.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Most learning tools are text-first: read the material, type the answer, read the feedback. That works well for a lot of people, but it puts a barrier in front of learners who are more comfortable speaking than typing, or who find a keyboard-and-screen interface slower to engage with than a conversation.&lt;/p&gt;

&lt;p&gt;Speaking is a lower-friction way to interact for a lot of learners. If a learner can talk through a problem, get evaluated on what they said out loud, and get a spoken response back, the interaction feels closer to being tutored than to filling out a form. That's the gap EduBuddy is aimed at — not a claim about literacy rates or any specific population, just a bet that voice removes friction that text adds for some learners.&lt;/p&gt;
&lt;h2&gt;
  
  
  What Is EduBuddy?
&lt;/h2&gt;

&lt;p&gt;EduBuddy is a voice-and-text learning companion built for anyone who wants to practice a subject through conversation rather than through a form. A learner can talk to it in the browser or receive an outbound practice call, work through an exercise, have their spoken answer evaluated, and pick up where they left off next time because the agent remembers them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsrp0n1elx4fsf0lju2pg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsrp0n1elx4fsf0lju2pg.png" alt=" " width="799" height="392"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The difference from a plain chatbot is that EduBuddy doesn't just answer questions — it runs a structured interaction: it can fetch an exercise, score an answer, decide whether to escalate to a human, or route a maths-specific question to a specialist agent that continues the same conversation. It's closer to a small system of cooperating agents and tools than to a single prompt.&lt;/p&gt;
&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Learner
   ↓
Frontend / SIP
   ↓
LiveKit
   ↓
Deepgram STT
   ↓
Google Gemini
   ↓
Tools / SQLite / Agent Handoff
   ↓
Murf Falcon TTS
   ↓
Learner
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontend / SIP&lt;/strong&gt; — entry point for the learner, either a Next.js browser session or an inbound/outbound phone call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LiveKit&lt;/strong&gt; — handles real-time audio transport between the learner and the agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deepgram (Nova-3)&lt;/strong&gt; — converts the learner's speech into text, with language configuration for English and Hindi/code-mixed input.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google Gemini&lt;/strong&gt; — the reasoning layer; decides whether to respond directly, call a tool, store memory, escalate, or hand off to the maths specialist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tools / SQLite / Agent Handoff&lt;/strong&gt; — the functional layer: exercise fetching, answer scoring, learner memory, escalation records, call analytics, and specialist routing all live here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Murf Falcon&lt;/strong&gt; — converts the agent's text response back into speech, described in the challenge as the fastest TTS API available.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  10-Day Build Journey
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Day&lt;/th&gt;
&lt;th&gt;What I Built&lt;/th&gt;
&lt;th&gt;Why It Mattered&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Set up and tested the base voice-agent pipeline&lt;/td&gt;
&lt;td&gt;Established a working STT → LLM → TTS loop before adding anything else&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Defined personality, objectives, and safety guardrails&lt;/td&gt;
&lt;td&gt;Gave the agent a consistent identity and boundaries before it had any real capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Personalized the frontend and made agent state visible&lt;/td&gt;
&lt;td&gt;Let learners (and me, debugging) see what the agent was doing during a call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Added persistent learner memory with SQLite&lt;/td&gt;
&lt;td&gt;Turned each session from a blank slate into a continuation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Added learning tools (&lt;code&gt;fetch_next_exercise&lt;/code&gt;, &lt;code&gt;score_spoken_answer&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Gave the agent something to actually &lt;em&gt;do&lt;/em&gt;, not just talk about&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Added outbound SIP calling&lt;/td&gt;
&lt;td&gt;Took EduBuddy out of the browser and into a real phone call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Added human escalation&lt;/td&gt;
&lt;td&gt;Gave the agent a way to recognize its limits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Added call analytics from real SQLite data&lt;/td&gt;
&lt;td&gt;Made performance visible instead of assumed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Added the Maths Practice Specialist and handoff logic&lt;/td&gt;
&lt;td&gt;Proved the agent could delegate instead of trying to do everything itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Documented the journey&lt;/td&gt;
&lt;td&gt;This article&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F89yr9vshnxw1tmwrsdvm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F89yr9vshnxw1tmwrsdvm.png" alt=" " width="799" height="388"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Core Features
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Voice + Personality
&lt;/h3&gt;

&lt;p&gt;Murf Falcon handles text-to-speech, and Day 2 was spent defining EduBuddy's personality, objectives, and safety guardrails before any real functionality existed — so every later feature had a consistent voice and boundaries to operate inside.&lt;/p&gt;
&lt;h3&gt;
  
  
  Multilingual / Code-Mixed Interaction
&lt;/h3&gt;

&lt;p&gt;Deepgram's STT is configured for multi-language detection to handle English, Hindi, and code-mixed speech. On the response side, correct script handling mattered — Hindi output needed to render in Devanagari rather than being Romanized, which took explicit attention in the STT config, the TTS config, and the system prompt, not just the LLM.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8a0rnmoftiqksrbi4n1t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8a0rnmoftiqksrbi4n1t.png" alt=" " width="799" height="383"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Learner Memory
&lt;/h3&gt;

&lt;p&gt;Learner information relevant to future sessions is stored in SQLite — not a full transcript log, just what's useful for continuity. This is what lets EduBuddy recognize a returning learner instead of starting cold every time.&lt;/p&gt;
&lt;h3&gt;
  
  
  Learning Tools
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;fetch_next_exercise&lt;/code&gt; and &lt;code&gt;score_spoken_answer&lt;/code&gt; are the two tools that turn the interaction into an actual practice loop: the agent hands the learner an exercise, the learner answers out loud, and the tool evaluates the response instead of the LLM eyeballing it.&lt;/p&gt;
&lt;h3&gt;
  
  
  Outbound SIP Calling
&lt;/h3&gt;

&lt;p&gt;Using LiveKit's SIP integration, EduBuddy can initiate a call rather than only responding to one. This is what makes a "daily practice call" possible — the agent reaches the learner instead of waiting for them to open the app.&lt;/p&gt;
&lt;h3&gt;
  
  
  Human Escalation
&lt;/h3&gt;

&lt;p&gt;Covered in detail below — this is the feature I'd point to first if someone asked what "responsible" looks like in this project.&lt;/p&gt;
&lt;h3&gt;
  
  
  Call Analytics
&lt;/h3&gt;

&lt;p&gt;A dashboard backed by real SQLite call records, not placeholder numbers, so I could actually see how sessions were going rather than assume.&lt;/p&gt;
&lt;h3&gt;
  
  
  Specialist Agent Handoff
&lt;/h3&gt;

&lt;p&gt;The main agent recognizes maths-specific requests and hands the conversation to a dedicated Maths Practice Specialist agent, which continues without asking the learner to repeat themselves.&lt;/p&gt;
&lt;h2&gt;
  
  
  Human Escalation
&lt;/h2&gt;

&lt;p&gt;This feature gets its own section because it's the clearest example of designing for the agent's limits rather than pretending it doesn't have any.&lt;/p&gt;

&lt;p&gt;Escalation triggers on two conditions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The learner shows severe frustration or emotional distress related to the learning session.&lt;/li&gt;
&lt;li&gt;The learner explicitly asks for a teacher, or stays stuck on the same concept after repeated attempts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Before anything is shared, the agent asks the learner for permission. What gets recorded is a short, human-readable summary — an escalation ID, the reason, relevant details, and a severity/status field — never sensitive information like OTPs, passwords, PINs, or account details. The learner gets a reference ID and a clear explanation of what happens next. The point isn't that the agent tries harder to solve the problem itself — it's that it knows when to stop trying and hand off cleanly.&lt;/p&gt;
&lt;h2&gt;
  
  
  Multi-Agent Handoff
&lt;/h2&gt;

&lt;p&gt;Instead of one large prompt trying to be good at everything, EduBuddy's main agent delegates maths-specific conversations to a separate specialist:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Learner: "Can you help me solve 3x + 9 = 24?"
Main Agent: "I will connect you to our maths specialist."
Maths Specialist: [continues the conversation directly, without asking
                    the learner to repeat the problem]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The benefit isn't just cleaner prompts — it's that each agent can be tuned, tested, and extended independently. Adding a second or third specialist later doesn't mean rewriting the main agent's entire personality and instruction set; it means adding a new handoff target.&lt;/p&gt;

&lt;h2&gt;
  
  
  Call Analytics
&lt;/h2&gt;

&lt;p&gt;The dashboard tracks total calls, successful calls, failed calls, and failure categories, all computed from real SQLite records rather than hardcoded placeholder data. For EduBuddy, a "successful" call is defined as one where the learner actively engages and completes at least one exercise or practice activity — a call that ends before that point counts as unsuccessful. Caller identifiers are masked in the dashboard to protect privacy. I'm not going to quote specific numbers here, since the point of this section is the mechanism, not a metric I'd be inventing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technical Challenges
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What Broke and What I Learned&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem&lt;/th&gt;
&lt;th&gt;Why It Happened&lt;/th&gt;
&lt;th&gt;What I Changed / Learned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Outbound SIP calls failed with &lt;code&gt;SipCallTo should be a phone number or SIP user, not a full SIP URI&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;The destination was being passed to LiveKit's SIP integration in the wrong format&lt;/td&gt;
&lt;td&gt;Corrected the SIP destination/configuration so it matched what LiveKit expected, rather than passing a full URI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent understood Hindi input but sometimes responded in the wrong language/voice&lt;/td&gt;
&lt;td&gt;The mismatch wasn't isolated to the LLM — it touched STT language settings, TTS voice/language config, and the system prompt&lt;/td&gt;
&lt;td&gt;Learned to treat language behavior as a cross-component problem, not something to fix by tweaking the prompt alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Specialist agent sometimes went silent after the main agent announced the handoff&lt;/td&gt;
&lt;td&gt;The specialist session wasn't reliably activating right after the handoff message&lt;/td&gt;
&lt;td&gt;Paid closer attention to LiveKit's agent/session lifecycle during handoff rather than assuming the framework would "just work"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern across all three: a bug in STT, TTS, session handling, or SIP config can present exactly like "the AI is behaving wrong," even when the LLM itself is doing its job correctly. Debugging a voice agent means checking the whole pipeline, not just the prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Build Your Own Voice Agent
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Set up a LiveKit project for real-time audio transport.&lt;/li&gt;
&lt;li&gt;Configure environment variables (see below) — never hardcode secrets.&lt;/li&gt;
&lt;li&gt;Connect Deepgram for speech-to-text.&lt;/li&gt;
&lt;li&gt;Connect Google Gemini as the reasoning layer.&lt;/li&gt;
&lt;li&gt;Connect Murf Falcon for text-to-speech.&lt;/li&gt;
&lt;li&gt;Get a basic STT → LLM → TTS conversation loop working end to end.&lt;/li&gt;
&lt;li&gt;Add function tools so the agent can do things, not just talk.&lt;/li&gt;
&lt;li&gt;Add SQLite for learner/session memory.&lt;/li&gt;
&lt;li&gt;Add a human escalation path with clear triggers.&lt;/li&gt;
&lt;li&gt;Add analytics from real usage data.&lt;/li&gt;
&lt;li&gt;Add specialist agent handoff for domain-specific tasks.&lt;/li&gt;
&lt;li&gt;Test through the browser first, then through SIP if you need telephony.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd7i73vpef9aw3spu0xei.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd7i73vpef9aw3spu0xei.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Example
&lt;/h2&gt;

&lt;p&gt;Here's the conceptual shape of the STT → LLM → TTS loop, written as pseudocode rather than a copy of any specific SDK call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Pseudocode — illustrates the flow, not a literal API
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;handle_turn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio_chunk&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio_chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="c1"&gt;# Deepgram
&lt;/span&gt;    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;TOOLS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# Gemini + function tools
&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_call&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fetch_next_exercise&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;exercise&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_next_exercise&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;learner_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exercise&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;synthesize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="c1"&gt;# Murf Falcon
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;audio&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Repository
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/deepbisen-06/Voice-Agents" rel="noopener noreferrer"&gt;https://github.com/deepbisen-06/Voice-Agents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repo contains the agent logic, the tool implementations, and the frontend used to run and test EduBuddy locally. Real credentials and caller data are excluded — you'll need your own API keys to run it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Improve Next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;More specialist agents beyond maths&lt;/li&gt;
&lt;li&gt;Stronger multilingual and code-mixed support&lt;/li&gt;
&lt;li&gt;More reliable specialist handoff (fixing the silence issue for good)&lt;/li&gt;
&lt;li&gt;Better telephony error recovery&lt;/li&gt;
&lt;li&gt;More personalized learning paths per learner&lt;/li&gt;
&lt;li&gt;Deeper, more actionable analytics&lt;/li&gt;
&lt;li&gt;Better interruption handling mid-conversation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Key Lessons
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A voice agent is a system, not a model — STT, LLM, TTS, transport, and session state all have to work together.&lt;/li&gt;
&lt;li&gt;Language correctness (script, voice, config) has to be checked across the whole pipeline, not patched at the prompt level.&lt;/li&gt;
&lt;li&gt;Giving an agent a way to say "I don't know, let me get a human" is a feature, not a limitation.&lt;/li&gt;
&lt;li&gt;Specialist sub-agents scale better than one prompt trying to do everything.&lt;/li&gt;
&lt;li&gt;Real data — for memory, analytics, or escalation — is worth the extra setup over hardcoded placeholders from day one.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;I started this challenge with a pipeline that could just about hold a conversation. Ten days later, EduBuddy remembers learners, runs practice sessions, makes outbound calls, knows when to ask for human help, tracks its own performance, and delegates maths questions to a specialist built for the job. None of that happened in one step — it was ten small, deliberate additions, each one breaking something the previous day's work didn't expose.&lt;/p&gt;

&lt;p&gt;Built during &lt;strong&gt;10 Days of Voice Agents — VoiceForBharat Edition&lt;/strong&gt;, using &lt;strong&gt;Murf Falcon&lt;/strong&gt;.&lt;/p&gt;




</description>
      <category>ai</category>
      <category>agents</category>
      <category>murfai</category>
      <category>voiceagent</category>
    </item>
  </channel>
</rss>
