<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 927tanmay</title>
    <description>The latest articles on DEV Community by 927tanmay (@927tanmay).</description>
    <link>https://dev.to/927tanmay</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F398519%2F8db6194f-b540-4c2c-952f-7a712ab56d65.jpeg</url>
      <title>DEV Community: 927tanmay</title>
      <link>https://dev.to/927tanmay</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/927tanmay"/>
    <language>en</language>
    <item>
      <title>I built my friend a Gemma-powered mock interviewer that talks back, entirely in the browser</title>
      <dc:creator>927tanmay</dc:creator>
      <pubDate>Mon, 05 Oct 2026 00:11:52 +0000</pubDate>
      <link>https://dev.to/927tanmay/i-built-my-friend-a-gemma-powered-mock-interviewer-that-talks-back-entirely-in-the-browser-1o94</link>
      <guid>https://dev.to/927tanmay/i-built-my-friend-a-gemma-powered-mock-interviewer-that-talks-back-entirely-in-the-browser-1o94</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hf26"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;My friend is a software engineer applying for backend roles. She's switching after a long time, so interviews feel new again. She has the skills. What makes her hesitant is the conversation itself: saying what she knows out loud, to a stranger, while they listen. She's been practicing on her own, and practicing interviews alone doesn't really work.&lt;/p&gt;

&lt;p&gt;You read a question, think of an answer in your head, and move on. You never say it out loud, nobody stops you halfway, and nobody asks "okay, but what did &lt;em&gt;you&lt;/em&gt; do there?" Asking a friend to play interviewer works once or twice, then it gets awkward for both of you.&lt;/p&gt;

&lt;p&gt;So I built her &lt;strong&gt;Interview Room&lt;/strong&gt;: a mock interviewer that actually talks.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You pick a track (Frontend, Backend or Machine learning), a round (behavioural, technical, HR, or a full loop of five questions), your level, how long your answers should be, and an interviewer: Ananya or Aarav, friendly, neutral or tough.&lt;/li&gt;
&lt;li&gt;It asks out loud. You answer out loud, and you can pause to think.&lt;/li&gt;
&lt;li&gt;It listens to what you said and asks one follow-up on it. If you said "we" the whole time, it asks what you did yourself. If you never said how it ended, it asks.&lt;/li&gt;
&lt;li&gt;At the end you review your own answers, one at a time, next to what a strong answer usually covers. Then you get a report built from your own marks and from what was measured: how long you spoke, your pace, your filler phrases, how often you said "I" vs "we". &lt;strong&gt;There are no scores.&lt;/strong&gt; The app never decides if your answer was good. You do.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can do it as a &lt;strong&gt;video interview&lt;/strong&gt; with a 3D interviewer who lip syncs, or as a &lt;strong&gt;phone screen&lt;/strong&gt; with just a voice, like a recruiter call. If you want it to feel like a real video call, you can turn your own camera on: a small mirror in the corner, off by default, never recorded or sent anywhere.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everything runs in the browser.&lt;/strong&gt; Hearing when you talk, turning your speech into text, the interviewer, its voice and the avatar all run on your own laptop. What you say is never uploaded. The network is only used to download the models the first time (they stay on your laptop after that), plus a few runtime files and a font.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Live:&lt;/strong&gt; &lt;a href="https://interview-room-iooj.onrender.com/" rel="noopener noreferrer"&gt;https://interview-room-iooj.onrender.com/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/pJPoltF6O6M" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Before you click:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It needs a laptop or desktop with WebGPU, so a recent Chrome or Edge.&lt;/li&gt;
&lt;li&gt;The first visit downloads about 1.6 GB. On my connection that took 4 to 8 minutes. After that it's ready in 2 to 3 seconds.&lt;/li&gt;
&lt;li&gt;The home page checks your device first and shows what will download and how big it is, before anything starts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F577jx4iptl92qap6znva.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F577jx4iptl92qap6znva.jpg" alt="The home page: what it does, Light or Heavy, and what each one downloads" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frb5xxmcrw1q0djqgr51d.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frb5xxmcrw1q0djqgr51d.jpg" alt="The self-review: your answer next to what a strong answer covers, each point marked covered, partly or missed" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj47nubfs2z3b5hzqr61x.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj47nubfs2z3b5hzqr61x.jpg" alt="The report: up to three things to work on, each saying where it came from" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/927tanmay" rel="noopener noreferrer"&gt;
        927tanmay
      &lt;/a&gt; / &lt;a href="https://github.com/927tanmay/interview-room" rel="noopener noreferrer"&gt;
        interview-room
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A spoken mock interviewer that runs entirely in your browser: Whisper, Gemma, Kokoro and a lip-synced avatar, all on-device.
    &lt;/h3&gt;
  &lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The pieces
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hears when you start and stop talking&lt;/td&gt;
&lt;td&gt;Silero VAD&lt;/td&gt;
&lt;td&gt;1.8 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Turns your speech into text&lt;/td&gt;
&lt;td&gt;Whisper base&lt;/td&gt;
&lt;td&gt;295 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The interviewer&lt;/td&gt;
&lt;td&gt;Gemma 3 1B (q4)&lt;/td&gt;
&lt;td&gt;880 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The interviewer's voice&lt;/td&gt;
&lt;td&gt;Kokoro 82M&lt;/td&gt;
&lt;td&gt;326 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runs it all&lt;/td&gt;
&lt;td&gt;ONNX Runtime Web, on WebGPU&lt;/td&gt;
&lt;td&gt;67 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All open-weight, all in the browser through transformers.js. The microphone, the voice, lip sync and interrupting the interviewer come from &lt;a href="https://github.com/927tanmay/react-ai-voice-avatar" rel="noopener noreferrer"&gt;react-ai-voice-avatar&lt;/a&gt;, an npm package I maintain. Interview Room itself is new. I started it on Saturday.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjw04sx46woogqhjif9gj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjw04sx46woogqhjif9gj.png" alt="How Interview Room works: from your answer to Gemma's follow-up, all in the browser" width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemma is the interviewer, but it's not in charge
&lt;/h3&gt;

&lt;p&gt;My first idea was to give Gemma the conversation and let it run the interview. With a 1B model that didn't work. In my first tests it said "That's a good start" before nearly every question, wrapped its lines in quote marks, sometimes answered &lt;em&gt;as the candidate&lt;/em&gt;, and once invented a "15% increase in session duration" that nobody had mentioned.&lt;/p&gt;

&lt;p&gt;So I split the job three ways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Plain code decides what to ask about.&lt;/strong&gt; Rules read the answer and pick an angle: what you did yourself if you only said "we", the outcome if you never gave one, specifics if it was vague, a missing key point, an edge case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemma only words it.&lt;/strong&gt; It gets the conversation as real turns, plus a short note like "ask what they did personally".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A guard checks every line before it's spoken.&lt;/strong&gt; It strips quotes and markdown, drops praise, drops anything copied from the question, and keeps exactly one question. If nothing usable is left, or Gemma takes more than 6 seconds, the interviewer says a written line for that angle instead.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;With that, 8 of 12 test follow-ups were usable straight from Gemma and the other 4 fell back cleanly, so you never hear a broken line. Each takes about a second. One real follow-up, after an answer that only said "we":&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What specific action did you take to improve the dashboard speed?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Two Gemmas, side by side
&lt;/h3&gt;

&lt;p&gt;Before building anything, I wrote a small eval page that runs the same prompts on any model in the browser, and ran it on Gemma 3 1B and Gemma 4 E2B on my M4 MacBook:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Gemma 3 1B&lt;/th&gt;
&lt;th&gt;Gemma 4 E2B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Download&lt;/td&gt;
&lt;td&gt;859 MB&lt;/td&gt;
&lt;td&gt;3.11 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First load, download included&lt;/td&gt;
&lt;td&gt;4 min&lt;/td&gt;
&lt;td&gt;14 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Follow-ups on the right angle&lt;/td&gt;
&lt;td&gt;8 of 12&lt;/td&gt;
&lt;td&gt;12 of 12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time per follow-up&lt;/td&gt;
&lt;td&gt;1.0 to 1.5 s&lt;/td&gt;
&lt;td&gt;1.0 to 1.6 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code review (2 cases)&lt;/td&gt;
&lt;td&gt;0 right&lt;/td&gt;
&lt;td&gt;2 right&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judging a STAR answer (2 cases)&lt;/td&gt;
&lt;td&gt;0 right&lt;/td&gt;
&lt;td&gt;1 right, by calling everything present&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things surprised me. Gemma 4 E2B was as fast as the 1B, because its replies are shorter. And it loads with the plain &lt;code&gt;text-generation&lt;/code&gt; pipeline in transformers.js, which only pulls the text parts of the model. The vision and audio encoders are never downloaded.&lt;/p&gt;

&lt;p&gt;The table also shaped the whole app. &lt;strong&gt;Neither model could tell me whether a STAR answer was complete.&lt;/strong&gt; If a model can't judge that reliably, I shouldn't hand my friend a score from it. That's where the "no scores" rule came from.&lt;/p&gt;

&lt;h3&gt;
  
  
  The report: you judge, the app measures
&lt;/h3&gt;

&lt;p&gt;Instead of a model grading answers, the report has two halves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your own marks.&lt;/strong&gt; When the interview ends you go through your answers one at a time. First, a quick "how did that one feel?": good, okay or rough. Then the points a strong answer usually covers appear next to your own words, and you mark each one covered, partly or missed. Every question in the bank has 3 to 5 of these points, 206 in total, each tagged as a result, your own role, an example, a trade-off and so on. It takes a minute or two, and works with the keyboard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What was measured.&lt;/strong&gt; Plain code, no model, so it's instant:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Number&lt;/th&gt;
&lt;th&gt;How&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Answer length&lt;/td&gt;
&lt;td&gt;Your first word to your last, from the voice detector, across thinking pauses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pace&lt;/td&gt;
&lt;td&gt;Words heard divided by the time you were actually talking, next to a rough 120 to 160 a minute&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to first word&lt;/td&gt;
&lt;td&gt;From the end of the question to when you started&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long pauses&lt;/td&gt;
&lt;td&gt;Gaps over 3 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Filler phrases&lt;/td&gt;
&lt;td&gt;"you know", "I mean", "basically", "kind of", "sort of", "like" next to a comma&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"I" vs "we"&lt;/td&gt;
&lt;td&gt;Counted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Numbers&lt;/td&gt;
&lt;td&gt;Figures and spoken numbers. Vague ones like "one or two" don't count&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Then simple rules put the two together. At the top you get up to three things to work on, and each says where it came from: "you left out the result in 3 of 4 stories", or "you felt good about this one but marked most points missed". Where your marks and the numbers disagree, it says so gently. If you marked your own role as covered but said "we" twelve times and "I" twice, it asks you to take another look. Every answer ends with a short &lt;strong&gt;Next time&lt;/strong&gt; list made from the points you missed, and a &lt;strong&gt;Practice this one again&lt;/strong&gt; button.&lt;/p&gt;

&lt;p&gt;One thing I had to be honest about: Whisper is trained to leave out "um" and "uh", so the app can't count them. It counts the filler phrases Whisper does keep, highlights every one in your transcript so you can see if it got one wrong, and says plainly that "um" and "uh" aren't counted.&lt;/p&gt;

&lt;h3&gt;
  
  
  Letting people think
&lt;/h3&gt;

&lt;p&gt;Real answers have long pauses. A voice assistant tuned for chat would cut you off mid-thought, which is the fastest way to make this feel fake.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A pause doesn't end your answer. It waits for about 5 seconds of quiet, or for you to press &lt;strong&gt;I'm done&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;If you say nothing at all, it nudges you after 12 seconds and moves on after 25.&lt;/li&gt;
&lt;li&gt;You can talk to it: "can you repeat that?", "what do you mean?", "I don't know", "can we skip this one?", "give me a minute". It handles those out loud instead of counting them as your answer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Keeping 1.6 GB in a browser
&lt;/h3&gt;

&lt;p&gt;transformers.js caches models in the browser's Cache API, but Chrome won't store a single file of 256 MB or more there, so the 859 MB Gemma file was downloaded again on every visit. My package already kept its own models in the Origin Private File System, so I exported that cache and used it for Gemma as well. Now a second visit is ready in 2 to 3 seconds.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fixing my own package along the way
&lt;/h3&gt;

&lt;p&gt;Building a real app on my own package found two bugs in it. Whisper was called without chunking, so if you spoke for more than 30 seconds without a pause, the rest was cut off. And if the app's reply was empty, the hook got stuck on "thinking" and never listened again. For an interview app both are deal breakers. I fixed them in the package, added the speaking time to &lt;code&gt;onSubmit&lt;/code&gt; for the pace, exported the model cache, and released it as 0.7.0 during the weekend.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hosting on Render
&lt;/h3&gt;

&lt;p&gt;It's a free static site on Render. There's no server because there's nothing to run on one. What matters is the cross-origin isolation headers, which ONNX Runtime needs for multithreaded WebAssembly. They live in &lt;code&gt;render.yaml&lt;/code&gt; next to the code, and every push to main deploys.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Letting Gemma run the interview.&lt;/strong&gt; Covered above. A small model is a good writer and a bad judge, so I gave it the writing and kept the judging away from it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trusting Whisper's transcript for fillers.&lt;/strong&gt; My first filler count was confidently low, because "um" and "uh" were never in the text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The first real run.&lt;/strong&gt; Speaking to it for real showed things no test had: a spoken "skip" didn't skip, "sorry, I don't know" wasn't understood, some follow-ups drifted off the point they were meant to ask about, "like" was counted as a filler when it wasn't one, and "one or two" was counted as a number. All fixed before submitting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Downloads fail sometimes.&lt;/strong&gt; At 1 am, testing the live site, Whisper and Kokoro lost their connection halfway through. The page sat at "Loading models 21%" with no message, and the voice quietly switched to a backup. Now it says which model failed and offers Try again, which only fetches what's missing.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I didn't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No scores, on purpose.&lt;/strong&gt; Not out of 10, not a percentage, not a tick grid.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No STAR judging.&lt;/strong&gt; Neither model got it right reliably.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heavy mode isn't there yet.&lt;/strong&gt; Gemma 4 E2B gave better follow-ups in my tests and reviewed code correctly, but running it alongside Whisper, Kokoro and the avatar needs more memory and more testing than I had time for. It shows as "coming soon".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System design questions&lt;/strong&gt; are written but parked. A system design round is a conversation, not a question and an answer, and it deserves its own design.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What my friend said
&lt;/h2&gt;

&lt;p&gt;I sent her the link while I was still building. She ran it on her own laptop, start to finish: everything local, no hiccups, and nothing to pay for.&lt;/p&gt;

&lt;p&gt;Her reaction was really positive. What she liked:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's good for practicing, and it helps with the anxiety.&lt;/strong&gt; Saying your answers out loud to an interviewer, before the real one, takes some of the fear out of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;She saw the gaps in her answers straight away.&lt;/strong&gt; Her favourite part of the report was the side-by-side view of what she covered and what she missed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The level of the questions&lt;/strong&gt; felt right to her.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A proper mock interview, free.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;She also gave me a list, which is the most useful thing a friend can do. Two of the four were small enough to add before submitting:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;She asked for&lt;/th&gt;
&lt;th&gt;What I think&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;More than one follow-up per question&lt;/td&gt;
&lt;td&gt;Fair. Real interviewers dig twice when the first answer is thin. The engine already picks the angle, so a second round is mostly a rule change.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The report should say what to improve, not only what was missed&lt;/td&gt;
&lt;td&gt;She was right, and this one was quick. &lt;strong&gt;I added it before submitting:&lt;/strong&gt; every answer card now ends with a short "Next time" list, built from the points she marked missed or partly, each turned into something to do.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Adding her own questions&lt;/td&gt;
&lt;td&gt;Her idea: upload her own list of questions and answers, and practice them in random order. That turns it from a demo into something she'd use every day.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Her camera on&lt;/td&gt;
&lt;td&gt;Seeing yourself makes it feel like a real call. &lt;strong&gt;Added before submitting:&lt;/strong&gt; a "show my camera" switch on the interview screen, off by default. It's a mirror in the corner, it only asks for the camera when you turn it on, and like everything else it stays on your laptop: never recorded, never sent.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;A mock interview is private. You stumble, you say the wrong thing, you talk about your current job and why you want to leave. I didn't want any of that going to a server, and with open-weight models it doesn't have to. Whisper, Gemma and Kokoro run on your own laptop, and the code is open, so you can check that nothing else leaves it.&lt;/p&gt;

&lt;p&gt;Open models also meant I could test them myself. I ran the same eval on two Gemma models, saw exactly where each one failed, and built the app around that. With a hosted API I'd be testing a model that could change under me next week.&lt;/p&gt;

&lt;p&gt;And it keeps working. There's no API key, no account and no bill that stops the demo when the challenge ends. It's a static page and some open models, so it's just a link my friend can keep using. It stands on open pieces (transformers.js, ONNX Runtime, Silero, Whisper, Kokoro, Gemma, three.js), and the app and my package are open too, so anyone can take it further.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;I did the planning and the first model tests in a private chat with Claude, because it has my personal notes in it. What came out of it is all in the repo: &lt;a href="https://github.com/927tanmay/interview-room/blob/main/docs/PLAN.md" rel="noopener noreferrer"&gt;&lt;code&gt;docs/PLAN.md&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/927tanmay/interview-room/blob/main/docs/MODEL-TESTS.md" rel="noopener noreferrer"&gt;&lt;code&gt;docs/MODEL-TESTS.md&lt;/code&gt;&lt;/a&gt; and the first two commits.&lt;/p&gt;

&lt;p&gt;The build itself was one Claude Code session, recorded with Entire: more than 30 checkpoints, each tied to a commit, with the prompts and reasoning behind it. I set the direction and the rules, Claude Code planned and wrote the code, and every step stopped for me to check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://entire.io/gh/927tanmay/interview-room/session/e0972c9d-f551-419e-8995-044a83f59546" rel="noopener noreferrer"&gt;The full session on Entire&lt;/a&gt;&lt;/strong&gt; (&lt;a href="https://entire.io/gh/927tanmay/interview-room" rel="noopener noreferrer"&gt;the repo on Entire&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;Three checkpoints worth opening:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://entire.io/gh/927tanmay/interview-room/commit/50ecef9e0d61796f9ecb01ec8ff31b1c498f6bae" rel="noopener noreferrer"&gt;Testing Gemma 4 E2B against Gemma 3 1B, and choosing the interviewer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://entire.io/gh/927tanmay/interview-room/commit/c293b226e973ca7849235a0e231f47de0e259212" rel="noopener noreferrer"&gt;Building the report: things to work on, second looks, charts, answer cards&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://entire.io/gh/927tanmay/interview-room/commit/58c196752de9b7f0ca1d45b6919c12b20c32ef69" rel="noopener noreferrer"&gt;Her feedback, shipped: the "Next time" list on every answer&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One tip for other Entire users: Claude Code includes your account email in the session context, and Entire's email redaction is off by default. Turn it on in &lt;code&gt;.entire/settings.json&lt;/code&gt; before you push to a public repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma:&lt;/strong&gt; Gemma 3 1B is the live interviewer, running in the browser on WebGPU. It words every follow-up, with rules choosing the angle and a guard checking the line. I tested Gemma 4 E2B against it on the same eval before choosing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Render:&lt;/strong&gt; Render hosts the front end as a static site, with the cross-origin isolation headers the in-browser models need set in &lt;code&gt;render.yaml&lt;/code&gt;. Every push deploys.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Entire:&lt;/strong&gt; the whole build is one recorded session, more than 30 checkpoints, each tied to its commit and public on Entire and in the repo. Linked above.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Her list comes first:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your own questions and answers, uploaded and practiced in random order.&lt;/li&gt;
&lt;li&gt;A second follow-up when the first answer is still thin.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then mine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Heavy mode with Gemma 4 E2B, and a deeper review on the device: code review and possible gaps, as suggestions, never scores.&lt;/li&gt;
&lt;li&gt;A small embedding model that, after you mark a point, highlights the sentence in your answer that comes closest, so you can check your own marks.&lt;/li&gt;
&lt;li&gt;Progress across sessions: pace, fillers and answer length over time, kept on your laptop.&lt;/li&gt;
&lt;li&gt;Hosting the voice detector and ONNX runtime files on the site, so it doesn't depend on a CDN.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>ai</category>
    </item>
    <item>
      <title>3D AI Voice Avatar That Runs Entirely in Your Browser — No Servers, No API Keys, No GPU Cloud</title>
      <dc:creator>927tanmay</dc:creator>
      <pubDate>Tue, 11 Aug 2026 17:05:53 +0000</pubDate>
      <link>https://dev.to/927tanmay/i-built-a-3d-ai-voice-avatar-that-runs-entirely-in-your-browser-no-servers-no-api-keys-no-gpu-i4p</link>
      <guid>https://dev.to/927tanmay/i-built-a-3d-ai-voice-avatar-that-runs-entirely-in-your-browser-no-servers-no-api-keys-no-gpu-i4p</guid>
      <description>&lt;p&gt;&lt;em&gt;Speech recognition, voice synthesis, and lip-synced 3D animation — running inside Web Workers on your desktop browser. Here's how.&lt;/em&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; I open-sourced a React component that drops a fully conversational, lip-syncing 3D avatar into any web app. Built primarily as a drop-in 3D frontend for your existing cloud LLMs (OpenAI, Claude, custom backends), it handles speech recognition (Whisper), voice synthesis (Kokoro TTS), and real-time ARKit facial blendshape animation in-browser via WebAssembly and WebGPU. It can also run completely offline with on-device models. One &lt;code&gt;npm install&lt;/code&gt;, one component, done.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;🔗 &lt;strong&gt;&lt;a href="https://react-ai-voice-avatar.vercel.app/" rel="noopener noreferrer"&gt;Live Demo&lt;/a&gt;&lt;/strong&gt; · 📦 &lt;strong&gt;&lt;a href="https://www.npmjs.com/package/react-ai-voice-avatar" rel="noopener noreferrer"&gt;NPM Package&lt;/a&gt;&lt;/strong&gt; · 🐙 &lt;strong&gt;&lt;a href="https://github.com/927tanmay/react-ai-voice-avatar" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem That Kept Bugging Me
&lt;/h2&gt;

&lt;p&gt;Every time I explored building a conversational AI interface — the kind where a character actually &lt;em&gt;talks back to you&lt;/em&gt; — I hit the same wall:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cloud TTS APIs&lt;/strong&gt; charge per character and add 200–500ms of round-trip latency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WebSocket video streaming&lt;/strong&gt; from GPU servers is fragile, expensive, and adds heavy infrastructure overhead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Existing avatar libraries&lt;/strong&gt; just give you a static 3D model. You still have to wire up speech, lip-sync, turn-taking, and microphone handling yourself (which takes months of integration).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I wanted something different: a single React component where I write &lt;code&gt;&amp;lt;AiVoiceAvatar/&amp;gt;&lt;/code&gt; and it just... works. The avatar listens, thinks, speaks with a natural voice, and moves its mouth in perfect sync — serving as the perfect visual layer for your AI backend.&lt;/p&gt;

&lt;p&gt;So, I built it.&lt;/p&gt;




&lt;h2&gt;
  
  
  What It Actually Does
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;react-ai-voice-avatar&lt;/code&gt;&lt;/strong&gt; is a React + React Three Fiber component that orchestrates an entire voice conversation pipeline inside the browser:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🎤 Microphone → Whisper ASR → LLM Reasoning → Kokoro TTS → 3D Lip-Sync → 🔊 Speaker&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flpvb8f6hxbpbsegwg2sw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flpvb8f6hxbpbsegwg2sw.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every heavy ML stage runs inside dedicated &lt;strong&gt;Web Workers&lt;/strong&gt;. This ensures the main thread stays buttery smooth at 60 FPS while the neural networks execute quietly in the background.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Two-Brain Architecture
&lt;/h3&gt;

&lt;p&gt;While the package supports fully offline execution, it was built first and foremost to plug into your existing cloud infrastructure:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Brain Mode&lt;/th&gt;
&lt;th&gt;How It Works&lt;/th&gt;
&lt;th&gt;Best Used For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;🧠 Connected Brain (Primary)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Route transcribed speech to your existing backend (OpenAI, Claude, custom FastAPI, Vercel AI SDK, etc.) using the &lt;code&gt;onSubmit&lt;/code&gt; prop.&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Production web apps, SaaS products, and enterprise AI assistants.&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;🔒 On-Device Brain (Offline)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A 0.5B parameter LLM (Qwen 2.5) runs via WebGPU locally. No API keys or network calls required.&lt;/td&gt;
&lt;td&gt;Privacy-sensitive web apps, kiosks, or demo environments with spotty WiFi.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the &lt;strong&gt;Connected Brain&lt;/strong&gt;, the avatar handles all the complex frontend tasks: listening, transcribing, audio playing, and lip animation. Your backend just streams the text back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Connected Brain: Your backend handles the thinking&lt;/span&gt;
&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;AiVoiceAvatar&lt;/span&gt;
  &lt;span class="na"&gt;avatarPreset&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"ananya"&lt;/span&gt;
  &lt;span class="na"&gt;ttsVoice&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"af_heart"&lt;/span&gt;
  &lt;span class="na"&gt;onSubmit&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userSpeech&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/chat&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userSpeech&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// Streams natively!&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. The 3D character loads over CDN, the TTS and ASR models download once (cached in IndexedDB permanently), and your app gets an interactive 3D character connected directly to your existing AI API.&lt;/p&gt;




&lt;h2&gt;
  
  
  Under the Hood: How It Works
&lt;/h2&gt;

&lt;p&gt;I spent weeks getting this architecture right. Here are the most interesting technical hurdles I had to cross.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Web Worker Isolation
&lt;/h3&gt;

&lt;p&gt;Running Whisper or Kokoro on the main thread would freeze the web app instantly. To fix this, every model runs in isolation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ML Pipeline Worker:&lt;/strong&gt; Handles Whisper ASR transcription and (optionally) the local Qwen LLM. Streams LLM tokens back to the main thread as they're generated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kokoro TTS Worker:&lt;/strong&gt; Converts text chunks into 24kHz Float32 audio. Runs the Kokoro-82M ONNX model using multi-threaded WASM or WebGPU acceleration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These workers are pre-bundled using esbuild at build time and stringified into the package. &lt;strong&gt;You don't have to host worker files or configure Webpack.&lt;/strong&gt; It works instantly with Vite, Next.js, or Create React App.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Real-Time Lip Synchronization (80 FPS)
&lt;/h3&gt;

&lt;p&gt;This was the hardest part. The avatar needs to move its mouth naturally and in perfect sync with the audio stream. I built a &lt;strong&gt;dual-source blending&lt;/strong&gt; approach:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Phoneme Timing Engine:&lt;/strong&gt; We extract phoneme-level timing data from Kokoro and map it to 15 standard viseme shapes. This gives us &lt;em&gt;predictive&lt;/em&gt; mouth shapes that lead the audio slightly, mimicking real human speech.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audio Amplitude Fallback:&lt;/strong&gt; A Web Audio API &lt;code&gt;AnalyserNode&lt;/code&gt; reads real-time frequency data, providing a secondary signal blended for amplitude-driven jaw movement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Procedural Facial Dynamics:&lt;/strong&gt; The avatar has continuous idle micro-animations (randomized Poisson interval blinks, breathing, head drift) so it feels alive even when silent.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of this targets the &lt;strong&gt;52 standard Apple ARKit blendshapes&lt;/strong&gt;, meaning any humanoid &lt;code&gt;.glb&lt;/code&gt; model rigged with these morph targets works out of the box.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Safari &amp;amp; WASM Limits: The Invisible OOM Safety Net
&lt;/h3&gt;

&lt;p&gt;Safari and restricted browser runtimes impose strict WebAssembly memory limits. Running heavier TTS models (~90MB) under tight WASM memory constraints can occasionally trigger &lt;code&gt;Out of memory&lt;/code&gt; errors.&lt;/p&gt;

&lt;p&gt;Instead of letting the component crash or fail silently, I built an &lt;strong&gt;automatic failover system&lt;/strong&gt;. If Kokoro initialization encounters memory restrictions, the engine transparently switches to a lightweight MMS TTS model (~30MB). The voice quality drops slightly, but the avatar keeps talking without breaking the user session. No crashes, no frozen UI, and zero developer intervention required.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Hard Interruption
&lt;/h3&gt;

&lt;p&gt;Real conversations are messy. If a user taps "Stop" mid-sentence, everything must halt instantly. I implemented a coordinated interrupt system across both Web Workers that clears all internal queues, aborts in-flight LLM generation via a sentinel error, and flushes the audio context.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Developer Experience
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Zero Configuration, Genuinely
&lt;/h3&gt;

&lt;p&gt;I'm allergic to "zero config" tools that require 14 setup steps. Here is the actual install process:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install &lt;/span&gt;react-ai-voice-avatar three @react-three/fiber @react-three/drei

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;AiVoiceAvatar&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react-ai-voice-avatar&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// Inside your R3F Canvas:&lt;/span&gt;
&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;AiVoiceAvatar&lt;/span&gt; &lt;span class="na"&gt;avatarPreset&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"ananya"&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No Vite &lt;code&gt;optimizeDeps&lt;/code&gt; overrides. No worker file hosting. The NPM footprint is only ~3.3 MB (including pre-bundled workers and viseme mapping). The heavy 3D models and neural network weights load on-demand over the network and cache in the browser.&lt;/p&gt;

&lt;h3&gt;
  
  
  Imperative Control
&lt;/h3&gt;

&lt;p&gt;For times when you need programmatic control, the component exposes a clean Ref API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;avatarRef&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;useRef&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;AiVoiceAvatarHandle&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// Make the avatar speak programmatically&lt;/span&gt;
&lt;span class="nx"&gt;avatarRef&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;speak&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Hello! How can I help you?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// Submit text as if the user spoke it&lt;/span&gt;
&lt;span class="nx"&gt;avatarRef&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;sendText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;What's the weather like?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// Hard-interrupt mid-speech&lt;/span&gt;
&lt;span class="nx"&gt;avatarRef&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;current&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;interrupt&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  What Surprised Me Building This
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Numbers break TTS models:&lt;/strong&gt; Kokoro chokes on symbols like &lt;code&gt;%&lt;/code&gt; or &lt;code&gt;$&lt;/code&gt;. I had to build a &lt;code&gt;sanitizeForSpeech&lt;/code&gt; preprocessor that translates "95% complete" to "ninety-five percent complete" before it hits the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Web Worker Bundling is a Nightmare:&lt;/strong&gt; Standard worker approaches break in library distribution. My solution (pre-compiling with esbuild and reconstructing as a Blob URL at runtime) isn't elegant, but it is universally compatible across all modern web bundlers.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The core engine is stable and production-ready for web interfaces. Next on the roadmap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hindi &amp;amp; Indic Language Voices:&lt;/strong&gt; The phoneme engine supports retroflex and aspirated consonants, and the &lt;code&gt;visemeTable&lt;/code&gt; has Devanagari mappings. We just need to train the TTS voices!&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ready Player Me Integration:&lt;/strong&gt; Official support for RPM avatars.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversation Memory:&lt;/strong&gt; APIs like &lt;code&gt;addContext()&lt;/code&gt; for injecting dynamic knowledge mid-conversation.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Try It Out
&lt;/h2&gt;

&lt;p&gt;If you want to see the 80FPS lip-sync and streaming TTS integration in action, check out the demo:&lt;/p&gt;

&lt;h3&gt;
  
  
  🌐 &lt;a href="https://react-ai-voice-avatar.vercel.app/" rel="noopener noreferrer"&gt;Live Demo — react-ai-voice-avatar.vercel.app&lt;/a&gt;
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://github.com/927tanmay/react-ai-voice-avatar" rel="noopener noreferrer"&gt;GitHub repo&lt;/a&gt; includes four complete example apps, ranging from a 30-line quickstart to a full hybrid-cloud OpenAI streaming integration.&lt;/p&gt;

&lt;p&gt;If you build something with this, I'd genuinely love to see it. Open an issue, tag me, or drop a comment below,  or &lt;strong&gt;&lt;a href="https://www.linkedin.com/in/927tanmay/" rel="noopener noreferrer"&gt;connect with me on LinkedIn&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;




</description>
      <category>ai</category>
      <category>webgpu</category>
      <category>webdev</category>
      <category>react</category>
    </item>
  </channel>
</rss>
