<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Omkar tripathi</title>
    <description>The latest articles on DEV Community by Omkar tripathi (@code_triggered_).</description>
    <link>https://dev.to/code_triggered_</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F785393%2F2b3d5d0a-549b-4d07-9586-e10252136fa4.png</url>
      <title>DEV Community: Omkar tripathi</title>
      <link>https://dev.to/code_triggered_</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/code_triggered_"/>
    <language>en</language>
    <item>
      <title>Vent 🙄- I built my friend someone to call when he just needs to vent (and it talks back in my voice)</title>
      <dc:creator>Omkar tripathi</dc:creator>
      <pubDate>Sun, 04 Oct 2026 21:34:54 +0000</pubDate>
      <link>https://dev.to/code_triggered_/vent-a-friend-that-you-call-2ilj</link>
      <guid>https://dev.to/code_triggered_/vent-a-friend-that-you-call-2ilj</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;I have a friend who has bad days, like all of us. The problem is not the bad day. The problem is when he tells someone about it, everyone becomes a life coach.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"You should talk to your manager."&lt;/p&gt;

&lt;p&gt;"Have you tried journaling?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Bro, he just wants to say out loud that some guy in a white SUV cut him off and then flipped &lt;em&gt;him&lt;/em&gt; off. That's it. He wants someone to go &lt;em&gt;"wait, what??"&lt;/em&gt; and let him finish.&lt;/p&gt;

&lt;p&gt;So I built him &lt;strong&gt;VENT&lt;/strong&gt;. It's a friend you can call.&lt;/p&gt;

&lt;p&gt;You open it, press the phone button, it rings, and it picks up. Then you just talk. VENT listens and reacts like a friend on a phone call would:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;You:&lt;/strong&gt; my manager blamed me for the delay and it wasn't even my part&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;VENT:&lt;/strong&gt; He blamed &lt;em&gt;YOU&lt;/em&gt;? For the backend's mess?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It doesn't give advice. If you're in the middle of a sentence, it waits.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two ways to talk it out
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;🌈 &lt;strong&gt;Breathe&lt;/strong&gt; is the default. A small cloud called &lt;strong&gt;Puff&lt;/strong&gt; breathes with you, opens its eyes when you talk, and VENT stays calm.&lt;/li&gt;
&lt;li&gt;🥊 &lt;strong&gt;Punch&lt;/strong&gt; is &lt;em&gt;devil mode&lt;/em&gt;, for the days you're actually angry. A little devil boxer punches a heavy bag every time &lt;strong&gt;your voice&lt;/strong&gt; gets loud. The bag tears a bit more with every hit and bursts on the 20th, a bell rings, and a new bag comes down. VENT gets angry &lt;em&gt;with&lt;/em&gt; you here ("Nope. NOT okay.").&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy4nqysd1giees2m76r12.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy4nqysd1giees2m76r12.gif" alt="Breathe: Puff listens and breathes with you" width="300" height="608"&gt;&lt;/a&gt; &lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzlpop3wyy6r8c7nfsbx7.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzlpop3wyy6r8c7nfsbx7.gif" alt="Punch: your loud words throw the punches" width="300" height="608"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you hang up, VENT asks &lt;strong&gt;&lt;em&gt;"Feel a little lighter?"&lt;/em&gt;&lt;/strong&gt; You get two choices:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Let it go.&lt;/strong&gt; The whole call dissolves and &lt;em&gt;nothing is saved&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep it as a page.&lt;/strong&gt; It goes into your journal.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa6mwyxg6f56yei3ond5p.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa6mwyxg6f56yei3ond5p.gif" alt="Let it go: the call dissolves and nothing is kept" width="360" height="729"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A journal you don't have to write
&lt;/h3&gt;

&lt;p&gt;The second part is for him also, but on normal days. He never keeps a journal, because writing feels like homework. Talking about your day is easy though.&lt;/p&gt;

&lt;p&gt;So there's a &lt;strong&gt;Make a journal page&lt;/strong&gt; call, where VENT is curious instead of quiet ("Ooh, where did you two go?"). When you hang up, Gemma turns the call into a proper page for the day:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;how the mood moved through the day&lt;/li&gt;
&lt;li&gt;the people, the food, the places&lt;/li&gt;
&lt;li&gt;the good moments and the hard ones&lt;/li&gt;
&lt;li&gt;things to remember for tomorrow&lt;/li&gt;
&lt;li&gt;the day told back &lt;em&gt;in your own words&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Later you can ask your own life questions, like &lt;em&gt;"When did I last talk about Greg?"&lt;/em&gt; It answers &lt;strong&gt;only from your pages&lt;/strong&gt; and shows which days it used. If something isn't in there, it says so instead of making it up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fihffkedwm9nr876aq3ji.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fihffkedwm9nr876aq3ji.gif" alt="Ask your life: answers only from your pages, with the days used" width="360" height="729"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/lmmsIo0QAdU" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;One small confession about the video.&lt;/strong&gt; My screen recorder only recorded my mic, not VENT's voice. So the captions at the bottom show what I said (transcribed with &lt;strong&gt;ElevenLabs Scribe&lt;/strong&gt;) and what VENT said.&lt;/p&gt;

&lt;p&gt;For the Breathe and journal calls these are the &lt;em&gt;real&lt;/em&gt; replies. I got them back from &lt;strong&gt;Temporal&lt;/strong&gt;, because every page VENT writes goes through a workflow that stores the full call. The Punch call I let go, so by design nothing was saved. For that one the replies are regenerated by the same model from my words, and the video says so.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/triggeredcode" rel="noopener noreferrer"&gt;
        triggeredcode
      &lt;/a&gt; / &lt;a href="https://github.com/triggeredcode/vent" rel="noopener noreferrer"&gt;
        vent
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A friend you can call when you just need to talk — open-weight Gemma listens, journals your days, and remembers. Runs locally.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;VENT&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;A friend you can call when you just need to talk.&lt;/strong&gt; VENT listens without trying to fix you, turns the days you talk through into a visual journal, and lets you ask about your own life later.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Call → talk it out → see your day → ask your life&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Every piece of intelligence in VENT is an open-weight model running on your own machine: these are the most private conversations a person has, so they shouldn't have to leave it.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What it does&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Vent&lt;/strong&gt; — a real phone call. VENT reacts like a friend would ("Wait, he blamed &lt;em&gt;you&lt;/em&gt;?"), nudges you on, stays quiet when you're mid-thought, and never gives advice. When you hang up, nothing is kept unless you ask.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Journal&lt;/strong&gt; — tell VENT about your day. When the call ends, Gemma writes a magazine-style page: mood arc, people, food, places, highlights, hard moments, things to…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/triggeredcode/vent" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Running it takes two commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pnpm &lt;span class="nb"&gt;install
&lt;/span&gt;pnpm vent            &lt;span class="c"&gt;# everything starts, the browser opens&lt;/span&gt;
pnpm voice:record    &lt;span class="c"&gt;# optional: make VENT talk in your own voice&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pnpm vent&lt;/code&gt; starts Ollama and pulls the Gemma models, starts the voice server, starts Temporal and the app, puts a few sample days in the journal and opens the browser. &lt;code&gt;pnpm connect&lt;/code&gt; sets up Sentry and ElevenLabs if you want them.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1qkcdi8g5flwcfh69xks.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1qkcdi8g5flwcfh69xks.png" alt="How a VENT call works" width="800" height="469"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The whole idea dies if it doesn't feel like a call. If you're upset and the thing on the other side sounds like a GPS lady, you hang up.&lt;/p&gt;

&lt;p&gt;So most of my weekend actually went into one question: &lt;strong&gt;how do I make it sound like a person?&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Version 1: the robot
&lt;/h3&gt;

&lt;p&gt;The first version worked on paper: mic, speech-to-text, a small model, and the browser's built-in text-to-speech. It was terrible. The voice was flat, and it said "I'm here" after almost everything I said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Me:&lt;/strong&gt; &lt;em&gt;(something really specific about my day)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;VENT:&lt;/strong&gt; I'm here. Take your time.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;My friend would last about ten seconds. Two things were wrong, and only one of them was the voice.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fixing &lt;em&gt;what&lt;/em&gt; it says
&lt;/h3&gt;

&lt;p&gt;I was using &lt;strong&gt;Gemma 3 1B&lt;/strong&gt; for replies because it was fast, about 80 ms. But it kept returning only this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"acknowledge"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There was no text at all, so the app fell back to "I'm here" every time. &lt;em&gt;A fast model saying nothing is still saying nothing.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;So I split the job between two Gemma 4 models in Ollama:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Warm time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hear you (audio → text)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Gemma 4 E2B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~150–400 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reply, with the &lt;em&gt;whole call&lt;/em&gt; as context&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Gemma 4 E4B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~400–550 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I ran E2B and E4B as the listener on the same scripted rants to compare:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;You said&lt;/th&gt;
&lt;th&gt;Gemma 4 E2B&lt;/th&gt;
&lt;th&gt;Gemma 4 E4B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;em&gt;"I literally labelled it"&lt;/em&gt; (roommate ate my biryani)&lt;/td&gt;
&lt;td&gt;"You labelled it out clearly, right?"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;"Labelled it, and still? Unbelievable."&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;em&gt;"whatever. I ordered pizza instead"&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;"Pizza sounds way better right now."&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;"Pizza fix. Good call."&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;E4B costs about 200 ms more, and it's completely worth it. The reply now comes back as a small action plus the words:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"follow_up"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"tone"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"fired_up"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Wait, he blamed YOU?"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The prompt has a few strict rules:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- Never generic: never say "That sounds rough", "I'm here", "I hear you", "Take your time".
- No advice, no tips, no "you should", no therapy talk.
- Choose silence ONLY when their line is obviously cut off mid-sentence.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One funny thing: when my test lines were too close to the examples in the prompt, the model just repeated my examples back to me 😅. I changed the examples to random situations (a landlord raising rent, a friend skipping your birthday), and it started reacting properly.&lt;/p&gt;

&lt;p&gt;I also pre-load the listener while the phone is &lt;strong&gt;ringing&lt;/strong&gt;, so the first reply isn't slower than the rest. One full turn, from you stopping to VENT starting to talk, is &lt;strong&gt;about a second&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fixing &lt;em&gt;how&lt;/em&gt; it sounds
&lt;/h3&gt;

&lt;p&gt;This is the part I'm happiest about. The voice went through three versions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Browser text-to-speech.&lt;/strong&gt; Flat and robotic. Nobody wants to vent to that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kokoro-82M&lt;/strong&gt;, an open-weight TTS running locally. Night and day: about &lt;strong&gt;100 ms&lt;/strong&gt; per line and it sounds like a real person. But it's &lt;em&gt;a&lt;/em&gt; person, some voice from a model card.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My own voice.&lt;/strong&gt; He's calling a friend, so why not make it sound like one?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I talk to my laptop a lot. I have my own little dictation app, so I had a folder full of recordings of my own voice. I took &lt;strong&gt;16 seconds&lt;/strong&gt; of me talking normally, cleaned it up with &lt;code&gt;ffmpeg&lt;/code&gt;, and cloned it locally with &lt;strong&gt;Chatterbox-Turbo&lt;/strong&gt; (MIT licence), running on my Mac through &lt;code&gt;mlx-audio&lt;/code&gt;. It takes around half a second per line.&lt;/p&gt;

&lt;p&gt;Now when he calls VENT, the voice that answers is &lt;strong&gt;mine&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Honestly, the first time I heard my own voice ask &lt;em&gt;"Wait, he said that to you?"&lt;/em&gt; it was a little weird. And then it was kind of perfect.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You can do the same with your own voice, it's one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pnpm voice:record
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It shows a short script, records you for 45 seconds, cleans up the take (trims silence, removes rumble, evens out the loudness), and restarts the voice server on your cloned voice. The recording stays on your machine and is never committed. If you skip it, VENT just uses Kokoro.&lt;/p&gt;

&lt;p&gt;Same voice, but the &lt;em&gt;mood&lt;/em&gt; should change with the mode. In Punch, VENT should sound fired up. Chatterbox-Turbo just ignores emotion controls (I checked its source code), so I shape the audio after it's generated:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Voice style&lt;/th&gt;
&lt;th&gt;What changes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Breathe&lt;/td&gt;
&lt;td&gt;&lt;code&gt;calm&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;nothing, as it is&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Punch&lt;/td&gt;
&lt;td&gt;&lt;code&gt;fired&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;~14% faster, small pitch lift, more presence, compression&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Journal&lt;/td&gt;
&lt;td&gt;&lt;code&gt;bright&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a bit livelier and warmer, more curious&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The words change too. In Punch the model is told to be angry &lt;em&gt;on your side&lt;/em&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;VENT (Punch):&lt;/strong&gt; FLIPPED you off?! ARE YOU SERIOUS?!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With my &lt;strong&gt;ElevenLabs&lt;/strong&gt; creator account, I also cloned my voice there (&lt;em&gt;Instant Voice Clone&lt;/em&gt;), so VENT can talk through &lt;strong&gt;ElevenLabs Flash v2.5&lt;/strong&gt; instead if you don't have a Mac. The punch sounds were generated with the &lt;strong&gt;ElevenLabs Sound Effects API&lt;/strong&gt;: heavy-bag thumps, the uppercut, the bag tearing, the bell, the chain.&lt;/p&gt;

&lt;h3&gt;
  
  
  Not talking over you
&lt;/h3&gt;

&lt;p&gt;The most annoying thing in early tests was VENT interrupting &lt;em&gt;itself&lt;/em&gt;. On laptop speakers its voice went back into the mic, it thought I was talking, and it stopped in the middle of its own sentence. Two fixes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the mic is &lt;strong&gt;deaf while VENT speaks&lt;/strong&gt;, plus a short echo tail&lt;/li&gt;
&lt;li&gt;the server &lt;strong&gt;ignores any "turn"&lt;/strong&gt; that's just VENT's own words coming back&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want to cut it off, you tap the character. Same as putting your hand up ✋.&lt;/p&gt;

&lt;h3&gt;
  
  
  The journal page, and why Temporal
&lt;/h3&gt;

&lt;p&gt;After a journal call, &lt;strong&gt;Gemma 4 E4B&lt;/strong&gt; writes the page against a JSON schema (Ollama structured output). Only what the caller said counts as fact, and VENT's lines are just context. Then the page is embedded locally with &lt;code&gt;nomic-embed-text&lt;/code&gt; and saved.&lt;/p&gt;

&lt;p&gt;On a laptop this part can fail. The model might still be loading, or Ollama might be busy. And I really didn't want my friend to talk for five minutes and &lt;em&gt;lose the page&lt;/em&gt;. So it runs as a &lt;strong&gt;Temporal&lt;/strong&gt; workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;writeJournalPage:  extractDay  →  embedPage  →  savePage
                   (each step with its own timeout + retries)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I tested it by killing the worker with &lt;code&gt;kill -9&lt;/code&gt; in the middle of writing. When the worker came back, Temporal just picked it up again and the page arrived. And the workflow input is the full call, which is how I got VENT's real lines back for the demo captions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb23y8nx79mrigk6a280j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb23y8nx79mrigk6a280j.png" alt="The writeJournalPage workflow in the Temporal UI: the full call as input, the page as result, and extractDay → embedPage → savePage on the timeline" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Remembering, and watching it &lt;em&gt;without reading it&lt;/em&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Memories&lt;/strong&gt; gives Gemma your pages, never the raw transcripts, and it has to tell me which pages it used. If an id doesn't exist I throw it away, so it can't invent a day.&lt;/li&gt;
&lt;li&gt;Pages live in a &lt;strong&gt;local file&lt;/strong&gt;. If you want sync, they go to &lt;strong&gt;MongoDB Atlas&lt;/strong&gt;, with &lt;strong&gt;Atlas Vector Search&lt;/strong&gt; over the same local embeddings (tested against Atlas Local).&lt;/li&gt;
&lt;li&gt;Every call turn is traced in &lt;strong&gt;Sentry&lt;/strong&gt;: hear → reply → speak → write the page, with timings and token counts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are probably the most private conversations someone has, so I switched off &lt;em&gt;everything&lt;/em&gt; in Sentry that collects content. The new SDK collects request bodies by default, and that would have quietly sent journal text out. Now it only knows &lt;strong&gt;how long&lt;/strong&gt; things took, never &lt;strong&gt;what&lt;/strong&gt; was said.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Because of &lt;em&gt;what people say to VENT&lt;/em&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Nobody tells a cloud API about their worst day if they think about it for even a second.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With open models, the listening happens on your laptop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemma&lt;/strong&gt; hears you, replies, writes the page and answers your questions, all through Ollama, all local.&lt;/li&gt;
&lt;li&gt;Your &lt;strong&gt;voice and journal never leave the machine&lt;/strong&gt; unless you turn something on.&lt;/li&gt;
&lt;li&gt;I could &lt;strong&gt;compare two Gemma sizes&lt;/strong&gt; on my own test calls and pick E2B for hearing and E4B for thinking.&lt;/li&gt;
&lt;li&gt;I could &lt;strong&gt;clone my own voice&lt;/strong&gt; with an open model, with no account and no per-minute bill.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A closed API would charge for every &lt;em&gt;"wait, seriously?"&lt;/em&gt;, and a thing you call at 11 pm on a bad day shouldn't be counting minutes. The hosted parts (ElevenLabs, Sentry, Atlas) are all &lt;strong&gt;opt-in&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemma:&lt;/strong&gt; Gemma 4 E2B hears every turn. Gemma 4 E4B replies on the call (and reads the mood that changes the scene), writes the journal page with a JSON schema, and answers memory questions. All local through Ollama.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ElevenLabs:&lt;/strong&gt; my voice as an &lt;em&gt;Instant Voice Clone&lt;/em&gt; with Flash v2.5, the Punch sound effects from the Sound Effects API, and Scribe for the demo captions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sentry:&lt;/strong&gt; &lt;code&gt;gen_ai&lt;/code&gt; traces of every call turn and journal write, with all content collection turned off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Temporal:&lt;/strong&gt; every journal page is written by a durable workflow with retries. It survived me killing the worker mid-write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MongoDB Atlas:&lt;/strong&gt; optional synced journal with Atlas Vector Search for Memories.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>Kizuna 絆 the Gemini-web to Local Environment Bridge</title>
      <dc:creator>Omkar tripathi</dc:creator>
      <pubDate>Wed, 04 Mar 2026 11:05:10 +0000</pubDate>
      <link>https://dev.to/code_triggered_/kizuna-5hc8</link>
      <guid>https://dev.to/code_triggered_/kizuna-5hc8</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/mlh-built-with-google-gemini-02-25-26"&gt;Built with Google Gemini: Writing Challenge&lt;/a&gt;&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;╔══════════════════════════════════════════════════════════════╗
║ ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ║
║                                                              ║
║   ██╗  ██╗ ██╗ ███████╗ ██╗   ██╗ ███╗   ██╗  █████╗         ║
║   ██║ ██╔╝ ██║ ╚══███╔╝ ██║   ██║ ████╗  ██║ ██╔══██╗        ║
║   █████╔╝  ██║   ███╔╝  ██║   ██║ ██╔██╗ ██║ ███████║        ║
║   ██╔═██╗  ██║  ███╔╝   ██║   ██║ ██║╚██╗██║ ██╔══██║        ║
║   ██║  ██╗ ██║ ███████╗ ╚██████╔╝ ██║ ╚████║ ██║  ██║        ║
║   ╚═╝  ╚═╝ ╚═╝ ╚══════╝  ╚═════╝  ╚═╝  ╚═══╝ ╚═╝  ╚═╝        ║
║                                                              ║
║          [ The Gemini to Local Environment Bridge ]          ║
║ ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ║
╚══════════════════════════════════════════════════════════════╝
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Integrating AI into our daily coding workflows is a recurring discussion. The discourse focuses heavily on context windows, reasoning models, and whether AI will replace or augment engineers. But I think centering this debate purely on the models themselves is reductive.&lt;/p&gt;

&lt;p&gt;The bigger question to me is the environment. How do we actually connect these floating, cloud-based brains to our physical work?&lt;/p&gt;

&lt;p&gt;For a long time, my workflow was &lt;em&gt;agonizing&lt;/em&gt;. I was stuck in what I call "&lt;strong&gt;&lt;em&gt;Copy-Paste Torture.&lt;/em&gt;&lt;/strong&gt;" I would give the AI context, copy a file from my IDE, paste it into Google Gemini, ask for a change, copy the resulting code, paste it back, run it, hit an error, and start over.&lt;/p&gt;

&lt;p&gt;Gemini was incredibly capable, but the friction of constant context-switching was killing the momentum. A brain in a browser jar has no hands.&lt;/p&gt;

&lt;p&gt;So, I decided to build it a nervous system. I called it &lt;strong&gt;Kizuna&lt;/strong&gt; (絆 - meaning &lt;em&gt;bond&lt;/em&gt; or &lt;em&gt;connection&lt;/em&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built with Google Gemini
&lt;/h2&gt;

&lt;p&gt;Kizuna is an end-to-end toolchain that transforms the standard Google Gemini web interface into a localized, agentic IDE companion. I didn't want to just build a wrapper; I wanted to create a system that felt intentional, granting the web-based LLM the ability to read, search, and safely patch a local codebase without compromising my machine.&lt;/p&gt;

&lt;p&gt;To achieve this, I broke the system down into three foundational pillars:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The Engine (Local Daemon) ⚙️
&lt;/h3&gt;

&lt;p&gt;Code is where software lives, but you can't just give an AI raw shell access—that's a massive security risk. I built a local backend service to act as a secure sandbox. It translates strict JSON intents from the browser into optimized local file reads, writes, and Git operations.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;╔══════════════════════════════════════════════════════════════╗
║ ░░░░░░░░░░░░░░ THE PATH JAIL (SANDBOX) ░░░░░░░░░░░░░░░░░░░░░ ║
╠══════════════════════════════════════════════════════════════╣
║                                                              ║
║  ╭── [ ALLOWED WORKSPACE ] ──────────────────────────────╮   ║
║  │                                                       │   ║
║  │   📂 /workspace/my-app/     (Sandbox Root)            │   ║
║  │    ├── 📄 src/main.js             [ ✓ ] OK            │   ║
║  │    └── 📄 package.json            [ ✓ ] OK            │   ║
║  │                                                       │   ║
║  ╰───────────────────────────────────────────────────────╯   ║
║                                                              ║
║  ╭── [ BLOCKED EXTERNALS ] ──────────────────────────────╮   ║
║  │                                                       │   ║
║  │   🚨 /etc/passwd                  [ ⨉ ] BLOCKED: 403  │   ║
║  │                                                       │   ║
║  ╰───────────────────────────────────────────────────────╯   ║
║                                                              ║
║  ▪ Engine drops all path traversal requests (../)            ║
║  ▪ Symlinks are resolved prior to boundary validation        ║
╚══════════════════════════════════════════════════════════════╝
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The Engine resolves all symlinks before executing. If an AI hallucinates a path traversal, the system structurally drops the request.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The Bridge (Chrome Extension)
&lt;/h3&gt;

&lt;p&gt;A web extension that sits quietly on the right side of the Gemini window. Building this was an architectural challenge. Because Gemini streams text, scraping the DOM naively crashes the parser with incomplete JSON. I had to build a &lt;code&gt;MutationObserver&lt;/code&gt; that waits for the absolute "completion" state of the chat UI before parsing. It captures the AI's outputs, relays them to my local engine via a Background Worker (to bypass strict browser CORS restrictions), and injects the results back into the chat.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The Protocol (Documentation &amp;amp; Rules)
&lt;/h3&gt;

&lt;p&gt;A deterministic system prompt fed to Gemini at the start of every chat. This acts as the "Operating Manual." LLMs are chaotic; they need boundaries. The protocol forces Gemini to use a strict JSON schema instead of conversational markdown.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Here is a look at the complete, fully-boxed architectural flow of Kizuna:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;╔══════════════════════════════════════════════════════════════╗
║ ▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒ BROWSER ENVIRONMENT ▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒ ║
║                                                              ║
║  ┌────────────────┐   [ DOM ]     ┌───────────────────────┐  ║
║  │  🧠 Gemini UI  │ ◀───────────▶ │  🧩 Chrome Extension   │  ║
║  └────────────────┘               └───────────────────────┘  ║
║          ▲                                    │              ║
║          │ (Injects UI Data)                  │ (WebSockets) ║
╠══════════╪════════════════════════════════════╪══════════════╣
║          │                                    ▼              ║
║  ┌───────┴────────────────────────────────────┴───────────┐  ║
║  │  💻 KIZUNA ENGINE (Local Daemon :8080)                 │  ║
║  │                                                        │  ║
║  │  ╭─────────────────╮          ╭─────────────────────╮  │  ║
║  │  │ 🛡️ Path Sandbox │ ───────▶ │ 🛠️ Tool Dispatcher   │  │  ║
║  │  ╰─────────────────╯          ╰───┬──────┬──────┬───╯  │  ║
║  │                                   │      │      │      │  ║
║  │                                [Read] [Write] [Git]    │  ║
║  │                                   │      │      │      │  ║
║  │                                 ╭─┴──────┴──────┴─╮    │  ║
║  │                                 │ 📂 Local Storage│    │  ║
║  │                                 ╰─────────────────╯    │  ║
║  └────────────────────────────────────────────────────────┘  ║
║                                                              ║
║ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ LOCAL ENVIRONMENT ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ ║
╚══════════════════════════════════════════════════════════════╝
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When I ask Gemini to "Update the database connection," it doesn't write me a markdown tutorial. It outputs its intent directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"patch_file"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"src/config.js"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"search"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"localhost"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"replace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"process.env.DB_HOST"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The extension picks this up, the engine verifies the path hasn't escaped the workspace, the file is safely patched, and the success message is injected directly back into the Gemini chat or I pasted it manually sometimes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;Through this work, I learned to question the constraints of LLMs first, not treat them as assumptions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Abstract Syntax Trees (ASTs) vs. Raw Text
&lt;/h3&gt;

&lt;p&gt;The naive approach to the context window is to just send the AI the whole file. But why send a 1,000-line file when the AI just needs to know what the file &lt;em&gt;does&lt;/em&gt;? Sending raw text is a massive waste of tokens and scatters the model's focus.&lt;/p&gt;

&lt;p&gt;I built a tool into the engine that parses code into an Abstract Syntax Tree (AST). Instead of returning raw code, it strips the implementation logic and returns a "Skeleton" of class names, imports, and function signatures.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;╔══════════════════════════════════════════════════════════════╗
║ ░░░░░░░░░░░░░░░ AST PARSER ROUTING LOGIC ░░░░░░░░░░░░░░░░░░░ ║
╠══════════════════════════════════════════════════════════════╣
║                 [ 📄 Gemini Requests File ]                  ║
║                              │                               ║
║                              ▼                               ║
║                  ╭───────────────────────╮                   ║
║                  │  Is File &amp;gt; 300 Lines? │                   ║
║                  ╰───────────┬───────────╯                   ║
║                              │                               ║
║            ┌───( YES )───────┴───────( NO )────┐             ║
║            │                                   │             ║
║            ▼                                   ▼             ║
║    ╭───────────────╮                   ╭───────────────╮     ║
║    │  🌳 AST Parse │                   │  📝 Raw Read   │    ║
║    ╰───────┬───────╯                   ╰───────┬───────╯     ║
║            │                                   │             ║
║    ╭───────▼───────╮                   ╭───────▼───────╮     ║
║    │ ▪ Classes     │                   │ ▪ Full Logic  │     ║
║    │ ▪ Signatures  │                   │ ▪ Implement.  │     ║
║    │ ▪ Docstrings  │                   │ ▪ Variables   │     ║
║    ╰───────┬───────╯                   ╰───────┬───────╯     ║
║            │                                   │             ║
║    █▓▒░ Token Cost: ~5%                Token Cost: 100% ░▒▓█ ║
╚══════════════════════════════════════════════════════════════╝
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To make this work, I learned to rely on highly descriptive function names and docstrings. With good docstrings, the AI could understand the codebase's intent just by pinging the AST, mapping out the architecture flawlessly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sandboxing and Heuristics
&lt;/h3&gt;

&lt;p&gt;I actually used Gemini itself to help create a dataset filtering hundreds of developer CLI commands into "safe" and "harmful" categories. This allowed me to build heuristic judgments into the local sandbox, entirely disabling &lt;code&gt;shell: true&lt;/code&gt; evaluations in Node.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Ultimate Safety Net (Auto-Commits)
&lt;/h3&gt;

&lt;p&gt;You can never trust an AI blindly. I didn't want it making untracked changes. I built a feature into the engine that automatically runs a local &lt;code&gt;git commit&lt;/code&gt; after &lt;em&gt;every (almost) single file change&lt;/em&gt; and kept git files backed up when messing around with git commits and refactoring. If the AI messed up, I had an instant, local undo button. Git became the AI's eyes and my safety net.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google Gemini Feedback
&lt;/h2&gt;

&lt;p&gt;I think of designing with AI in two stages: when the system hums, and when the material fights back.&lt;/p&gt;

&lt;h3&gt;
  
  
  ✅ The 70% Magic
&lt;/h3&gt;

&lt;p&gt;When it worked, it felt like magic. About 70% of the time, the system operated perfectly. &lt;strong&gt;Reading code worked phenomenally well.&lt;/strong&gt; Gemini 2.5 Pro could absorb the AST skeletons, navigate the directory, and reason about system design with striking clarity.&lt;/p&gt;

&lt;h3&gt;
  
  
  ⚠️ When the material fights back (The 30% Chaos)
&lt;/h3&gt;

&lt;p&gt;The remaining 30% of the time was a battle against the realities of the medium: hallucinations, syntax errors, and context degradation.&lt;/p&gt;

&lt;h4&gt;
  
  
  1. The Writing Problem: Surgical Diffs vs. End-to-End
&lt;/h4&gt;

&lt;p&gt;While reading was elegant, writing was clumsy. Gemini 2.5 Pro struggled heavily with surgical, line-by-line insertions. Standard &lt;code&gt;diff&lt;/code&gt; patching is notoriously flaky with LLMs—they hallucinate line numbers or forget indentation. When I asked for complex changes across a large file, it would misplace the code entirely.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Workaround:&lt;/strong&gt; The solution wasn't to push harder; it was to change the abstraction. I stopped asking for line manipulations. Instead, I updated the protocol to force verbatim block patching or end-to-end rewrites of the entire file. The &lt;code&gt;search&lt;/code&gt; block had to match the local file &lt;em&gt;exactly&lt;/em&gt;, down to the space, or the Engine rejected it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  2. Amnesia and The Autonomous Self-Healing Loop
&lt;/h4&gt;

&lt;p&gt;Sometimes, the AI would just forget its own rules. It would hallucinate bad JSON syntax (trailing commas, unescaped quotes) or start outputting raw markdown. If I fixed it manually in my IDE, the sync between the "brain" and the "hands" broke.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Workaround:&lt;/strong&gt; I built an &lt;strong&gt;Autonomous Self-Correction Loop&lt;/strong&gt;. If the Chrome extension failed to &lt;code&gt;JSON.parse()&lt;/code&gt; the output, it automatically generated an error payload and injected it directly back into the chat. Gemini would immediately apologize, fix the syntax, and re-emit the tool call without me typing a single word and if this was too often then I inserted Protocol docs again.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;╔══════════════════════════════════════════════════════════════╗
║ ░░░░░░░░░░░░░ AUTONOMOUS SELF-HEALING LOOP ░░░░░░░░░░░░░░░░░ ║
╠══════════════════════════════════════════════════════════════╣
║                                                              ║
║   [🧠 Gemini UI]         [🧩 Extension]          [💻 Engine] ║
║         │                      │                      │      ║
║         │ 1. Invalid JSON      │                      │      ║
║         │ ───────────────────▶ │                      │      ║
║         │                      │ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓   │      ║
║         │ 2. Catch Error       │ ▓ JSON Parse Fails ▓ │      ║
║         │ ◀─────────────────── │ ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓   │      ║
║         │                      │                      │      ║
║         │ 3. Fixes Syntax      │                      │      ║
║         │ ───────────────────▶ │                      │      ║
║         │                      │ 4. Valid Request     │      ║
║         │                      │ ━━━━━━━━━━━━━━━━━━━▶ │      ║
║         │                      │                      │      ║
║         │                      │  ╭────────────────╮  │      ║
║         │                      │  │ ⚙️ Validates   │  │      ║
║         │                      │  │ ⚡ Executes    │  │      ║
║         │                      │  ╰────────────────╯  │      ║
║         │                      │ 5. Returns Data      │      ║
║         │                      │ ◀━━━━━━━━━━━━━━━━━━━ │      ║
║         │ 6. Inject Status     │                      │      ║
║         │ ◀━━━━━━━━━━━━━━━━━━━ │                      │      ║
║                                                              ║
╚══════════════════════════════════════════════════════════════╝
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  3. Context Degradation (The Final Boss)
&lt;/h4&gt;

&lt;p&gt;This compounds when the conversation gets too long. When dealing with detailed codebases, the LLM eventually loses its understanding of the current situation. The context window fills up, attention mechanisms drift, and it makes bad decisions.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Master Workaround:&lt;/strong&gt; To tackle this, I created a specific &lt;strong&gt;"Summarization Prompt"&lt;/strong&gt; within the Chrome extension. When a chat got too long and the AI started losing the plot, I didn't try to salvage it.&lt;/p&gt;

&lt;p&gt;I would run the prompt, instructing Gemini to condense its current understanding of the architecture, the problem, and our progress into one cohesive document. I would then open a brand new chat, paste that summary along with my base JSON rules, and resume.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;╔══════════════════════════════════════════════════════════════╗
║ ░░░░░░░░░ CONTEXT HANDOVER PROTOCOL (SUMMARIZATION) ░░░░░░░░ ║
╠══════════════════════════════════════════════════════════════╣
║                                                              ║
║  [ ⏳ STAGE 1: Degradation ]                                 ║
║   │                                                          ║
║   ╰─▶ 💬 Long Chat Session ──▶ ⚠️ Hallucinations Begin       ║
║                                                              ║
║  [ 💉 STAGE 2: The Handoff ]                                 ║
║   │                                                          ║
║   ├─▶ 1. Inject "Summarization Prompt" via Extension         ║
║   ╰─▶ 2. Gemini Generates Cohesive 'State Document' 🗂️        ║
║                                                              ║
║  [ ✨ STAGE 3: Resurrection ]                                ║
║   │                                                          ║
║   ├─▶ 1. Close Degraded Session 🗑️                            ║
║   ├─▶ 2. Open Empty Chat Session 🆕                          ║
║   ├─▶ 3. Inject [ Base Rules + State Document ] 📥           ║
║   ╰─▶ 4. Fresh AI Instance with Perfect Context 🚀           ║
║                                                              ║
╚══════════════════════════════════════════════════════════════╝
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;It was like giving the AI a fresh cup of coffee and a perfect handover document.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Elevating the craft
&lt;/h3&gt;

&lt;p&gt;The Kizuna system did not work flawlessly out of the box. But building it taught me that modern LLMs are not a pipeline; they are a search. They hallucinate, they forget, and they make errors.&lt;/p&gt;

&lt;p&gt;But by building the right scaffolding—strict constraints, local safety nets, AST parsers, and clever workarounds like the self-healing loop and summarization prompt—you can harness them to do incredible things. It forced me to be intentional about how software is written, and it laid a profound foundation for the systems I want to build next.&lt;/p&gt;

&lt;h1&gt;
  
  
  GeminiChallenge #AI #WebDev #Productivity #Agents #Automation #SoftwareDesign
&lt;/h1&gt;

</description>
      <category>devchallenge</category>
      <category>geminireflections</category>
      <category>gemini</category>
    </item>
    <item>
      <title>DialogueAI: Interactive Playground for assemblyai with automatic code generation</title>
      <dc:creator>Omkar tripathi</dc:creator>
      <pubDate>Sat, 23 Nov 2024 13:47:18 +0000</pubDate>
      <link>https://dev.to/code_triggered_/dialogueai-interactive-playground-for-assemblyai-speech-to-text-api-and-lemur-api-and-generate-30de</link>
      <guid>https://dev.to/code_triggered_/dialogueai-interactive-playground-for-assemblyai-speech-to-text-api-and-lemur-api-and-generate-30de</guid>
      <description>&lt;p&gt;This is a submission for the &lt;a href="https://dev.to/challenges/assemblyai"&gt;AssemblyAI Challenge:&lt;/a&gt; : Sophisticated Speech-to-text and No More Monkey Business.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built : &lt;a href="https://github.com/triggeredcode/DialogueAI" rel="noopener noreferrer"&gt;DialogueAI ( GITHUB )&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;I built DialogueAI, an interactive platform that leverages the powerful capabilities of AssemblyAI's sophisticated speech-to-text API and their LeMUR summarization model. The primary goal of this platform is to simplify the process for users who are new to these APIs, helping them overcome the steep learning curve typically associated with diving into new documentation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features of the Platform&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Interactive Playground&lt;/strong&gt;: Users can explore and experiment with various API functionalities through an intuitive interface. Input boxes, selection options, model selection, and summary types are all easily adjustable.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Instant Results&lt;/strong&gt;: With a single click, users can execute API calls and see the results immediately. This feature helps bridge the gap between learning and actual implementation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Code Generation&lt;/strong&gt;: For those who prefer to handle API calls manually, the platform generates the necessary code snippets, which can be directly run on their systems. This feature significantly reduces the time and effort required to understand and use the API.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Smart Summary Page&lt;/strong&gt;: Similar to the main playground, this page offers various configuration options and examples to help users generate summaries of transcripts quickly. Users can also get the generated code to use by themselves.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;By providing these features, the platform ensures that users can quickly and efficiently learn how to use AssemblyAI's APIs, reducing the frustration and time typically spent navigating complex documentation. This makes it an invaluable tool for developers and anyone looking to incorporate speech-to-text and summarization capabilities into their projects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Journey
&lt;/h2&gt;

&lt;p&gt;The inspiration for this platform came from my own experience when I first encountered AssemblyAI's API. I found it a bit confusing to get started with the documentation and the API usage. So, I set out to solve this problem not just for myself but for everyone else who might face the same challenge.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tech Used
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontend&lt;/strong&gt;: React, TypeScript, Tailwind CSS&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API&lt;/strong&gt;: AssemblyAI Speech-to-Text, LeMUR LLM model summary API&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Animations&lt;/strong&gt;: Framer Motion&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Working Features
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Interactive Speech-to-Text Configurations&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;Users can easily configure and experiment with various settings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single Click Run&lt;/strong&gt;: Execute the configuration and see results immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single Click Code Generation&lt;/strong&gt;: Generates the code based on the configuration for users to use directly.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Configurations Available&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API Key&lt;/li&gt;
&lt;li&gt;Speech Model&lt;/li&gt;
&lt;li&gt;Word Boost&lt;/li&gt;
&lt;li&gt;Profanity Filter&lt;/li&gt;
&lt;li&gt;Audio Range&lt;/li&gt;
&lt;li&gt;Audio Intelligence&lt;/li&gt;
&lt;li&gt;Summary Model&lt;/li&gt;
&lt;li&gt;Summary Type&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqdy3i857awn7gcagyd66.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqdy3i857awn7gcagyd66.png" alt="Interactive Speech-to-Text Configurations" width="800" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnocbuprkv0chetc8oksv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnocbuprkv0chetc8oksv.png" alt="Interactive Speech-to-Text Configurations" width="800" height="782"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fx02awcqmd34mypjg7b1e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fx02awcqmd34mypjg7b1e.png" alt="Interactive Speech-to-Text Configurations" width="800" height="531"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fx71uvjj0gqdk68us9wfa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fx71uvjj0gqdk68us9wfa.png" alt="Interactive Speech-to-Text Configurations" width="800" height="759"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fgv348kbgt4e2fi0vtgr9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fgv348kbgt4e2fi0vtgr9.png" alt="Interactive Speech-to-Text Configurations" width="800" height="740"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flmpst5ngfabyrn1jpbky.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flmpst5ngfabyrn1jpbky.png" alt="Generted code and summary" width="800" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9xrqfn5u7mzg5s8wre01.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9xrqfn5u7mzg5s8wre01.png" alt="Generted code and summary" width="800" height="842"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Interactive Summary Generation with LeMUR&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;Users can generate summaries with various options and configurations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single Click Run&lt;/strong&gt;: Instantly generate summaries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single Click Code Generation&lt;/strong&gt;: Provides the code for generating summaries.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Configurations Available&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API Key&lt;/li&gt;
&lt;li&gt;Summary Type (Basic, Custom)&lt;/li&gt;
&lt;li&gt;Transcript ID&lt;/li&gt;
&lt;li&gt;Model&lt;/li&gt;
&lt;li&gt;Prompt&lt;/li&gt;
&lt;li&gt;Custom Prompt&lt;/li&gt;
&lt;li&gt;Max Output Tokens (Example Pre-coded)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzom6koz439t2yeyymp3p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzom6koz439t2yeyymp3p.png" alt="Image description" width="800" height="817"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fb41b3qm3srzvokf5d8h9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fb41b3qm3srzvokf5d8h9.png" alt="Image description" width="800" height="790"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fs71aqh9fddn9eory9zn4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fs71aqh9fddn9eory9zn4.png" alt="Image description" width="800" height="826"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fk1smmt2kluqj1giglqee.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fk1smmt2kluqj1giglqee.png" alt="Image description" width="800" height="758"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  In Development
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chat with the Transcript&lt;/strong&gt;: Using LeMUR API to enable interactions with the generated transcript.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interactive Quiz Generation&lt;/strong&gt;: Generate quizzes based on the transcript.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff2o3u1j1vuli16ygw7yb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff2o3u1j1vuli16ygw7yb.png" alt="Image description" width="800" height="468"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Journey
&lt;/h3&gt;

&lt;p&gt;So far, I've successfully addressed the initial problem statements for the speech-to-text API and LeMUR summary model. This project has been incredibly exciting to work on, pushing the boundaries of what can be done with API interactions and user interface design.&lt;/p&gt;

&lt;p&gt;Looking ahead, I plan to expand the platform to include interactive playgrounds and code generation capabilities for real-time APIs and more sophisticated use cases of LeMUR. This will further streamline the learning and implementation process for developers and enhance the overall user experience.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>assemblyaichallenge</category>
      <category>ai</category>
      <category>developertool</category>
    </item>
  </channel>
</rss>
