<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: moonbird</title>
    <description>The latest articles on DEV Community by moonbird (@yangpeng802).</description>
    <link>https://dev.to/yangpeng802</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161547%2F72a48087-df65-4329-be4f-ae798b507bab.png</url>
      <title>DEV Community: moonbird</title>
      <link>https://dev.to/yangpeng802</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yangpeng802"/>
    <language>en</language>
    <item>
      <title>SpeakBuddy: a patient AI speaking partner that runs 100% on your laptop</title>
      <dc:creator>moonbird</dc:creator>
      <pubDate>Sun, 04 Oct 2026 12:11:58 +0000</pubDate>
      <link>https://dev.to/yangpeng802/speakbuddy-a-patient-ai-speaking-partner-that-runs-100-on-your-laptop-38lc</link>
      <guid>https://dev.to/yangpeng802/speakbuddy-a-patient-ai-speaking-partner-that-runs-100-on-your-laptop-38lc</guid>
      <description>&lt;p&gt;&lt;em&gt;Built for the DEV Hacktoberfest 2026 "Build for a Friend" weekend challenge.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The friend
&lt;/h2&gt;

&lt;p&gt;This one's personal: I'm learning English myself — speaking, specifically. Not vocabulary quizzes or grammar worksheets; the actual terrifying part: opening your mouth and making sounds in front of another human.&lt;/p&gt;

&lt;p&gt;The classic beginner trap: you need reps to get less shy, but you're too shy to get reps. What I needed wasn't another course. It was a patient partner who's available at 11pm on a Sunday, never rolls their eyes, and never repeats what I said to anyone.&lt;/p&gt;

&lt;p&gt;So I built one. Over a weekend.&lt;/p&gt;

&lt;h2&gt;
  
  
  What SpeakBuddy does
&lt;/h2&gt;

&lt;p&gt;SpeakBuddy is a web app where you pick an everyday scenario, roleplay it with an AI, get gentle corrections as you go, and finish with a personal feedback report. A session looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick a scenario.&lt;/strong&gt; Six to choose from, all things a beginner actually needs: ordering coffee, hotel check-in, airport small talk, ordering at a restaurant, asking for directions, and a job-interview self-introduction. Each one tells you the AI's role, your goal, and gives you starter phrases if you freeze up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Roleplay.&lt;/strong&gt; The AI stays in character — a friendly barista, a front-desk receptionist, a fellow passenger at the gate. Replies are short (1–3 sentences), because a wall of text is the last thing a beginner needs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gentle corrections.&lt;/strong&gt; When you write something imperfect — say, &lt;em&gt;"I want a big coffee"&lt;/em&gt; — the AI keeps the conversation flowing and drops an occasional 💡 tip: &lt;em&gt;"You could say 'a large coffee, please' — sounds more natural!"&lt;/em&gt; Never a lecture. Nobody learns while being embarrassed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;End session &amp;amp; get feedback.&lt;/strong&gt; Hit the button and you get a structured report: new vocabulary with Chinese glosses, your mistakes corrected, what went well, one thing to practice next time, and some encouragement.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two things that make it feel real: a 🎤 mic button (browser speech recognition, works in Chrome/Edge) so you can actually &lt;em&gt;talk&lt;/em&gt;, and a 🔊 read-aloud toggle so you can &lt;em&gt;hear&lt;/em&gt; the AI's replies. Speaking practice should involve your mouth and ears, not just your keyboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  The open-source AI part (and why local matters here)
&lt;/h2&gt;

&lt;p&gt;The challenge requires open-source AI at the core, and honestly this project turned out to be a case where local inference isn't just a checkbox — it's the whole point.&lt;/p&gt;

&lt;p&gt;SpeakBuddy talks to &lt;strong&gt;Ollama's OpenAI-compatible API&lt;/strong&gt; running on your own machine, using &lt;strong&gt;Gemma 3&lt;/strong&gt; (Google's open-weights model). Default is &lt;code&gt;gemma3:1b&lt;/code&gt; — tiny, fast, good enough to start; the 4B version (&lt;code&gt;gemma3&lt;/code&gt;) is noticeably better at English — recommended if your machine has the headroom. Any Ollama model works via the &lt;code&gt;OLLAMA_MODEL&lt;/code&gt; setting.&lt;/p&gt;

&lt;p&gt;Three reasons local was the right call for &lt;em&gt;this&lt;/em&gt; app:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Privacy.&lt;/strong&gt; A shy beginner practicing out loud and making mistakes — the last thing they need is their stammering sent to a cloud API and logged somewhere. The conversation text is handled entirely by the local model: no cloud API, no account, no per-token billing. (One caveat: the mic button uses your browser's built-in speech recognition, which in Chrome/Edge may send audio to the browser vendor's servers. Typing keeps everything fully local.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline.&lt;/strong&gt; The text conversation works fully offline — on a plane, in a subway, anywhere. (Mic input depends on your browser's speech recognition, which may need a connection.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero cost.&lt;/strong&gt; No API keys, no accounts, no per-token billing. A friend learning English shouldn't need a credit card to practice saying "a latte, please."&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The stack
&lt;/h2&gt;

&lt;p&gt;Short version: &lt;strong&gt;Python + Gradio + Ollama (OpenAI-compatible API) + Web Speech API&lt;/strong&gt;. That's it. &lt;code&gt;app.py&lt;/code&gt; holds the Gradio UI and session logic, &lt;code&gt;scenarios.json&lt;/code&gt; holds the six scenarios, &lt;code&gt;prompts.py&lt;/code&gt; builds the roleplay and feedback prompts. Three dependencies: &lt;code&gt;gradio&lt;/code&gt;, &lt;code&gt;openai&lt;/code&gt;, &lt;code&gt;python-dotenv&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it in 60 seconds
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull gemma3:1b
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
python app.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open &lt;a href="http://127.0.0.1:7860" rel="noopener noreferrer"&gt;http://127.0.0.1:7860&lt;/a&gt;, pick ☕ &lt;strong&gt;Ordering Coffee&lt;/strong&gt;, press &lt;strong&gt;Start session&lt;/strong&gt;, and try typing something imperfect like &lt;code&gt;I want a big coffee&lt;/code&gt;. Chat a few turns, notice the barista stays in character and the tips stay kind, then press &lt;strong&gt;🏁 End session &amp;amp; get feedback&lt;/strong&gt; to see your report. If Ollama isn't running, the app tells you exactly what to do instead of crashing.&lt;/p&gt;

&lt;p&gt;🔗 Repo: &lt;a href="https://github.com/yangpeng802/llmd" rel="noopener noreferrer"&gt;https://github.com/yangpeng802/llmd&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Honest list, roughly in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pronunciation scoring.&lt;/strong&gt; Right now corrections are text-based. Comparing the learner's actual speech against the target phrase — even roughly — would close the loop between "what I said" and "what I meant to say."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Progress tracking.&lt;/strong&gt; A simple per-scenario history: words learned, recurring mistakes, sessions completed. Beginners quit when they can't see themselves improving.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;More scenarios, and harder ones.&lt;/strong&gt; The current six cover survival English. Next: small talk at work, phone calls, disagreeing politely — the stuff that's actually hard.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma&lt;/strong&gt; ($200 featured category): SpeakBuddy runs on Gemma 3 — Google's open-weight model — served locally via Ollama (&lt;code&gt;gemma3:1b&lt;/code&gt; by default, &lt;code&gt;gemma3&lt;/code&gt; 4B recommended). No cloud, no API key; the open model is the whole brain of the app.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Built for the DEV Hacktoberfest 2026 "Build for a Friend" weekend challenge.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
