Built for the DEV Hacktoberfest 2026 "Build for a Friend" weekend challenge.
The friend
This one's personal: I'm learning English myself — speaking, specifically. Not vocabulary quizzes or grammar worksheets; the actual terrifying part: opening your mouth and making sounds in front of another human.
The classic beginner trap: you need reps to get less shy, but you're too shy to get reps. What I needed wasn't another course. It was a patient partner who's available at 11pm on a Sunday, never rolls their eyes, and never repeats what I said to anyone.
So I built one. Over a weekend.
What SpeakBuddy does
SpeakBuddy is a web app where you pick an everyday scenario, roleplay it with an AI, get gentle corrections as you go, and finish with a personal feedback report. A session looks like this:
- Pick a scenario. Six to choose from, all things a beginner actually needs: ordering coffee, hotel check-in, airport small talk, ordering at a restaurant, asking for directions, and a job-interview self-introduction. Each one tells you the AI's role, your goal, and gives you starter phrases if you freeze up.
- Roleplay. The AI stays in character — a friendly barista, a front-desk receptionist, a fellow passenger at the gate. Replies are short (1–3 sentences), because a wall of text is the last thing a beginner needs.
- Gentle corrections. When you write something imperfect — say, "I want a big coffee" — the AI keeps the conversation flowing and drops an occasional 💡 tip: "You could say 'a large coffee, please' — sounds more natural!" Never a lecture. Nobody learns while being embarrassed.
- End session & get feedback. Hit the button and you get a structured report: new vocabulary with Chinese glosses, your mistakes corrected, what went well, one thing to practice next time, and some encouragement.
Two things that make it feel real: a 🎤 mic button (browser speech recognition, works in Chrome/Edge) so you can actually talk, and a 🔊 read-aloud toggle so you can hear the AI's replies. Speaking practice should involve your mouth and ears, not just your keyboard.
The open-source AI part (and why local matters here)
The challenge requires open-source AI at the core, and honestly this project turned out to be a case where local inference isn't just a checkbox — it's the whole point.
SpeakBuddy talks to Ollama's OpenAI-compatible API running on your own machine, using Gemma 3 (Google's open-weights model). Default is gemma3:1b — tiny, fast, good enough to start; the 4B version (gemma3) is noticeably better at English — recommended if your machine has the headroom. Any Ollama model works via the OLLAMA_MODEL setting.
Three reasons local was the right call for this app:
- Privacy. A shy beginner practicing out loud and making mistakes — the last thing they need is their stammering sent to a cloud API and logged somewhere. The conversation text is handled entirely by the local model: no cloud API, no account, no per-token billing. (One caveat: the mic button uses your browser's built-in speech recognition, which in Chrome/Edge may send audio to the browser vendor's servers. Typing keeps everything fully local.)
- Offline. The text conversation works fully offline — on a plane, in a subway, anywhere. (Mic input depends on your browser's speech recognition, which may need a connection.)
- Zero cost. No API keys, no accounts, no per-token billing. A friend learning English shouldn't need a credit card to practice saying "a latte, please."
The stack
Short version: Python + Gradio + Ollama (OpenAI-compatible API) + Web Speech API. That's it. app.py holds the Gradio UI and session logic, scenarios.json holds the six scenarios, prompts.py builds the roleplay and feedback prompts. Three dependencies: gradio, openai, python-dotenv.
Try it in 60 seconds
ollama pull gemma3:1b
pip install -r requirements.txt
python app.py
Open http://127.0.0.1:7860, pick ☕ Ordering Coffee, press Start session, and try typing something imperfect like I want a big coffee. Chat a few turns, notice the barista stays in character and the tips stay kind, then press 🏁 End session & get feedback to see your report. If Ollama isn't running, the app tells you exactly what to do instead of crashing.
🔗 Repo: https://github.com/yangpeng802/llmd
What's next
Honest list, roughly in order:
- Pronunciation scoring. Right now corrections are text-based. Comparing the learner's actual speech against the target phrase — even roughly — would close the loop between "what I said" and "what I meant to say."
- Progress tracking. A simple per-scenario history: words learned, recurring mistakes, sessions completed. Beginners quit when they can't see themselves improving.
- More scenarios, and harder ones. The current six cover survival English. Next: small talk at work, phone calls, disagreeing politely — the stuff that's actually hard.
Prize Categories
-
Best Use of Gemma ($200 featured category): SpeakBuddy runs on Gemma 3 — Google's open-weight model — served locally via Ollama (
gemma3:1bby default,gemma34B recommended). No cloud, no API key; the open model is the whole brain of the app.
Built for the DEV Hacktoberfest 2026 "Build for a Friend" weekend challenge.
Top comments (0)