This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My wife Eugenia moved to San Francisco and is learning English. She's at an A2 (elementary) level, and the thing that trips her up most isn't grammar, it's idioms. People here say "it's a piece of cake", "let's call it a day" and "I'm under the weather" all the time, and a textbook doesn't teach you how to say them.
So I built her IdiomsMasterGirl: a patient practice partner that teaches one common idiom a day and then listens to her use it.
A daily session takes under 20 minutes (her first full session did) and has six steps:
- Learn β the idiom, its meaning in simple English, a "when to use it" note, a literal-vs-figurative note and 3 example sentences. Every line has a βΆ button with a natural American voice.
- Speak (try 1) β she records the idiom and each example. She gets accuracy, clarity and a combined score, plus 1β3 warm tips that name the exact word that went wrong.
- Speak (try 2) β she records everything again and sees +N points next to each line.
- Write β she writes 3 sentences of her own. Each one gets an idiom-use grade, a grammar grade, a corrected version with the change highlighted, and a short why.
- Read aloud β she listens to her corrected sentences, then reads them and gets scored.
- Score β a daily score, a 30-day chart and a π₯ streak.
The part I care most about is the feedback. It's never "wrong". It's "I heard βbananaβ for βschoolβ. Say 'school' slowly. Make sure you say the 'oo' sound clearly."
Demo
The app runs on our laptop on purpose: the speech recognition, the tutor model and her recordings all stay there. So there's no hosted link. The screenshots below are the real app on a demo account. To keep anyone's real voice out of the post, the recordings were made with the Mac's built-in voice, with one word swapped on purpose so you can see a tip.
Learn β today's idiom, with a play button on every line:
Speak β every recording is scored, and the tips name the actual words:
Write β her own sentences, corrected with the smallest change and a reason:
Progress β a streak and a 30-day chart, to keep her coming back:
Code
gusanchefullstack
/
idioms-master-girl
One English idiom a day β learn it, say it, write it. A local-first practice partner built for my wife (Gemma via Ollama + faster-whisper).
IdiomsMasterGirl π¬
One English idiom a day: learn it, hear it, say it, write it, and read it aloud. A local-first practice partner for an A2 English learner. It listens to her speak, tells her kindly which exact words to fix, and corrects her own sentences. The AI runs on the laptop so her voice and mistakes stay private.
Table of Contents
- Why this exists
- How it works
- Installation
- Quick start
- Configuration
- API reference
- Project structure
- Tests
- Contributing
- Roadmap
- What I learned
- Open models, privacy and why open matters
- Known limitations
- License
- Credits
- Screenshots
- Author
Why this exists
I built this for Eugenia, my wife. She is learning English in San Francisco at an A2 (elementary) level. Idioms are everywhere in American life ("it's a piece of cake", "let's call it a day"), and they are hard to learn from a textbook.
General apps don't fit her well. Flashcardβ¦
It's MIT-licensed. Setup is four commands (nvm use, npm install, npm run setup, npm run dev). The README covers configuration, the API and what I learned along the way.
How I Built It
Chrome (localhost)
β records webm/opus Β· falls back to the Mac's on-device voice
βΌ
Next.js 16 API routes
βββ scoring ββββΊ pure TypeScript: word alignment (WER), clarity, problem words
βββ tutor ββββββΊ Ollama + Gemma 4 E2B (4.6B, Apache 2.0), JSON-schema output
βββ listening ββΊ Python sidecar: FastAPI + faster-whisper small.en
βββ voice ββββββΊ ElevenLabs (opt-in, text only, cached on disk)
βββ storage ββββΊ SQLite via Node 26's built-in node:sqlite
Gemma is the tutor. It writes the 3 example sentences for each day's idiom, the coaching hints, and the grades and corrections for her sentences. Every call uses Ollama's JSON-schema format, and the code then validates the answer: an example must contain the idiom and have at most 15 words, a correction must keep the idiom, and an explanation must be 40 words or fewer. A failed answer is retried once, then falls back honestly. The app never shows an invented grade.
Whisper is the ear. faster-whisper gives a confidence for every word, which is the whole scoring model:
-
Accuracy =
100 Γ (1 β word error rate)against the target sentence (case, punctuation and contractions ignored). - Clarity = the average word confidence.
- Combined = 70% accuracy + 30% clarity.
- Problem words = words that were missed, replaced, or heard with less than 0.6 confidence.
With temperature=0.0 and condition_on_previous_text=False, the same recording gets the same score every time. In my tests it was identical.
The lesson that changed the design
My first version asked Gemma to write the tips. The tips were warm but vague, full of other idioms ("don't let tricky words get under your skin!"). In my first test, 0 of 4 named the word that was actually missed.
The fix was to split the job. The code writes the facts it already knows, and Gemma writes only the coaching:
/** The factual part of a tip is written in code, so every tip names a real word. */
export function leadFor(pw: ProblemWord): string {
switch (pw.type) {
case "substituted":
return `I heard β${pw.heard}β for β${pw.word}β.`;
case "missing":
return `I didn't hear β${pw.word}β.`;
case "unclear":
return `β${pw.word}β was a little unclear.`;
}
}
// final tip = leadFor(word) + " " + gemmaHint(word) // e.g. "Make the 'oo' sound long."
After that change, in a check of 20 recordings, 20 of 20 tips named the right word, and 18 of 20 hints were fully written by Gemma. When a recording is clean, there's no model call at all, just praise, so it never invents a problem.
Measured on a MacBook (M4):
- Score and tips after a recording: 1.1β2.9 s.
- Today's examples: about 2.6 s.
- Saved audio replay: 17 ms.
- Grammar corrections: 13 of 15 known errors fixed.
I planned and built it spec-first with GitHub Spec Kit. The spec, plan, research, contracts and tasks are all in the repo under specs/, with Claude Code as my pair programmer.
Why Does Open Innovation Matter?
For this project, open isn't a nice extra. It's the reason it works for her.
-
Her mistakes stay hers. A learner's recordings are personal: the hesitations, the wrong words, the accent. With an open-weight model and open speech recognition running locally, her voice never leaves our laptop. Recordings go to a helper on
127.0.0.1, are scored in memory and are never saved. There's no telemetry. - It's built to work offline. The tutor, the speech recognition and the database all run locally, and the speech helper is started in offline mode. Only new natural-voice audio needs the internet, and saved clips plus the Mac's own voice cover that. (A full run with Wi-Fi switched off is next on my list.)
-
It's tuned to one person. The prompts are Markdown files, the idioms are a curated JSON list I control, and the model is one line in
config/settings.json. Changing the tutor's personality ("warm, patient, playful big-sister tutor") is editing a text file. Grounding it in a curated list of 100 real A2βB1 idioms means it can't invent fake idioms, which is a real risk with a 4.6B model. - It costs nothing to run. There's no per-session API bill for something meant to be used every day.
- Where open beat closed: the scoring needs per-word confidence and repeatable results. Open Whisper gives both directly. A closed speech API would hide the confidences, or change them under me without notice. Being able to read and change every prompt also let me fix the vague-tips problem in an evening.
The one non-open piece is the optional ElevenLabs voice, and I kept it honest. It's off by default, it only ever receives the sentence to speak (never her audio or answers), and each sentence is generated once and cached.
My Agent Session
I built this with Claude Code, following the Spec Kit flow from constitution to spec, plan, tasks and implementation.
Prize Categories
- Best Use of Gemma β Gemma 4 E2B (via Ollama) is the tutor. It writes the daily examples, the pronunciation coaching and the writing grades and corrections, all with schema-constrained output that the code validates.
- Best Use of ElevenLabs β the natural American voice she imitates. It's used in a privacy-respecting way: opt-in, text only, cached per sentence, with an on-device fallback when it's off or offline.
This post was written with AI assistance (Claude Code) and checked by me before publishing.





Top comments (0)