DEV Community

Adpirs
Adpirs

Posted on

Speech Practice Partner

What I Built

Most voice assistants decide you've finished speaking after about a second of silence. For someone who stammers, is nervous about speaking, or is searching for words (after a stroke, for example), that means being cut off again and again. It makes practising out loud feel worse, not better.

So I built Speech Practice Partner, a patient conversation partner that never interrupts. You pick a topic you enjoy, press one big button and just talk. It waits as long as you need (you choose anywhere from 2 to 15 seconds of quiet), then replies in one or two short, gentle sentences. It never finishes your sentences, never corrects your grammar and never comments on pauses or repetition.

I built it for anyone and everyone, who gets anxious speaking to people and wants a low-pressure place to practise.

Things I cared about:

  • Calm and simple: one screen, one button, a soft orb that shows whether it is listening, hearing you, thinking or speaking.
  • Private: everything runs on the laptop and nothing goes to the cloud. Only the written transcript and a few numbers are saved, and your voice is never recorded. One button deletes everything.
  • Progress without pressure: a Progress page shows personal trends (pause length, pace, words per session). They are framed as trends, not scores, and the numbers are hidden during a session by default.

Demo

Video Link :https://youtu.be/UPSXFH25tTc?si=QE1RDLcW4migRK6D

The app has a --demo mode that previews the whole interface with synthetic sample data and a pretend conversation. Any trend charts in the demo come from that synthetic data and are not real people's results.

How I Built It

The whole pipeline runs locally:

Mic → Silero VAD → faster-whisper → Gemma (via Ollama) → Piper voice

  • Silero VAD detects when you start and stop talking. I set the end-of-speech silence window myself, which is the heart of the project.
  • faster-whisper (small.en) transcribes locally. I prompt it to keep repetitions and fillers instead of "cleaning them up".
  • Gemma (gemma3:4b, open weights, run through Ollama) is the conversation partner. A short system prompt keeps it warm, brief and non-corrective.
  • Piper speaks the replies, with an adjustable slower voice.
  • TabPFN (optional, local) forecasts the next session's pause length and speaking rate from past sessions. It is backtested against a naive baseline, and it needs at least 8 real sessions before it will forecast. [Add your backtest result here once you have real sessions, or delete this line.]
  • Flask + plain HTML/CSS/JS make up the interface. There are no frontend frameworks and no CDN, so it works fully offline. The server only listens on 127.0.0.1.

I started from a terminal-only prototype and turned it into a simple browser app: settings inside the app, friendly error messages (like "Open the Ollama app and try again"), a session you can end at any moment, and dark mode.

Why Does Open Innovation Matter?

  • Control over the one setting that matters. With a closed voice API, I can't change how long it waits before deciding you're done. With open components I could set that window to match the person, and tune it with them.
  • Privacy. A person practising something vulnerable shouldn't have to send their voice to someone else's servers. Because every model runs locally, their speech never leaves the laptop.
  • Free, unlimited practice. Confidence builds with repetition, and nobody has to worry about per-minute API costs.
  • Tunable to the person. The prompt, voice speed and wait time can all be adjusted after each session.

Prize Categories

  • Gemma

Top comments (0)