DEV Community

Cover image for I Built a Private AI Pronunciation Coach So My Friend Wouldn't Give Up on Her Placement Dreams
Triveni K O
Triveni K O

Posted on AI-assisted

I Built a Private AI Pronunciation Coach So My Friend Wouldn't Give Up on Her Placement Dreams

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

Prize Categories

  • Overall Winner
  • Best Use of Gemma

đź’ˇThe Raw Truth: Why I Built This

We are currently sitting for our final-year campus placement drives. The pressure in the air is suffocating, but for my closest friend, it has been downright terrifying.

She comes from a rural vernacular background. In our computer science labs, she is a rockstar—her code is elegant, her logic is sharp, and her problem-solving skills put most of us to shame. But the moment she has to speak in English, her voice trembles.

A few weeks ago, she took an automated AI assessment required by an visiting tech company. She knew every single answer. But when the results came back, she was rejected. The machine flagged her response as "unclear" simply because of her regional accent and vowel shifts—saying "tablet" like "toblet" or "WhatsApp" as "whyatsup".

She came back to the hostel room, sat down quietly, and said: "I know the answers, but my mouth just fails me. Maybe corporate isn't for people like me."

That broke my heart.

Our college gave us access to a commercial communication training tool, but it made her anxiety ten times worse. It dumped a cold, robotic score like "Fluency: 58/100" without telling her which specific words were wrong or how to shape her mouth to say them correctly. It felt like being graded by a wall.

Naive AI wrappers don't help either—Speech-to-Text models like Whisper automatically "autocorrect" mispronunciations into perfect text, causing standard LLMs to hallucinate fake errors when they can't actually hear the audio.

I promised her I would build something better: a safe, private, non-judgmental space that actually hears acoustic speech variations, gives gentle phonetic guidance, and never leaves her laptop.


🛠️ What I Built: Placement Pronunciation Coach

Placement Pronunciation Coach is an acoustic-backed, locally-run AI speech evaluation dashboard designed specifically for students facing interview anxiety.

⚙️ How the Hybrid Acoustic Pipeline Works:

  1. Target Prompt Alignment: Evaluates expected interview responses against recorded user audio.
  2. Acoustic Ambiguity Detection (Whisper Timestamps): Uses OpenAI Whisper with word_timestamps=True to extract per-word confidence probabilities (< 0.75 threshold) to catch where speech clarity drops.
  3. Phonetic Sound Alignment (g2p_en): Converts expected and spoken text into IPA/Arpabet phonemes to detect real acoustic sound mismatches without relying on LLM guesses.
  4. Empathetic Coaching via Gemma 2:2b: Verified acoustic flags are passed to Google Gemma 2 (via Ollama). Using an empathetic system prompt, Gemma generates a supportive communication scorecard, natural interview rewrites, and simple syllable sound-outs (e.g., "TAB-lit" like 'cab').
  5. Listen & Repeat Reference Engine: Generates offline reference audio (pyttsx3) for target words so candidates can listen to clean model pronunciations and repeat them immediately.

📸 Demo & Screenshots

Placement Pronunciation Coach Dashboard

Acoustic Analysis & Phonetic Breakdown

Model Reference Audio & Practice

💡 Real-Time Feedback: The system highlights speech confidence scores, breaks down mispronounced words phonetically, and plays back clean reference audio—all running 100% locally.

Example Feedback: When practicing technical sentences, the tool flags low-confidence words, provides simple syllable breakdowns, and plays back the target sound without judging the speaker.


đź”’ Why Open-Source AI (Gemma 2) Matters Here

When you're terrified of mispronouncing words, uploading your vulnerable voice recordings to a cloud server or a paid API feels terrifying:

  1. Zero Judgment & Total Privacy: Running Gemma 2 locally via Ollama guarantees that her voice recordings never leave her laptop. No telemetry, no cloud logging, no fear.
  2. Offline Accessibility: Students from rural backgrounds often face spotty internet access in hostel rooms or hometowns. Local inference means she can practice anywhere, anytime, completely offline.
  3. Customized Empathy: Open-weight models allow us to craft tailored system prompts that prioritize gentle encouragement over cold automated scoring.

đź”— Code Repository

GitHub Repository: https://github.com/triveniko/placement-pronunciation-coach

Top comments (0)