This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built Uchaar, a personalized speech-practice coach for my brother who struggles with articulation and wants a better way to practice.
The problem I wanted to solve was not simply a lack of speech exercises. The bigger problem was how difficult it is to practice consistently when you don't know what to focus on or how to structure the next practice session.
A typical practice session can become:
Say a word → repeat it → move to another word → repeat again.
There is very little connection between one session and the next. The user has to decide what to practice, how much to practice, and how to hear the target before attempting it.
I wanted to change that.
Uchaar turns speech practice into a guided loop built around the user's own recordings.
The user first completes a controlled assessment by recording words and sentences. Uchaar analyzes those recordings and produces structured findings.
Those findings are then given to Gemma 3, running locally through Ollama.
Gemma's job is to take the application's findings and turn them into something the user can actually act on:
- What does the assessment mean?
- What stood out?
- What should I practice next?
- What cues can I use while practicing? The result is a personalized practice plan rather than a generic list of exercises.
Then Uchaar uses ElevenLabs to generate natural spoken examples for those practice targets.
This creates the complete experience:
The user speaks → Uchaar analyzes → Gemma creates the practice plan → ElevenLabs provides a spoken reference → the user practices again.
So the solution is not just an AI-generated report.
It is a closed practice loop where the user's assessment influences what they practice next.
The goal is to make practice feel less like repeatedly saying random words and more like following a guided session that adapts to the user's own recordings.
Demo
Watch the Demo Video
The demo shows the complete journey:
Assessment → Recording → Analysis → Gemma Explanation → Personalized Practice Plan → ElevenLabs Reference → Practice
The most important part is seeing the transition from the user's recording to an actual practice session.
Code
Github Repository
The repository contains the FastAPI backend, React frontend, deterministic audio-analysis layer, assessment system, Gemma/Ollama integration, ElevenLabs integration, practice flow, progress persistence, and tests.
How I Built It
I deliberately did not make one AI model responsible for the entire application.
Uchaar has three main intelligence layers.
1. Deterministic Audio Analysis
- The user's recording is first processed by Uchaar's own audio-analysis layer.
- It extracts structured signal-level information such as duration, RMS, dBFS, peak amplitude, spectral information, and recording-quality information.
- This gives the application structured evidence before an AI model is involved.
2. Gemma 3 + Ollama
- Gemma 3 runs locally through Ollama.
- It receives the structured findings from the assessment rather than raw audio.
- Gemma then turns those findings into:
assessment explanation → observations → personalized practice targets → practice cues
- Its output is validated by the backend before being shown to the user.
- This makes Gemma the reasoning and personalization layer of Uchaar.
3. ElevenLabs
- ElevenLabs solves a different problem.
- Once Uchaar has determined what the user should practice, ElevenLabs generates a natural spoken reference.
The user can:
Listen → Repeat → Practice
So:
- Gemma decides what to practice.
- ElevenLabs helps the user hear the practice target.
The application itself is built with FastAPI and React, with local JSON persistence and automated tests around the core backend functionality.
Why Does Open Innovation Matter?
For this project, open innovation was important because I didn't want the entire application to depend on a closed AI API.
Gemma 3 runs locally through Ollama.
That gives me control over the reasoning layer: what information is provided to the model, what structure it must return, how its output is validated, and how the model fits into the rest of the application.
More importantly, it allowed me to build the system around multiple specialized components instead of asking one proprietary model to do everything.
The architecture is:
Deterministic Audio Analysis → Structured Evidence → Gemma 3 → Reasoning & Personalization → ElevenLabs → Natural Spoken Reference
Each component has a clear job.
That is what open innovation made possible for Uchaar: I could experiment with the AI layer as part of my own system rather than making my entire product a wrapper around a single closed API.
My Agent Session
I used AI-assisted development throughout the project to help implement features, debug integration problems, test the backend, and iterate on the Gemma and ElevenLabs workflows.
The important engineering decisions remained under my control, including the separation between deterministic analysis and AI reasoning, the validation of Gemma's output, and the overall practice workflow.
Prize Categories
- Best Use of Gemma
Gemma 3 is used as the reasoning and personalization engine of Uchaar.
It takes structured assessment findings and turns them into a user-friendly explanation and personalized practice plan.
- Best Use of ElevenLabs
ElevenLabs is integrated directly into the practice experience.
It generates the spoken reference examples that the user listens to before attempting the practice target themselves.
Top comments (0)