This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My brother is 15 and sometimes has to speak in front of people. I asked him what he thought about a small app that could help him practise. He said: "It's good."
So I made one.
Second Take is a small web app that runs on my laptop. You record yourself talking for a short time, and it gives you a few things to look at: speaking pace, pauses of 0.7 seconds or more, and filler words like "um", "uh", "like", "basically", and "you know".
The filler words and pauses have timestamps, so you can click on one and go back to that part of the recording.
It also sends the transcript and the measurements to Gemma through Ollama. Gemma gives practice tips based on them. Each take is saved locally, and the app has a chart showing fillers per minute and long pauses across the takes.
The important part is that Gemma isn't calculating the numbers. Python does that first. The model gets the finished measurements and the transcript and writes the advice.
I also added a check after Gemma responds. It finds numbers in the response and compares them with the numbers calculated by the app. It caught numbers in several of my test responses, although it also produced some false alarms.
Demo
I made five real takes.
The first take was 23.5 seconds long. It measured 99.5 words per minute, 12.8 fillers per minute, and 3 long pauses.
The next take was 27.2 seconds long. It measured 119.2 words per minute, 4.4 fillers per minute, and 5 long pauses.
The app also shows where those pauses happened. In the second recording, for example, it found pauses between "scalability." and "When", "server..." and "What", and "copy." and "It".
The transcription isn't perfect. One part came back as:
ThesecondpointIwanttomakeisaboutscalability
So there is clearly still work to do there. The timestamps are more reliable than the transcript formatting, which is why the measurements are based on the timestamps.
My brother actually tried it. I asked him what he thought.
He said:
"It's good."
That's the whole review.
Code
https://github.com/srishtipandey1/second-take
How I Built It
Transcription uses faster-whisper with the base.en model and word timestamps.
Whisper sometimes removes things like "um" and "uh", which is a problem because those are exactly the things I wanted to count. I turned off the voice-activity filter and gave it a short prompt containing some disfluencies.
That helped. The test recordings included "uh" and "um" in the transcript, although the transcription itself still wasn't always clean.
The metrics are just a Python function that takes (word, start, end) values.
For pace, it counts words in each 30-second window and converts that to words per minute.
For pauses, it looks at the gap between one word ending and the next word starting. Anything at least 0.7 seconds counts as a long pause.
For fillers, it checks the transcript against two lists of words.
I also wrote a small test using made-up timestamps where I already knew what the answers should be. The test passed.
For the AI part, I used Ollama with gemma3:1b. I went with the 1B model because I wanted something that would actually run on my laptop without a GPU.
The prompt gives Gemma the measurements and transcript and asks for three practice tips with timestamps.
The app stores takes in a local JSON file. The chart reads from the same file.
I built the app with GitHub Copilot in agent mode. I gave it the requirements, let it make the first version, then ran it and fixed things that didn't behave the way I wanted.
One thing I caught was the recording button still looking like it was recording after I stopped a take. I also changed the history so it doesn't make old takes look playable when their audio isn't stored.
Results
Here are the five takes currently in the app:
| Take | Fillers per minute | Long pauses |
|---|---|---|
| 1 | 4.4 | 5 |
| 2 | 12.8 | 3 |
| 3 | 2.2 | 4 |
| 4 | 0.0 | 1 |
| 5 | 0.0 | 1 |
The numbers don't move neatly downward, which is probably what you'd expect from five short recordings.
The 10:30 take had 12.8 fillers per minute and 3 long pauses. The 10:33 take had 4.4 fillers per minute but 5 long pauses. So there were fewer fillers, but more long pauses.
The tips from Gemma were mostly about pauses and explaining things more clearly.
What my brother said after:
"It's good."
What didn't work: the 1B model isn't particularly good at following the coaching format every time. Some responses are generic, and sometimes it puts numbers in the answer that aren't part of the measurements. In the latest take, the number check flagged 1 and 3.
The other problem is transcription. Whisper occasionally runs words together or misses things, so the filler count depends on what makes it through the transcript.
Why Does Open Innovation Matter?
The recordings are a teenager's voice, so I wanted to keep them on my laptop.
Whisper runs locally. Gemma runs locally through Ollama. There is no account, API key, or audio upload involved. The take history is just a file in the project folder.
That also made testing easier. I could record something, run it through the app again, change the prompt, and try again without paying for an API call.
The model name is an environment variable too, so I can swap gemma3:1b for a bigger model later if the computer can handle it.
Prize Categories
- Best Use of Gemma: Gemma writes the practice tips from the measurements calculated by the app, using Ollama locally.
- Best Use of GitHub Copilot: I built the app with Copilot in agent mode and used it to generate the first version from my specification.





Top comments (0)