This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Some time back I worked with a team where two of the teammates couldn't hear. In meetings, they would use the captions. Well, after the meetings, what's next? We still had to communicate.
When the rest of us wanted to tell them something, we typed and showed them the screen or the paper we wrote on.
It worked. So in this hackathon I built M'aso, which means my ear in Akan, a Ghanaian language. Two people open a private room:
- Live captions. What one person says appears on both screens about a second later, with their name on it.
- Typing for everyone. Anyone can type, and a typed message shows just as large as speech.
- Show on screen. One button turns the laptop into a full-screen board of huge text: type it, turn the laptop around.
- Both agree first. Captions don't start until everyone in the room has agreed.
- A note of what was agreed. At the end, Gemma writes a short summary of decisions, action items and dates that you can correct, save or delete.
Demo
Code
m’aso
Make room for every voice.
I used to work on a team where one or two teammates couldn’t hear. In meetings they used captions. When the rest of us wanted to tell them something, we typed what we meant and showed them the screen. That’s where m’aso came from.
Two people open a private room:
- Live captions. What one person says appears on both screens about a second later, with their name on it.
- Typing for everyone. Anyone can type a message, and it shows just as large as speech. Show on screen turns the laptop into a full-screen board of huge, high-contrast text: type it, turn it around.
- Both of you agree first. Captions don’t start until everyone in the room has agreed on the consent screen.
- A note of what was agreed. At the end, Gemma writes a short summary — decisions, action items, dates — that…
scripts/setup.sh # env files, dependencies, Gemma model (one time)
scripts/dev.sh # Ollama + AI service + web app → http://localhost:3000
https://m-aso.onrender.com/
## How I Built It
Two parts: a Next.js 16 app for the screens, and a small Python service (FastAPI) that holds the rooms and runs the models.
Browser (Next.js) AI service (FastAPI)
microphone → AudioWorklet → 16 kHz PCM ──▶ WebSocket /rooms/{code}/ws
├─ segmenter: splits speech on pauses
captions, typed messages, presence ◀───── ├─ faster-whisper (small.en, int8, CPU)
└─ room hub: broadcasts to everyone
end of conversation ── POST /summary ────▶ Gemma 3 4B via Ollama → decisions, actions, dates
## Why Does Open Innovation Matter?
For one teammate the transcript is how they take part, so faster-whisper and Gemma run on my own laptop instead of a paid third-party API that bills every minute and hears every word. Because the models are open, I could also shape them to the job Gemma always returns the same structured summary and typed lines count as much as speech.
## Prize Categories
**Best Use of Gemma**: Gemma 3 4B, served locally through Ollama, turns the conversation into structured decisions, action items, dates and open questions, with spoken and typed lines treated equally.

Top comments (0)