What I Built
I built my friend a voice-note assistant that keeps their audio on their laptop
My friend has a habit I think a lot of us have: they record voice notes whenever something comes to mind, then almost never come back to them.
assignments. People to call. Things to buy. Meetings. Random details they know they will forget.
So for this Hacktoberfest weekend, I built VoxBuddy β a small local AI app that turns a messy voice note into a clear summary, actionable tasks, priorities, and important details.
Demo
Code
VoxBuddy ποΈ
Turn messy voice notes into things you can act on β privately, with local open AI.
VoxBuddy was built for a friend who records lots of voice notes but rarely revisits them. Instead of sending those recordings to a cloud AI API, VoxBuddy keeps the AI pipeline on the user's machine:
Voice note β Whisper β transcript β Qwen3 β summary + tasks + important details
Why this project exists
A voice note is easy to record and surprisingly easy to forget.
VoxBuddy takes the few seconds your friend already spends speaking and turns them into something actionable:
- A short summary
- Clear tasks
- Priority levels
- Deadlines when they were actually mentioned
- The original transcript for verification
The goal is not to create another general-purpose chatbot. It is a tiny tool built around one real person's habit.
Why open AI matters here
The most important design choice is thatβ¦
How I Built It
I built VoxBuddy as a local-first AI application using open-source AI models and local inference.
The app starts with a voice note recorded or uploaded through the browser. The audio is sent to a Flask backend, where Whisper runs locally to convert the speech into text. This gives VoxBuddy an accurate transcript without sending the recording to a cloud speech API.
The transcript is then passed to Qwen3 (4B) running locally through Ollama. I prompt the model to turn the unstructured voice note into structured information such as a summary, actionable tasks, priorities, and deadlines. The backend validates the model's response and extracts the JSON before sending the results back to the frontend.
The main flow is:
Voice note β Local Whisper β Transcript β Local Qwen3/Ollama β Structured tasks & summary β Web UI
I used HTML, CSS, and JavaScript for the frontend and Python + Flask for the backend. Everything is designed to run on the user's own laptop once the models and dependencies are installed.
The open-source/local approach is an important part of the project rather than just a technical choice: voice notes can contain personal or sensitive information, so VoxBuddy can process them locally without requiring a third-party AI server. It also means the models can be swapped, modified, or upgraded without redesigning the entire application.
Why Does Open Innovation Matter?
- Privacy: Voice notes can contain personal information, and VoxBuddy processes them locally instead of sending them to a cloud AI service.
- Offline use: After setup, the core AI processing can run without an internet connection.
- No per-request API costs: Local inference avoids ongoing cloud API charges.
- Model freedom: I can swap or experiment with different open models instead of being locked into one provider.
- Control: I can change the prompts, processing pipeline, and application behavior to fit the user's needs.
Prize Categories
Overall Winner
Top comments (0)