This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built Voice2Note, a local-first AI application that turns voice recordings into structured, searchable knowledge.
I built it for a friend who frequently records voice notes to capture ideas, plans, reminders, and thoughts, but often struggles to find or act on that information later.
The problem isn't recording a voice note.
The problem is what happens after the recording.
Important ideas get buried. Tasks get forgotten. Finding something mentioned several days ago becomes difficult.
Voice2Note turns that unstructured voice data into something useful:
Voice Recording
↓
Local Speech-to-Text
↓
Open-Weight LLM
↓
Summary + Tasks + Ideas + Decisions
↓
Local Knowledge Base
↓
Semantic Search
↓
Ask My Notes
Instead of simply transcribing a recording, Voice2Note helps answer:
"What did I say about this before?"
Key Features
- 🎙️ Record voice notes directly from the browser
- 📁 Upload existing audio recordings
- 📝 Local speech-to-text transcription
- 🧠 AI-generated summaries
- ✅ Automatic task extraction
- 💡 Idea and decision extraction
- 🔎 Semantic search across previous notes
- 💬 Ask questions about your own voice notes
- 📌 Source-backed answers
- 🔒 Local-first/private AI architecture
- ⚙️ Swappable open-weight AI models
The goal was to build something small enough to actually use, but meaningful enough to solve a real problem for one person.
Demo
🌐 Live Demo: https://voice2note-one.vercel.app
The demo shows the complete workflow:
Record → Transcribe → Understand → Search → Ask
Code
💻 GitHub: https://github.com/affanraza84/voice2note
The repository contains the complete application, local AI setup, architecture documentation, evaluation tests, and setup instructions.
How I Built It
Voice2Note is built around open-source/open-weight AI, rather than using a proprietary AI API as the core intelligence layer.
The architecture consists of several AI components:
1. Local Speech Recognition
Voice recordings are processed using an open speech-recognition model such as Whisper/faster-whisper.
Audio
↓
Local Speech Model
↓
Transcript
This means the application doesn't need to send a private voice recording to a third-party speech API for its core transcription workflow.
2. Open-Weight LLM
The transcript is passed to a locally running open-weight language model.
The model extracts structured information such as:
- summaries
- tasks
- ideas
- decisions
- people
- topics
- important dates
For example:
{
"summary": "Discussed the college project and deployment plan.",
"tasks": [
"Finish the landing page",
"Deploy the backend"
],
"ideas": [
"Add semantic search to the project"
]
}
3. Local Knowledge Base
The transcripts are chunked and converted into embeddings.
Transcript
↓
Chunking
↓
Embeddings
↓
Vector Store
This allows Voice2Note to understand semantic relationships instead of relying only on keyword matching.
4. Retrieval-Augmented Generation
When the user asks:
"What did I say about my project?"
Voice2Note retrieves relevant sections from previous voice notes and provides them as context to the local LLM.
Question
↓
Embedding
↓
Semantic Retrieval
↓
Relevant Voice Notes
↓
Local LLM
↓
Grounded Answer
The response also points back to the source note so the user can verify where the information came from.
Why Does Open Innovation Matter?
This is the most important part of Voice2Note.
Voice recordings can contain extremely personal information.
A voice note might contain:
- personal plans
- private conversations
- project ideas
- financial information
- family information
- unfinished thoughts
I didn't want the fundamental value of the application to depend on sending that information to a proprietary AI provider.
With an open/local AI architecture, Voice2Note can keep the core AI pipeline under the user's control:
Microphone
↓
Local Speech Model
↓
Local LLM
↓
Local Embeddings
↓
Local Knowledge Base
Open AI made several things possible.
Privacy
The application can process sensitive voice data locally instead of requiring a cloud AI API for its core intelligence.
Model freedom
The application isn't permanently tied to one model provider. The AI provider is abstracted so compatible models can be swapped.
Customization
Because the AI pipeline is under our control, we can customize transcription, extraction, retrieval, prompting, and model behavior around the actual user's workflow.
Lower recurring cost
Once the required models are available locally, the core inference pipeline doesn't require paying a per-request proprietary AI API.
Transparency
I can understand and control the components responsible for processing the user's data instead of treating the AI layer as a black box.
There are trade-offs, of course.
Local AI can require more hardware, model setup can be more complicated, and some closed models may provide better performance for particular tasks.
But for this application, control and privacy are more important than simply choosing the strongest available API.
That's where open innovation made the most sense.
What I Learned
The biggest lesson from building Voice2Note was that AI doesn't have to mean building another chatbot.
The interesting part was designing the pipeline around a real person's workflow:
Raw Voice
↓
Information
↓
Structure
↓
Memory
↓
Action
The AI is valuable because it transforms something the user already does—recording voice notes—into something much more useful.
I also learned how different AI components can work together:
- speech recognition
- structured LLM extraction
- embeddings
- vector search
- RAG
- local inference
- source attribution
rather than treating a single LLM call as the entire application.
Final Thoughts
Voice2Note started with a simple observation:
My friend already had the information. They just couldn't easily find or use it.
So instead of building another place to create more information, I built a tool that helps turn the information they already create into something searchable, structured, and actionable.
That's what Build for a Friend meant to me.
Not building the biggest AI application.
Building something that one real person would actually want to keep using.
Top comments (0)