This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My friend sends a lot of voice messages.
The problem isn't receiving them. The problem is remembering what was actually said later.
A voice note might contain a task, a person's name, a plan for next week, an important date, or something that needs to be done later.
But once the conversation gets buried under dozens of other messages, finding that information again becomes surprisingly difficult.
So I built Voice → Life.
- Voice → Life turns voice messages into structured, searchable memories using open-source tools, an open-weight language model, and local AI inference. Instead of manually listening through old voice notes, you can give Voice → Life a recording and let it:
Voice Note
↓
Speech-to-Text
↓
AI Understanding
↓
Memory Extraction
↓
SQLite
↓
Timeline / Search / Questions
For example, a voice note like:
"Kal 11 baje Rahul ko internship ke documents bhej dena. Aur Monday ko uska interview hai."
can become memories such as:
- Task: Send internship documents
- Person: Rahul
- Time: 11:00
- Event: Interview
- Person: Rahul
- Date: Monday
And instead of remembering which voice note contained the information, you can ask:
"What did I need to send Rahul?"
and Voice → Life answers using the stored memories.
Why I built it
I wanted to build something that solves an actual everyday problem for a friend, rather than making another generic chatbot.
Voice messages are convenient for communicating.
They're terrible as a database.
Voice → Life is my attempt to bridge that gap.
Demo
Demo Video: Watch the Voice → Life demo
The demo shows the complete flow:
- Upload a voice message
- Transcribe it
- Let the AI understand the transcript
- Extract useful memories
- Store them
- View them on the timeline
- Search memories
- Ask questions about them
Code
The complete project is open source:
GitHub:
🎙️ Voice → Life
Turn your voice into useful memories.
Voice → Life is an open-source, local-first AI application that transforms voice notes into structured, searchable memories and actionable information.
Instead of letting important information get buried inside voice recordings, Voice → Life uses open-source AI to understand what was said, extract useful information, and let the user decide what should be remembered.
✨ What It Does
Give Voice → Life a voice note like:
"Bro kal 11 baje Rahul ko internship ke documents bhej dena. Aur usko bol dena ki interview Monday ko hai."
It can turn that into:
🔔 Task
Send internship documents
👤 Person
Rahul
📅 Event
Interview — Monday
The user can review the extracted information and choose what to save.
Later:
"What did I need to send Rahul?"
Voice → Life can retrieve the relevant memory.
🎯 Core Idea
🎙 Voice
↓
📝 Transcription…
The project is built so that the core AI pipeline can run locally rather than depending on a proprietary AI API.
How I Built It
The stack is intentionally simple.
Architecture:
┌──────────────────────┐
│ Frontend │
│ Next.js + TypeScript │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ FastAPI │
│ Backend │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ faster-whisper │
│ Speech-to-Text │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Ollama + Open Model │
│ AI Understanding │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Memory Extraction │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ SQLite │
└──────────┬───────────┘
│
▼
Timeline / Search
/ Q&A
Speech-to-Text
For transcription, I use faster-whisper.
This became particularly interesting when I started testing the project with real voice messages.
Normal benchmark-style English isn't the only thing a personal assistant needs to understand.
Real people mix languages.
My voice-note tests included Hinglish (Hindi + English) and code-switched speech, which exposed weaknesses in the initial lightweight Whisper configuration.
That forced me to treat transcription quality as an actual engineering problem rather than assuming:
audio → text = solved.
Open-Weight AI
The understanding layer uses an open-weight language model through Ollama.
The model receives the transcript and extracts useful information such as:
- tasks
- events
- people
- places
- plans
- facts
- reminders The model is explicitly instructed not to invent information that isn't present in the transcript.
Memory Storage
The extracted information is stored in SQLite.
I deliberately didn't start with a complicated vector database or large RAG pipeline.
For the MVP, structured memories are enough.
That also keeps the project easy to run and understand.
Asking Your Memories
The final layer lets users ask natural-language questions about their memories.
For example:
"What did I need to send Rahul?"
The relevant stored memories are passed to the language model, which generates the answer.
The model isn't supposed to make up an answer when the information isn't present.
Why Does Open Innovation Matter?
This project deals with something much more personal than a normal chatbot:
people's voice messages
Voice notes can contain conversations, names, plans, schedules, personal information and things people simply don't want uploaded somewhere else.
That's one of the reasons open-source AI mattered to me while building Voice → Life.
The core AI components can run locally:
- faster-whisper handles speech recognition
- Ollama provides local model inference
- an open-weight LLM handles understanding
- SQLite stores the memories locally That changes the relationship between the user and the AI. Instead of:
My private voice
↓
Third-party API
↓
AI service
the goal is:
My private voice
↓
My machine
↓
Open AI models
↓
My memories
Open models also give developers much more freedom to experiment.
If transcription isn't good enough for Hinglish, I can change the model or configuration.
If the language model isn't performing well, I can replace it.
I'm not locked into one proprietary API or one provider's ecosystem.
For a project built around personal memories, that control matters.
What I Learned
The biggest lesson from this project was that building an AI application isn't just about putting an LLM in the middle of an architecture.
The interesting problems appeared at the boundaries:
Audio → Speech
Is the transcription actually accurate?
Speech → Meaning
Can the model distinguish a task from casual conversation?
Meaning → Memory
What information is actually worth keeping?
Memory → Answer
Can the system answer a question without hallucinating?
And most importantly:
Does the final result solve the original person's problem?
That's what I wanted to explore with Voice → Life.
What's Next?
The current version is intentionally an MVP.
Some ideas for future versions include:
- better multilingual and Hinglish support
- WhatsApp voice-message integration
- task/reminder generation
- notifications
- stronger semantic search
- local-first privacy controls
- better confirmation/editing before memories are saved But I intentionally stopped the first version before turning it into a giant AI assistant. The core idea had to work first. Voice in → useful memory out.
Top comments (0)