DEV Community

Kushagra Bhardwaj
Kushagra Bhardwaj

Posted on

I Built an AI That Turns My Friend's Voice Notes Into Searchable Memories

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My friend sends a lot of voice messages.
The problem isn't receiving them. The problem is remembering what was actually said later.
A voice note might contain a task, a person's name, a plan for next week, an important date, or something that needs to be done later.
But once the conversation gets buried under dozens of other messages, finding that information again becomes surprisingly difficult.
So I built Voice → Life.

  • Voice → Life turns voice messages into structured, searchable memories using open-source tools, an open-weight language model, and local AI inference. Instead of manually listening through old voice notes, you can give Voice → Life a recording and let it:
Voice Note
    ↓
Speech-to-Text
    ↓
AI Understanding
    ↓
Memory Extraction
    ↓
SQLite
    ↓
Timeline / Search / Questions
Enter fullscreen mode Exit fullscreen mode

For example, a voice note like:

"Kal 11 baje Rahul ko internship ke documents bhej dena. Aur Monday ko uska interview hai."
Enter fullscreen mode Exit fullscreen mode

can become memories such as:

  • Task: Send internship documents
  • Person: Rahul
  • Time: 11:00
  • Event: Interview
  • Person: Rahul
  • Date: Monday

And instead of remembering which voice note contained the information, you can ask:

"What did I need to send Rahul?"
Enter fullscreen mode Exit fullscreen mode

and Voice → Life answers using the stored memories.

Why I built it

I wanted to build something that solves an actual everyday problem for a friend, rather than making another generic chatbot.
Voice messages are convenient for communicating.
They're terrible as a database.
Voice → Life is my attempt to bridge that gap.

Demo

Demo Video: Watch the Voice → Life demo

The demo shows the complete flow:

  1. Upload a voice message
  2. Transcribe it
  3. Let the AI understand the transcript
  4. Extract useful memories
  5. Store them
  6. View them on the timeline
  7. Search memories
  8. Ask questions about them

Code

The complete project is open source:
GitHub:

🎙️ Voice → Life

Turn your voice into useful memories.

Voice → Life is an open-source, local-first AI application that transforms voice notes into structured, searchable memories and actionable information.

Instead of letting important information get buried inside voice recordings, Voice → Life uses open-source AI to understand what was said, extract useful information, and let the user decide what should be remembered.


✨ What It Does

Give Voice → Life a voice note like:

"Bro kal 11 baje Rahul ko internship ke documents bhej dena. Aur usko bol dena ki interview Monday ko hai."

It can turn that into:

🔔 Task
Send internship documents

👤 Person
Rahul

📅 Event
Interview — Monday

The user can review the extracted information and choose what to save.

Later:

"What did I need to send Rahul?"

Voice → Life can retrieve the relevant memory.


🎯 Core Idea

🎙 Voice
   ↓
📝 Transcription
…



The project is built so that the core AI pipeline can run locally rather than depending on a proprietary AI API.

How I Built It

The stack is intentionally simple.

Architecture:

┌──────────────────────┐
│      Frontend        │
│ Next.js + TypeScript │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│       FastAPI        │
│      Backend         │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│    faster-whisper    │
│   Speech-to-Text     │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│ Ollama + Open Model  │
│   AI Understanding   │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│ Memory Extraction    │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│       SQLite         │
└──────────┬───────────┘
           │
           ▼
    Timeline / Search
        / Q&A
Enter fullscreen mode Exit fullscreen mode

Speech-to-Text

For transcription, I use faster-whisper.
This became particularly interesting when I started testing the project with real voice messages.
Normal benchmark-style English isn't the only thing a personal assistant needs to understand.
Real people mix languages.
My voice-note tests included Hinglish (Hindi + English) and code-switched speech, which exposed weaknesses in the initial lightweight Whisper configuration.
That forced me to treat transcription quality as an actual engineering problem rather than assuming:

audio → text = solved.
Enter fullscreen mode Exit fullscreen mode

Open-Weight AI

The understanding layer uses an open-weight language model through Ollama.
The model receives the transcript and extracts useful information such as:

  • tasks
  • events
  • people
  • places
  • plans
  • facts
  • reminders The model is explicitly instructed not to invent information that isn't present in the transcript.

Memory Storage

The extracted information is stored in SQLite.
I deliberately didn't start with a complicated vector database or large RAG pipeline.
For the MVP, structured memories are enough.
That also keeps the project easy to run and understand.

Asking Your Memories

The final layer lets users ask natural-language questions about their memories.
For example:

"What did I need to send Rahul?"
Enter fullscreen mode Exit fullscreen mode

The relevant stored memories are passed to the language model, which generates the answer.
The model isn't supposed to make up an answer when the information isn't present.

Why Does Open Innovation Matter?

This project deals with something much more personal than a normal chatbot:
people's voice messages
Voice notes can contain conversations, names, plans, schedules, personal information and things people simply don't want uploaded somewhere else.
That's one of the reasons open-source AI mattered to me while building Voice → Life.
The core AI components can run locally:

  • faster-whisper handles speech recognition
  • Ollama provides local model inference
  • an open-weight LLM handles understanding
  • SQLite stores the memories locally That changes the relationship between the user and the AI. Instead of:
My private voice
       ↓
Third-party API
       ↓
AI service
Enter fullscreen mode Exit fullscreen mode

the goal is:

My private voice
       ↓
My machine
       ↓
Open AI models
       ↓
My memories
Enter fullscreen mode Exit fullscreen mode

Open models also give developers much more freedom to experiment.
If transcription isn't good enough for Hinglish, I can change the model or configuration.
If the language model isn't performing well, I can replace it.
I'm not locked into one proprietary API or one provider's ecosystem.
For a project built around personal memories, that control matters.

What I Learned

The biggest lesson from this project was that building an AI application isn't just about putting an LLM in the middle of an architecture.
The interesting problems appeared at the boundaries:
Audio → Speech
Is the transcription actually accurate?
Speech → Meaning
Can the model distinguish a task from casual conversation?
Meaning → Memory
What information is actually worth keeping?
Memory → Answer
Can the system answer a question without hallucinating?
And most importantly:
Does the final result solve the original person's problem?
That's what I wanted to explore with Voice → Life.

What's Next?

The current version is intentionally an MVP.
Some ideas for future versions include:

  • better multilingual and Hinglish support
  • WhatsApp voice-message integration
  • task/reminder generation
  • notifications
  • stronger semantic search
  • local-first privacy controls
  • better confirmation/editing before memories are saved But I intentionally stopped the first version before turning it into a giant AI assistant. The core idea had to work first. Voice in → useful memory out.

Top comments (0)