<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kushagra Bhardwaj</title>
    <description>The latest articles on DEV Community by Kushagra Bhardwaj (@kush05bhardwaj).</description>
    <link>https://dev.to/kush05bhardwaj</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2213235%2F3a7b64c3-f858-4278-88bb-ec71989eb23e.jpg</url>
      <title>DEV Community: Kushagra Bhardwaj</title>
      <link>https://dev.to/kush05bhardwaj</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kush05bhardwaj"/>
    <language>en</language>
    <item>
      <title>I Built an AI That Turns My Friend's Voice Notes Into Searchable Memories</title>
      <dc:creator>Kushagra Bhardwaj</dc:creator>
      <pubDate>Sun, 04 Oct 2026 17:12:38 +0000</pubDate>
      <link>https://dev.to/kush05bhardwaj/i-built-an-ai-that-turns-my-friends-voice-notes-into-searchable-memories-mld</link>
      <guid>https://dev.to/kush05bhardwaj/i-built-an-ai-that-turns-my-friends-voice-notes-into-searchable-memories-mld</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;My friend sends a lot of voice messages.&lt;br&gt;
The problem isn't receiving them. The problem is remembering what was actually said later.&lt;br&gt;
A voice note might contain a task, a person's name, a plan for next week, an important date, or something that needs to be done later.&lt;br&gt;
But once the conversation gets buried under dozens of other messages, finding that information again becomes surprisingly difficult.&lt;br&gt;
So I built Voice → Life.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Voice → Life turns voice messages into structured, searchable memories using open-source tools, an open-weight language model, and local AI inference.
Instead of manually listening through old voice notes, you can give Voice → Life a recording and let it:
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Voice Note
    ↓
Speech-to-Text
    ↓
AI Understanding
    ↓
Memory Extraction
    ↓
SQLite
    ↓
Timeline / Search / Questions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;For example, a voice note like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Kal 11 baje Rahul ko internship ke documents bhej dena. Aur Monday ko uska interview hai."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;can become memories such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Task: Send internship documents&lt;/li&gt;
&lt;li&gt;Person: Rahul&lt;/li&gt;
&lt;li&gt;Time: 11:00&lt;/li&gt;
&lt;li&gt;Event: Interview&lt;/li&gt;
&lt;li&gt;Person: Rahul&lt;/li&gt;
&lt;li&gt;Date: Monday&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And instead of remembering which voice note contained the information, you can ask:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"What did I need to send Rahul?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;and Voice → Life answers using the stored memories.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why I built it
&lt;/h2&gt;

&lt;p&gt;I wanted to build something that solves an actual everyday problem for a friend, rather than making another generic chatbot.&lt;br&gt;
Voice messages are convenient for communicating.&lt;br&gt;
They're terrible as a database.&lt;br&gt;
Voice → Life is my attempt to bridge that gap.&lt;/p&gt;
&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Demo Video:&lt;/strong&gt; &lt;a href="https://drive.google.com/file/d/18HTWVgH5B1oJxqPePSpBBwUPjxsMUMNs/view?usp=sharing" rel="noopener noreferrer"&gt;Watch the Voice → Life demo&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The demo shows the complete flow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Upload a voice message&lt;/li&gt;
&lt;li&gt;Transcribe it&lt;/li&gt;
&lt;li&gt;Let the AI understand the transcript&lt;/li&gt;
&lt;li&gt;Extract useful memories&lt;/li&gt;
&lt;li&gt;Store them&lt;/li&gt;
&lt;li&gt;View them on the timeline&lt;/li&gt;
&lt;li&gt;Search memories&lt;/li&gt;
&lt;li&gt;Ask questions about them&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;The complete project is open source:&lt;br&gt;
GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Kush05Bhardwaj" rel="noopener noreferrer"&gt;
        Kush05Bhardwaj
      &lt;/a&gt; / &lt;a href="https://github.com/Kush05Bhardwaj/Voice-Life" rel="noopener noreferrer"&gt;
        Voice-Life
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;🎙️ Voice → Life&lt;/h1&gt;
&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Turn your voice into useful memories.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Voice → Life is an open-source, local-first AI application that transforms voice notes into &lt;strong&gt;structured, searchable memories and actionable information&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead of letting important information get buried inside voice recordings, Voice → Life uses open-source AI to understand what was said, extract useful information, and let the user decide what should be remembered.&lt;/p&gt;




&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;✨ What It Does&lt;/h2&gt;
&lt;/div&gt;

&lt;p&gt;Give Voice → Life a voice note like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Bro kal 11 baje Rahul ko internship ke documents bhej dena. Aur usko bol dena ki interview Monday ko hai."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It can turn that into:&lt;/p&gt;

&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;🔔 Task
Send internship documents

👤 Person
Rahul

📅 Event
Interview — Monday
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The user can review the extracted information and choose what to save.&lt;/p&gt;
&lt;p&gt;Later:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"What did I need to send Rahul?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Voice → Life can retrieve the relevant memory.&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;🎯 Core Idea&lt;/h2&gt;

&lt;/div&gt;

&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;
&lt;pre class="notranslate"&gt;&lt;code&gt;🎙 Voice
   ↓
📝 Transcription&lt;/code&gt;&lt;/pre&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Kush05Bhardwaj/Voice-Life" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;&lt;br&gt;&lt;br&gt;
The project is built so that the core AI pipeline can run locally rather than depending on a proprietary AI API.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The stack is intentionally simple.&lt;/p&gt;

&lt;h3&gt;
  
  
  Architecture:
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌──────────────────────┐
│      Frontend        │
│ Next.js + TypeScript │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│       FastAPI        │
│      Backend         │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│    faster-whisper    │
│   Speech-to-Text     │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│ Ollama + Open Model  │
│   AI Understanding   │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│ Memory Extraction    │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│       SQLite         │
└──────────┬───────────┘
           │
           ▼
    Timeline / Search
        / Q&amp;amp;A
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Speech-to-Text
&lt;/h4&gt;

&lt;p&gt;For transcription, I use faster-whisper.&lt;br&gt;
This became particularly interesting when I started testing the project with real voice messages.&lt;br&gt;
Normal benchmark-style English isn't the only thing a personal assistant needs to understand.&lt;br&gt;
Real people mix languages.&lt;br&gt;
My voice-note tests included Hinglish (Hindi + English) and code-switched speech, which exposed weaknesses in the initial lightweight Whisper configuration.&lt;br&gt;
That forced me to treat transcription quality as an actual engineering problem rather than assuming:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;audio → text = solved.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Open-Weight AI
&lt;/h4&gt;

&lt;p&gt;The understanding layer uses an open-weight language model through Ollama.&lt;br&gt;
The model receives the transcript and extracts useful information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tasks&lt;/li&gt;
&lt;li&gt;events&lt;/li&gt;
&lt;li&gt;people&lt;/li&gt;
&lt;li&gt;places&lt;/li&gt;
&lt;li&gt;plans&lt;/li&gt;
&lt;li&gt;facts&lt;/li&gt;
&lt;li&gt;reminders
The model is explicitly instructed not to invent information that isn't present in the transcript.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  Memory Storage
&lt;/h4&gt;

&lt;p&gt;The extracted information is stored in SQLite.&lt;br&gt;
I deliberately didn't start with a complicated vector database or large RAG pipeline.&lt;br&gt;
For the MVP, structured memories are enough.&lt;br&gt;
That also keeps the project easy to run and understand.&lt;/p&gt;
&lt;h4&gt;
  
  
  Asking Your Memories
&lt;/h4&gt;

&lt;p&gt;The final layer lets users ask natural-language questions about their memories.&lt;br&gt;
For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"What did I need to send Rahul?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The relevant stored memories are passed to the language model, which generates the answer.&lt;br&gt;
The model isn't supposed to make up an answer when the information isn't present.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;



&lt;p&gt;This project deals with something much more personal than a normal chatbot:&lt;br&gt;
&lt;strong&gt;people's voice messages&lt;/strong&gt;&lt;br&gt;
Voice notes can contain conversations, names, plans, schedules, personal information and things people simply don't want uploaded somewhere else.&lt;br&gt;
That's one of the reasons open-source AI mattered to me while building Voice → Life.&lt;br&gt;
The core AI components can run locally:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;faster-whisper handles speech recognition&lt;/li&gt;
&lt;li&gt;Ollama provides local model inference&lt;/li&gt;
&lt;li&gt;an open-weight LLM handles understanding&lt;/li&gt;
&lt;li&gt;SQLite stores the memories locally
That changes the relationship between the user and the AI.
Instead of:
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;My private voice
       ↓
Third-party API
       ↓
AI service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;the goal is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;My private voice
       ↓
My machine
       ↓
Open AI models
       ↓
My memories
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open models also give developers much more freedom to experiment.&lt;br&gt;
If transcription isn't good enough for Hinglish, I can change the model or configuration.&lt;br&gt;
If the language model isn't performing well, I can replace it.&lt;br&gt;
I'm not locked into one proprietary API or one provider's ecosystem.&lt;br&gt;
For a project built around personal memories, that control matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;The biggest lesson from this project was that building an AI application isn't just about putting an LLM in the middle of an architecture.&lt;br&gt;
The interesting problems appeared at the boundaries:&lt;br&gt;
Audio → Speech&lt;br&gt;
Is the transcription actually accurate?&lt;br&gt;
Speech → Meaning&lt;br&gt;
Can the model distinguish a task from casual conversation?&lt;br&gt;
Meaning → Memory&lt;br&gt;
What information is actually worth keeping?&lt;br&gt;
Memory → Answer&lt;br&gt;
Can the system answer a question without hallucinating?&lt;br&gt;
And most importantly:&lt;br&gt;
Does the final result solve the original person's problem?&lt;br&gt;
That's what I wanted to explore with Voice → Life.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next?
&lt;/h2&gt;

&lt;p&gt;The current version is intentionally an MVP.&lt;br&gt;
Some ideas for future versions include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;better multilingual and Hinglish support&lt;/li&gt;
&lt;li&gt;WhatsApp voice-message integration&lt;/li&gt;
&lt;li&gt;task/reminder generation&lt;/li&gt;
&lt;li&gt;notifications&lt;/li&gt;
&lt;li&gt;stronger semantic search&lt;/li&gt;
&lt;li&gt;local-first privacy controls&lt;/li&gt;
&lt;li&gt;better confirmation/editing before memories are saved
But I intentionally stopped the first version before turning it into a giant AI assistant.
The core idea had to work first.
Voice in → useful memory out.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
