This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built Voice2Memory for my friend Alex, a busy founder who records dozens of audio voice notes every week while commuting and moving between meetings.
The Problem
Voice notes are write-only memory. Finding a specific task, deadline, or promise later required Alex to replay minutes of audio at 1x speed. As a result, critical commitments got forgotten, deadlines were missed, and useful ideas remained trapped inside unsearchable audio files.
What It Does & How It Solves It
Voice2Memory automatically converts raw voice notes into structured, searchable personal memories:
- Actionable Tasks: Extracts clear to-do items from spoken thoughts.
- Dates & Deadlines: Identifies exact meeting times and commitments.
- People & Topics: Detects names mentioned and categorizes notes with tags.
- Searchable Vault: Lets Alex search across past recordings instantly instead of listening back to audio files.
Demo
- Live Application: https://voice2memory.onrender.com
- Video Walkthrough: https://drive.google.com/file/d/1dZRGLBKB8y0aG1TXI1AOgyJMCSVXrNZY/view?usp=sharing
Code
hitesh-kumar123
/
Voice2Memory
Transform raw voice notes into structured, actionable personal memories with Google Gemma 2 open-weight AI, Faster-Whisper, and MongoDB Atlas.
Voice2Memory
Transform raw voice notes into structured, actionable personal memories using open-weight AI.
Built for the Hacktoberfest 2026 DEV Weekend Challenge: "Build for a Friend".
Overview
The Friend Problem
My friend records dozens of voice notes throughout the week—while commuting, walking between meetings, or capturing sudden startup ideas and daily errands.
A typical spoken recording sounds like this:
"Hey Sarah, I'm heading over to the office now. Please remember to finalize the Q4 investor pitch deck with Alex by Thursday at 4 PM. We also need to schedule the follow-up meeting with Priya for next Monday afternoon. On my way back I have to order a new microphone for our podcast setup and email the contract to Tom."
The problem? Voice notes are write-only memory.
Days later, finding a specific deadline, promise, or action item requires replaying minutes of audio at 1x speed. Crucial commitments get lost, deadlines slip…
How I Built It
Voice2Memory is built around local open-weight AI models to process voice notes privately and deterministically:
-
Open-Weight AI Brain (Google Gemma 2): Used Google's
gemma2:2binstruction-tuned open-weight model running locally via Ollama. It takes raw transcripts and extracts structured JSON containing executive summaries, action items, dates, people, and topics. -
Speech-to-Text (Faster-Whisper): Used
faster-whisper(CTranslate2int8backend) with Voice Activity Detection (VAD) for fast, offline transcription directly from audio buffers. - Database & Storage (MongoDB Atlas): Used MongoDB Atlas with weighted text search indexes to store and instantly query structured memories by topic, name, or keyword, with live task completion persistence.
- Frontend & App Framework: Built with Next.js 16 (App Router), React 19, TypeScript, and Tailwind CSS.
Why Does Open Innovation Matter?
Voice notes are deeply personal—they contain unfiltered thoughts, private work updates, financial decisions, and personal appointments.
Open innovation made three things possible that closed proprietary APIs could not:
- Absolute Privacy: By pairing local Faster-Whisper with Google's open-weight Gemma 2, audio files and transcripts never need to leave the user's infrastructure or be fed into third-party AI training pipelines.
- Zero Recurring API Costs: Closed APIs charge per-minute for audio transcription and per-token for LLM calls, making continuous daily voice logging expensive. Open-weight models run locally at near-zero incremental cost.
- Reliability & Independence: Open-weight models ensure the application will continue running without vendor lock-in, rate limits, or sudden API deprecations.
My Agent Session
Prize Categories
1. Best Use of Gemma ($200)
-
Model Used: Google Gemma 2 (
gemma2:2bopen-weight model). -
How It Was Used: Powers the core intelligence engine in
src/lib/ollama.tsand/api/analyze. It processes raw voice transcripts through custom Gemma instruction templates (<start_of_turn>user ...) and outputs deterministic JSON containing summaries, actionable tasks, temporal dates, people, and topics.
2. Best Use of MongoDB Atlas ($100)
-
How It Was Used: Serves as the central persistent memory store in
src/lib/db.ts. Utilizes weighted compound text search indexes across titles, summaries, transcripts, topics, and people, and supports real-time task checklist synchronization viaPATCH /api/memories/[id].
Thanks for checking out Voice2Memory! Feel free to explore the repository, try the live demo, and share any feedback or thoughts in the comments below.
Top comments (0)