What I Built
I built the Nostalgia Cookbook, a privacy-first web application designed specifically for my grandparent, who loves cooking but only preserves recipes in their head or via lengthy, disjointed voice memos.
The application provides a seamless path to digital archiving. Users can either upload raw voice files (.mp3, .wav, .m4a) or record their voices directly inside a vintage, kitchen-card-styled UI. The application transcribes the audio and uses a hybrid AI system to separate the recipe (ingredients, metric/imperial measurements, structured steps) from the invaluable personal stories, historical context, and anecdotes woven into the recording.
Demo
Code
How I Built It
The application is built using a modern Next.js (App Router) frontend and backend architecture containerized with Docker, backed by MongoDB Atlas for data persistence.
To maximize accuracy while strictly protecting data privacy, I engineered a Hybrid AI Pipeline:
-
Multimodal Ingestion Layer: The raw audio file is initially parsed using the Gemini API (
gemini-2.5-flash) via the@google/genaiSDK to produce a raw, verbatim stream of consciousness text block from the voice recording. - Open-Source Processing Core: This unedited text transcript is then piped locally into Gemma 2 (9B) running on an open-source inference runner (Ollama). Gemma acts as our structured data extraction engine, isolating the recipe formatting from the personal narrative elements.
Why Does Open Innovation Matter?
Open innovation and open-weight models like Gemma 2 change the paradigm of personal applications for three reasons:
- Absolute Family Data Privacy: Voice recordings and transcriptions contain deeply sensitive family history, names, dates, and locations. Processing these texts through our self-hosted open-weight model means these intimate family memories never live on an external server or get consumed by corporate public training loops.
- Permanent Determinism: Closed third-party APIs constantly deprecate or alter their internal alignments, breaking downstream parsing code. By using Gemma 2, the core functionality of this heirloom tool remains structurally unchanged indefinitely.
- Zero Ingestion Costs: Processing long, rambling family voice recordings creates massive token lengths. Running an open-weight model on cloud infrastructure eliminates variable token fee volatility.
My Agent Session
My complete coding, debugging, and environment setup journey can be inspected via my development workflow updates.
Prize Categories
- Built with Gemma: Uses Google's Gemma 2 (9B) open-weight model as the primary formatting engine.
- DigitalOcean Deployment: Fully dockerized and deployed across the DigitalOcean App Platform and a custom GPU Droplet.
- MongoDB Atlas Integration: Leverages MongoDB Atlas for robust storage of parsed family recipes and narrative timeline logs.
Top comments (0)