DEV Community

Pratik
Pratik

Posted on

Recipe-maker

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built Family Recipe Memory for a friend whose family cooks almost entirely from memory.

Her grandmother and parents never write recipes down or measure much. They explain them out loud: add some of this, cook it until it smells right, you'll know when it's done.

Those little details are the actual recipe, and they're very easy to lose. Right now, a lot of them live in voice notes buried somewhere in a chat app, often in a mix of Nepali and English, with no practical way to search them.

So the app lets her just talk.

She uploads a voice note, and it:

  • Transcribes it on her own machine
  • Lets her fix any misheard words
  • Turns it into a clean recipe with ingredients, steps, tips, and the family story
  • Saves it to a small recipe library
  • Lets her ask things like "What did Grandma say about making momo?"

One rule I cared about: if something wasn't said in the recording, the app leaves it blank instead of guessing. I didn't want a model quietly rewriting someone's grandmother's recipe.

Demo

[https://www.youtube.com/watch?v=x8TJe7HyIew]

The flow is:

Upload voice note → Transcribe → Review transcript → Extract and save → Browse or ask questions

Each recipe is shown as a card with the story, an ingredients table, numbered steps, and tips. Recipes can also be downloaded as Markdown.

Code


🍲 Family Recipe Memory

Preserve voice recipes and stories from loved ones using Local AI + Backboard Persistent Memory.

Problem

Family culinary heritage is often trapped in unwritten spoken traditions. Audio notes from grandmothers or parents contain precious nuances, but searching through raw voice recordings is frustrating and inefficient.

Solution

Family Recipe Memory turns unstructured voice notes into a searchable family knowledge base. Audio transcription and AI structured parsing run 100% locally, ensuring raw voice recordings never leave your machine. Only extracted text is persisted in Backboard memory for cross-session natural language retrieval.

Architecture & Privacy Model




How I Built It

Everything runs locally.

The stack is Python and Streamlit, with two open models doing the heavy lifting:

  • OpenAI Whisper for speech-to-text. The voice notes mix Nepali and English, and the smallest Whisper model gets shaky on that, so the app defaults to small, lets you switch to medium, and lets you pick the spoken language.
  • Ollama for everything language-related. I tested with llama3.2, gemma2, and qwen2.5, and the app lists whichever models are installed.

The extraction step was the part that took the most care.

The transcript goes to the model with strict instructions: only use what was said, keep dish and ingredient names as spoken, and leave anything that wasn't mentioned empty.

The output is JSON, which I then clean up, because small local models love to return a quantity as an object or write "N/A" instead of leaving a field empty.

Two design choices came directly from using it:

  1. A review step between transcription and extraction. Speech-to-text can mishear a Nepali word now and then, and one wrong word can ruin an ingredient. Letting my friend fix the transcript first made the recipes noticeably better.

  2. Memory is just a local JSON file. Recipes are saved to data/recipes.json. For questions, the app passes the most relevant saved recipes to the local model and tells it to answer only from them, and to say so if the answer isn't there.

I also had a bug early on where the results disappeared the moment I clicked the download button, because Streamlit reruns the script on every click. Moving the state into st.session_state fixed it.

My friend's reaction: [Add your friend's actual feedback here, or remove this line.]

Why Does Open Innovation Matter?

A family voice note isn't just audio. It's someone's voice, their stories, names, and private conversations, often from people who are getting older. I wasn't comfortable sending that to a cloud API just because it was convenient.

Open models meant the whole pipeline—from audio to transcript to recipe to answers—runs on a laptop with no account, no API key, and no per-request cost.

It also meant I could test the extraction prompt as many times as I liked on real recordings without worrying about a bill, and I could swap models freely when one handled Nepali words better than another.

A closed API could have done the extraction too. What it couldn't give me is the same level of control over keeping the original recordings on my friend's computer.

For this project, that mattered more than anything else.

Prize Categories

  • Build for a Friend / Loved One

hf26challenge

Top comments (0)