DEV Community

Elev πŸ’₯
Elev πŸ’₯

Posted on

I turned my sister in law's voice notes into a cookbook that answers in her voice

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I have known Purity for ten years. She was my friend first. Today she is my younger brother's wife.

She is also the cook everyone waits for, and almost everything she makes, she learned from her mother. None of it is written down. It lives in her hands and in the way she explains it: half English, half Igbo, full of small warnings like pour away the first water and fry the fresh tomato first.

So this weekend I asked her for voice notes. One dish per note, explained the way she would if I were standing next to her in the kitchen. She sent eight: jollof rice, egusi soup with uziza seed, okra soup with uziza leaf and ugba, beans and ripe plantain cooked together, bitter leaf soup, abacha, porridge yam with ogiri, and egg sauce the way her elder sister taught her. Forty minutes of her voice.

Legacy Loom turns those voice notes into something the family can use:

  • Every note is written down word for word, with a timestamp on every word.
  • Every note is filed: the dish, the ingredients, the steps, her tricks, and the lines worth keeping exactly as she said them.
  • You can ask it anything. "What is the secret to her jollof?" "Why does she pour away the first water from the beans?" The answer uses only what she actually said, and each fact carries a small number. Tap the number and you hear Purity say it, from that exact second.
  • Her recipes become a printable cookbook. Every page has a QR code. Scan it and she explains the dish herself.
  • Ask about something she never mentioned and it says so plainly, then suggests a question to ask her next time.

When it was ready I sent her the link. She asked it about her own jollof, tapped a number and heard herself. This is what she sent back, exactly as she typed it:

wooow meeeenhn this is good, onyi you did a great job, i love this

What got her was hearing her own voice and seeing her mother's tricks written down as tips. That was the whole point.

Demo

Purity's archive (listen and ask, no login): https://legacy-loom-ashen.vercel.app/app

Try it with your own voice: https://legacy-loom-ashen.vercel.app/app?space=try
Record yourself or upload any voice note. No passcode. It is kept apart from Purity's archive and cleared after 24 hours.

Her printable cookbook: https://legacy-loom-ashen.vercel.app/book

Five minute guide: https://legacy-loom-ashen.vercel.app/guide

Three questions worth asking her archive. Each link opens the app and asks it for you:

  1. What did Purity learn from her mother?
  2. What is the secret to her jollof rice? The answer quotes her in Igbo.
  3. What did she say about her wedding? She never talked about it, so watch it say exactly that.

The Legacy Loom landing page: Keep their voice. Ask it anything.

On a phone: asking the secret to her jollof rice. The answer quotes Purity in Igbo, with numbers you tap to hear her.

Her okra soup note, filed, with her own words kept exactly: Everything I know today is my mom that taught me.

The first pages of Purity's Kitchen, her printable cookbook.

Code

GitHub logo angelraph / legacy-loom

A private archive of one person's voice notes, built on open models (Gemma, Whisper, TabPFN) with MongoDB Atlas

Legacy Loom

Live: https://legacy-loom-ashen.vercel.app
App: https://legacy-loom-ashen.vercel.app/app
Docs: https://legacy-loom-ashen.vercel.app/docs

Legacy Loom keeps one person's voice notes and makes them useful. You drop in the voice notes a friend sends you (the recipes, the stories, the "this is how my mum did it"), and it turns them into a private archive you can ask questions of. Every answer links to the second they said it. Recipes become a printable cookbook, and each page has a QR code that plays the cook explaining the dish.

Open models do the core work: Gemma reads, files and answers, Whisper listens, and an open embedding model searches. Gemma and Whisper can run on a laptop with no GPU, so the voice notes don't have to leave the machine.

What it does

Ask. Gemma picks a tool (search the memories, look up a recipe, read the timeline, or ask the call planner), runs it, and answers only…

How I Built It

One voice note, start to finish

  1. Store. The audio goes into MongoDB Atlas (GridFS) and is streamed back with byte ranges. That is what lets the player jump straight to a cited second.
  2. Listen. faster-whisper, running the open Whisper weights on my laptop's CPU, writes every word down with its timing. The hosted version uses ElevenLabs Scribe instead. Both return the same shape, so nothing downstream cares which one ran.
  3. File. Gemma reads the transcript and returns JSON: a title, what kind of memory it is, a summary, the people and places, the year if one is said, up to three exact quotes, and a recipe card with ingredients, steps and tips.
  4. Index. The transcript is cut into passages of about 45 seconds that keep their timestamps. Each passage, plus one English overview of the note, is embedded with BAAI/bge-small-en-v1.5 (open, running through fastembed on ONNX) and stored in Atlas.
  5. Ask. A small agent loop. Gemma picks a tool (search_memories, find_recipe, timeline or plan_calls). Search runs Atlas Vector Search and Atlas Search side by side and merges them with reciprocal rank fusion, so a passage that ranks well for meaning and for exact words rises to the top. Gemma answers only from the numbered passages and cites them.
  6. Check. Every quote in the answer is compared with the real transcripts before you see it. More on why below.

Gemma, in two places

On my laptop (16 GB of RAM, no GPU) it runs Gemma 3 4B through Ollama. Filing a five minute note takes a minute or two on the CPU, which is fine in the background. The hosted demo uses Gemma 4 26B A4B through Google AI Studio, because a free serverless function cannot hold a 4B model. Same prompts, same code, one setting switches between them.

Two things I learned the hard way:

  • Gemma 4 thinks out loud by default. Its reasoning used around 300 tokens per answer and once used up the whole output budget before it wrote the recipe card. Setting its thinking level to minimal brought answers down to about three seconds.
  • The small model sometimes labels a note "recipe" and then leaves the recipe card empty. When that happens, Legacy Loom asks Gemma a second, narrower question for just the card.

The rule it never bends: no invented quotes

Early on, testing with a short clip, the 4B model wrote this:

She always said, "It's a good, warming soup."

Nobody ever said that. The model was trying to be warm. In an archive of someone's voice, that is the one mistake you cannot make, so now every quoted span is checked against what was actually said. A sentence quoting words that are not in any transcript is removed, and the answer shows Removed 1 quote that was never said.

Then Purity's real notes found a sneakier version. Gemma 4 joined two true sentences into one quote:

"Everything I know today is my mom that taught me that uziza leaf and the uziza seed gives it a very sweet flavor"

She said both halves, minutes apart. She never said them together. My first checker glued all her saved quotes into one block of text, so the stitched quote passed. Now each passage is checked on its own, Gemma is told to quote one sentence at a time, and there are tests for a real quote, an invented one and a stitched one.

TabPFN: the best time to call

Every recording is a session: what hour, what day, who asked, what topic, how long she spoke. I rate each session from one to five. TabPFN, the tabular foundation model from Prior Labs, is built for exactly this size of table: a few dozen rows, no training loop of my own. It learns which days, hours and topics tend to give the recordings we treasure, scores every hour of the coming week, and a second TabPFN model estimates how many minutes she will talk.

When I first published this, Purity had sent five notes, all rated five stars, and the planner refused to guess. It needs at least eight rated sessions, with some that went less well. She sent three more that afternoon. I rated them honestly (egg sauce 3, abacha 4, porridge yam 3) and TabPFN switched on.

Its first real forecast: call her around 1 PM, with a 91% chance of a keeper recording, and expect about five minutes of talk. That lines up with what happened. Her five notes from one o'clock all got five stars, and the later afternoon ones I rated lower. It cannot tell weekdays apart yet, because every note so far was recorded on the same Sunday. It runs through the Prior Labs API, and the open tabpfn package works too.

Making it work on free hosting

  • The bundle was too big. The full app came to 554 MB against Vercel's 500 MB limit, because the TabPFN client brings pandas, pyarrow, scipy and scikit learn with it. The planner now runs as its own small Vercel project, and the main site forwards /api/planner to it.
  • Serverless functions stop when they reply, so on Vercel a voice note is transcribed, filed and indexed inside the upload request. A five minute note takes about 30 seconds.
  • Two spaces. Purity's archive is read only without a passcode, so nobody can delete her voice. A separate Try it yourself space is open to everyone, kept apart in every query and search filter, and cleared after 24 hours.
  • Read aloud uses ElevenLabs. Each clip is cached, and new clips are limited to three per visitor per day so the free credits last through judging. Purity's own recordings are never limited.

Why Does Open Innovation Matter?

Because these are a family's recipes in a woman's own voice. On my laptop, with Ollama, Whisper and the open embedding model, the listening, filing and answering all happen on the machine. Her recordings never have to leave it. The hosted demo trades that for convenience so you can try it in a browser, and the code runs either way.

Because Purity speaks Igbo. She moves between English and Igbo in the middle of a sentence. Whisper does not list Igbo among its languages. Because every piece of the pipeline is open and separate, I could put ElevenLabs Scribe in for her notes without touching anything else. The transcript, Gemma, the search and the quote check all work the same. When a model does not fit the person, you change the model, not the person.

Because I could see what went wrong. When the small model invented a quote, I could read the prompt, the output and the transcript side by side, change my own code, and write a test so it never comes back. Running the same prompts on Gemma 3 4B and Gemma 4 26B also showed me which mistakes were the model's size and which were my prompt.

Because it costs nothing to keep. Locally it is free. Hosted, it fits inside free tiers. A family archive should not need a subscription to stay alive.

Where the open path cost me something: speed. On my CPU, Gemma 3 4B takes a minute or two to file one note, and the hosted Gemma 4 does it in seconds. For a demo the hosted path is the right call. For Purity's own copy, I would pick the laptop.

My Agent Session

I built Legacy Loom with Claude Code as my pair programmer over the weekend: planning, the pipeline, the debugging stories above, and the deploys.

Prize Categories

  • Best Use of Gemma
  • Best Use of TabPFN
  • Best Use of MongoDB Atlas
  • Best Use of ElevenLabs

Update, Sunday evening: after this went live, Purity sent three more voice notes: egg sauce, abacha and porridge yam. Her archive now holds eight recordings and forty minutes of her voice, and the TabPFN planner is forecasting from real ratings (see above).

Thank you, Purity, for trusting me with your mother's recipes and your voice. And thank you, Silver, who sent five recordings of her own the same afternoon and reminded me this is not a one person problem.

Top comments (13)

Collapse
 
uzoechi_franklin_51aab075 profile image
Uzoechi Franklin •

This is the kind of AI project that actually makes sense.

Turning something as personal as a family voice note into a living archive, while keeping the original voice and exact context attached to every answer, is a beautiful idea.

The β€œno invented quotes” rule is especially important. You’re preserving memories, not generating content.

Really well done.

Collapse
 
raphelevator profile image
Elev πŸ’₯ •

Thank you great mind

Collapse
 
atugbokoh_charles_3eb94bd profile image
Atugbokoh Charles •

Really nice

Collapse
 
ugwueze_deborah_1a81a2b79 profile image
ugwueze Deborah •

Nice one

Collapse
 
uzoma_onwusiribe_5a548d85 profile image
Uzoma Onwusiribe •

❀️❀️❀️ Lovely

Collapse
 
corleone_ profile image
Corleone •

Solid one.
Keep it up πŸ‘

Collapse
 
mrnetwork001 profile image
MrNetwork •

This is a very good one, bro. Keep it up

Collapse
 
gozie_uzoechi_7027fedb414 profile image
Gozie Uzoechi •

I really appreciate the work you put into developing this Ai web app.

Collapse
 
amaka_tessy_11285f281ba82 profile image
Amaka Tessy •

Thank you for creating such a useful AI web app.

Collapse
 
wireless1 profile image
Wireless •

This is great bro. I've seen a secret learning environment. Thanks for this.

Collapse
 
boyante_83bc599bca3d0925a profile image
Boyante •

Definitely not a one person problem. GG

Collapse
 
raphelevator profile image
Elev πŸ’₯ •

Thank you big name