DEV Community

Cover image for Salt & Story: My Maasi's Recipes, Saved in Her Own Voice
Aditya Kumar
Aditya Kumar

Posted on AI-assisted

Salt & Story: My Maasi's Recipes, Saved in Her Own Voice

Hacktoberfest: Maintainer Spotlight

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

⭐ If this story means something to you, please star the repo: github.com/ADITYA-KUMAR-2358/salt-and-story. It helps more families find it.

What I Built

Who I built it for

My Maasi (my mom's sister, my aunt) has cancer. A few weeks ago we learned she probably has two to three years left.

She makes the best food I know. Not fancy food: comforting, tasty, the kind you remember for the rest of your life.

A few days back I saw her writing in her diary. When I asked, she told me it was her famous chili and lemon pickle recipe. She had written a few others too, each with a lovely message and the name of the person it was for. My favourite thing she cooks is her Dahi Wada, and guess what: there it was, a recipe addressed to me.

She had spent hours writing those pages by hand. So I thought: what if she could just talk, the way she does in the kitchen, and her recipe came out exactly in her words?

That became Salt & Story. Maasi has now recorded those recipes with her voice, and they are live, written the way she said them. I also added a way to turn them into a personalised book, and a lot more, because every recipe has a story.

The problem it solves for her

Writing recipes by hand is slow and tiring, and the way Maasi really cooks ("a pinch of this, cook it till it looks right") never fits neatly on a page anyway. Salt & Story lets her simply speak, and it keeps her exact words instead of flattening them into a generic recipe. For our family, it means her food and her voice stay with us.

What it does

Salt & Story is a web app that turns spoken recipes into a printable family cookbook.

You record a voice memo in the browser (or upload one, or paste the words), and Salt & Story:

  • transcribes it,
  • structures it into ingredients, steps and tips,
  • turns vague measures like "a fistful" into rough quantities while keeping the original phrase right next to them,
  • keeps one memorable line, verbatim, as a pull-quote (checked against the transcript in code),
  • never invents an ingredient, amount or step that wasn't said,
  • exports everything as a designed A5 PDF cookbook.

The goal was never a recipe database. It was a book that sounds like Maasi.

What's in it

Area What you can do
Capture Record in the browser with a live level meter, upload a memo (mp3, m4a, wav, webm, ogg), paste a transcript, or write a recipe by hand
Recipe Read it with an ingredient checklist and tappable steps, edit everything (with an unsaved-changes guard), add a photo, keep the original transcript one tap away
Cook mode One step at a time in large type, the screen kept awake, arrow keys or big buttons
Listen ElevenLabs narration of the whole recipe, cached so it is paid for once
Book Pick and order recipes, a live cover preview, a dedication, an editable "Recipes told by" line, and an A5 PDF with contents and page numbers
Organise Search, favourites, delete with confirmation, backup and restore as JSON
Accounts Sign in with Google; each family's recipes are private to their account; delete your account and everything in it
Start fast Six American and European family classics with openly licensed photos, added in one tap
Comfort Light and dark mode, 17px base text, 44px tap targets, keyboard and screen reader friendly, reduced-motion support

Demo

Live app: https://salt-and-story-web.onrender.com/

Salt & Story landing page with the tagline

A recipe page in Salt & Story showing the photo, a pull-quote in the cook's own words, an ingredient checklist and numbered steps

Cook mode on a phone: a single recipe step in large type with big Previous and Next buttons

A two-page spread from the printed A5 PDF cookbook: photo, title, story and the cook's words on the left; ingredients, method and tips on the right

Code

Salt & Story

Because Every Recipe Has A Story

Salt & Story turns voice memos of a relative cooking into clean, printable recipes, while keeping the speaker's own words beside them. Record or upload a memo, an open-weight model (Gemma 4) writes it up, and you export a cookbook PDF that looks like a gift.

It is built for one real family: the person who tells recipes out loud and has never written them down, and the grandchild who wants to keep them.

Built by Aditya Kumar (@ADITYA-KUMAR-2358) for the DEV Hacktoberfest Weekend Challenge, "Build for a Friend".

Agents and contributors: start with docs/ (PRD, Architecture, Rules, Design, Tasks, Memory).

What it does

  • Capture: record in the browser (live level meter), upload a memo, paste a transcript, or write a recipe by hand.
  • Write it up: ElevenLabs…

⭐ A star on the repo would genuinely make my day.

How I Built It

Open-source AI at the core: Gemma 4 (gemma-4-26b-a4b-it, Google's open-weight model) structures every recipe. I call it through the Gemini API's OpenAI-compatible endpoint on the free tier, so the same code also runs against local Ollama (for example gemma3:4b) by changing three environment variables. The model turns a rambling transcript into structured JSON, and it is the heart of Salt & Story.

Architecture

+--------------------------------------+
| BROWSER                              |
| React 19, Vite, TypeScript 7         |
| Tailwind v4, GSAP                    |
+--------------------------------------+
    |             |              |
    | app files   | JSON +       | Sign in
    |             | Bearer       | with Google
    v             v              v
+--------+  +-------------+  +----------+
| Static |  | FastAPI     |  | Google   |
| site   |  | Python 3.14 |  | Identity |
+--------+  +-------------+  +----------+
                   |
                   +--> SQLite: accounts, recipes
                   +--> Files: photos, audio
                   +--> ElevenLabs: STT + TTS
                   +--> Gemma 4: Gemini API
                   |             or Ollama
                   +--> WeasyPrint: HTML -> PDF

  Hosted on Render (or docker compose locally)
Enter fullscreen mode Exit fullscreen mode
Layer Choice Why
Frontend React 19, Vite, TypeScript 7, Tailwind v4, GSAP Fast to build, strict types, a few deliberate animations
Backend FastAPI, Pydantic v2, SQLModel on Python 3.14 Typed models end to end; background tasks for slow AI work
AI writer Gemma 4 26B (A4B) via an OpenAI-compatible API Open weights, swappable host, free tier
Speech ElevenLabs Scribe (speech-to-text) and multilingual TTS Handles messy kitchen audio; natural narration
PDF WeasyPrint with Jinja2 templates Real print CSS: page counters, running headers, leaders
Accounts Google Identity Services + signed sessions No passwords to store
Hosting Render Blueprint: Docker web service + static site One render.yaml describes both

From a voice memo to a recipe

 Cook records or uploads a voice memo
        |
        v
 ElevenLabs Scribe ----> transcript
        |
        v
 Gemma 4 <---- rules: never invent,
        |             keep their words
        v
 recipe as JSON
        |
        v
 Pydantic check --bad JSON--> retry,
        |                     error fed back
        v
 Is the quote really in the transcript?
        |   no --> use closest real sentence
        v
 Ready to read, cook, edit and print
Enter fullscreen mode Exit fullscreen mode

The status moves uploaded → transcribing → structuring → ready (or failed, always with a reason and a next step). If the server restarts mid-memo, it marks that memo as interrupted instead of leaving a spinner forever.

Things I had to get right

  • Not inventing things. The prompt forbids adding ingredients or steps that weren't said, the JSON is validated with Pydantic, and bad output is retried with the validation error fed back. For Maasi's recipes this mattered more than anything.
  • Keeping her voice. The pull-quote must be an exact substring of the transcript. The code checks this and falls back to the closest real sentence if the model paraphrases.
  • Swappable models. The model endpoint is three environment variables, so the same code runs against local Ollama or a hosted open-weight model.
  • A book that looks like a gift. Each recipe is a two-page spread: photo, title, story and their words on the left; ingredients, method and tips on the right. Steps never split across pages, headings never strand at the bottom, and continuation pages carry the recipe name in a running header. I checked it by rendering real books and looking at every page.
  • A free API key that strangers can't drain. See the next section.

Protecting a free API key

A public demo on free tiers needs guardrails, so every request passes through several layers:

 Sign in with Google
   |-- throwaway email domain?  --> refused
   |   (about 9,200 blocked)
   |-- 2+ new accounts from this
   |   network today?           --> refused
   v
 Ask for a recipe
   |-- account used its 5 today? --> try tomorrow
   |-- network used its 10?      --> refused
   |-- same transcript before?   --> reused, free
   v
 Throttle + queue (one call at a time)
   |-- 429 or 5xx --> wait, retry, then explain
   v
 Gemma 4 writes the recipe
Enter fullscreen mode Exit fullscreen mode
Guard Default Purpose
Recipes (or narrations) per account per day 5 Fair share of the free quota
New accounts per network per day 2 Stops one device making many accounts
AI requests per network per day 10 Stops one device spending many accounts' quota
Throwaway email domains refused about 9,200 No disposable sign-ups
Requests per account per 10 minutes 8 Stops bursts
Model calls per minute / per day (whole server) 10 / 300 Last line of defence for the key
Identical transcripts reused "Try a sample" costs nothing the second time

Only real AI calls count against an account's 5: writing by hand, editing, photos, classics and books are free. IP addresses are stored only as salted hashes, and the visitor's IP is read from Cloudflare's True-Client-IP header on Render, never from a header the client can fake.

Numbers

What Result
Time to write one recipe, Gemma 4 default reasoning 80 to 190 seconds
Same, with minimal thinking (thinking_level: minimal) 12 to 44 seconds in most runs (one slow free-tier run took about 2 minutes)
Backend tests 83, all mocked (no quota used)
Backend branch coverage 92% (CI requires at least 90%)
Browser checks against the Docker build 123: every page, every link, the book PDF, a real Gemma recipe
Responsive audit 14 pages at 5 widths (320 to 768px) plus 6 sizes up to 1920px: no text overflow
A six-recipe book 19 A5 pages including cover, contents and closing page
Python versions tested 3.13 and 3.14

Quality checks on every pull request

 Pull request
   |
   +--> Lint + format (ruff, Oxlint) -----+
   |                                      |
   +--> Security (CodeQL, Gitleaks,       |
   |    pip-audit, npm audit) ------------+
   |                                      |
   +--> Tests (Py 3.13 + 3.14,            |
   |    coverage >= 90%) ---+             |
   |                        |             |
   +--> Type check + build -+             |
                            |             |
                            v             |
              Docker build + 123          |
              browser checks -------------+
                                          |
                                          v
                              All checks passed
                                          |
                                          v
                                   Ready to merge
Enter fullscreen mode Exit fullscreen mode

Dependabot keeps Python, npm, Docker images and GitHub Actions up to date, and every update goes through the same checks.

What was hard

  • Speed. Gemma 4 reasons before it answers, and my first recipes took up to three minutes. Passing a minimal thinking level brought that down to well under a minute with the same quality.
  • Reasoning text in the output. Gemma 4 prefixes its answer with a <thought> block, which broke JSON parsing until I stripped it.
  • Free-tier limits. 429s and the occasional 500 from the provider needed polite retries, a one-at-a-time queue and clear "try again in a minute" messages instead of crashes.
  • The PDF engine everywhere. WeasyPrint needs native libraries (Pango): easy in Docker and on Render, painful on Windows until I scripted a private install.
  • The real visitor IP. Behind Render's proxy, the usual X-Forwarded-For header can be faked by the client, so the per-network limits read Cloudflare's True-Client-IP instead.
  • Tooling on the newest versions. TypeScript 7 is so new that typescript-eslint doesn't support it yet, so I lint with Oxlint instead.

What I'd do next

Hindi and Hinglish support (with a Devanagari PDF font), so Maasi's recipes can be printed in the language she actually speaks; shared family libraries so several relatives can add to one book; playback of the original voice memo, so we can always hear her say it; and a printed hardcover option.

Why Does Open Innovation Matter?

Salt & Story handles something deeply personal: a family member's voice and memories.

  • Control. With an open-weight model, the structuring step can run on my own machine or server. In the hosted demo, audio goes to ElevenLabs for transcription and transcripts go to Google's Gemini API, which runs Gemma; switching LLM_BASE_URL to local Ollama keeps transcripts on my own computer.
  • Swappability. Because the model sits behind a standard endpoint, I can change models without rewriting the app.
  • Cost. It's free to run, on the free tier or locally, which matters when the "user" is a family member and not a customer.
  • Tunability. I could shape the prompt and behaviour to protect what mattered most: faithfulness to what was actually said.

With a closed model, Maasi's words would only ever be processed wherever one company decided. With open weights, they can stay with our family.

Prize Categories

  • Best Use of Gemma: Gemma 4 structures every recipe, with its output validated and its quotes checked against the transcript.
  • Best Use of Render: Salt & Story is deployed on Render from a Blueprint: a Docker web service for the API and a static site for the app.
  • Best Use of ElevenLabs: speech-to-text turns voice memos into transcripts, and text-to-speech reads recipes aloud.

Somewhere in that app is a Dahi Wada recipe with my name on it, in Maasi's own words. That's the whole reason it exists.

If you have someone whose recipes live only in their head, record them while you can. Try it: https://salt-and-story-web.onrender.com/

And if this story touched you, please ⭐ star Salt & Story on GitHub. Thank you for reading. ❤️

Top comments (2)

Collapse
 
sizzlebop profile image
Jessica Doering •

This is such a thoughtful project. I really like that you didn’t just turn the recordings into generic cleaned-up recipes, but actually made preserving her wording part of the design. The fact that the pull quotes are checked against the transcript and the model is explicitly prevented from inventing ingredients or steps is a really nice touch too. It feels like you built around what actually matters here, not just what the AI could do. Really lovely idea.

Collapse
 
respect17 profile image
Kudzai Murimi •

This is the kind of build that actually matters, not just a demo. Keeping her real voice in it instead of turning it into plain text is a really nice choice.