<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Piyush Anand</title>
    <description>The latest articles on DEV Community by Piyush Anand (@piyush_anand_9d508e5ee7af).</description>
    <link>https://dev.to/piyush_anand_9d508e5ee7af</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4157515%2F676cecbb-c907-4b4f-842d-e0475fc1f19c.jpg</url>
      <title>DEV Community: Piyush Anand</title>
      <link>https://dev.to/piyush_anand_9d508e5ee7af</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/piyush_anand_9d508e5ee7af"/>
    <language>en</language>
    <item>
      <title>EchoBook: Turning Grandpa's Rambling Voice Memos into a Family Recipe Book, Fully Offline</title>
      <dc:creator>Piyush Anand</dc:creator>
      <pubDate>Sat, 03 Oct 2026 04:46:37 +0000</pubDate>
      <link>https://dev.to/piyush_anand_9d508e5ee7af/echobook-turning-grandpas-rambling-voice-memos-into-a-family-recipe-book-fully-offline-oon</link>
      <guid>https://dev.to/piyush_anand_9d508e5ee7af/echobook-turning-grandpas-rambling-voice-memos-into-a-family-recipe-book-fully-offline-oon</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The person
&lt;/h3&gt;

&lt;p&gt;I built this for my grandfather.&lt;/p&gt;

&lt;p&gt;For years he has recorded voice memos about the food our family grew up on: the dal my grandmother made every Sunday, the banana bread our neighbour Mrs. Fernandes taught him in the seventies, the chai he picked up from a railway-station chaiwala over forty years of mornings before work. Whenever someone asks "how did you make that?", he records another memo.&lt;/p&gt;

&lt;h3&gt;
  
  
  The problem
&lt;/h3&gt;

&lt;p&gt;None of it is written down. The memos ramble: "Okay, is this thing recording?", a story about the day my father was born, back to the lentils, "four whistles, not three, she was very strict about that." There are no ingredient lists and no numbered steps. The measurements are "a big pinch" and "the ugly black bananas nobody wants to eat." The recipes exist &lt;strong&gt;only in his voice&lt;/strong&gt;, scattered across audio files on an old phone, and once those files are lost the recipes go with them.&lt;/p&gt;

&lt;p&gt;Typing them up by hand means listening to hours of audio and pausing every few seconds. Sending them to a cloud transcription or AI service means uploading my grandfather's voice and our family stories to somebody else's servers. I didn't want either.&lt;/p&gt;

&lt;h3&gt;
  
  
  The solution: EchoBook
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;EchoBook turns raw voice memos into a clean, searchable, shareable family recipe book, with every bit of processing done on the family's own computer using open-weight models.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You drop in a recording, and EchoBook:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Listens.&lt;/strong&gt; It transcribes the memo locally with an open-weight Whisper model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Writes the recipe.&lt;/strong&gt; A local open-weight LLM turns the ramble into a structured recipe card: title, servings, every ingredient with its quantity &lt;em&gt;exactly as he said it&lt;/em&gt;, numbered steps in the order he does them, and his own tips and warnings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keeps his voice.&lt;/strong&gt; Every recipe carries an &lt;strong&gt;"In Grandpa's words"&lt;/strong&gt; section: a verbatim quote of the memory behind the dish (&lt;em&gt;"You know, we made this the day your father was born. The whole hostel floor could smell it."&lt;/em&gt;), right above an audio player with the original recording.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adds it to the book.&lt;/strong&gt; It's stored with the full transcript and a link to the source audio.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpnj90okui3tfhojggdpl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpnj90okui3tfhojggdpl.png" alt="A recipe page with Grandpa's quote and the original recording" width="800" height="833"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Features
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;🎙️ &lt;strong&gt;Drop in a voice memo&lt;/strong&gt; (mp3, m4a, wav, ogg, webm…) or pick a sample, and watch each stage run live: saving → listening (progress bar plus live transcript) → writing the recipe (live token count) → adding to the book.&lt;/li&gt;
&lt;li&gt;📖 &lt;strong&gt;Recipe cards&lt;/strong&gt; with tick-off ingredient checklists, numbered steps, "Grandpa's tips", and the original transcript one click away.&lt;/li&gt;
&lt;li&gt;🗣️ &lt;strong&gt;"In Grandpa's words"&lt;/strong&gt;: an authentic quote plus the original recording on every recipe.&lt;/li&gt;
&lt;li&gt;🔍 &lt;strong&gt;Search&lt;/strong&gt; by dish or ingredient (&lt;code&gt;ghee&lt;/code&gt;, &lt;code&gt;cardamom milk&lt;/code&gt;), with matching ingredients highlighted on each card.&lt;/li&gt;
&lt;li&gt;✏️ &lt;strong&gt;Review &amp;amp; edit&lt;/strong&gt;: small models sometimes mishear, so every recipe can be checked against the recording and corrected before the book is shared. The family gets the final say, not the AI.&lt;/li&gt;
&lt;li&gt;🖨️ &lt;strong&gt;Export&lt;/strong&gt;: download one recipe or the whole book as Markdown, or open a print layout (cover, table of contents, one recipe per page) and save it as a PDF to print and bind.&lt;/li&gt;
&lt;li&gt;🌐 &lt;strong&gt;A shareable link on Render&lt;/strong&gt;: relatives anywhere can browse the book on their phones and hear Grandpa tell the story, while recordings are processed only at home.&lt;/li&gt;
&lt;li&gt;📴 &lt;strong&gt;Works offline&lt;/strong&gt;: once the models are downloaded, transcription and structuring run with the network unplugged.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The recipes are cleaned up. His voice is left as it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;🔗 &lt;strong&gt;Live recipe book: &lt;a href="https://echobook-8fpj.onrender.com/" rel="noopener noreferrer"&gt;https://echobook-8fpj.onrender.com/&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(It's on Render's free tier, so if it has been idle the first load can take ~30–50 seconds to wake up.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to click:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Open "Sunday Dal".&lt;/strong&gt; Read the &lt;em&gt;In Grandpa's words&lt;/em&gt; quote and press ▶ to hear the memo it came from. Notice the quantities are his ("a big pinch" of asafoetida, "4 whistles"). Scroll down and expand &lt;strong&gt;Original voice memo transcript&lt;/strong&gt; to compare the raw ramble with the finished card.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search&lt;/strong&gt; &lt;code&gt;ghee&lt;/code&gt;, then &lt;code&gt;cardamom&lt;/code&gt;. Only the matching recipes remain, with the matched ingredient highlighted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open "Masala Chai"&lt;/strong&gt; for the railway-station story and the "three times" trick.&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;Print book&lt;/strong&gt;, then &lt;strong&gt;Print / Save as PDF&lt;/strong&gt;, to get the whole family cookbook as a PDF.&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;Add a memo&lt;/strong&gt;. On the public site this explains that processing happens at home. That's deliberate (see &lt;em&gt;How I Built It&lt;/em&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;The local pipeline in action&lt;/strong&gt; (this is the part that runs on the family's computer):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fswrsretyxaca1knwrnn1.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fswrsretyxaca1knwrnn1.gif" alt="EchoBook demo: a voice memo becomes a recipe card, then search and the printable book" width="560" height="387"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A real, unedited run on a CPU-only machine. The two waiting stages are sped up (labelled in the GIF); the full memo-to-recipe run took about 2½ minutes. Note the raw output still says "asan" (Assam); that's what the Review &amp;amp; edit screen is for.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A note on the sample audio:&lt;/strong&gt; to avoid publishing private family recordings, the demo uses three stand-in memos I wrote in the style of Grandpa's real ones (false starts, side stories, loose measurements, a bit of family history) and voiced with &lt;a href="https://github.com/rhasspy/piper" rel="noopener noreferrer"&gt;Piper&lt;/a&gt;, an open-source local text-to-speech engine. Real recordings go through exactly the same pipeline.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Anandpiyush21" rel="noopener noreferrer"&gt;
        Anandpiyush21
      &lt;/a&gt; / &lt;a href="https://github.com/Anandpiyush21/EchoBook" rel="noopener noreferrer"&gt;
        EchoBook
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Turn family voice memos into a living recipe book, entirely on-device with open-weight models (faster-whisper + Ollama). Built for the Hacktoberfest 'Build for a Friend' challenge.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;EchoBook&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Turn family voice memos into a living recipe book, entirely on your own computer.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;My grandfather has years of voice memos describing family recipes from memory. They ramble, there are no
ingredient lists and no steps, and the recipes only exist in his voice, scattered across audio files
EchoBook takes those recordings and turns them into a clean, searchable, printable recipe book. It keeps
the important part: his own words and the original recording, attached to every recipe.&lt;/p&gt;
&lt;p&gt;All processing runs locally on open-weight models. No recording, transcript or family story is sent to a
cloud API, and once the models are downloaded the pipeline works with no internet connection at all.&lt;/p&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/Anandpiyush21/EchoBook/docs/demo.gif"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FAnandpiyush21%2FEchoBook%2FHEAD%2Fdocs%2Fdemo.gif" alt="EchoBook demo"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/Anandpiyush21/EchoBook/docs/book.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FAnandpiyush21%2FEchoBook%2FHEAD%2Fdocs%2Fbook.png" alt="The recipe book"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;A recipe, with Grandpa's own words and the original recording&lt;/th&gt;
&lt;th&gt;The local pipeline, stage by stage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/Anandpiyush21/EchoBook/docs/recipe.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FAnandpiyush21%2FEchoBook%2FHEAD%2Fdocs%2Frecipe.png" alt="Recipe page"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;&lt;a rel="noopener noreferrer" href="https://github.com/Anandpiyush21/EchoBook/docs/pipeline.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FAnandpiyush21%2FEchoBook%2FHEAD%2Fdocs%2Fpipeline.png" alt="Pipeline progress"&gt;&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Features&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Drop in a voice memo&lt;/strong&gt; (mp3, m4a, wav, ogg, webm…) or pick a sample, and watch each stage run live…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Anandpiyush21/EchoBook" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;The repo includes setup instructions, an architecture write-up, &lt;code&gt;render.yaml&lt;/code&gt;, &lt;code&gt;.env.example&lt;/code&gt;, the sample memos, a prompt-evaluation script, and a step-by-step &lt;code&gt;demo_script.md&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Architecture
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;voice memo ─► faster-whisper ─► Ollama + qwen2.5:7b ─► SQLite ─► web UI
  (.m4a)     (Whisper small.en,   (JSON-schema-        (recipe +     (browse, search,
              on-device)           constrained output)  transcript +  edit, export)
                                                        audio link)
└──────────────── all on the family's computer, works offline ───────────────┘
                                   │ git push (finished recipes only)
                                   ▼
                     Render: same app in read-only "viewer" mode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Speech-to-text&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;faster-whisper&lt;/strong&gt; (open-weight Whisper &lt;code&gt;small.en&lt;/code&gt;, CTranslate2, int8)&lt;/td&gt;
&lt;td&gt;Accurate, fast on a plain CPU, fully local&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structuring&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Ollama + &lt;code&gt;qwen2.5:7b-instruct&lt;/code&gt;&lt;/strong&gt; (Apache-2.0)&lt;/td&gt;
&lt;td&gt;Strong instruction-following at 7B; supports schema-constrained output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;SQLite&lt;/strong&gt;, mirrored to &lt;code&gt;data/recipes.json&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Zero setup; the JSON snapshot is diffable and deployable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backend&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;FastAPI&lt;/strong&gt; + a background worker&lt;/td&gt;
&lt;td&gt;Long jobs don't block the UI; per-stage progress API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frontend&lt;/td&gt;
&lt;td&gt;Plain &lt;strong&gt;HTML/CSS/JS&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;No framework, no build step; mobile, dark mode and print styles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosting&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Render&lt;/strong&gt; (free web service, &lt;code&gt;render.yaml&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;A shareable link for the family, auto-deploy on push&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Stage 1: Listening (faster-whisper)
&lt;/h3&gt;

&lt;p&gt;faster-whisper runs OpenAI's open-weight Whisper model on CTranslate2: int8 on CPU, float16 on a GPU. Two settings made a real difference for this use case:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Voice activity detection&lt;/strong&gt; skips the long pauses that old memos are full of.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An initial prompt seeded with kitchen vocabulary&lt;/strong&gt; ("ghee, cumin, turmeric, cardamom, asafoetida, toor dal, Fahrenheit…") biases recognition toward the words that matter most in a recipe.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A one-minute memo transcribes in about 15 seconds on a 12-core CPU. The transcript is still imperfect ("hing pia saffitida", "the tor dal", "still worn" for "still warm"), which is exactly why the next stage has to be careful.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stage 2: Writing the recipe (Ollama + Qwen 2.5)
&lt;/h3&gt;

&lt;p&gt;Most of my time went here. The transcript goes to the local model with an "archivist" system prompt, and Ollama's &lt;strong&gt;structured outputs&lt;/strong&gt; feature constrains decoding to a JSON schema:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"servings"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"total_time"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ingredients"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"quantity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"item"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"note"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"steps"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"tips"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"story_quote"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"tags"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because generation is constrained to this schema, the response always parses, and I never have to scrape JSON out of chatty text. Getting it to &lt;em&gt;parse&lt;/em&gt; was easy. Getting it to be &lt;em&gt;faithful&lt;/em&gt; was the real work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The 3B model invented things.&lt;/strong&gt; I started with &lt;code&gt;qwen2.5:3b&lt;/code&gt; for speed. It added a "30 minutes" total time nobody mentioned, gave the dal a "don't over-mix" tip that belonged to the banana bread, and turned "one teaspoon of baking soda" into "½ tsp". For a family recipe, a wrong quantity is worse than a missing one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stricter rules plus a bigger model fixed most of it.&lt;/strong&gt; I rewrote the prompt around faithfulness: &lt;em&gt;use only what is said; copy every amount exactly as spoken; never add a unit he didn't say ("four cardamom pods" → quantity &lt;code&gt;4&lt;/code&gt;, not &lt;code&gt;4 tsp&lt;/code&gt;); leave a field empty rather than guess; include ingredients mentioned in passing ("cinnamon if you like it", "if you have walnuts").&lt;/em&gt; With those rules, &lt;code&gt;qwen2.5:7b&lt;/code&gt; copied amounts exactly and left unknown fields blank.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Examples in a prompt leak.&lt;/strong&gt; I gave one example of fixing a misheard word involving "Assam", and the model started naming unrelated dishes "Assam Dal". I replaced it with a neutral example.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Verbatim" isn't always verbatim.&lt;/strong&gt; Small models quietly paraphrase quotes, and the whole point of &lt;em&gt;In Grandpa's words&lt;/em&gt; is that they are &lt;em&gt;his&lt;/em&gt; words. So after generation, EchoBook fuzzy-matches the model's quote against the transcript and snaps it to the closest real span of one to three sentences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Humans get the last word.&lt;/strong&gt; Even the 7B model occasionally slips (one run wrote "4 tsp" of cardamom pods). Rather than pretend the AI is perfect, I built a &lt;strong&gt;Review &amp;amp; edit&lt;/strong&gt; screen with the recording and transcript right next to the form, so a family member can check each recipe before the book is shared.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To iterate quickly I wrote &lt;code&gt;scripts/eval_structuring.py&lt;/code&gt;. It caches transcripts, re-runs only the LLM step, and &lt;strong&gt;flags any ingredient quantity whose numbers never appear in the transcript&lt;/strong&gt;, plus any quote that isn't verbatim. That made comparing prompts and models a two-minute loop instead of a re-recording session.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A performance surprise:&lt;/strong&gt; my development machine has only a small, unsupported GPU, and Ollama was quietly offloading a sliver of the model onto it. Prompt processing crawled at ~4 tokens/s. Forcing pure CPU (&lt;code&gt;OLLAMA_NUM_GPU=0&lt;/code&gt;) made it &lt;strong&gt;6x faster&lt;/strong&gt; (~27 tokens/s) and brought a full one-minute memo down to &lt;strong&gt;about 2 minutes end to end on CPU&lt;/strong&gt;. On a proper GPU the whole pipeline takes seconds, and you can step up to &lt;code&gt;large-v3&lt;/code&gt; and a 14B model through &lt;code&gt;.env&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stage 3: Storage
&lt;/h3&gt;

&lt;p&gt;SQLite stores each recipe alongside its &lt;strong&gt;full original transcript&lt;/strong&gt;, the &lt;strong&gt;source audio file&lt;/strong&gt;, and &lt;strong&gt;which models produced it&lt;/strong&gt;, so provenance is never lost. Every write is also mirrored to &lt;code&gt;data/recipes.json&lt;/code&gt;, a human-readable, diffable snapshot that doubles as the deployment artifact.&lt;/p&gt;

&lt;h3&gt;
  
  
  The web app
&lt;/h3&gt;

&lt;p&gt;FastAPI serves a small JSON API and the static frontend. Processing runs in a single background worker (the models are large, so one job at a time keeps memory predictable). The UI polls &lt;code&gt;/api/jobs/{id}&lt;/code&gt; for stage, transcription progress, the live transcript and the token count. The frontend is one page of plain JavaScript with hash routing, works on phones, follows dark mode, and has a print stylesheet that turns the book into a clean PDF with a cover, a table of contents and page breaks between recipes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deploying with Render: inference stays local, viewing goes to the cloud
&lt;/h3&gt;

&lt;p&gt;Whisper plus a 7B LLM won't run on a free web instance. More importantly, &lt;strong&gt;family recordings shouldn't be processed in the cloud at all&lt;/strong&gt;, because that's the whole point. So the same codebase runs in two modes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;At home (&lt;code&gt;ECHOBOOK_MODE=local&lt;/code&gt;)&lt;/th&gt;
&lt;th&gt;On Render (&lt;code&gt;viewer&lt;/code&gt;)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Upload &amp;amp; process new memos&lt;/td&gt;
&lt;td&gt;✅ faster-whisper + Ollama&lt;/td&gt;
&lt;td&gt;❌ disabled (read-only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browse, search, listen, export&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dependencies&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;requirements.txt&lt;/code&gt; (includes faster-whisper)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;requirements-web.txt&lt;/code&gt; (FastAPI + uvicorn only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data&lt;/td&gt;
&lt;td&gt;SQLite, mirrored to &lt;code&gt;recipes.json&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;SQLite rebuilt from &lt;code&gt;recipes.json&lt;/code&gt; on boot&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The publishing workflow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Process memos at home and review or fix them in the UI.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;git add data &amp;amp;&amp;amp; git commit &amp;amp;&amp;amp; git push&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Render auto-deploys and rebuilds the book from the snapshot.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The viewer installs no ML dependencies, so it builds in seconds and fits the free tier. The repo includes a &lt;code&gt;render.yaml&lt;/code&gt; blueprint, and the app &lt;strong&gt;automatically falls back to read-only mode when it detects it's running on Render&lt;/strong&gt;, so a missing setting can never expose an upload endpoint. Relatives get a link they can open on their phones, and the cloud only ever sees the finished recipes the family chose to publish.&lt;/p&gt;

&lt;h3&gt;
  
  
  What's next
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Batch-import a whole folder of old memos overnight.&lt;/li&gt;
&lt;li&gt;Multilingual memos: Whisper's multilingual models can handle Hindi and mixed Hindi-English recordings, with the recipe written out in English.&lt;/li&gt;
&lt;li&gt;Let family members add photos of the finished dish to each recipe.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;These recordings are my grandfather's voice and our family's history, and uploading them to a cloud API to save some typing never felt right. Because Whisper and Qwen are open-weight and run locally, &lt;strong&gt;no audio, transcript or family story ever leaves the machine during processing&lt;/strong&gt;, and the whole pipeline works with the network cable unplugged. There's &lt;strong&gt;no per-recording API cost&lt;/strong&gt;, so digitizing years of memos costs nothing but electricity. Because the model runs on our own machine, I could &lt;strong&gt;tune the structuring prompt freely&lt;/strong&gt;, compare models side by side, and swap in a bigger one on a better GPU without anyone's permission or a pricing page. And a family archive should outlive any single product: in twenty years, when today's cloud APIs have changed or shut down, these model weights and this code will still turn Grandpa's voice into recipes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Render&lt;/strong&gt;: the live, shareable family recipe book runs on Render (&lt;a href="https://echobook-8fpj.onrender.com/" rel="noopener noreferrer"&gt;https://echobook-8fpj.onrender.com/&lt;/a&gt;), deployed from a &lt;code&gt;render.yaml&lt;/code&gt; blueprint with auto-deploy on every push. Inference is kept local on purpose, and only finished, family-approved recipes are deployed: a lightweight viewer with no ML dependencies that builds in seconds on the free tier.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Thanks for reading. If you have a grandparent with stories in their voice memos, record a few more this weekend.&lt;/em&gt; 🍲&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hacktoberfest</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
