<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Win Aung</title>
    <description>The latest articles on DEV Community by Win Aung (@win_aung_c84205e70ae9cba3).</description>
    <link>https://dev.to/win_aung_c84205e70ae9cba3</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4160627%2F46e1927e-0f6a-4ce9-a8eb-ade1b32a5725.png</url>
      <title>DEV Community: Win Aung</title>
      <link>https://dev.to/win_aung_c84205e70ae9cba3</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/win_aung_c84205e70ae9cba3"/>
    <language>en</language>
    <item>
      <title>RecipeTape: My Friend's Voice Memos, Now a Family Recipe Book</title>
      <dc:creator>Win Aung</dc:creator>
      <pubDate>Sun, 04 Oct 2026 07:20:14 +0000</pubDate>
      <link>https://dev.to/win_aung_c84205e70ae9cba3/recipetape-my-friends-voice-memos-now-a-family-recipe-book-4m31</link>
      <guid>https://dev.to/win_aung_c84205e70ae9cba3/recipetape-my-friends-voice-memos-now-a-family-recipe-book-4m31</guid>
      <description>

&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;RecipeTape&lt;/strong&gt; — a tool that turns a loved one's voice memos into a printable family recipe book.&lt;/p&gt;

&lt;p&gt;My friend &lt;strong&gt;Win Ko Aung&lt;/strong&gt;'s mother has never written down a single recipe. She cooks the way her mother taught her: by voice, by feel, by "a spoon of this, until it smells right." The family has years of her voice memos — half-remembered curries narrated over a sizzling pan — and every one of them was one lost phone away from disappearing forever. I built this for Win Ko Aung, so his mother's recipes outlive her phone.&lt;/p&gt;

&lt;p&gt;RecipeTape takes one of those voice memos and gives back a recipe book page: a title, the ingredients with their quantities, the steps in the order she spoke them, her own tips, and a dedication line in her words. Open the page in a browser, hit print, and it becomes a page in a book the family keeps.&lt;/p&gt;

&lt;p&gt;Here's what came out of a 90-second demo memo (a mother narrating her chicken curry):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Okay sweetheart, I am going to tell you my chicken curry recipe, the way my own mother taught me, so you never lose it. First, take two pounds of chicken, cut into small pieces, wash it well, and set it aside..."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the page RecipeTape made from it — &lt;strong&gt;Chicken Curry&lt;/strong&gt;, Serves 4, 5 ingredients, 13 steps, "Serve hot with rice":&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkvc38zeusplt42wgja4g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkvc38zeusplt42wgja4g.png" alt=" " width="800" height="1136"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;The whole thing runs as a Kaggle notebook on a free GPU — press run, get a recipe page:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.kaggle.com/code/winkoaung/notebookbe7ed42789" rel="noopener noreferrer"&gt;https://www.kaggle.com/code/winkoaung/notebookbe7ed42789&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The notebook accepts any voice memo (mp3/m4a/wav): attach a Kaggle dataset or upload your own with the widget. Out come three files: &lt;code&gt;transcript.txt&lt;/code&gt;, &lt;code&gt;recipe.json&lt;/code&gt;, and &lt;code&gt;recipe.html&lt;/code&gt; — the printable book page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Kaggle notebook (full code + demo):&lt;/strong&gt; &lt;a href="https://www.kaggle.com/code/winkoaung/notebookbe7ed42789" rel="noopener noreferrer"&gt;https://www.kaggle.com/code/winkoaung/notebookbe7ed42789&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The pipeline is four small Python modules:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;voice memo
  │  faster-whisper (open source, runs locally)
  ▼
transcript
  │  Gemma 3 1B (open weights) → strict JSON, validated, never invents
  ▼
recipe.json
  │  renderer
  ▼
recipe.html → print to PDF
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The open-source pieces are the product, not a garnish.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Speech&lt;/strong&gt; is &lt;code&gt;faster-whisper&lt;/code&gt; (the &lt;code&gt;small&lt;/code&gt; model), running entirely on the machine — on Kaggle's free T4 it transcribes with &lt;code&gt;float16&lt;/code&gt;, on a CPU-only laptop it drops to &lt;code&gt;int8&lt;/code&gt;. A voice-activity filter skips the silences, and if GPU initialisation ever fails it falls back to CPU instead of dying. The 90-second demo transcribed near-perfectly, including "a squeeze of half a lemon."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Understanding&lt;/strong&gt; is &lt;strong&gt;Gemma 3 1B Instruct&lt;/strong&gt;, Google's open-weight model, pulled from Kaggle Models (no tokens, no API keys). It gets the transcript plus a strict JSON schema and five rules — the most important being &lt;em&gt;never invent an ingredient, quantity, or step that isn't in the transcript&lt;/em&gt;. The reply is parsed with a &lt;code&gt;JSONDecoder.raw_decode()&lt;/code&gt; scan (so stray prose can't corrupt it) and validated hard: ingredients must be well-formed entries, steps must be real sentences. Deterministic decoding (&lt;code&gt;do_sample=False&lt;/code&gt;) — a recipe should not improvise.&lt;/p&gt;

&lt;p&gt;One honest wart: Gemma nailed the structure (13 steps in spoken order) but echoed my schema's placeholder text for the dedication line instead of writing one. The prompt needs one more iteration there — the page above uses the transcript's own closing line instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The page&lt;/strong&gt; is plain HTML/CSS with a print stylesheet: warm paper tones, a dedication line, ingredients in two columns, numbered method, the cook's tips, and a pull-quote of the original voice memo ("In Their Own Words"). No framework, no build step — &lt;code&gt;Ctrl+P&lt;/code&gt; gives you the book.&lt;/p&gt;

&lt;p&gt;I built the pipeline with Muse (code, notebook, demo) and had ChatGPT do a full adversarial code review — it caught 15 issues, including a real one: my first version requested &lt;code&gt;bfloat16&lt;/code&gt; on a T4, which has no native BF16 support. The review earned its keep before a single GPU minute was spent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;A voice memo of your mother describing her curry is about as personal as data gets. The closed alternative — upload it to a transcription API, then to a chat API — means an account, a per-minute bill, and a copy of her voice on someone else's server, plus an app that dies the moment the network does.&lt;/p&gt;

&lt;p&gt;RecipeTape runs fully offline after the one-time model download. The audio never leaves the laptop. It costs exactly $0 — the demo ran on Kaggle's free GPUs. And every open piece can be inspected and changed: the transcriber, the weights, the prompt, the JSON schema, the validation rules. When Gemma misheard "a spoon of" as a precise gram weight in an early test, I could read the prompt, see the rule that allowed it, and tighten it. A closed API would have handed me a confident paragraph and no way to demand it quote the audio I already had.&lt;/p&gt;

&lt;p&gt;That's the whole bet: the people whose recipes these are should never need anyone's permission — or anyone's server — to keep them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;h2&gt;
  
  
  - &lt;strong&gt;Best Use of Gemma.&lt;/strong&gt; Gemma 3 1B Instruct (open weights, via Kaggle Models) is the recipe editor: it turns a rambling spoken transcript into validated, structured recipe JSON without inventing a single ingredient. Deterministic decoding, chat-template prompting, FP16 on the T4.
&lt;/h2&gt;

</description>
      <category>hacktoberfest</category>
      <category>ai</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?</title>
      <dc:creator>Win Aung</dc:creator>
      <pubDate>Sat, 03 Oct 2026 23:15:08 +0000</pubDate>
      <link>https://dev.to/win_aung_c84205e70ae9cba3/myanmarchemcalc-bench-do-llms-do-chemistry-better-in-english-than-in-burmese-1iab</link>
      <guid>https://dev.to/win_aung_c84205e70ae9cba3/myanmarchemcalc-bench-do-llms-do-chemistry-better-in-english-than-in-burmese-1iab</guid>
      <description>&lt;p&gt;&lt;strong&gt;Benchmark:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/winkoaung/myanmar-chem-calc-bench" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/winkoaung/myanmar-chem-calc-bench&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What tasks did I run?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;MyanmarChemCalc-Bench&lt;/strong&gt; — 30 tasks built from 15 chemistry calculation questions drawn from Myanmar's Grades 9–12 chemistry textbook worked examples (ground-truth answers verified against the textbooks). Each question exists in &lt;strong&gt;two versions: English and Burmese&lt;/strong&gt; — the same problem, the same numbers, the same expected answer.&lt;/p&gt;

&lt;p&gt;Topics span the full calculation curriculum: mole concept, Avogadro's number, isotopes, percentage composition, combustion analysis, stoichiometry, gas volumes, ideal gas law, molarity, dilution, titration, pH, Faraday's laws, percentage yield, and equilibrium constants.&lt;/p&gt;

&lt;p&gt;Scoring is strict and automatic: the model must show its work, then write &lt;code&gt;FINAL ANSWER: &amp;lt;number&amp;gt;&lt;/code&gt; on its own line. A Python checker extracts the number and compares it against the ground truth within a per-question tolerance.&lt;/p&gt;

&lt;p&gt;The creative hook: &lt;strong&gt;does a model score lower when the exact same chemistry question is asked in Burmese?&lt;/strong&gt; Burmese is a low-resource language — if the multilingual gap is real, it should show up here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which models did I run it against?
&lt;/h2&gt;

&lt;p&gt;Five models from Kaggle's suite:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.5&lt;/td&gt;
&lt;td&gt;100.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-R1&lt;/td&gt;
&lt;td&gt;93.33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Pro&lt;/td&gt;
&lt;td&gt;93.33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;93.33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 235B A22B Instruct&lt;/td&gt;
&lt;td&gt;46.67&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What are the main insights?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The most interesting finding isn't chemistry — it's run reliability
&lt;/h3&gt;

&lt;p&gt;Qwen's 46.67 looks like a chemistry failure. It isn't — not mostly. Of Qwen's runs that actually produced a scored answer, the vast majority passed. The missing cells were execution errors: 429 rate-limit errors from provider overload, not wrong answers. The leaderboard score divides passes by all 30 tasks, so errors masquerade as incompetence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson: on a small benchmark, run reliability can distort the leaderboard more than model capability.&lt;/strong&gt; I report passes, fails, and errors separately — and so should you.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Paired translations make failures inspectable
&lt;/h3&gt;

&lt;p&gt;The EN/MM pairing paid off in a concrete case: on the mole question ("How many moles are in 8 g of NaOH?"), Qwen passed in English but wrote &lt;strong&gt;23 + 16 + 1 = 50&lt;/strong&gt; in the Burmese version, getting 0.16 mol instead of 0.2. Same numbers, same question — a genuine arithmetic slip in the Burmese run. That's a single case, not proof of a systematic language gap, but it's exactly the kind of failure this benchmark design surfaces.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Small benchmarks demand restraint
&lt;/h3&gt;

&lt;p&gt;With 30 tasks, &lt;strong&gt;one task swings the score by 3.33 points&lt;/strong&gt;. The gap between Claude (100.00) and the 93.33 group is just two tasks. Don't over-interpret the ordering — it needs completed evaluations and repeated runs to be meaningful.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Burmese didn't break the top models — and the numeral fix is validated
&lt;/h3&gt;

&lt;p&gt;After updating the Burmese prompts to use Arabic numerals (8 instead of ၈), &lt;strong&gt;every completed run on the new Burmese tasks was a PASS — 100%&lt;/strong&gt; across all five models. The numeral change didn't degrade anything; models handle Arabic numerals in Burmese prompts reliably. Claude reached a perfect &lt;strong&gt;100.00&lt;/strong&gt; overall. For textbook-style calculations, the current frontier models handle Burmese about as well as English.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where can we see it?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Benchmark (public):&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/winkoaung/myanmar-chem-calc-bench" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/winkoaung/myanmar-chem-calc-bench&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All 30 tasks, the leaderboard, and per-model run traces are public. The task files are generated from a single Python script with ground truths recomputed against the Myanmar textbook worked examples.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Fix the remaining error cells with bounded retries and re-run for a clean comparison&lt;/li&gt;
&lt;li&gt;Expand beyond 15 questions — stoichiometry and multi-step problems are where I'd expect the real gaps&lt;/li&gt;
&lt;li&gt;Test whether strict Burmese explanations (not just prompts) change anything&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Built for the Kaggle Benchmarking Challenge. All questions verified against Myanmar basic-education chemistry textbooks (Grades 9–12).&lt;/em&gt;&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>llm</category>
      <category>benchmarking</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
