<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ShalyX</title>
    <description>The latest articles on DEV Community by ShalyX (@shalyx).</description>
    <link>https://dev.to/shalyx</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4162386%2F6d666c4c-3b29-4672-bf4f-0c6b63baf550.jpg</url>
      <title>DEV Community: ShalyX</title>
      <link>https://dev.to/shalyx</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shalyx"/>
    <language>en</language>
    <item>
      <title>Screenshot Graveyard: I built my friend a memory for the screenshots they never revisit</title>
      <dc:creator>ShalyX</dc:creator>
      <pubDate>Mon, 05 Oct 2026 04:16:01 +0000</pubDate>
      <link>https://dev.to/shalyx/screenshot-graveyard-i-built-my-friend-a-memory-for-the-screenshots-they-never-revisit-2729</link>
      <guid>https://dev.to/shalyx/screenshot-graveyard-i-built-my-friend-a-memory-for-the-screenshots-they-never-revisit-2729</guid>
      <description>&lt;p&gt;Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝&lt;/p&gt;

&lt;p&gt;This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;My friend saves everything as a screenshot.&lt;/p&gt;

&lt;p&gt;A restaurant they want to try. A product they might buy. A tweet worth rereading. A design idea. A technical explanation. A recommendation somebody sent them.&lt;/p&gt;

&lt;p&gt;At the moment they save it, the screenshot feels useful.&lt;/p&gt;

&lt;p&gt;A week later, it is just image number 6,842 in the camera roll.&lt;/p&gt;

&lt;p&gt;That was the problem I wanted to solve this weekend.&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;Screenshot Graveyard&lt;/strong&gt;, a tiny memory layer for screenshots.&lt;/p&gt;

&lt;p&gt;You drop in screenshots. It reads what is actually visible, turns each image into a compact memory of what it is and why you probably saved it, and stores that memory alongside the original screenshot in the browser.&lt;/p&gt;

&lt;p&gt;Later you can ask something vague like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;what was that private access thing I saved?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You do not need the filename. You do not need the exact title. You do not even need to remember the proper noun.&lt;/p&gt;

&lt;p&gt;Screenshot Graveyard searches the recovered memories, finds the likely match, and brings the original screenshot back up.&lt;/p&gt;

&lt;p&gt;That is the whole product.&lt;/p&gt;

&lt;p&gt;I deliberately kept it narrow because the useful moment is not “AI organized my photos.”&lt;/p&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I remembered the idea, but not enough of it to find the screenshot — and the app found it anyway.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The friend I built it for is real, but I am keeping them anonymous. Their screenshot habit is the product brief.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Live app:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://screenshot-graveyard.vercel.app" rel="noopener noreferrer"&gt;https://screenshot-graveyard.vercel.app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Try this flow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Drop in a few screenshots.&lt;/li&gt;
&lt;li&gt;Let Screenshot Graveyard recover a memory for each one.&lt;/li&gt;
&lt;li&gt;Search with a fuzzy description instead of the exact words in the screenshot.&lt;/li&gt;
&lt;li&gt;Open the result and confirm it surfaced the original image.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The archive itself lives in the browser using IndexedDB.&lt;/p&gt;

&lt;p&gt;There is no account and no server-side screenshot library.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/ShalyX/screenshot-graveyard" rel="noopener noreferrer"&gt;https://github.com/ShalyX/screenshot-graveyard&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repository includes the Flask backend, browser UI, RapidOCR pipeline, Gemma integration, structured-output recovery logic, IndexedDB archive, and smoke tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The final architecture looks simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Screenshot
   ↓
RapidOCR
   ↓
literal text
   ↓
Gemma 3 4B
   ↓
title / summary / why saved / entities / search terms
   ↓
IndexedDB
   ↓
fuzzy natural-language retrieval
   ↓
original screenshot
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It did not start that way.&lt;/p&gt;

&lt;h3&gt;
  
  
  Version one: let Gemma do everything
&lt;/h3&gt;

&lt;p&gt;My first version ran &lt;strong&gt;Gemma 3 4B locally through Ollama&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I sent the screenshot to Gemma and asked it to both read the image and understand it.&lt;/p&gt;

&lt;p&gt;Architecturally, that was beautifully simple.&lt;/p&gt;

&lt;p&gt;In practice, it exposed two problems immediately.&lt;/p&gt;

&lt;p&gt;The first was speed. Vision inference on my machine took long enough that a small batch felt broken even when it was technically still processing.&lt;/p&gt;

&lt;p&gt;The second problem was worse because it attacked trust.&lt;/p&gt;

&lt;p&gt;A screenshot contained the name:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;VEYRA&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Gemma read it as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;VETRA&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then the retrieval layer repeated the same wrong spelling back to me.&lt;/p&gt;

&lt;p&gt;The app had successfully remembered the screenshot — incorrectly.&lt;/p&gt;

&lt;p&gt;For this product, that is not a cosmetic bug.&lt;/p&gt;

&lt;p&gt;If Screenshot Graveyard quietly changes the names inside the things you saved, it is manufacturing memories instead of recovering them.&lt;/p&gt;

&lt;h3&gt;
  
  
  The architecture changed because of one typo
&lt;/h3&gt;

&lt;p&gt;That VEYRA → VETRA failure made the separation of responsibilities obvious.&lt;/p&gt;

&lt;p&gt;An LLM is useful for questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What is this screenshot about?&lt;/li&gt;
&lt;li&gt;Why might somebody have saved it?&lt;/li&gt;
&lt;li&gt;Which future query should retrieve it?&lt;/li&gt;
&lt;li&gt;Does this fuzzy question refer to this memory?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It should not be the only authority for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;exact product names&lt;/li&gt;
&lt;li&gt;usernames&lt;/li&gt;
&lt;li&gt;acronyms&lt;/li&gt;
&lt;li&gt;project names&lt;/li&gt;
&lt;li&gt;visible quoted text&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So I split the pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RapidOCR owns literal text. Gemma owns meaning.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For text-heavy screenshots, RapidOCR first extracts the on-screen text. That transcript is then sent to Gemma with an explicit rule: distinctive names and proper nouns are evidence, not material to “improve.”&lt;/p&gt;

&lt;p&gt;I also added a small grounding pass. If Gemma produces a suspicious near-match to a distinctive OCR token, the stored title is corrected back toward the literal evidence.&lt;/p&gt;

&lt;p&gt;If OCR finds very little useful text, the system falls back to Gemma's multimodal image understanding.&lt;/p&gt;

&lt;p&gt;That gave me a much better division of labor:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OCR: what does the screenshot literally say?
Gemma: what does it mean to this person?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Gemma creates a memory, not just a caption
&lt;/h3&gt;

&lt;p&gt;Each screenshot becomes a small structured record:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"category"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ideas"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"summary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"whySaved"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"entities"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"searchTerms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The field I care about most is &lt;code&gt;whySaved&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Generic image classification can tell me there is a restaurant menu in a screenshot.&lt;/p&gt;

&lt;p&gt;Screenshot Graveyard is trying to reconstruct something closer to intent:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You probably saved this because you wanted to compare this ramen place later.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That makes the archive searchable by the future thought, not only by the pixels.&lt;/p&gt;

&lt;h3&gt;
  
  
  Retrieval is deliberately small
&lt;/h3&gt;

&lt;p&gt;I did not add a vector database.&lt;/p&gt;

&lt;p&gt;When you search, the browser sends Gemma a compact catalog of recovered metadata — not every original image again.&lt;/p&gt;

&lt;p&gt;Gemma ranks plausible matches and returns screenshot IDs with short reasons.&lt;/p&gt;

&lt;p&gt;The browser then resolves those IDs back to the original screenshot already stored in IndexedDB.&lt;/p&gt;

&lt;p&gt;That means the retrieval loop is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;vague memory → semantic match → original evidence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;vague memory → generated answer that replaces the screenshot
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The screenshot remains the source of truth.&lt;/p&gt;

&lt;h3&gt;
  
  
  I had to make structured output fail gracefully too
&lt;/h3&gt;

&lt;p&gt;Dense screenshots exposed another model failure mode: Gemma occasionally returned useful content wrapped in slightly malformed JSON.&lt;/p&gt;

&lt;p&gt;The first implementation treated that as a total failure.&lt;/p&gt;

&lt;p&gt;That was silly.&lt;/p&gt;

&lt;p&gt;The server now handles common structured-output mistakes such as code fences, trailing commas, and truncated closing brackets. If a response is still unusable, it retries once with a stricter instruction.&lt;/p&gt;

&lt;p&gt;Batch processing is also incremental. If screenshot 4 of 5 fails, the first three recovered memories remain visible instead of the whole interface looking stuck.&lt;/p&gt;

&lt;p&gt;For a weekend app, those boring failure states mattered more than adding another feature.&lt;/p&gt;

&lt;h3&gt;
  
  
  From Ollama to hosted Gemma
&lt;/h3&gt;

&lt;p&gt;The original build was fully local: Gemma ran through Ollama.&lt;/p&gt;

&lt;p&gt;That made the privacy story strong, but it created a judging problem. A live demo should not depend on my laptop staying online, and local vision inference was too slow on my hardware anyway.&lt;/p&gt;

&lt;p&gt;So for the submitted version, I moved the same Gemma family to &lt;strong&gt;Hugging Face Inference Providers&lt;/strong&gt;, currently serving &lt;code&gt;google/gemma-3-4b-it&lt;/code&gt;, and deployed the app on Vercel.&lt;/p&gt;

&lt;p&gt;The browser archive is still local.&lt;/p&gt;

&lt;p&gt;The hosted backend does not maintain a screenshot database. It receives screenshot content or OCR-derived text for inference and returns the recovered memory.&lt;/p&gt;

&lt;p&gt;I want to be explicit about that distinction: the submitted version is &lt;strong&gt;not fully on-device&lt;/strong&gt; anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;The most useful thing about building this around Gemma was not a slogan about open models.&lt;/p&gt;

&lt;p&gt;It was &lt;strong&gt;portability&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I started with Gemma 3 4B on my own machine through Ollama.&lt;/p&gt;

&lt;p&gt;When local inference turned out to be the wrong deployment choice for judging, I did not have to redesign the product around a completely different intelligence layer.&lt;/p&gt;

&lt;p&gt;I could keep the same model family and move inference to a hosted provider.&lt;/p&gt;

&lt;p&gt;That mattered because the product's behavior had already been shaped around Gemma:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;structured screenshot memories&lt;/li&gt;
&lt;li&gt;semantic “why saved” reconstruction&lt;/li&gt;
&lt;li&gt;fuzzy retrieval&lt;/li&gt;
&lt;li&gt;multimodal fallback&lt;/li&gt;
&lt;li&gt;strict proper-noun handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Open-weight AI let the model be part of the architecture instead of a dependency on one closed endpoint.&lt;/p&gt;

&lt;p&gt;It also made the local version real, not hypothetical. I was able to run the same core intelligence on my own hardware while iterating on the product.&lt;/p&gt;

&lt;p&gt;And the biggest lesson from the build was that “use AI for everything” was the wrong approach anyway.&lt;/p&gt;

&lt;p&gt;Open tools gave me the freedom to compose the system:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RapidOCR for literal evidence. Gemma for interpretation. Browser storage for memory.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Each part does the job it is actually good at.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Improve Next
&lt;/h2&gt;

&lt;p&gt;I am intentionally not turning the weekend build into a giant photo-management platform.&lt;/p&gt;

&lt;p&gt;The core loop I care about is still:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;save screenshot → recover meaning → forget the exact details → find it again&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The next improvements would stay inside that loop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;stronger OCR grounding for unusual names&lt;/li&gt;
&lt;li&gt;faster batch ingestion&lt;/li&gt;
&lt;li&gt;duplicate screenshot detection&lt;/li&gt;
&lt;li&gt;an optional fully local mode for people who prefer on-device inference&lt;/li&gt;
&lt;li&gt;better confidence signals when the model is uncertain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No accounts. No social layer. No giant dashboard.&lt;/p&gt;

&lt;p&gt;The graveyard only needs to remember.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Best Use of Gemma
&lt;/h3&gt;

&lt;p&gt;I am entering &lt;strong&gt;Best Use of Gemma&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Gemma is not an optional assistant attached to Screenshot Graveyard. It performs the semantic work that makes the product useful: turning screenshots into personal memories, reconstructing likely intent, handling multimodal fallback, and matching fuzzy future queries back to saved evidence.&lt;/p&gt;

&lt;p&gt;The build also exercised one of the reasons an open-weight model matters in practice: I moved the same core model from local Ollama inference to hosted Hugging Face inference without changing the product into something else.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Screenshot Graveyard&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Your camera roll remembers everything.&lt;/p&gt;

&lt;p&gt;You remember none of it.&lt;/p&gt;

</description>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>VoiceDebt: the voice-note inbox for people with good intentions</title>
      <dc:creator>ShalyX</dc:creator>
      <pubDate>Mon, 05 Oct 2026 02:43:20 +0000</pubDate>
      <link>https://dev.to/shalyx/voicedebt-the-voice-note-inbox-for-people-with-good-intentions-539</link>
      <guid>https://dev.to/shalyx/voicedebt-the-voice-note-inbox-for-people-with-good-intentions-539</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h1&gt;
  
  
  devchallenge #weekendchallenge #hf26challenge
&lt;/h1&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;My friend sends voice notes like podcasts.&lt;/p&gt;

&lt;p&gt;Eight minutes. Twelve minutes. Long enough that I can listen with every intention of replying properly — and still finish with three questions unanswered, one link unsent, and a detail I absolutely meant to acknowledge.&lt;/p&gt;

&lt;p&gt;That became VoiceDebt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;VoiceDebt turns the voice notes you keep meaning to reply to into the things you actually owe the sender.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not just a summary.&lt;/p&gt;

&lt;p&gt;It separates a voice note into four useful layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;THE ACTUAL STORY&lt;/strong&gt; — what happened&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;YOU OWE THEM&lt;/strong&gt; — the questions, promises, decisions, calls, or things you actually need to send back&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DO NOT FORGET&lt;/strong&gt; — dates, personal details, sender commitments, and context worth remembering&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SUGGESTED REPLY&lt;/strong&gt; — a concise reply that covers the important bits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key word is &lt;strong&gt;owe&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A friend can tell me they quit their job, ask if I am free Saturday, remind me to send an Airbnb link, mention an interview on Tuesday, and then spend another five minutes on unrelated lore.&lt;/p&gt;

&lt;p&gt;A normal summary compresses that into one paragraph.&lt;/p&gt;

&lt;p&gt;VoiceDebt tries to preserve the social obligations hiding inside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bug That Defined the Product
&lt;/h2&gt;

&lt;p&gt;The synthetic demo worked.&lt;/p&gt;

&lt;p&gt;The first real voice note was more useful.&lt;/p&gt;

&lt;p&gt;I uploaded a 23-second message from a sales rep explaining that an address could not be edited after payment and that &lt;strong&gt;they&lt;/strong&gt; would process a refund the next morning.&lt;/p&gt;

&lt;p&gt;VoiceDebt understood the story, but the first version produced this under &lt;strong&gt;YOU OWE THEM&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Confirm when the refund will be processed tomorrow morning.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That was backwards.&lt;/p&gt;

&lt;p&gt;The sales rep owed &lt;strong&gt;me&lt;/strong&gt; the refund.&lt;/p&gt;

&lt;p&gt;That one mistake forced me to define the product more carefully:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Reply debt is always from the listener's perspective.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;After the fix, the exact same note correctly showed &lt;strong&gt;NO DEBT&lt;/strong&gt;, while the promised refund timing moved into &lt;strong&gt;DO NOT FORGET&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A second real note exposed another boundary. Someone explained transport options, prices, drop-off points, and warnings. VoiceDebt initially turned "you'll need to find your way from the gate" into a task.&lt;/p&gt;

&lt;p&gt;But that is not reply debt either.&lt;/p&gt;

&lt;p&gt;It is logistics.&lt;/p&gt;

&lt;p&gt;So the final rule became:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;VoiceDebt tracks social and communication obligations to the sender — not every task implied by a conversation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction is the heart of the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Live app:&lt;/strong&gt; &lt;a href="https://voicedebt.onrender.com" rel="noopener noreferrer"&gt;https://voicedebt.onrender.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Source:&lt;/strong&gt; &lt;a href="https://github.com/ShalyX/voicedebt" rel="noopener noreferrer"&gt;https://github.com/ShalyX/voicedebt&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The core flow is deliberately small:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Drop in a voice note.&lt;/li&gt;
&lt;li&gt;Whisper transcribes it.&lt;/li&gt;
&lt;li&gt;Gemma separates the story from the listener's actual reply debt.&lt;/li&gt;
&lt;li&gt;VoiceDebt surfaces the things worth remembering.&lt;/li&gt;
&lt;li&gt;It drafts a reply.&lt;/li&gt;
&lt;li&gt;Only notes with real listener obligations enter the Reply Debt inbox.&lt;/li&gt;
&lt;li&gt;Send the reply and clear the debt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The public app also includes a synthetic Amaka example so the interaction can be explored without exposing anyone else's private conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;VoiceDebt is intentionally simple.&lt;/p&gt;

&lt;p&gt;The frontend is plain HTML, CSS, and JavaScript. The backend is a small Node server using built-in APIs. There is no framework-heavy application layer and no server-side conversation database.&lt;/p&gt;

&lt;p&gt;The AI path is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;voice note
   ↓
Whisper large-v3
   ↓
transcript
   ↓
Gemma 3
   ↓
structured reply debt
   ├── actual story
   ├── listener obligations
   ├── do-not-forget context
   └── suggested reply
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Whisper for transcription
&lt;/h3&gt;

&lt;p&gt;Audio is sent to &lt;code&gt;openai/whisper-large-v3&lt;/code&gt; through Hugging Face Inference Providers.&lt;/p&gt;

&lt;p&gt;The transcript remains visible in the UI because speech recognition is not perfect. Real testing already surfaced examples such as "Airbnb" becoming "urban" and "Lekki" becoming "Leckie."&lt;/p&gt;

&lt;p&gt;Instead of pretending those errors do not happen, VoiceDebt keeps the source transcript inspectable and tells the reasoning model not to turn garbled or uncertain phrases into confident facts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemma for reply-debt reasoning
&lt;/h3&gt;

&lt;p&gt;The transcript is then sent to &lt;code&gt;google/gemma-3-12b-it&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Gemma returns a strict structure containing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;summary&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;debt[]&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;remember[]&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;reply&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;vibe&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The interesting part was not getting an LLM to write a reply.&lt;/p&gt;

&lt;p&gt;The interesting part was teaching it what &lt;strong&gt;counts as debt&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The final reasoning rules are conservative:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an explicit question can be debt&lt;/li&gt;
&lt;li&gt;"send me X" can be debt&lt;/li&gt;
&lt;li&gt;"call me when you can" can be debt&lt;/li&gt;
&lt;li&gt;a promise the listener previously made can be debt&lt;/li&gt;
&lt;li&gt;a sender promising a refund is &lt;strong&gt;not&lt;/strong&gt; listener debt&lt;/li&gt;
&lt;li&gt;advice is not debt&lt;/li&gt;
&lt;li&gt;a price is not debt&lt;/li&gt;
&lt;li&gt;a transport recommendation is not debt&lt;/li&gt;
&lt;li&gt;something the listener may need to do for themselves is not automatically debt&lt;/li&gt;
&lt;li&gt;when the obligation is ambiguous, omit it instead of manufacturing one&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And zero debt is a completely valid result.&lt;/p&gt;

&lt;h3&gt;
  
  
  A local Reply Debt inbox
&lt;/h3&gt;

&lt;p&gt;The analyzed inbox lives in the browser with &lt;code&gt;localStorage&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;VoiceDebt itself does not persist uploaded audio or transcripts in a server-side history.&lt;/p&gt;

&lt;p&gt;Only notes with actual listener obligations count toward the Reply Debt inbox. Informational notes can still be analyzed and remembered without making the user feel artificially behind.&lt;/p&gt;

&lt;h3&gt;
  
  
  Render
&lt;/h3&gt;

&lt;p&gt;Render hosts the public Node application and the server-side inference flow.&lt;/p&gt;

&lt;p&gt;The repository includes a &lt;code&gt;render.yaml&lt;/code&gt; Blueprint with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the Node runtime&lt;/li&gt;
&lt;li&gt;the health endpoint&lt;/li&gt;
&lt;li&gt;Whisper and Gemma model configuration&lt;/li&gt;
&lt;li&gt;a secret Hugging Face token supplied through Render rather than committed to source&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That made it easy to keep the repo reproducible while still deploying a real working inference path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Open Innovation Matters
&lt;/h2&gt;

&lt;p&gt;Voice notes are unusually personal data.&lt;/p&gt;

&lt;p&gt;They contain family updates, work problems, money, addresses, relationships, jokes, plans, names, and all the context people do not usually put into polished text.&lt;/p&gt;

&lt;p&gt;Using open models matters here for more than putting an "open" badge on the stack.&lt;/p&gt;

&lt;p&gt;Whisper and Gemma keep the intelligence-heavy parts of VoiceDebt replaceable and inspectable.&lt;/p&gt;

&lt;p&gt;I can change the speech model.&lt;/p&gt;

&lt;p&gt;I can swap the reasoning model.&lt;/p&gt;

&lt;p&gt;I can change the exact semantics of "reply debt" in my own prompt and application logic.&lt;/p&gt;

&lt;p&gt;And the product can evolve toward more local inference without needing to be redesigned around one proprietary assistant.&lt;/p&gt;

&lt;p&gt;The current demo uses Hugging Face-hosted inference, so I do &lt;strong&gt;not&lt;/strong&gt; claim that audio stays entirely on-device.&lt;/p&gt;

&lt;p&gt;VoiceDebt keeps no server-side conversation history, while the analyzed inbox is stored locally in the browser.&lt;/p&gt;

&lt;p&gt;That distinction matters to me.&lt;/p&gt;

&lt;p&gt;"Open" should describe the actual architecture, not become vague privacy marketing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;The best part of this build was that the hardest problem was not technical plumbing.&lt;/p&gt;

&lt;p&gt;Whisper worked.&lt;/p&gt;

&lt;p&gt;Gemma worked.&lt;/p&gt;

&lt;p&gt;Render worked.&lt;/p&gt;

&lt;p&gt;The harder problem was deciding what the product actually means by &lt;strong&gt;debt&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Real voice notes forced that definition to get sharper:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;someone else's promise is not your task&lt;/li&gt;
&lt;li&gt;useful advice is not automatically an obligation&lt;/li&gt;
&lt;li&gt;context and action are different&lt;/li&gt;
&lt;li&gt;a model should be allowed to say "nothing you owe them right now"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is what made VoiceDebt feel less like an audio summarizer and more like a product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current Limitations
&lt;/h2&gt;

&lt;p&gt;Whisper can still mishear proper nouns, accents, or casual speech.&lt;/p&gt;

&lt;p&gt;The inbox is browser-local, so it does not sync across devices.&lt;/p&gt;

&lt;p&gt;And VoiceDebt is intentionally not trying to ingest WhatsApp, Telegram, iMessage, contacts, or entire messaging histories in this weekend version.&lt;/p&gt;

&lt;p&gt;The goal was to get one small behavior right:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I listened to your voice note. What do I actually need to reply to?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Best Use of Gemma
&lt;/h3&gt;

&lt;p&gt;Gemma is the reasoning core of VoiceDebt. It turns an unstructured transcript into listener-specific reply debt, memorable context, and a reply draft while enforcing the directionality and ambiguity rules that emerged from real-world testing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Best Use of Render
&lt;/h3&gt;

&lt;p&gt;Render hosts the complete public VoiceDebt application and its server-side inference path, with environment-managed secrets and the repository's Blueprint configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Live demo:&lt;/strong&gt; &lt;a href="https://voicedebt.onrender.com" rel="noopener noreferrer"&gt;https://voicedebt.onrender.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/ShalyX/voicedebt" rel="noopener noreferrer"&gt;https://github.com/ShalyX/voicedebt&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
