<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sakethbalijepalli</title>
    <description>The latest articles on DEV Community by sakethbalijepalli (@sakethbalijepalli).</description>
    <link>https://dev.to/sakethbalijepalli</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F650540%2Fda9fa51d-c890-4811-a165-f42d9c1f1979.png</url>
      <title>DEV Community: sakethbalijepalli</title>
      <link>https://dev.to/sakethbalijepalli</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sakethbalijepalli"/>
    <language>en</language>
    <item>
      <title>NannaDesk: An AI Assistant for My Dad That's Not Allowed to Guess</title>
      <dc:creator>sakethbalijepalli</dc:creator>
      <pubDate>Sat, 03 Oct 2026 13:27:22 +0000</pubDate>
      <link>https://dev.to/sakethbalijepalli/nannadesk-an-ai-assistant-for-my-dad-thats-not-allowed-to-guess-242l</link>
      <guid>https://dev.to/sakethbalijepalli/nannadesk-an-ai-assistant-for-my-dad-thats-not-allowed-to-guess-242l</guid>
      <description>&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;NannaDesk is a personal AI assistant built for one real person: my dad. He's a government employee in Telangana, India — comfortable talking to AI, not comfortable with technology in general. He needs help with four specific things: remembering whether he took his medicine, understanding his blood reports and government paperwork, following up on office emails that have gone quiet, and getting small administrative things &lt;em&gt;done&lt;/em&gt; instead of just explained.&lt;/p&gt;

&lt;p&gt;It runs on his laptop. A small open-weight model (Qwen3.5-0.8B) handles the conversation. Whisper.cpp handles his voice, including when he mixes Telugu and English mid-sentence, which he does constantly. ElevenLabs reads answers back to him when he'd rather listen than read. None of that is the interesting part, though.&lt;/p&gt;

&lt;p&gt;The interesting part is the rule I built everything else around:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The LLM decides how to help. Deterministic software decides what actually happened.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;NannaDesk's model is not allowed to tell my dad he took a medicine he didn't confirm taking. It's not allowed to say an email sent if the send endpoint didn't return success. It's not allowed to invent a government policy, or guess at a lab value, or decide on its own that a report is "probably fine." Every one of those is a database write that only happens after an explicit, logged, human action — and the model only ever narrates what already happened, never what it thinks probably happened.&lt;/p&gt;

&lt;p&gt;For an assistant that's going to manage someone's medication schedule, I didn't want "probably."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fth83ge8te2j3wbjrl5o4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fth83ge8te2j3wbjrl5o4.png" alt="NannaDesk home screen — today's confirmed routines in a ledger, not a chat bubble" width="480" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Code:&lt;/strong&gt; &lt;a href="https://github.com/sakethbalijepalli/NannaDesk" rel="noopener noreferrer"&gt;github.com/sakethbalijepalli/NannaDesk&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Screenshots throughout this post are from the actual running app — the "Venkata Rao" profile is sanitized demo data (the repo ships with a seed script for exactly this reason); my dad's real profile goes in locally through the Settings screen below and never leaves his machine.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Ask a question, get a drafted email, reviewed and sent — not auto-sent:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq8tja8qjok3t7k07z7kl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq8tja8qjok3t7k07z7kl.png" alt="Chat-driven email draft, reviewed inline, sent only after an explicit tap" width="480" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real data entry, not a script someone has to edit for him:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhr8fd69dknrgf070qrth.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhr8fd69dknrgf070qrth.png" alt="Settings screen for profile and medications" width="480" height="2022"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The open-source AI layer:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.5-0.8B&lt;/strong&gt; (open-weight, via unsloth's GGUF release) running locally through &lt;strong&gt;llama.cpp&lt;/strong&gt;. No API key, no per-token cost, no network call for the core assistant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;whisper.cpp&lt;/strong&gt; for local speech-to-text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ElevenLabs&lt;/strong&gt; for the one piece I didn't build locally — text-to-speech, so he can hear an answer instead of reading it. This is optional and additive: with no API key configured, the app degrades silently to text-only. Nothing else depends on it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The deterministic layer that the LLM is not allowed to touch:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A SQLite-backed state machine for medication and meal confirmations — &lt;code&gt;SCHEDULED → REMINDER_SENT → TAKEN&lt;/code&gt;, with explicit &lt;code&gt;SNOOZED&lt;/code&gt;, &lt;code&gt;NOT_CONFIRMED&lt;/code&gt;, and &lt;code&gt;SKIPPED&lt;/code&gt; states. A reminder fires on a real schedule (APScheduler), sends a real macOS notification, and only a user-confirmed action moves the state machine forward.&lt;/li&gt;
&lt;li&gt;A keyword/regex router instead of LLM tool-calling, on purpose. A 0.8B model cannot be trusted to reliably emit well-formed tool-call JSON — that's not a guess, I tested it. So routing ("is this a medication confirmation, a government question, an email request?") is deterministic and unit-tested, and the model is only ever asked to &lt;em&gt;phrase&lt;/em&gt; a response from a result that's already been decided and fetched.&lt;/li&gt;
&lt;li&gt;A Pydantic-validated email pipeline where the model can create a draft but has no tool that can send one. Sending is an HTTP endpoint that requires &lt;code&gt;confirm: true&lt;/code&gt; from a human tap, full stop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where open-source AI genuinely made the build better, not just cheaper:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Midway through, I needed NannaDesk to answer "any update on PRC?" (Pay Revision Commission — a live topic for Indian government employees) from an actual official source instead of guessing. The obvious URL for this — &lt;code&gt;telangana.gov.in/government-orders/&lt;/code&gt; — turned out, when I actually fetched and read it, to not be a government-orders listing at all. It's a generic info page. The real repository is &lt;code&gt;goir.telangana.gov.in&lt;/code&gt;, an ASP.NET site whose search is a form postback with view-state — not something a plain HTTP request can drive at all. I ended up scripting a real headless browser (Playwright) to fill in that form and read the results table directly, because that was the only way to get &lt;em&gt;real&lt;/em&gt; government data instead of a plausible-sounding guess.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1oxyf1ekpocthh2xsdmb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1oxyf1ekpocthh2xsdmb.png" alt="A real government-order search result — department, order number, date, and subject, not a guess" width="480" height="900"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That's the kind of problem that doesn't show up until you refuse to let the model paper over it with a confident-sounding answer. Open, local, inspectable tooling is what let me go find and fix the actual root cause instead of prompt-engineering around a bad source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It runs on a laptop with no internet.&lt;/strong&gt; Medicine reminders, meal confirmations, report explanations, "what's pending" — all of it works with Wi-Fi off, because I built and tested it that way deliberately. Only two things need a connection: sending an email and looking up live government information, and the app says so plainly when it can't reach either.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It keeps his data off a server he doesn't control.&lt;/strong&gt; My dad's medication history, his blood report values, his PF case details — all of it lives in one SQLite file on his laptop. Gitignored, never uploaded, never seen by me unless he shows me his screen. For someone who isn't going to read a privacy policy, "it doesn't leave this laptop" is a promise I can actually keep, because the architecture makes it true rather than asking him to trust a toggle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's swappable.&lt;/strong&gt; The app talks to an &lt;code&gt;LLMProvider&lt;/code&gt; interface, not to Qwen specifically. When a better small open model ships, I change one class, not the application.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It costs nothing to run.&lt;/strong&gt; No inference bill for the actual assistant. The only metered dependency (ElevenLabs) is optional, and the app works completely without it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where open beat closed, concretely:&lt;/strong&gt; I could not have built the "never guess" architecture the same way on top of a closed hosted agent API. Being able to see — and test — exactly why the model did or didn't narrate a confirmation, because that decision lives in code I wrote and can read, not inside someone else's black box, is the whole reason this assistant is safe to actually hand to my dad.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;I built this with Claude Code and saved the real session rather than writing this post from memory afterward. It includes the part I'm proudest of: the moment I found out the government-lookup URL I'd built against wasn't even the right page, and the fix that followed.&lt;/p&gt;


&lt;div class="ltag-agent-session"&gt;
  &lt;div class="agent-session-header"&gt;
    
    &lt;span class="agent-session-tool-icon-badge" title="Claude Code"&gt;
&lt;/span&gt;
    &lt;span class="agent-session-title"&gt;Building NannaDesk: A Local AI Assistant for My Dad&lt;/span&gt;
  &lt;/div&gt;

  &lt;div class="agent-session-scroll"&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;start building the NannaDesk project from the spec pdf in the directory&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Read the spec. It's a 25-section hackathon build document for a local-first AI assistant for one specific non-technical person — the user's dad. The one line that sets the whole architecture: 'LLM decides how to help; deterministic software decides what actually happens.' Before touching code, consulted an advisor on approach given the scope.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-read"&gt;Read&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;NannaDesk_Full_Project_Specification.pdf&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;NannaDesk_Full_Project_Specification.pdf&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;25-section spec: medication state machine, document understanding, government retrieval, email workflows, local LLM (Qwen3.5-0.8B via llama.cpp), voice (whisper.cpp), SQLite schema, hackathon MVP = 5 flows&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Advisor's key call: a 0.8B model can't be trusted to reliably emit tool-call JSON, so route deterministically with keyword/regex and only ask the model to phrase an already-decided result. Also: build the deterministic core first since it needs no model to test.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Install everything now&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Installed llama.cpp, whisper.cpp, ffmpeg via brew, downloaded the Qwen3.5-0.8B GGUF and a whisper model, scaffolded the FastAPI backend, SQLite schema (events/reminders/cases/documents/lab_results with two separate FTS5 tables so personal and government content can never cross-contaminate), the medication state machine, PDF report parsing, a mock email provider, and a React/Vite frontend with a custom 'household ledger' visual design — not a generic SaaS dashboard, since the whole product is built around one specific government-employee father.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-bash"&gt;Bash&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;llama-server -m models/qwen/... --reasoning-budget 128&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre agent-session-terminal"&gt;&lt;code&gt;$ llama-server -m models/qwen/... --reasoning-budget 128&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;Caught mid-build: the model would ramble for tens of thousands of tokens before answering a one-line question without a reasoning budget cap.&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Ran two self-review passes plus a /code-review pass before calling it done. Found and fixed real bugs each time: reminders firing in UTC instead of local time, snooze not actually delaying, a government query matching page footer boilerplate instead of real content.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;go after the government lookup gap next&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;The government-lookup flow was the weakest of the five demo flows — it always said 'nothing found.' Checked the actual page being fetched, with real JS rendering via Playwright to rule out a client-rendering problem.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-bashplaywright"&gt;Bash (Playwright)&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;render telangana.gov.in/government-orders/ and search for 'prc'&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;render telangana.gov.in/government-orders/ and search for 'prc'&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;False, False, False — the page isn't a GO listing at all, JS-rendered or not. It's a generic 'order of precedence' info page. Wrong URL from the start.&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Searched for how Telangana actually publishes government orders and found the real repository: goir.telangana.gov.in, an ASP.NET site whose search is a view-state form postback — not something a plain HTTP GET can drive at all.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-bashplaywright"&gt;Bash (Playwright)&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;fill #txtSearchText with 'Pay Revision', click #BtnSearch&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;fill #txtSearchText with 'Pay Revision', click #BtnSearch&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;Real results: GO RT No.10, ITE&amp;amp;C Department, 30/03/2022 — Pay Revision arrears orders. Genuine government data for the first time.&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Built a persistent headless Chromium in the FastAPI lifespan to drive that real search form, with a static-page fetch as fallback. Verified live through the actual chat UI, not just curl.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;/code-review&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;The review found two real safety bugs — the kind that matter most in an app managing someone's medication. First: the negation guard I'd added for named-medicine confirmations wasn't applied to the generic confirmation path, so 'I had not taken my tablet' recorded TAKEN. Second: naming an already-confirmed medicine when exactly one other was still pending silently confirmed that unrelated one instead.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-bashcurl"&gt;Bash (curl)&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;POST /api/chat {message: 'I had not taken my tablet'}&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;POST /api/chat {message: 'I had not taken my tablet'}&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;Before fix: recorded TAKEN. After fix: routed to clarify, nothing recorded, confirmed via direct DB query.&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Fixed both, added regression tests reproducing the exact failing phrases, verified live against the running app before and after. Also fixed two substring-matching bugs ('da' matching inside 'update') the same review caught.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;whats the progress here? Is the application done end to end?&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Gave an honest answer: yes for the hackathon MVP scope, all 5 demo flows verified through the real UI with a real local LLM, real voice transcription, 67 passing tests at that point. But flagged the real gap plainly — this still only runs as three hand-started dev processes against fictional demo data, not something the actual target user (a non-technical father) could use without someone setting it up for him.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;/devrelay-mlh-submissions&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Found the exact matching DEV Challenge by its judging criteria and theme text the user pasted — 'Build for a Friend,' which is literally the spec's own tagline. Checked git commit timestamps against the challenge window to confirm eligibility before going further. Then researched every partner-tech option listed (Render, Tinker, Gemma, TabPFN, Backboard, ElevenLabs) against what was actually available and what genuinely fit the product.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-listeventofferslisteventagentskills"&gt;list_event_offers / list_event_agent_skills&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;Hacktoberfest 2026 event&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;Hacktoberfest 2026 event&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;Only ElevenLabs had both a real claimable credit and an official integration skill for this event. Nothing for the others.&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Recommended ElevenLabs specifically over Backboard (cloud memory) and Render (cloud hosting) because those would have actively contradicted the entry's own core claim — that Dad's data never leaves his laptop. ElevenLabs closes a real gap instead: the app could only listen (whisper.cpp), never speak back.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Proceed with ElevenLabs text-to-speech&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Claimed the credit, installed ElevenLabs' own official skill bundle and followed its documented convention exactly — a plain ELEVENLABS_API_KEY env var via their own SDK, not translated into the app's own prefixed settings scheme. Added an on-demand speaker button per chat reply that degrades silently to text-only with no key configured, same pattern already used for Outlook and the local-LLM-unavailable path. Verified the fallback live in the actual browser before the key was even set up.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Let's also add real data, instead of mock data.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Before building anything, surfaced the real gap: there was no way to enter real data except hand-editing the demo seed script, which isn't something the target user could ever do himself. Built a proper Settings screen — real profile and medication CRUD through the browser — plus a --wipe-only flag to clear the fictional seed first. Verified the full save -&amp;gt; reload -&amp;gt; persisted round-trip live.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Push to github&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;SSH auth to GitHub was already working, but creating a new public repo is outward-facing enough that it belonged to the user, not a script — asked them to create the empty repo rather than trying to automate around it. Pushed via the working SSH credential once it existed, then installed and authenticated the GitHub CLI to set the repo description.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;start drafting dev post next.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Drafted the submission post leaning on the real, specific material from the build rather than generic hackathon-post language — the GOIR wrong-URL story, the two safety bugs, the architecture thesis — since Writing Quality is the heaviest-weighted judging criterion. Committed and hosted real screenshots from the actual running app via raw.githubusercontent.com, verified each one actually loads before embedding. Left it as a draft for review rather than publishing unprompted.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;
  &lt;/div&gt;

  &lt;div class="agent-session-footer"&gt;
    &lt;span class="agent-session-meta"&gt;
        20 of 13 messages
    &lt;/span&gt;
  &lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Best Use of ElevenLabs&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Status
&lt;/h2&gt;

&lt;p&gt;My dad hasn't used this yet — the Settings screen that lets real data in (not the sanitized demo data in these screenshots) just went live. The actual "hand it over and see what he says" part comes next, and I'll update here when it happens.&lt;/p&gt;




&lt;p&gt;81 automated tests, all deterministic — the suite stubs the LLM as unavailable so nothing depends on a model actually being loaded to verify the medication state machine, the email safety gate, or the government-source separation behave correctly.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
