DEV Community

sakethbalijepalli
sakethbalijepalli

Posted on

NannaDesk: An AI Assistant for My Dad That's Not Allowed to Guess

What I Built

NannaDesk is a personal AI assistant built for one real person: my dad. He's a government employee in Telangana, India — comfortable talking to AI, not comfortable with technology in general. He needs help with four specific things: remembering whether he took his medicine, understanding his blood reports and government paperwork, following up on office emails that have gone quiet, and getting small administrative things done instead of just explained.

It runs on his laptop. A small open-weight model (Qwen3.5-0.8B) handles the conversation. Whisper.cpp handles his voice, including when he mixes Telugu and English mid-sentence, which he does constantly. ElevenLabs reads answers back to him when he'd rather listen than read. None of that is the interesting part, though.

The interesting part is the rule I built everything else around:

The LLM decides how to help. Deterministic software decides what actually happened.

NannaDesk's model is not allowed to tell my dad he took a medicine he didn't confirm taking. It's not allowed to say an email sent if the send endpoint didn't return success. It's not allowed to invent a government policy, or guess at a lab value, or decide on its own that a report is "probably fine." Every one of those is a database write that only happens after an explicit, logged, human action — and the model only ever narrates what already happened, never what it thinks probably happened.

For an assistant that's going to manage someone's medication schedule, I didn't want "probably."

NannaDesk home screen — today's confirmed routines in a ledger, not a chat bubble

Demo

  • Code: github.com/sakethbalijepalli/NannaDesk
  • Screenshots throughout this post are from the actual running app — the "Venkata Rao" profile is sanitized demo data (the repo ships with a seed script for exactly this reason); my dad's real profile goes in locally through the Settings screen below and never leaves his machine.

Ask a question, get a drafted email, reviewed and sent — not auto-sent:

Chat-driven email draft, reviewed inline, sent only after an explicit tap

Real data entry, not a script someone has to edit for him:

Settings screen for profile and medications

How I Built It

The open-source AI layer:

  • Qwen3.5-0.8B (open-weight, via unsloth's GGUF release) running locally through llama.cpp. No API key, no per-token cost, no network call for the core assistant.
  • whisper.cpp for local speech-to-text.
  • ElevenLabs for the one piece I didn't build locally — text-to-speech, so he can hear an answer instead of reading it. This is optional and additive: with no API key configured, the app degrades silently to text-only. Nothing else depends on it.

The deterministic layer that the LLM is not allowed to touch:

  • A SQLite-backed state machine for medication and meal confirmations — SCHEDULED → REMINDER_SENT → TAKEN, with explicit SNOOZED, NOT_CONFIRMED, and SKIPPED states. A reminder fires on a real schedule (APScheduler), sends a real macOS notification, and only a user-confirmed action moves the state machine forward.
  • A keyword/regex router instead of LLM tool-calling, on purpose. A 0.8B model cannot be trusted to reliably emit well-formed tool-call JSON — that's not a guess, I tested it. So routing ("is this a medication confirmation, a government question, an email request?") is deterministic and unit-tested, and the model is only ever asked to phrase a response from a result that's already been decided and fetched.
  • A Pydantic-validated email pipeline where the model can create a draft but has no tool that can send one. Sending is an HTTP endpoint that requires confirm: true from a human tap, full stop.

Where open-source AI genuinely made the build better, not just cheaper:

Midway through, I needed NannaDesk to answer "any update on PRC?" (Pay Revision Commission — a live topic for Indian government employees) from an actual official source instead of guessing. The obvious URL for this — telangana.gov.in/government-orders/ — turned out, when I actually fetched and read it, to not be a government-orders listing at all. It's a generic info page. The real repository is goir.telangana.gov.in, an ASP.NET site whose search is a form postback with view-state — not something a plain HTTP request can drive at all. I ended up scripting a real headless browser (Playwright) to fill in that form and read the results table directly, because that was the only way to get real government data instead of a plausible-sounding guess.

A real government-order search result — department, order number, date, and subject, not a guess

That's the kind of problem that doesn't show up until you refuse to let the model paper over it with a confident-sounding answer. Open, local, inspectable tooling is what let me go find and fix the actual root cause instead of prompt-engineering around a bad source.

Why Does Open Innovation Matter?

It runs on a laptop with no internet. Medicine reminders, meal confirmations, report explanations, "what's pending" — all of it works with Wi-Fi off, because I built and tested it that way deliberately. Only two things need a connection: sending an email and looking up live government information, and the app says so plainly when it can't reach either.

It keeps his data off a server he doesn't control. My dad's medication history, his blood report values, his PF case details — all of it lives in one SQLite file on his laptop. Gitignored, never uploaded, never seen by me unless he shows me his screen. For someone who isn't going to read a privacy policy, "it doesn't leave this laptop" is a promise I can actually keep, because the architecture makes it true rather than asking him to trust a toggle.

It's swappable. The app talks to an LLMProvider interface, not to Qwen specifically. When a better small open model ships, I change one class, not the application.

It costs nothing to run. No inference bill for the actual assistant. The only metered dependency (ElevenLabs) is optional, and the app works completely without it.

Where open beat closed, concretely: I could not have built the "never guess" architecture the same way on top of a closed hosted agent API. Being able to see — and test — exactly why the model did or didn't narrate a confirmation, because that decision lives in code I wrote and can read, not inside someone else's black box, is the whole reason this assistant is safe to actually hand to my dad.

My Agent Session

I built this with Claude Code and saved the real session rather than writing this post from memory afterward. It includes the part I'm proudest of: the moment I found out the government-lookup URL I'd built against wasn't even the right page, and the fix that followed.

Building NannaDesk: A Local AI Assistant for My Dad
You

start building the NannaDesk project from the spec pdf in the directory

Agent

Read the spec. It's a 25-section hackathon build document for a local-first AI assistant for one specific non-technical person — the user's dad. The one line that sets the whole architecture: 'LLM decides how to help; deterministic software decides what actually happens.' Before touching code, consulted an advisor on approach given the scope.

Advisor's key call: a 0.8B model can't be trusted to reliably emit tool-call JSON, so route deterministically with keyword/regex and only ask the model to phrase an already-decided result. Also: build the deterministic core first since it needs no model to test.

You

Install everything now

Agent

Installed llama.cpp, whisper.cpp, ffmpeg via brew, downloaded the Qwen3.5-0.8B GGUF and a whisper model, scaffolded the FastAPI backend, SQLite schema (events/reminders/cases/documents/lab_results with two separate FTS5 tables so personal and government content can never cross-contaminate), the medication state machine, PDF report parsing, a mock email provider, and a React/Vite frontend with a custom 'household ledger' visual design — not a generic SaaS dashboard, since the whole product is built around one specific government-employee father.

Ran two self-review passes plus a /code-review pass before calling it done. Found and fixed real bugs each time: reminders firing in UTC instead of local time, snooze not actually delaying, a government query matching page footer boilerplate instead of real content.

You

go after the government lookup gap next

Agent

The government-lookup flow was the weakest of the five demo flows — it always said 'nothing found.' Checked the actual page being fetched, with real JS rendering via Playwright to rule out a client-rendering problem.

Searched for how Telangana actually publishes government orders and found the real repository: goir.telangana.gov.in, an ASP.NET site whose search is a view-state form postback — not something a plain HTTP GET can drive at all.

Built a persistent headless Chromium in the FastAPI lifespan to drive that real search form, with a static-page fetch as fallback. Verified live through the actual chat UI, not just curl.

You

/code-review

Agent

The review found two real safety bugs — the kind that matter most in an app managing someone's medication. First: the negation guard I'd added for named-medicine confirmations wasn't applied to the generic confirmation path, so 'I had not taken my tablet' recorded TAKEN. Second: naming an already-confirmed medicine when exactly one other was still pending silently confirmed that unrelated one instead.

Fixed both, added regression tests reproducing the exact failing phrases, verified live against the running app before and after. Also fixed two substring-matching bugs ('da' matching inside 'update') the same review caught.

You

whats the progress here? Is the application done end to end?

Agent

Gave an honest answer: yes for the hackathon MVP scope, all 5 demo flows verified through the real UI with a real local LLM, real voice transcription, 67 passing tests at that point. But flagged the real gap plainly — this still only runs as three hand-started dev processes against fictional demo data, not something the actual target user (a non-technical father) could use without someone setting it up for him.

You

/devrelay-mlh-submissions

Agent

Found the exact matching DEV Challenge by its judging criteria and theme text the user pasted — 'Build for a Friend,' which is literally the spec's own tagline. Checked git commit timestamps against the challenge window to confirm eligibility before going further. Then researched every partner-tech option listed (Render, Tinker, Gemma, TabPFN, Backboard, ElevenLabs) against what was actually available and what genuinely fit the product.

Recommended ElevenLabs specifically over Backboard (cloud memory) and Render (cloud hosting) because those would have actively contradicted the entry's own core claim — that Dad's data never leaves his laptop. ElevenLabs closes a real gap instead: the app could only listen (whisper.cpp), never speak back.

You

Proceed with ElevenLabs text-to-speech

Agent

Claimed the credit, installed ElevenLabs' own official skill bundle and followed its documented convention exactly — a plain ELEVENLABS_API_KEY env var via their own SDK, not translated into the app's own prefixed settings scheme. Added an on-demand speaker button per chat reply that degrades silently to text-only with no key configured, same pattern already used for Outlook and the local-LLM-unavailable path. Verified the fallback live in the actual browser before the key was even set up.

You

Let's also add real data, instead of mock data.

Agent

Before building anything, surfaced the real gap: there was no way to enter real data except hand-editing the demo seed script, which isn't something the target user could ever do himself. Built a proper Settings screen — real profile and medication CRUD through the browser — plus a --wipe-only flag to clear the fictional seed first. Verified the full save -> reload -> persisted round-trip live.

You

Push to github

Agent

SSH auth to GitHub was already working, but creating a new public repo is outward-facing enough that it belonged to the user, not a script — asked them to create the empty repo rather than trying to automate around it. Pushed via the working SSH credential once it existed, then installed and authenticated the GitHub CLI to set the repo description.

You

start drafting dev post next.

Agent

Drafted the submission post leaning on the real, specific material from the build rather than generic hackathon-post language — the GOIR wrong-URL story, the two safety bugs, the architecture thesis — since Writing Quality is the heaviest-weighted judging criterion. Committed and hosted real screenshots from the actual running app via raw.githubusercontent.com, verified each one actually loads before embedding. Left it as a draft for review rather than publishing unprompted.

Prize Categories

  • Best Use of ElevenLabs

Status

My dad hasn't used this yet — the Settings screen that lets real data in (not the sanitized demo data in these screenshots) just went live. The actual "hand it over and see what he says" part comes next, and I'll update here when it happens.


81 automated tests, all deterministic — the suite stubs the LLM as unavailable so nothing depends on a model actually being loaded to verify the medication state machine, the email safety gate, or the government-source separation behave correctly.

Top comments (0)