What I Built
NannaDesk is a personal AI assistant built for one real person: my dad. He's a government employee in Telangana, India — comfortable talking to AI, not comfortable with technology in general. He needs help with four specific things: remembering whether he took his medicine, understanding his blood reports and government paperwork, following up on office emails that have gone quiet, and getting small administrative things done instead of just explained.
It runs on his laptop. A small open-weight model (Qwen3.5-0.8B) handles the conversation. Whisper.cpp handles his voice, including when he mixes Telugu and English mid-sentence, which he does constantly. ElevenLabs reads answers back to him when he'd rather listen than read. None of that is the interesting part, though.
The interesting part is the rule I built everything else around:
The LLM decides how to help. Deterministic software decides what actually happened.
NannaDesk's model is not allowed to tell my dad he took a medicine he didn't confirm taking. It's not allowed to say an email sent if the send endpoint didn't return success. It's not allowed to invent a government policy, or guess at a lab value, or decide on its own that a report is "probably fine." Every one of those is a database write that only happens after an explicit, logged, human action — and the model only ever narrates what already happened, never what it thinks probably happened.
For an assistant that's going to manage someone's medication schedule, I didn't want "probably."
Demo
- Code: github.com/sakethbalijepalli/NannaDesk
- Screenshots throughout this post are from the actual running app — the "Venkata Rao" profile is sanitized demo data (the repo ships with a seed script for exactly this reason); my dad's real profile goes in locally through the Settings screen below and never leaves his machine.
Ask a question, get a drafted email, reviewed and sent — not auto-sent:
Real data entry, not a script someone has to edit for him:
How I Built It
The open-source AI layer:
- Qwen3.5-0.8B (open-weight, via unsloth's GGUF release) running locally through llama.cpp. No API key, no per-token cost, no network call for the core assistant.
- whisper.cpp for local speech-to-text.
- ElevenLabs for the one piece I didn't build locally — text-to-speech, so he can hear an answer instead of reading it. This is optional and additive: with no API key configured, the app degrades silently to text-only. Nothing else depends on it.
The deterministic layer that the LLM is not allowed to touch:
- A SQLite-backed state machine for medication and meal confirmations —
SCHEDULED → REMINDER_SENT → TAKEN, with explicitSNOOZED,NOT_CONFIRMED, andSKIPPEDstates. A reminder fires on a real schedule (APScheduler), sends a real macOS notification, and only a user-confirmed action moves the state machine forward. - A keyword/regex router instead of LLM tool-calling, on purpose. A 0.8B model cannot be trusted to reliably emit well-formed tool-call JSON — that's not a guess, I tested it. So routing ("is this a medication confirmation, a government question, an email request?") is deterministic and unit-tested, and the model is only ever asked to phrase a response from a result that's already been decided and fetched.
- A Pydantic-validated email pipeline where the model can create a draft but has no tool that can send one. Sending is an HTTP endpoint that requires
confirm: truefrom a human tap, full stop.
Where open-source AI genuinely made the build better, not just cheaper:
Midway through, I needed NannaDesk to answer "any update on PRC?" (Pay Revision Commission — a live topic for Indian government employees) from an actual official source instead of guessing. The obvious URL for this — telangana.gov.in/government-orders/ — turned out, when I actually fetched and read it, to not be a government-orders listing at all. It's a generic info page. The real repository is goir.telangana.gov.in, an ASP.NET site whose search is a form postback with view-state — not something a plain HTTP request can drive at all. I ended up scripting a real headless browser (Playwright) to fill in that form and read the results table directly, because that was the only way to get real government data instead of a plausible-sounding guess.
That's the kind of problem that doesn't show up until you refuse to let the model paper over it with a confident-sounding answer. Open, local, inspectable tooling is what let me go find and fix the actual root cause instead of prompt-engineering around a bad source.
Why Does Open Innovation Matter?
It runs on a laptop with no internet. Medicine reminders, meal confirmations, report explanations, "what's pending" — all of it works with Wi-Fi off, because I built and tested it that way deliberately. Only two things need a connection: sending an email and looking up live government information, and the app says so plainly when it can't reach either.
It keeps his data off a server he doesn't control. My dad's medication history, his blood report values, his PF case details — all of it lives in one SQLite file on his laptop. Gitignored, never uploaded, never seen by me unless he shows me his screen. For someone who isn't going to read a privacy policy, "it doesn't leave this laptop" is a promise I can actually keep, because the architecture makes it true rather than asking him to trust a toggle.
It's swappable. The app talks to an LLMProvider interface, not to Qwen specifically. When a better small open model ships, I change one class, not the application.
It costs nothing to run. No inference bill for the actual assistant. The only metered dependency (ElevenLabs) is optional, and the app works completely without it.
Where open beat closed, concretely: I could not have built the "never guess" architecture the same way on top of a closed hosted agent API. Being able to see — and test — exactly why the model did or didn't narrate a confirmation, because that decision lives in code I wrote and can read, not inside someone else's black box, is the whole reason this assistant is safe to actually hand to my dad.
My Agent Session
I built this with Claude Code and saved the real session rather than writing this post from memory afterward. It includes the part I'm proudest of: the moment I found out the government-lookup URL I'd built against wasn't even the right page, and the fix that followed.
Prize Categories
- Best Use of ElevenLabs
Status
My dad hasn't used this yet — the Settings screen that lets real data in (not the sanitized demo data in these screenshots) just went live. The actual "hand it over and see what he says" part comes next, and I'll update here when it happens.
81 automated tests, all deterministic — the suite stubs the LLM as unavailable so nothing depends on a model actually being loaded to verify the medication state machine, the email safety gate, or the government-source separation behave correctly.




Top comments (0)