DEV Community

Cover image for I built my own JARVIS for Windows. The hardest part wasn't the AI
platonkuch-dev
platonkuch-dev

Posted on AI-assisted

I built my own JARVIS for Windows. The hardest part wasn't the AI

I wanted something closer to JARVIS than to a chatbot: I say a sentence, and something actually happens on my computer. No typing, no copy-pasting answers from a chat window.

After a few months of evenings, it works. It opens and closes apps, moves windows, types into programs, looks at the screen when it needs to, gives me a morning briefing, sets timers, plays music, controls the smart home and messages people on Telegram. It's about 25,000 lines of Python now, and most of what I learned had nothing to do with picking the smartest model.

Here's how it's built, and the three lessons that changed the design the most.

The pipeline

microphone → speech-to-text → LLM picks a tool → tool runs on the PC → text-to-speech → speakers
Enter fullscreen mode Exit fullscreen mode
  • Voice transport: LiveKit Agents handles the audio loop, turn detection and interruptions.
  • Speech-to-text: Deepgram, streaming.
  • Brain: Claude (Haiku for normal conversation, a bigger model only for "look at the screen and do it" tasks).
  • Voice: Edge TTS (free) or ElevenLabs.
  • Hands: about 40 tool modules: Windows UI Automation for windows and buttons, a DOM-based browser tool, files, media keys, Home Assistant, email, Telegram, timers, and more.

A small supervisor process (app.py) starts everything: the voice worker, a status bar, the Telegram bridge, a health monitor and a local control panel. When something crashes, it restarts it. If something crashes five times in five minutes, it backs off for half an hour and tells me on Telegram, instead of spinning forever.

Lesson 1: most commands should never reach the LLM

When I looked at what I actually say to the assistant, the bulk of it was boring: "volume 30", "pause", "timer for five minutes", "open Telegram", "what time is it".

Sending each of those through the LLM meant a full round trip with roughly 11k tokens of (cached) system prompt and tool descriptions, plus about a second of extra latency, only for the model to call one obvious tool.

So there's a fast path in front of the model. It's plain regex, with no model and no network:

import re

# whole-utterance matches only: anything longer or ambiguous still goes to the LLM
RULES = [
    (re.compile(r"^(set )?volume (to )?(?P<n>\d{1,3})( percent)?$"), "set_volume"),
    (re.compile(r"^(pause|stop the music)$"),                         "media_pause"),
    (re.compile(r"^(set a )?timer for (?P<n>\d+) minutes?$"),         "start_timer"),
]

def parse(utterance: str):
    text = utterance.lower().strip(" .!?")
    for pattern, tool in RULES:
        if m := pattern.match(text):
            return tool, m.groupdict()
    return None          # not sure → let the LLM handle it
Enter fullscreen mode Exit fullscreen mode

The important design rule: a false positive is worse than a miss. If the fast path does the wrong thing without asking the model, that's a bug the user feels. If it misses, the LLM handles the command exactly as before. So the rules match whole utterances only and stay conservative.

The fast path calls the same tool functions the LLM would call, so there's one implementation of "set volume", not two. The result: the most frequent commands are instant and cost nothing.

Lesson 2: latency beats intelligence

A slightly dumber answer in one second feels much better than a perfect answer in four. Things that helped the most:

  • Prompt caching. The system prompt and tool schemas are big and stable, so they're cached. A normal turn costs a fraction of a cent.
  • A small model by default. The large model is only used for multi-step "look at my screen" tasks.
  • Reading the page, not looking at it. For web tasks I replaced screenshots with a tool that reads the page's DOM. It's faster, cheaper, and the model makes fewer mistakes than when it reads pixels.
  • A daily spending cap set in the control panel, so a bug in a loop can't burn money overnight.

Lesson 3: giving an AI your PC is scary, so design for that

An assistant with real access to the file system and apps needs guard rails that don't depend on the model behaving.

  • Unattended actions need a "yes". When I'm talking to it, I'm the safety check: I asked for it and I hear what happens. Background tasks and triggers run with nobody watching, so every tool call they make is classified first: safe (read-only) runs, confirm (sends something, closes or changes something, drives the screen) waits for my "yes" in a Telegram message I can answer from my phone, and block is never done autonomously. Unknown tools default to confirm.
  • The rules live in code, not in the prompt. A prompt can be argued with; a lookup table can't.
  • Deletes go to the Recycle Bin, never straight to /dev/null.
  • The control panel listens on 127.0.0.1 only, rejects foreign Host headers and requires a token. API keys are never sent back to the browser, only their last four characters.
  • Notifications don't need a live connection. Any process can drop a line into a small JSONL outbox, and one process that holds the Telegram session delivers it. A reminder fired by a background task still reaches me even if the voice process is restarting.

Small things that made it feel like JARVIS

  • Wake word, offline. "Hey Jarvis" is detected locally with openWakeWord, so nothing is streamed until I actually talk to it. F10 puts it to sleep and wakes it up, and it falls asleep on its own after a few minutes of silence.
  • Memory. It remembers facts and preferences between sessions, plus voice macros like "when I say X, do Y".
  • Triggers. "When this app starts / a file lands in this folder / it's 8:30 / the battery is low, do something", with results sent by voice or to Telegram if I'm away from the PC.
  • A morning briefing (weather, calendar, to-dos, headlines) that runs with no LLM at all, so it's free.

Shipping it

Nobody wants to install Python to try a voice assistant, so it ships as a one-click Windows installer: Python and all dependencies bundled, per-user install, no admin rights, no console windows. You paste two API keys (speech-to-text and the LLM) into a local panel, and the panel checks each key before you start.

One fun side effect: I can ask the assistant to change its own code. It runs a coding agent on its source, and if that succeeds, a script builds a new installer, commits, publishes a GitHub release and silently installs the update over itself.

What's next

  • Better local fallbacks (an optional local LLM via Ollama already handles basic commands offline).
  • More reliable screen understanding for apps without accessibility info.
  • Making it easier to add your own tools without touching the core.

The code is open source: github.com/platonkuch-dev/jarvis-ai. There's a demo and the installer on the project site.

I'd love to hear what you'd want a desktop voice assistant to do that current ones don't, and I'm happy to answer questions about the architecture or the costs in the comments.

Top comments (0)