DEV Community

Cover image for Project JARVIS: The Offline AI Lab Assistant
Abhinav krishna N
Abhinav krishna N

Posted on

Project JARVIS: The Offline AI Lab Assistant

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Ask a new student in a college lab where the vernier is, and most will say nothing. They'll wander, give up, and not come back.

The IoT club lab at my college is an open space where anyone can walk in and learn about hardware. Nobody is taught from a podium. Students learn from each other, share what they know, and get their hands on real components, and there are a lot of them. The club also hosts workshops. When I was there I hosted workshops too, and I mentored students and taught them.

Now my junior, John, coordinates the lab. With so many tools and components, and so many people passing through, finding the right one and keeping track of what has been borrowed is a job in itself. John loves automation and cool tech, and he is a big Iron Man fan, so I thought: why not give him his own JARVIS for his lab? That's what I built: an offline AI assistant you talk to, which points a laser at the tool you need, lights up component boxes, and keeps the lending log.

It runs entirely on an Arduino UNO Q with 2 GB of RAM: no internet, no cloud account, no API key.

Ask out loud, and the dot lands on the hook.

New students are often too shy to ask where things are. With JARVIS they just ask the wall. It takes the friction out of a first visit, and it makes the lab a bit more fun, so people get curious and start poking around.

How you use it. You say the wake command, "Hey Jarvis", and the assistant comes alive. It answers "Yes, sir?" and waits for you to say what you need. Then you just ask in plain words, and it replies out loud, points, or lights up whatever you asked for. Say "No, thank you" and it goes quiet again until the next "Hey Jarvis".

Here is what it does today. These are examples of what it can do:

  • "Where's the vernier?" It answers "The vernier caliper is on hook 2, sir. I've highlighted it for you." A servo tilts a small laser pointer onto that hook. The laser lights with the first word of the reply and switches off when the reply ends. Code on both processors also cuts it after 10 seconds regardless.
  • "I want to measure the size of this groove." It works out the tool for the job (the vernier) and points at it. If you just say "I have a job, which tool should I use?", it asks "What's the job?" first.
  • "Lend two ESP32s to Arjun till Friday." It reads the loan back and saves it only after a spoken "yes" or a tap on the dashboard.
  • "What's overdue?" / "Who has the soldering iron?" It answers from the loan log.
  • After each answer it asks "Anything else?", and "No, thank you" ends the conversation.

It talks just like JARVIS in the Iron Man movies, with lines inspired by the films themselves, made for my Iron Man fan John.

Demo

JARVIS running on the UNO Q in the lab: wake word, a spoken request, and the reply, all offline.

Code

GitHub logo Abhinavkrishna3211 / jarvis-lab-assistant

Offline voice assistant on an Arduino UNO Q that finds tools, lights component boxes and keeps the lending log for a lab. Gemma 3 1B, whisper.cpp and Piper, all on-device

JARVIS: an offline lab assistant on the Arduino UNO Q

A voice assistant for an open IoT-club lab. Ask it where a tool is and a laser marks the hook; ask for a component and the LED under its box lights; lend and return hardware by voice and it keeps the log Everything runs on the board: no internet, no cloud account.

How it works

mic -> "Hey Jarvis" (App Lab keyword spotting, always listening) -> "Good evening, sir."
    -> record until you pause (6 s max) -> whisper.cpp tiny.en -> fuzzy candidate match (about 10 items)
    -> keyword rules; if they give up on a known item, Gemma 3 1B (App Lab llama.cpp runner,
       JSON-schema constrained) -> validator -> SQLite
    -> Piper reply and Bridge calls to the STM32: point / laser / box_led / ring_state

The model never touches facts it could invent. It returns one short JSON…

What I Used

Hardware

  • Arduino UNO Q, 2 GB version
  • Servo, 180 degrees (an SG90), to tilt the laser
  • Laser module salvaged from an old keychain pointer, switched through a logic-level MOSFET (IRLZ44N) with a series resistor
  • Arducam B0490 camera (2 MP, IMX462 sensor, day and IR night vision, USB), watching the tool wall. Its built-in microphone is the microphone JARVIS listens through
  • Bluetooth speaker, for JARVIS's voice
  • USB-C hub, to connect the camera and other USB devices to the board
  • A separate 5 V supply for the servo, sharing ground with the board
  • Status ring and LEDs for the component boxes

Software and models

  • Arduino App Lab (Bridge between the Linux side and the STM32)
  • Gemma 3 1B through llama.cpp
  • whisper.cpp tiny.en
  • Piper
  • A "Hey Jarvis" wake-word model trained in Edge Impulse
  • A FOMO object-detection model trained in Edge Impulse (public project), which finds the tools on the wall
  • SQLite

How I Built It

mic -> "Hey Jarvis" (my Edge Impulse keyword model) -> cached greeting
    -> record until you pause -> whisper.cpp tiny.en
    -> keyword rules; only if they give up on a known item: Gemma 3 1B -> one JSON action
    -> validate() -> SQLite / servo + laser / box LED -> Piper voice, sentence by sentence
Enter fullscreen mode Exit fullscreen mode

Open models, all running locally on the board:

  • Gemma 3 1B (open weights), served by App Lab's llama.cpp runner.
  • whisper.cpp tiny.en for speech to text.
  • Piper for the voice.
  • A "Hey Jarvis" wake-word model I trained in Edge Impulse (MFE plus a small classifier, classes JARVIS and NOISE, about 10 ms per window). It runs inside App Lab's keyword-spotting brick, swapped in with one variable in app.yaml. The wake word alone never goes to a model: "Hey Jarvis" by itself gets "Yes, sir?" straight from code.


On its validation set the wake-word model reaches 100% accuracy with a loss of 0.08. That set is small and has only the JARVIS and NOISE classes, so I don't read it as proof the model never wakes by mistake. Testing it with more voices is the first item under Still To Do

  • A FOMO object-detection model, also trained in Edge Impulse, which finds each tool on the wall in the camera image so the laser can follow it (see The camera below).


On the validation set the model scores an F1 of 96.6%. The multimeter and the wire stripper were found every time, and the vernier 80% of the time; the misses were counted as background. Inference takes 22 ms, with 305.5K of peak RAM and 81.3K of flash. The validation set is small, so a single missed vernier moves that score a lot. I treat these numbers as a first check, not as an accuracy claim.

JARVIS is a physical agent. It hears a request, decides what to do, and acts in the room: a servo swings the laser onto the right hook.

The UNO Q's two processors split the work:

  • Linux side (Qualcomm MPU): the wake word, whisper.cpp, the keyword rules and Gemma, the validator, SQLite and Piper.
  • STM32 side (microcontroller): the servo, the laser, the status ring and the box LEDs, called from Python over App Lab's Bridge in about 7 ms per call.

Wiring


The main wiring schematic, drawn in KiCad.

The pointer

It's built from scrap. The laser module and the servo came from old electronic parts I had at home, not bought.

What I planned, and what I built. I designed a pan-tilt head: two servos, one to swing the laser sideways and one to tilt it, inside a 3D-printed enclosure. I wrote the wiring and the code for both axes. But I had only one servo and no 3D printer. So the prototype in the demo is a simpler version: one SG90 servo on a cardboard mount, glued together, that moves the laser up and down. The hooks hang in a single column, so one axis is enough to point at the tool, and the camera tells JARVIS where along that column the tool is.


The old keychain pointer, from a drawer at home.


Opened up: the board and contacts I reused. The cells and case stayed behind.

The laser is the module from a keychain pointer with its cells removed. I never measured its output power, so I treat it as an unrated pointer and keep it away from people. It's switched from pin D3 through a logic-level MOSFET (IRLZ44N) and fed from 5 V through a series resistor that limits the current. The servo runs from a separate 5 V supply that shares ground with the board, and the sketch detaches it after each move so it doesn't jitter.

I started with fixed hook angles (50°, 65° and 79° for the three tools on the wall), found by jogging the servo until the dot sat on the hook. I used them for the first validation, to check that the laser pointed at the right place. Later I trained a FOMO object-detection model in Edge Impulse, and now the laser follows the tool wherever it is placed.

The camera

A camera watches the tool wall. I trained a FOMO object-detection model of the wall in Edge Impulse (public project). It finds each tool in the camera image, JARVIS turns that position into a servo angle, and the laser points at the tool wherever it hangs. If someone puts the vernier on the wrong hook, the laser follows the tool and JARVIS says "Someone has moved the vernier caliper, sir."

Safety

The laser's limits live on the microcontroller as well as in Python. The sketch caps every laser-on at 10 seconds and switches it off by itself, even if Linux crashes mid-reply. The module is current-limited, and I keep it aimed at the wall, never at people.

The model never gets the last word

It returns one small JSON action. Code checks that the item exists, the quantity is in stock, the borrower is a name and the date is real, and code does all the date and stock arithmetic. Nothing the model says reaches the database or the laser without that check.

What I Measured

All numbers come from the board; full tables are in docs/measurements.md. The 60-command evaluation uses typed commands. The speed numbers come from spoken requests through the mic.

A 1B model lost to keyword rules. On 60 typed test commands, Gemma 3 1B got 22 right and simple keyword rules got 40. Gemma handled loose wording better, but it labelled almost every loan and return as a search. It also invented answers to small talk: "what's the weather" came back as "the multimeter". On top of that, it takes 28-56 s per request on this board.

So JARVIS uses both. The rules answer first, in under a second. Gemma is asked only when the rules give up and the request mentions a known item, and JARVIS says "One moment, sir" before it starts. In a second run on the board, that hybrid scored 45/60, against 41/60 for the rules alone in the same run (the rules scored 40 in the first run; the one-command difference is run-to-run variation). Gemma was called on 18 of the 60 commands, and the net gain was 4 commands the rules missed. When it was wrong, it almost always answered with a search, never a loan or a return, so a mistake costs a wrong pointer, not a wrong record.

Since then I've added rules for asking about a tool by its job, and the rules alone now score 51/60. One caveat: I wrote the rules while looking at the same 60 commands, so these numbers are optimistic.

Speed. In the first version the reply took a while to start. Three changes cut that down:

Change Before After
Greetings cached at start-up 2.2-3.6 s from wake to first word about 0 s of synthesis
Replies spoken sentence by sentence (next one synthesised while this one plays) 2.4-5.5 s before the first word 1.2-2.3 s
A short filler ("Right, sir.") once whisper has heard words silence something to hear while it thinks

whisper.cpp takes about 1.7 s for a short request. The Bridge call is about 7 ms.

What Went Wrong (and Stays In)

  • Whisper heard "wire stripper" as "wise triple". The rules gave up, and Gemma recovered the right tool. That one example is the reason for the hybrid design.
  • Whisper also transcribed the wake word: "Hey Aradino" matched the Arduino Uno and lit its box. The wake word is now stripped before parsing.
  • A desk fan pushed the mic to full scale for whole recordings. Whisper heard nothing, and JARVIS kept saying "Could you repeat that?" until I found the fan.
  • The laser pointed at the floor for a tool whose hook only existed in placeholder data. Now a tool without a measured hook angle gets a spoken answer and no laser.
  • I tried to calibrate the laser by having the camera find the red dot. It never saw the dot on the white foam board. In the end I jogged the servo a few degrees at a time while watching the dot, and wrote the three hook angles down by hand.
  • At full volume the Bluetooth speaker distorted. 85% volume and a slower, calmer voice fixed that.
  • The Bluetooth speaker paired without bonding and forgot the board on every reboot. It needed pairable on before pairing.
  • JARVIS could wake itself: its own voice came back through the mic. Wake words are now ignored while it speaks and for 1.5 s after (Bluetooth audio lags about a second).

Still To Do

  • Retrain "Hey Jarvis" with more voices and an "other speech" class. It only knows JARVIS and NOISE today, so other words can wake it.
  • Wire the LED strip under the component boxes and test it on the board.
  • Measure the laser's pointing error.
  • Load the lab's full inventory: three tools hang on the wall today.
  • Record spoken commands from other club members and report accuracy on those. -Build the two-servo pan-tilt head. It needs a second servo and a 3D-printed enclosure.

Why Does Open Innovation Matter?

A college open lab can't run on an API key. Someone has to pay for it, it needs Wi-Fi that works, and it stops the day the free tier ends. Open-weight models and open runtimes let the whole assistant live on a board we own, work with the router unplugged, and cost nothing per question.

Because everything is open, I could also measure each piece honestly. When the 1B model lost to plain keyword rules, I could see it, and change the design to use it only where it actually helps.

My Agent Session

I built JARVIS with Claude as my coding agent. It wrote most of the Python and the sketch. I did the hardware, the wiring, the Edge Impulse wake-word training, the Edge Impulse Vision model, the testing in the lab and the design decisions. The rules it worked under are in AGENTS.md in the repo: the model only returns one JSON action, loans need a spoken "yes", the laser always times out, and nothing goes to the cloud. Every number in this post was measured on the board, not estimated.

Prize Categories

  • Best Use of Arduino: JARVIS runs entirely on the Arduino UNO Q: wake word, FOMO model, speech to text, Gemma and the voice on the Linux side, and the servo, laser and LEDs on the STM32 over the Bridge. It senses (it listens), decides, and acts physically (the laser marks the hook), with the laser's safety limits enforced on the microcontroller.
  • Best Use of Gemma: Gemma 3 1B runs locally on the board through App Lab's llama.cpp runner. It's constrained to emit one JSON action, and code does the stock, dates and validation. I measured it honestly against keyword rules on the same 60 commands and used it where it actually helps.

Top comments (0)