This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
I built RoomTidy for my roommate: a small Windows desktop application that offers a gentle reminder when visible clutter remains in an area they choose.
The user selects a fixed-camera area, explicitly saves a tidy reference, and sets preferences such as allowing books while asking for dishes to be removed. RoomTidy compares later pictures with that reference locally. It describes visible items; it does not claim to detect dust, prove someone has not cleaned, or judge the person.
The user controls monitoring, snooze, resets and reference updates. References and observation history stay on the computer. The app does not automatically treat persistent clutter as the new tidy baseline. This is a local MVP built for my roommate; recipient feedback and a full timed room trial are still pending.
Demo
Watch the 60-second RoomTidy demo on YouTube
The demo uses generated photographs of a fictional desk, not images of our private room. It runs seven manually staged checks through the actual local comparison and persistence pipeline:
- Save and check an explicit tidy reference.
- Introduce visible cups, dishes and loose papers.
- Show a dark/blocked view; the count remains unchanged.
- Observe clutter again.
- Reach three valid clutter observations and activate a reminder.
- Observe more clutter without repeating the same episode's notification.
- Return to the reference and reset the count.
The accelerated demo threshold is three valid observations. It is not a recording of thirty elapsed minutes or thirty hourly checks. Normal monitoring checks hourly while the app and computer are awake, with a default threshold of thirty valid clutter observations.
Code
RoomTidy source code on GitHub
The application source is MIT-licensed. The README explains setup on the tested Windows computer, local storage, model/runtime provenance and known limits. This MVP is not yet a distributable installer for other computers.
How I Built It
RoomTidy uses Python and PySide6 for the desktop interface, OpenCV for camera handling and image checks, SQLite for observation history, and Google Gemma 4 E4B Instruct for understanding changed scenes.
The quantized Q4_0 model and matching BF16 vision projector run through an application-owned llama.cpp process on authenticated loopback. Both artifacts were pinned and SHA256-verified. Google and the official GGUF conversion declare Apache-2.0 licensing.
Gemma receives the user-approved reference and current selected-area picture with allowed-object preferences. It returns a structured status: tidy, clutter_present or unable_to_assess, with a short description of visible evidence. The app validates that response before counting it.
OpenCV rejects dark, overexposed and low-detail images first. Exact readable pixel matches return a factual reference match without asking the model to invent differences. Changed valid scenes use Gemma. Application logic, rather than the model, owns timestamps, counters and reminder deduplication. Unknown and missed checks do not count as clutter evidence.
I used AI coding assistance and Opus/Astra source reviews. The generated demonstration assets came from OpenAI image generation; they were not generated by Gemma. Runtime room-image analysis stays local, with no cloud image uploads or fallback AI API.
Testing mattered more than choosing a model by name. The smaller Gemma E2B candidate failed comparison checks. Initial E4B tests also exposed preference mistakes and false added-object claims on identical pictures. Those findings led to clearer preference instructions and the exact-image safeguard.
The final regression suite passed 71 tests. The seven-step controlled demo produced counters 0, 1, 1, 2, 3, 4, 0. Additional local model checks recognized a dim tidy scene, dim clutter, and a scene whose added objects were explicitly permitted.
These tests do not establish general room accuracy. Camera capture/release was tested separately; classification across genuinely changed real-camera scenes, a full timed monitoring trial, and independently observed Windows toast delivery remain unverified. Windows can suppress toasts, so an active reminder also remains visible in the application.
Why Does Open Innovation Matter?
A room camera captures a private space. Running openly licensed model weights locally lets the app compare pictures without sending them to an image-analysis provider. The user keeps the references, settings and history on their own computer and can pause or delete the app's data.
The open runtime also made model evaluation possible: I could try different weights, inspect failure behavior, and strengthen validation without rebuilding the product around a hosted API. There is no per-request AI-provider billing for local analysis, though hardware, storage and electricity still have costs.
Prize Categories
Best Use of Gemma. Gemma performs the application's central image-understanding task: comparing changed selected-area scenes with an explicit tidy reference while interpreting permitted objects. It is part of the working product, rather than a tool used only to produce the presentation.
Top comments (0)