This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
We spend most of our day in front of screens. Deep work matters, but so do the moments when we stand up, look away, and step outside. And when we're focused, those are exactly the moments we skip.
So I built Moss: a small, deadpan fern creature that lives on your desk, works beside you while you focus, and occasionally makes you leave your desk.
Moss is a focus timer, a Tamagotchi, and an outdoor quest game in one:
- Focus. You start a session; Moss sits beside its tiny laptop and works with you.
- Moss develops a need. When the timer ends, the deterministic game engine looks at Moss's stats (energy, nature, movement, rest…) and picks what it needs most.
- A real-world quest. "Find me a tree." "Show me the sky." "Find some grass." Local Gemma words the quest in Moss's voice.
- Take your phone outside. Scan a QR code and the phone becomes Moss's eyes. Moss literally walks out the door of the 3D room ("Moss is outside with you").
- Photograph it. The photo goes over your local Wi-Fi to your laptop, where Gemma 3 4B checks it with vision.
- Moss recovers and the world grows. Fronds spring up, a leaf blows in through the window, and the desk ecosystem grows: a tiny tree after your first tree, a sod tray after your first grass, a flower pot, a sun crystal. Achievements like Actually Touched Grass unlock along the way.
- Back to work, refreshed.
The screen is the shortest part of the loop. The point of the app is the two minutes you spend outside.
It's for anyone who works long stretches at a laptop (developers, students, writers) and knows they should take real breaks but never does. Moss makes the break the game.
Demo
The video is the app's built-in demo mode (/?demo=1): a scripted 90-second replay of the full loop. It runs the real game engine on a scripted clock, so the quest choice, XP and the tiny-tree unlock are genuine engine output. The "Nature found ✓" verdict on the phone is a recorded run of local Gemma 3 4B on that exact photo (confidence 0.98, 1.2 s on my laptop), not a made-up answer. There's a script in the repo to re-run it.
| Break quest on the phone | The "why" |
|---|---|
![]() |
![]() |
Code
Moss
A local-first AI productivity companion: Tamagotchi × focus timer × outdoor quests.
Moss is a small forest creature that works beside you on your laptop. When a focus session ends, Moss develops a need that can only be met in the physical world. You take your phone outside, photograph a tree (or grass, or a flower), and local Gemma 3 4B checks the photo. Moss recovers, and its tiny desk ecosystem grows.
Everything runs on your machine. Photos travel only over your local Wi-Fi, and no cloud AI is involved.
▶ Demo video: youtu.be/xXZvVzqoxFI · Code: github.com/Reterics/hacktoberfest_touch_grass
Or watch the whole loop in about 90 seconds…
Quick start (Node 22.13+, pnpm 9, Ollama):
ollama pull gemma3:4b
pnpm install
pnpm dev # laptop UI on :5173, phone pairs via QR on the same Wi-Fi
How I Built It
Open-source AI at the core: Gemma 3 4B, running locally through Ollama. It does three jobs, and only these three:
- Quest wording. The engine picks the need and the quest template (requirement, rewards, timer). Gemma rewrites the title, description and Moss's line in character. The output is constrained to a JSON schema and validated with zod. If Gemma drops the actual requirement (e.g. the word "tree"), the wording is rejected and the built-in text is used instead.
-
Vision validation. The phone photo is checked by Gemma's vision with a strict verdict schema (
success,confidence,detected,reason). The model never awards anything. The engine applies the verdict, and it only passes at ≥ 0.6 confidence. Two rejections, or any technical failure, unlock manual completion at half XP, so a flaky model can never trap you. - Moss's reactions. Short, deadpan one-liners after a quest. Moss is "unimpressed but secretly delighted".
Everything that matters for fairness (XP, stats, levels, unlocks, streaks, achievements) lives in a pure, deterministic TypeScript game engine with 35 tests. The AI is personality and perception, never authority. If Ollama isn't running, the whole game still works with fallback wording and manual completion.
The rest of the stack:
-
Laptop: Node + Fastify 5 + WebSockets + SQLite (
node:sqlite). The laptop is the hub, and the phone connects over the local network with a QR pairing token. - Desktop scene: React 19 + React Three Fiber + Three.js. A cozy 3D desk with Moss as a Meshy-generated, rigged model (walk and hop clips). The face, blinking and frond-droop expressions are rendered live in shaders, driven by a visual-state controller.
- Phone: a lightweight React companion with no Three.js. Moss's walk and hop on the phone are sprite strips baked from the rigged 3D model.
- Design-time AI, shipped as local files: Higgsfield for the character design, expression sheet and achievement art, Meshy for 3D, and ElevenLabs for Moss's voice, the narration, sound effects and the music bed in the demo. None of it is called at runtime.
The whole thing was built in pair sessions with Claude Code (session below), from the empty repo to the vertical slice: focus timer → need → Gemma quest → phone → photo → Gemma validation → Moss reaction → ecosystem growth → next focus session.
Why Does Open Innovation Matter?
Because the core promise of Moss, "go outside, show me what you found", involves photos of wherever you are. With an open-weight model running locally:
- Your photos never leave your network. They travel from your phone to your own laptop over Wi-Fi, get checked, and that's it. A closed vision API would mean shipping pictures of your street, your garden or your park to someone else's server, several times a day.
- It costs nothing to run. A habit app that fires a vision request on every break would rack up API bills for something you're supposed to use daily. Locally, it's free forever.
- No account, no key, no internet requirement for the AI part. A tiny companion app shouldn't need a cloud subscription to tell a tree from a lamp post.
-
Swappable. The model is an env var (
OLLAMA_VISION_MODEL). Because every output is schema-validated and the engine makes the decisions, a different open model can drop in without touching the game rules.
The bigger reason is what open innovation made possible for a one-person project: an open-weight model for perception, open web standards to connect two devices, Three.js for the world, and community tools for 3D, voice and art, combined into something none of them was designed to build. An AI companion designed to get you away from AI.
My Agent Session
A curated view of the Claude Code session behind this build: every prompt I gave, the tools the agent used for each step, and its report. Secrets, signed URLs, emails and local paths are redacted.
Prize Categories
- Best Use of Gemma. Gemma 3 4B (local, via Ollama) does all of Moss's runtime AI work: quest wording, photo validation with vision, and reactions, each schema-validated, with the deterministic engine making every decision.
- Best Use of ElevenLabs. Moss's deadpan voice (Eleven v3), the intro and end-card narration, the sound effects (shutter, success chime, leaf gust, frond spring) and the cozy music bed are all ElevenLabs, generated at design time and shipped as local audio, so the app stays offline-capable.








Top comments (1)
tr.ee/dev-to