DEV Community

Attila Reterics
Attila Reterics

Posted on

Moss: a tiny AI desk companion that makes you go outside and touch grass

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

We spend most of our day in front of screens. Deep work matters, but so do the moments when we stand up, look away, and step outside. And when we're focused, those are exactly the moments we skip.

So I built Moss: a small, deadpan fern creature that lives on your desk, works beside you while you focus, and occasionally makes you leave your desk.

Moss is a focus timer, a Tamagotchi, and an outdoor quest game in one:

  1. Focus. You start a session; Moss sits beside its tiny laptop and works with you.
  2. Moss develops a need. When the timer ends, the deterministic game engine looks at Moss's stats (energy, nature, movement, rest…) and picks what it needs most.
  3. A real-world quest. "Find me a tree." "Show me the sky." "Find some grass." Local Gemma words the quest in Moss's voice.
  4. Take your phone outside. Scan a QR code and the phone becomes Moss's eyes. Moss literally walks out the door of the 3D room ("Moss is outside with you").
  5. Photograph it. The photo goes over your local Wi-Fi to your laptop, where Gemma 3 4B checks it with vision.
  6. Moss recovers and the world grows. Fronds spring up, a leaf blows in through the window, and the desk ecosystem grows: a tiny tree after your first tree, a sod tray after your first grass, a flower pot, a sun crystal. Achievements like Actually Touched Grass unlock along the way.
  7. Back to work, refreshed.

The screen is the shortest part of the loop. The point of the app is the two minutes you spend outside.

It's for anyone who works long stretches at a laptop (developers, students, writers) and knows they should take real breaks but never does. Moss makes the break the game.

Moss celebrating on the desk after a quest: a new tiny tree, a sunbeam through the window

Demo

The video is the app's built-in demo mode (/?demo=1): a scripted 90-second replay of the full loop. It runs the real game engine on a scripted clock, so the quest choice, XP and the tiny-tree unlock are genuine engine output. The "Nature found ✓" verdict on the phone is a recorded run of local Gemma 3 4B on that exact photo (confidence 0.98, 1.2 s on my laptop), not a made-up answer. There's a script in the repo to re-run it.

Break quest on the phone The "why"
The break quest End card about open innovation

Code

Moss

CI Node 22.13+ pnpm 9 TypeScript React 19 Three.js Gemma 3 4B via Ollama No cloud AI

Moss, a small fern creature, sitting on a mossy stone next to a glowing phone in an autumn forest

A local-first AI productivity companion: Tamagotchi × focus timer × outdoor quests.

Moss is a small forest creature that works beside you on your laptop. When a focus session ends, Moss develops a need that can only be met in the physical world. You take your phone outside, photograph a tree (or grass, or a flower), and local Gemma 3 4B checks the photo. Moss recovers, and its tiny desk ecosystem grows.

Everything runs on your machine. Photos travel only over your local Wi-Fi, and no cloud AI is involved.

Moss working beside a tiny laptop on a cozy desk, with a subtitle about screen time The break quest "Find me a tree" on the phone while Moss heads outside
Focus: Moss works beside you while you work. Break: Moss needs a tree, so you take the phone outside.
Moss celebrating after the quest, with a new tiny tree on the desk and a sunbeam through the window End card: open innovation lets us combine what already exists to create what doesn't
Reward: local Gemma confirmed the tree; Moss recovers and the desk ecosystem grows. Why: an AI companion designed to get you away from AI.

▶ Demo video: youtu.be/xXZvVzqoxFI · Code: github.com/Reterics/hacktoberfest_touch_grass

Or watch the whole loop in about 90 seconds…

Quick start (Node 22.13+, pnpm 9, Ollama):

ollama pull gemma3:4b
pnpm install
pnpm dev        # laptop UI on :5173, phone pairs via QR on the same Wi-Fi
Enter fullscreen mode Exit fullscreen mode

How I Built It

Open-source AI at the core: Gemma 3 4B, running locally through Ollama. It does three jobs, and only these three:

  • Quest wording. The engine picks the need and the quest template (requirement, rewards, timer). Gemma rewrites the title, description and Moss's line in character. The output is constrained to a JSON schema and validated with zod. If Gemma drops the actual requirement (e.g. the word "tree"), the wording is rejected and the built-in text is used instead.
  • Vision validation. The phone photo is checked by Gemma's vision with a strict verdict schema (success, confidence, detected, reason). The model never awards anything. The engine applies the verdict, and it only passes at ≥ 0.6 confidence. Two rejections, or any technical failure, unlock manual completion at half XP, so a flaky model can never trap you.
  • Moss's reactions. Short, deadpan one-liners after a quest. Moss is "unimpressed but secretly delighted".

Everything that matters for fairness (XP, stats, levels, unlocks, streaks, achievements) lives in a pure, deterministic TypeScript game engine with 35 tests. The AI is personality and perception, never authority. If Ollama isn't running, the whole game still works with fallback wording and manual completion.

The rest of the stack:

  • Laptop: Node + Fastify 5 + WebSockets + SQLite (node:sqlite). The laptop is the hub, and the phone connects over the local network with a QR pairing token.
  • Desktop scene: React 19 + React Three Fiber + Three.js. A cozy 3D desk with Moss as a Meshy-generated, rigged model (walk and hop clips). The face, blinking and frond-droop expressions are rendered live in shaders, driven by a visual-state controller.
  • Phone: a lightweight React companion with no Three.js. Moss's walk and hop on the phone are sprite strips baked from the rigged 3D model.
  • Design-time AI, shipped as local files: Higgsfield for the character design, expression sheet and achievement art, Meshy for 3D, and ElevenLabs for Moss's voice, the narration, sound effects and the music bed in the demo. None of it is called at runtime.

The whole thing was built in pair sessions with Claude Code (session below), from the empty repo to the vertical slice: focus timer → need → Gemma quest → phone → photo → Gemma validation → Moss reaction → ecosystem growth → next focus session.

Why Does Open Innovation Matter?

Because the core promise of Moss, "go outside, show me what you found", involves photos of wherever you are. With an open-weight model running locally:

  • Your photos never leave your network. They travel from your phone to your own laptop over Wi-Fi, get checked, and that's it. A closed vision API would mean shipping pictures of your street, your garden or your park to someone else's server, several times a day.
  • It costs nothing to run. A habit app that fires a vision request on every break would rack up API bills for something you're supposed to use daily. Locally, it's free forever.
  • No account, no key, no internet requirement for the AI part. A tiny companion app shouldn't need a cloud subscription to tell a tree from a lamp post.
  • Swappable. The model is an env var (OLLAMA_VISION_MODEL). Because every output is schema-validated and the engine makes the decisions, a different open model can drop in without touching the game rules.

The bigger reason is what open innovation made possible for a one-person project: an open-weight model for perception, open web standards to connect two devices, Three.js for the world, and community tools for 3D, voice and art, combined into something none of them was designed to build. An AI companion designed to get you away from AI.

My Agent Session

A curated view of the Claude Code session behind this build: every prompt I gave, the tools the agent used for each step, and its report. Secrets, signed URLs, emails and local paths are redacted.

Building Moss with Claude Code: a desk companion that sends you outside
You

Start implement the Moss game based on the documentation

Agent

The core Moss loop is built and works end to end with real Gemma, and it also completes when Ollama is down. One gap: I haven't tested it with a real phone yet. All phone tests ran from this laptop.
How to run it: pnpm install && pnpm dev, then open http://localhost:5173. Pick Demo 30s / 2m for a quick loop, and when the break starts, scan the QR code with your phone on the same Wi-Fi.
What's built:
- Desktop: a 3D desk scene with Moss, a laptop, an autumn …

You

Commit this work, then Use Higgsfield and/or Meshy MCP to design the real Moss character and character sheet, as CLAUDE.md asks. The current 3D Moss is a placeholder built in code and can be swapped for a proper model without touching the game logic.

Agent

Both commits are in, and Moss now has a real character design that the app uses on both the desktop and the phone.
The commits:
- 0a5c094: the core game loop from the last session.
- 345b225: the new Moss design and the 3D model.
Typecheck, the 32 tests and the production build all pass. I checked the new Moss in the browser, but haven't watched every animated transition closely.
The design:
- Concepts: six concepts are in visuals/concepts/. You picked 05's …

You

Start focus and stop focus and too frequent clicks are shaking the Moss

Agent

Fixed: Moss no longer shakes when you click Start and Stop quickly.
Cause: Moss's bob, breathing, hop and wiggle were computed from the total time since the page loaded, multiplied by the animation speed of its current state. Start and Stop switch Moss between idle (slower) and working (faster), and the speed eases between the two. After a few minutes on the page, even a small speed change jumped the animation to a different point in its cycle every frame. The longer the …

You

Lets to these: - A rigged walk or hop animation from Meshy, for when Moss heads outside on the phone.
- Higgsfield art for the first achievements, such as "Actually Touched Grass".

Agent

Moss now has a rigged walk and hop on both the laptop and the phone, and the first five achievements are in the game with their own art. Tests (34), typecheck and the production build pass. Nothing is committed yet; that includes last turn's shake fix.
Walk and hop (Meshy, 8 credits):
- Rig: Meshy's auto-rigger handled Moss despite its round shape. That gave the walk, plus the hop you picked ("Hop with Arms Raised").
- On the laptop:
- Moss hops when a break …

You

[Image #1] what are the overlapping grass-like object on the table and why?
[screenshot omitted]

Agent

questsCompleted is 0, so the running backend isn't reading the test save I meant. I'll check which databases exist and what each one contains:

You

Default jumping moss is detailed and light, while the rigged moving version is dark and differs. When switch to break time, it is a real bug

Agent

I fixed both problems, the darker rigged Moss and the grass spikes. Tests (34), typecheck and the production build pass. Nothing is committed yet.
The grass-like objects were the grass-patch reward you earn for completing your first grass quest; it was unlocked during my testing. I had drawn it as about 220 bare spikes in a wide ring around Moss, with no ground under them, cutting into the laptop. It's now a small terracotta tray of grass on its own patch of soil, placed …

You

yes, conventional commits

Agent

I committed everything as five conventional commits on top of your own 85eb73a. The working tree is clean and nothing is pushed.
```
031e02e docs: document Moss motion, achievements and asset pipeline
7206e76 feat(web): animate Moss walking and hopping, show achievements
b4455d2 feat(assets): add rigged Moss, motion sprite strips and achievement art
66cdc51 build: add script to merge Meshy animation clips into one GLB
fc956a3 feat(engine): add deterministic achievements …

You

./docs/DEMO-SCENE.md and implement demo via utilizing elevenlabs and higgsfield MCP

Agent

Tidying the Desktop wrapper indentation:

You

It is perfect, one note: after Open innovation text all sentences are sudden after each other, no breathing break, so it seems rushed. Can we correct this? everything else is perfect, do not change

Agent

I fixed the rushed end card: there's now about a 1.1 s breathing pause after each narrated sentence, and nothing else changed. I didn't re-watch the end card after this change.
How I did it: I didn't regenerate the narration. I took the same voice take and inserted silence at each sentence boundary. The on-screen lines now appear at 0, 4.9, 8.25, 11.75 and 15.4 s, matching the voice. The narration grew from 17.5 s to 21.1 s, so "Replay" now appears about 3.5 s later. The …

You

AI is too much in the text, so instead of We are using AI we can "We are using it" or something else

Agent

I changed the end-card closing line, the one that appears after the narration, to:

We are using it to make people spend less time staring at computers.
That line is on-screen text only, not part of the narration, so no audio changed. The narrated line before it, "…an AI companion that gets you away from AI.", now gives "it" its meaning. Nothing else changed, and it's still uncommitted.

You

Use this instead
Open innovation lets us combine what already exists to create what doesn’t.
Local AI understands the world.
Open technologies connect devices.
Community tools accelerate creation.

The result: an AI companion designed to get you away from AI.
Moss uses technology to reconnect you with the world outside your screen.
and change the sound as well

Agent

The end card now uses your six lines, and the narrator speaks them in the same voice as before (Carolyn). I watched the demo through to the Replay button, and tests still pass.
- Voice: I generated two takes and picked the one with a natural 0.7–1.3 s breathing pause after each sentence. It didn't need the silence I spliced in last time. Your lines go in as written, apart from adding the closing punctuation the script needs. I can't listen to audio, so I chose the take …

You

during the start instead of 9 sec, 7 is enough for countdown

Agent

The cues shifted correctly, but the command exited with an error and no message showed. Running typecheck and tests separately:

You

[Image #3] popup overlap
[screenshot omitted]

Agent

The countdown now opens at 7 s. My screenshot landed after the caption window, so I'm re-running and sampling around it:

You

Commit current state, and find possible solutions for problem statement: Demo is good, but does not explain to the user what it is. like "
We spend too much of the day sitting in front of screens.
Deep work is valuable, but people also need moments to move, recover, look away, and reconnect with the physical world.
The problem is that when we are focused, we often ignore those …

Agent

Committed as four conventional commits, and the working tree is clean:

259931b docs: add the demo scene script and document the demo
f87d214 feat(web): add scripted "Moss Needs a Tree" demo at ?demo=1
4311714 feat(assets): add demo voice, sound, music and art
5a93a9b feat(web): blow a leaf onto the desk and light a sunbeam on quest rewards

- Your staged docs/DEMO-SCENE.md went into the docs commit. Nothing is pushed.
- The 7-second countdown is in this commit, and …

You

A. and softened

Agent

The demo now opens with the problem statement, using your softened wording (option A), before the gameplay starts. I watched it play: subtitles, push-in, tagline and the hand-off to the countdown all landed on cue.
How the intro plays:
1. Problem: the camera starts wide on the desk in dim, late-afternoon light, with the HUD hidden and Moss working beside the glowing laptop. The narrator speaks the three problem lines, with each one shown as a subtitle.
2. Idea: …

You

[Image #6] for a video intro, it does feel a bit too empty. Not catastrophically empty, but empty enough that the scene reads more like a prototype room than a lived-in world.
A few reasons:
- Huge uninterrupted desk surface
- Large plain wall areas
- Only a few props, spaced far apart
- The scene has a center character, but not enough surrounding “life”
- In motion, …

Agent

The room now looks lived-in rather than like a prototype. I added every item on your list to the scene, so the real game shows them too, not just the demo. I checked the result in the browser during the intro, during gameplay and while Moss is outside.
| Your suggestion | What I added |
|---|---|
| desk mat or fabric runner | a sage felt mat with a stitched border under Moss and the laptop, which breaks up the big wood surface |
| mouse or small desk accessory | a cream …

You

Make it in video format as well

Agent

I'll record the demo as an MP4 with its full soundtrack. The plan:
- Drive a real Chrome window at 1920×1080 with Playwright and capture every frame through Chrome's screencast.
- Log the exact moment each voice line, SFX and music change plays.
- Rebuild the soundtrack from the original audio files at those timestamps with ffmpeg, so picture and sound stay in sync.
Checking the tools first:

You

Open Innovation text has no fade in so sound and text are too sudden. texts are jumping when new line added. Add smooth transitions or animations to avoid jumping and sudden movements

Agent

I've smoothed the end card and the intro text. Nothing should jump or pop in any more, and new lines fade in gently. I confirmed in the browser that the end card stays in place now, but I haven't seen the new fades play through on screen.
What I changed:
- No more jumping: all the end-card lines are laid out from the start but invisible, so adding a line can't push the others around. The headline's measured position stayed the same for the whole end card.
- **Line …

You

continue

Agent

Continuing with the video. Heads-up: a separate Chrome window will open and play the demo for about 1.5 minutes. It runs muted, and you can work around it.

You

too much fade, reduce it half

Agent

I halved the fades:
| | Before | Now |
|---|---|---|
| Each line reveal (end card, intro subtitles, Meet Moss card) | 1.4 s, drifts up 12 px, starts 3 px blurred | 0.7 s, drifts up 6 px, starts 1.5 px blurred |
| End card fade-in | 2 s | 1 s |
| Music dip under the narration | 2.5 s | 1.25 s |
| Narration start after the cut to the end card | 2.3 s | 1.3 s |
Typecheck and tests pass, but I haven't watched it play since this change. Nothing is committed. I've left the video …

You

After Play demo, fade out ply demo and About a minute - sound on, wait 3sec and then go next slide

Agent

It works. After Play, the button and hint fade out and the title holds; about 3 seconds later the card fades into the intro, with the first subtitle on screen. Cleaning up:

You

finalize application, commit changes, add github actions, readme badges if missing. (image, screenshot, or other visuals must have in readme) and stop before we continue with dev submission

Agent

Everything is committed and the working tree is clean. Nothing is pushed.

c74e9e9 docs: add README badges, hero image and screenshot gallery
20090e3 ci: typecheck, test and build on push and pull requests
8c1d1f7 feat(web): dress the desk so the room feels lived in
b2b3d50 feat(web): open the demo with the problem and the idea

Checks: 35 tests, typecheck across all four packages and the production build all pass.
GitHub Actions (.github/workflows/ci.yml):
- …

You

Demo recorded: https://youtu.be/jfAM9tN1Dsg Github: https://github.com/Reterics/hacktoberfest_touch_grass Product Categories: ElevenLabs, Gemma Must be added to Readme and using devrelay-gateway create a Draft to dev.to with the following template:
*This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch …

Agent

Added the demo video, repository link and prize categories (Gemma, ElevenLabs) to the README, sanitized this session's transcript (secrets, signed URLs, emails and local paths redacted; screenshots stripped), saved it with DevRelay, and created the DEV submission draft from the challenge template with the #hf26challenge tag and a cover image.

Prize Categories

  • Best Use of Gemma. Gemma 3 4B (local, via Ollama) does all of Moss's runtime AI work: quest wording, photo validation with vision, and reactions, each schema-validated, with the deterministic engine making every decision.
  • Best Use of ElevenLabs. Moss's deadpan voice (Eleven v3), the intro and end-card narration, the sound effects (shutter, success chime, leaf gust, frond spring) and the cozy music bed are all ElevenLabs, generated at design time and shipped as local audio, so the app stays offline-capable.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to