This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Touch Grass Quest is a daily outdoor photo task that is designed to be over in seconds. You open the page and read one small task, like "Find a berry on a bush." Then you put the laptop down and go outside. You take one photo, and a vision model tells you in one line whether it matches. After a pass, the app says "Done for today. Go touch more grass." and stops.
There is no feed, no notifications, and no leaderboard. The most the app does after a pass is show a streak counter. I wanted the screen to be the shortest part of the experience, and the best way to do that was to give the app nothing else to offer.
It's for people like me who spend most of the day at a laptop and keep meaning to go outside.
Demo
I haven't deployed a hosted version. The app runs locally, and the repo includes a Dockerfile for anyone who wants to host it. The screenshots here are from my own testing.
Code
GhanshyamJha05
/
touch_grass_quest
A daily outdoor photo quest: one small task, one photo, checked by Gemma (open-weight vision). Built in Go for the DEV Hacktoberfest Touch Grass challenge.
πΏ Touch Grass Quest
Touch Grass Quest is a daily, mobile-first web app built for the DEV Challenge. It encourages users to put down their screens and interact with the real world by giving them one unique outdoor photography task every day.
Upload your photo, and an open-weight Vision AI will determine if you successfully found the item outdoors. Build your streak, and go touch some grass!
β¨ Features
- Daily Quests: A new deterministic, outdoor-focused task generated for you every day.
- Open-Weight AI Verification: Verification is powered by open-source vision models via any OpenAI-compatible API.
- Mobile-First UX: Sleek, responsive, dark-mode UI designed to feel like a native mobile app.
- Privacy & Speed: Image downscaling happens locally on your device before upload, saving bandwidth and protecting raw photo data.
- Streak Tracking: Tracks your daily success streak locally on your device.
π Quick Start
Prerequisites
- Go 1.22+
- An API Key from anβ¦
How I Built It
-
Backend: Go with only the standard library (
net/http), plus a small front end in plain HTML, CSS, and JavaScript. There's no framework and no database. -
Tasks: a pool of short outdoor tasks in
tasks.json. The task of the day is picked deterministically from the date, so everyone gets the same one. Each task carries a safety hint. The berry task says "look for wild berries, but don't eat them." -
Model: Gemma 4 31B (open-weight), called through Google AI Studio's OpenAI-compatible endpoint. A separate
VISION_MODELsetting exists because a text-only model can silently ignore an image, so image checks can be routed to a model that actually sees them. I wrote a small test tool (cmd/visiontest) to confirm that Gemma really reads the photo and isn't answering from the task text alone. - Light uploads: the browser downscales the photo to at most 1024px and re-encodes it as a JPEG before upload, which also drops its metadata and keeps the upload small.
- Streak: stored in the browser only. There are no accounts.
The model gives evidence, and my code decides
The first design decision was to keep the verdict out of the model. The model returns one strict JSON object, with a match flag, an outdoors flag, and a short reason. Go then makes the call: a photo passes only if it matches the task and looks outdoors. If it matches but looks indoors, the reply is gentle ("Looks like it's indoors or a screen, take it outside").
The prompt asks the model to be fair but honest, to accept reasonable interpretations, to never invent objects, and to treat a photo of a screen or a printed picture as not outdoors. I wrote that last rule after thinking about the most obvious way to cheat, which is photographing a nice forest on your monitor.
The failure paths are handled in code too:
- Gemma 4 returns its reasoning inside
<thought>tags ahead of the real answer. I strip those before parsing, and I treat an unclosed tag as a failure instead of showing half a thought process. - If the model ignores the JSON format and replies in prose, the app retries once. If that fails, it falls back safely.
- If a check fails for technical reasons, the page says "Couldn't check that, try again." I chose that on purpose. A flaky API shouldn't hand out free passes, and it shouldn't unfairly fail someone who really did go outside.
The check I was most curious about
The obvious way to cheat is to photograph a nice picture on your monitor, so I tried exactly that. I pointed the camera at a picture of a berry bush on my laptop screen and submitted it.
In my first test runs, which used Gemini 2.5 Pro before I switched to Gemma, the app rejected photos like this. The model recognized the berries and still refused to pass them, because it noticed laptop bezels in the frame. That's the behavior I wanted. A couple of rejected photos isn't an accuracy number, though, and I changed models afterward, so I treat it as a sanity check on the design and not a benchmark.
What Went Wrong
- Spoofing isn't solved. A realistic 4K screen or a printed photo held outdoors might still fool the model. The prompt tells it to look for screens, but no vision model can prove someone is outside. This is a soft check and I'd rather say so than oversell it.
-
Garbled error text. One early rejection read "Not quite. Not quite." and showed a raw
'where an apostrophe should be. My code was building the message badly and escaping it wrongly. - A false rejection. In the same round of testing, a clear photo of berries on a bush was told it doesn't look like the requested item. A vision model being wrong on an easy case is exactly what an evaluation set is for, so this photo is a good test case for the harness in the repo.
-
Text-only models fail quietly. Without a vision model, the app would happily "check" a photo it can't see. That's why the project has a separate vision test and a
VISION_MODELsetting.
Why Does Open Innovation Matter?
I'll be careful here, because I didn't compare against a closed model on the same photos, so I can't claim open was more accurate. What open gave me:
-
No lock-in. The app talks to any OpenAI-compatible endpoint. Moving to a different provider or vision model means changing environment variables (
LLM_BASE_URL,LLM_MODEL,VISION_MODEL), not rewriting code. That wasn't hypothetical. I started this project on a closed model (Gemini 2.5 Pro) and moved to Gemma by changing my.envsettings. - Control over behavior. The rules, the strictness, and what counts as "outdoors" live in a prompt and in Go code I can read and change. The most useful decision I made was keeping the pass/fail logic out of the model entirely.
- A path to local. I used a hosted API, so the downscaled photo does leave the device while it's being checked. The app never asks for location and has no accounts. If I wanted photos to stay on the device, Gemma's open weights are what would let me run the same check locally.
- Small footprint. The whole thing is one Go binary and one API key. There's no GPU to rent and no database to run.
What's Next
A proper field test on real walks, with more people than just me and a measured hit rate. A larger photo set in the evaluation harness, including the failing cases above. And a hosted deployment so it doesn't depend on my laptop.
My Agent Session
I built this with Antigravity, working from a detailed spec that asked for tests, strict parsing, and an evaluation harness. I then ran it, tested it by hand, and found the bugs above myself.
Prize Categories
- Best Use of Gemma: Gemma 4 31B does the photo check, served through Google AI Studio. The category allows serving through "Google Cloud or another provider."



Top comments (1)
tr.ee/dev-to