This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
I spend most of my day on a screen. So I built Touch Grass, an app that gives me a reason to leave it.
Every morning, you get one outdoor task. "Find moving water." "Say hello to three people you pass." "Pick up ten pieces of litter." You can do it any time that day. Then you take a live photo as proof. A small AI on my laptop looks at the photo, and plain code decides if you passed.
The tasks are not only about nature. Some are about people: ask a shop owner how long they've been open, or ask someone for their favorite place to eat. The goal is to go out, look around, and talk to someone, not just count steps.

Today's task. Anytime today. No timer.
What it does:
- 40 hand-written tasks in six kinds: nature, movement, exploration, social, creative and community
- Live camera only. No gallery, so the proof comes from really being there
- If the AI gets it wrong, tap "Mark done anyway" for half points. The app never blocks your streak
- Works with no signal. The photo waits on your phone and goes through when you're back online
- XP, levels, a daily streak and a monthly count
- You choose which kinds of tasks you want more or less of, and how bold the social ones are
- Weather-aware tasks, a streak freeze, a daily reminder, a bad-day mode and a monthly recap
- It counts its own screen time, because the screen should be the shortest part of the day
It's for people like me who want a small push to go outside, and a little game to keep it going.

No signal? The photo is saved on the phone and waits.

Back online. The laptop checked the photo and the task is done.
Demo
Code
Touch Grass
Single-user app that gives you one daily outdoor task (nature, movement, exploration, social, creative, community). You submit a live photo as proof from a mobile app; a laptop server verifies it with a local VLM served by Ollama.
See AGENTS.md for the full spec, hard rules and game/verdict rules.
Prerequisites
- uv (Python package manager)
- Python 3.11+
- Node.js 20+ (for the Expo app in
app/) -
Ollama running locally with a vision model
ollama pull qwen2.5vl:3b
- Git
- Optional: Tailscale so the phone can reach the laptop over HTTPS-grade WireGuard while the app still uses HTTP
- Optional: an Expo account if you want a standalone APK (
eas build)
Setup
# 1. Install dependencies (creates .venv and uv.lock)
uv sync
# 2. Configure environment
copy .env.example .env # Windows; use `cp .env.example .env` on macOS/Linux
# Edit .env and set API_TOKEN to a long random string. Never commit .env.ā¦How I Built It
The AI in the app is open and runs on my own laptop. It's an everyday machine: an RTX 3050 with 4 GB of video memory and 16 GB of RAM.
The AI. I used Qwen2.5-VL (3B) through Ollama. It looks at the photo and answers small yes/no questions, like "Is a leaf visible?" It also writes a short note about what it sees, and that note becomes the hint you read in the app.
The rule that made it work: the AI never decides pass or fail. Small AI models like to say yes. So plain code makes the call. You pass only if every required question gets a "yes". That also means nobody can sweet-talk the AI into a pass.
Choosing the model. I ran a quick test on 7 photos:
| Model | Right answers | "Not sure" answers | Speed per photo |
|---|---|---|---|
| Qwen2.5-VL 3B | 12 of 12 | 2 | 3 to 6 seconds |
| Gemma 3 4B | 8 of 10 | 4 | 4.5 to 6.5 seconds, but 71 seconds on the first photo, and it crashed once |
Gemma also gave "no" answers while its own note said the thing was visible. That's a tiny test, not a benchmark, but it was enough to pick Qwen. Swapping the model took one line.
The rest of the build:
- Photo checks come first. Blurry, dark, tiny or repeated photos are stopped before the AI sees them.
- Server: Python (FastAPI) and SQLite, running on my laptop.
- App: Expo (React Native), built as an Android app.
- The link between phone and laptop: Tailscale, a private network between my own devices.
- Offline: the phone keeps a queue. It sends photos when it can reach the laptop, and it never sends the same photo twice.
- Tasks are written by hand, because a small AI writes vague tasks. The AI only rewrites the wording a little.
- Streaks and XP are worked out from my history each time, so they can't get out of sync.
- The demo voice-over was made with ElevenLabs.
Being honest about how I built it: I used coding agents (OpenCode and Cursor) to write most of the code, from detailed prompts and a rules file. Those run on hosted models. The AI that runs inside the app is the local one. I built the Android file with Expo's cloud build.
What doesn't work well yet:
- Social tasks are checked by the place only. The AI can see the cafe, but not the chat.
- The AI is small, so it's sometimes unsure or wrong.
- My laptop has to be on for photos to get checked.
- Android only.

The app counts its own screen time. That 7 minutes was a testing day.

You decide which kinds of tasks you want more or less of.
Bugs I hit along the way
A tiny photo was stopped in 0.6 seconds. My test photo was only 30 KB. The photo check caught it before the AI even looked, and that's the point of checking first. The message said "move a little closer", which was wrong, because the problem was resolution, not distance. I changed it to say the resolution is too low.
Gemma disagreed with itself. One of the two models I tried said "no" to "is this outdoors?" while its own note said "sky is visible". It also took 71 seconds on the first photo and crashed once. Qwen got every question it answered right and was much faster. That's why I picked Qwen. It was a small test, but the difference was clear.
I typed the model name wrong, twice. My first try at a third model failed with a "not found" error. The tag had a typo. I fixed it and moved on, because Qwen was already doing the job.
A photo of a screen passed the "outdoors" question. I photographed a monitor showing a tree, and the AI said "outdoors". When I asked directly "is this a screen?", it got it right. So the model can tell, I just wasn't asking. For now, spotting screens is only logged. It's a clear next step.
The same note appeared for every question. The small model describes the photo once and pastes that note under each answer. So some hints are less useful than I'd like. A fix would be to ask one question at a time, which costs a few more seconds.
One extra retry made a check take 13 seconds instead of 5. My rule that every note must be at least four words sometimes made the model try again. It's a fair trade for better hints, but I watch it.
Uploads failed with "Unsupported FormDataPart". Expo's network code didn't accept the usual way of sending a photo. The server never saw a request. The app also said "couldn't reach your server" for every kind of error, which was misleading. I switched the upload to a different method and made the error messages say what really went wrong.
My history thumbnails were blank. The image requests didn't carry my token, so my own server said "401 unauthorized" to my own app. I now download each thumbnail with the token and keep a copy on the phone.
The app showed "done" when the server said it wasn't. When I deleted a test photo on the laptop, the phone still remembered it. Today's task showed as finished with no way to upload. I made the phone check with the server and drop items the server doesn't know about, and I added a button to clear the local cache.
The Close button did nothing after an "unsure" result. I tapped it three times. It only went away when I switched tabs. The cause was the screen being redrawn from fresh data, which brought the card back. Closing now takes effect immediately and stays closed.
Expo Go can't open the app in airplane mode. It loads the app from my laptop, so with no connection it can't start. The offline queue works once the app is open. For real use I needed a standalone app that doesn't depend on my laptop.
Why Does Open Innovation Matter?
- My photos stay mine. They go from my phone to my laptop and nowhere else. No cloud service ever sees them.
- It costs nothing per photo. I can check as many photos as I want, and nobody sends me a bill.
- I could swap models in one line. I tried two, looked at the results, and picked the better one.
- I set the rules myself: the questions, the exact format the model must answer in, and the code that decides. That's why the AI can't just agree with everything.
- The small model made the design better. A bigger closed model might see more detail. But because mine is small, I had to break the job into simple yes/no questions and let code decide. That made the app fairer and easier to trust.
An app that tells you to go outside shouldn't send your photos to a company on the way. This one doesn't.

Top comments (1)
tr.ee/dev-to