This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Side Quests: Go Touch Grass 🍁 is a mobile-first web app that gives you tiny real-world quests. Gemma 4 writes each quest and also checks it, running locally on a laptop. You spend about 20 seconds on the screen: you get a quest, the app tells you to put your phone away, and when you're done you take one photo as proof.
The loop:
- Nudge. The app reads your local weather and sunset time from Open-Meteo, and Gemma writes a short, friendly push to go outside. For example: "38 minutes of golden hour left. Spend 20 of them outside." If it's dark or stormy, it suggests a doorstep or balcony moment instead and never guilt-trips you.
- Quest. You choose how much time you have (10, 20 or 45 min) and a difficulty (Chill, Curious or Explorer). Gemma designs a quest that fits the season, weather and daylight. It can send you to a real park or pond nearby, pulled from OpenStreetMap. Example: "Find three leaves with clearly different shapes."
- Go outside. The phone goes back in your pocket.
- Proof. You take one photo. Gemma 4's vision model checks it against the quest's criteria, then tells you whether you passed, explains why, and adds a nature fact about what you photographed.
- Crew. Streaks, a weekly XP leaderboard and join codes for run clubs, families and friend groups. Every finished quest is saved in your Field Journal.
Who it's for: anyone who opens their phone for a "quick break" and comes out 40 minutes later. It's also for groups that want a light, playful reason to get outside together, like families, run clubs and office teams. There's no sign-up: a random device token lives in your browser.
Demo
Not deployed to any cloud service since the premise is on device. Attaching few photos for the demo.
Code
This is the heart of the app: checking the photo. Gemma 4 looks at the picture and returns a verdict in a fixed JSON shape. Then plain Python has the final say, so the model can't be talked into a pass by a photo of a screen, an old photo from the camera roll, or a low-confidence guess.
# backend/src/ai/prompts.py: the exact shape Gemma must answer in
VERIFY_SCHEMA = {
"type": "object",
"properties": {
"passed": {"type": "boolean"},
"confidence": {"type": "number", "minimum": 0, "maximum": 1},
"criteria_met": {"type": "array", "items": {"type": "string"}},
"reasoning": {"type": "string"},
"fun_fact": {"type": "string"},
"looks_like_screen_or_stock_photo": {"type": "boolean"},
},
"required": ["passed", "confidence", "criteria_met", "reasoning",
"fun_fact", "looks_like_screen_or_stock_photo"],
}
# backend/src/ai/ollama_client.py: every Gemma call goes through one local endpoint
payload = {
"model": model, # gemma4:e4b, running on the laptop
"messages": [{"role": "system", "content": system}, user_msg], # photo rides along as base64
"format": schema, # Ollama constrains Gemma's output to the JSON Schema
"stream": False,
# Gemma 4 "thinks" by default, which takes minutes on CPU; the JSON schema does the work.
"think": settings.ollama_think,
"options": {"temperature": temperature},
}
resp = await client.post(f"{settings.ollama_url}/api/chat", json=payload)
# backend/src/ai/verify.py: Gemma judges, plain code has the final say
async def verify_photo(photo, *, title, description, criteria, quest_created_local) -> Verdict:
raw = await chat_json(
model=settings.vision_model,
system=VERIFY_SYSTEM,
user=build_verify_prompt(title, description, criteria), # the criteria Gemma itself wrote
schema=VERIFY_SCHEMA,
images_b64=[base64.b64encode(photo.jpeg_bytes).decode()],
temperature=0.2, # a referee should be boring and consistent
)
return apply_rules(raw, min_confidence=settings.verify_min_confidence,
taken_at=photo.taken_at, quest_created_local=quest_created_local)
def apply_rules(raw, *, min_confidence, taken_at, quest_created_local) -> Verdict:
model_passed = bool(raw.get("passed"))
confidence = max(0.0, min(1.0, float(raw.get("confidence", 0))))
flags: list[str] = []
if raw.get("looks_like_screen_or_stock_photo"):
flags.append("screen_or_stock")
if model_passed and confidence < min_confidence:
flags.append("low_confidence")
# Camera-local EXIF time vs quest creation. 10 min of slack covers clock drift;
# anything older was clearly taken before the quest existed.
if taken_at and quest_created_local and (quest_created_local - taken_at).total_seconds() > 600:
flags.append("photo_older_than_quest")
passed = model_passed and not flags
...
The quest generator writes 2–4 verification_criteria for every quest, and this referee checks each one against the photo. So Gemma writes the rubric and then grades the photo against it, all on the same machine.
My Agent Session
I am not adding my coding agent session as it has signatures which could be reverse engineered. But a Claude code was my coding agent and buddy.
How I Built It
Model: Gemma 4 (gemma4:e4b) served locally through Ollama. One model on one laptop handles three jobs:
| Job | How Gemma is used | Temperature |
|---|---|---|
| Nudge copy | Text model gets weather, sunset time, time budget and streak, and returns headline, message and a 0-10 go_outside_score
|
0.9 |
| Quest generation | Text model gets conditions, OSM places and recent quest titles, and returns title, description, 2–4 verification criteria, hint, est. minutes and safety note | 1.0 |
| Photo verification | Vision model gets the quest, its criteria and the photo, and returns pass/fail, confidence, which criteria were met, reasoning, a fun fact and a screen/stock-photo flag | 0.2 |
A huge thanks to claude code :).
Why Does Open Innovation Matter?
Open innovation matters because, the users don't need to be afraid their location is leaked to a closed model or the pictures you take are also not sent to some remote server for analysis. Everything is on device which makes this interesting.
Prize Categories
Best Use of Gemma - Gemma is the heart of this project from creating the quest to image analysis.




Top comments (0)