This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Side Quest is a tiny game where an AI running on your own laptop gives you one small real-world mission, and the only way to win is to go outside.
- Open it on your phone. Choose how long you have (10, 20, 45 or 90 minutes) and your mood (curious, calm, with kids, need to clear my head…).
- A local Gemma model acts as game master. It invents one quest that fits your time of day, season and current weather. One it gave me on a 35°C October morning in Arizona: > Autumn Leaf Color Echo. Find the most intensely colored autumn leaf you can, then trace the shapes of its veins with your finger on a nearby tree. What do you notice?
- Pocket the phone. The screen just shows a timer and a single "📷 Snap proof" button.
- Take one photo as proof. A local vision model referees it ("passed", plus one warm sentence that mentions a real detail it saw) and adds it to your field journal, which tracks quests done and minutes spent outside.
The screen part is about 30 seconds. Everything else happens outside.
It's for people who want to go outside but get stuck at "and do what?": kids who need a mission, someone on a walk-to-clear-my-head break, or a friend group that wants a silly challenge. Quests are about noticing things like a heart-shaped leaf, a reflection in a puddle, or three kinds of bark. They aren't workouts, and they never ask you to photograph strangers.
Demo
Video Created with assistance of Brag Skill
Tested end to end on my laptop:
| Photo submitted | Quest proof required | Referee (qwen3-vl:8b) |
|---|---|---|
| Big Sur coastline | "water meeting land or rocks" | "calm blue water meeting rocky outcrops and a shore lined with reddish vegetation" (9.1s) |
| Solid teal image | same | "a solid color, let's head outside to find that water-meets-land spot!" (10.8s) |
Code
The whole thing is about 180 lines of Python (standard library only, no pip installs) plus one HTML file.
ollama pull gemma3:4b && ollama pull qwen3-vl:8b
python3 app.py # then open http://<laptop-ip>:8080 on your phone
How I Built It
Models (all open weights, via Ollama):
-
gemma3:4b: the game master. It's small enough to answer in about 4 seconds on a laptop, and creative enough to write a whimsical quest. -
qwen3-vl:8b: the referee. It's a vision-language model that looks at your proof photo and returns a verdict. - You can swap either one with an env var (
QUEST_MODEL=… JUDGE_MODEL=…). For an all-Gemma setup,gemma3:4balso has vision and can be the referee.
Structured outputs everywhere. Both calls pass a JSON schema in Ollama's format field, so the game master must return {title, mission, photo_proof, why} and the referee must return {passed, saw, verdict, outdoors}. No regex parsing of chatty model output. The photo_proof field matters most: the game master writes a checkable success criterion, and the referee is graded against exactly that sentence. The referee also has to confirm the photo was taken outdoors, so a photo of your screen doesn't count.
Context without tracking. The quest prompt gets:
- the local time (no sunset quests at noon)
- the season, flipped for the southern hemisphere (October is spring in Sydney)
- the live weather from Open-Meteo, which is free and keyless, queried with coordinates rounded to about 1 km, and only if you share location
- your last 8 completed quests, so it doesn't repeat itself
With no internet, the weather is simply "unknown" and everything else still works.
Phone-first with no app store. The laptop serves one page over Wi-Fi. <input type="file" accept="image/*" capture="environment"> opens the rear camera directly. The photo is downscaled to 1024px on the phone before upload, which keeps the vision model fast.
The bug that ate an hour. My first referee calls all came back as JSONDecodeError. Calling Ollama directly showed why: with think: false and a JSON-schema format, qwen3-vl puts its answer into the thinking field, malformed, and leaves content empty. Leaving thinking enabled for the referee returns clean, schema-valid JSON in about 9 seconds. It's a good reminder that "structured output" depends on the model and the runtime flags. With open weights and a local runtime I could inspect the raw response and work out what was happening. A closed API would have just been a black box returning errors.
Why Does Open Innovation Matter?
- Your photos and location stay on your machine. Side Quest sees where you walk, when you go out, and pictures of your neighborhood, your kids and your backyard. That's exactly the data I don't want sent to someone else's server. Here the photos and journal sit in a folder on your laptop.
- It costs nothing to play. A closed vision API charges per image. A game you're supposed to play every day shouldn't have a meter running. Local inference makes the 100th quest as free as the first.
- You can tune the game. Swapping in a bigger referee, or an all-Gemma setup for a smaller laptop, takes one env var. You can also edit the prompt to make it a birding game or a kids' scavenger hunt. A tool you're meant to customize should run on models you can actually control.
- Known Current Limitation: Need to be deployed Online for Complete End-To-End Deployment and Usage on Mobile Devices. Local Inference will require porting to Android Code and using Local ML models. The current implementation is desktop based, Web-Powered Application. Or setup a tunnel from your remote machine where the code runs to your Mobile Device.
Prize Categories
-
Best Use of Gemma:
gemma3:4bis the game master that writes every quest.
Top comments (0)