This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Grass Quest is a photo scavenger hunt where the game master is a local AI.
You press one button and get a mission, like "find a leaf with a hole in it" or "three different shades of green in one frame." Then the phone goes in your pocket and you go look for it. When you find it, you take one photo, and the same model looks at the photo and decides whether it counts.
The screen is only there at the start and the end. Everything in between happens outside, which is the point. It's for anyone who goes on a walk and ends up staring at their phone the whole time: the app gives you a reason to look at the world instead.
Demo
Demo Image from: https://in.pinterest.com/pin/97460779408386405/
Code
Grass Quest
An offline photo scavenger hunt. A local Gemma model gives you a mission to do outside ("three different shades of green in one frame"), you go find it, and the same model looks at your photo and decides if it counts.
Runs an open-weight Gemma model through Ollama. No API keys, no cloud, your photos never leave your machine.
Run it
# 1. Install Ollama from https://ollama.com, then:
ollama pull gemma3:4b
# 2. Install and start the app
pip install -r requirements.txt
python app.py
Open the local URL Gradio prints. To use it from your phone, start it with SHARE=1 python app.py and open the gradio.live link Gradio prints (this tunnels through Gradio's servers while it's running).
Use a bigger model with MODEL=gemma3:12b python app.py (needs more VRAM), or point at another Ollama host with OLLAMA_URL.
Limits
The judge is a 4B vision model. It's good…
How I Built It
One Python file, around 130 lines.
-
Model: Gemma 3 (
gemma3:4b), Google's open-weight model, running locally through Ollama. Gemma 3 understands images as well as text, so one small model does both jobs: writing missions and judging photos. - UI: Gradio. A difficulty picker, a mission button, a photo box (upload or camera) and a judge button.
- Missions: the app picks a random theme (textures, light and shadow, something tiny, the sky...) and asks Gemma for one mission at a high temperature, so you don't get the same one twice. The prompt rules out anything unsafe, illegal, or involving photographing strangers' faces or private property.
-
Judging: the photo is shrunk to 1024px (phone photos are huge and it doesn't change the verdict), sent to Gemma with the mission, and Gemma returns JSON: passed or not, what it sees, and a one-line verdict. Ollama's
format: "json"mode keeps the output parseable, and a low temperature keeps the judge consistent.
The judge is told to be fair but not a pushover, so a blurry guess or a photo of a screen doesn't pass. It's a 4B model, so it's a referee, not ground truth: it handles colours, objects and scenes well and can get fiddly missions wrong.
Why Does Open Innovation Matter?
- Your photos stay yours. Photos from a walk show where you live, your street and your routine. With a local model, none of that gets uploaded to a company's server to be stored or trained on.
- It costs nothing per photo. A scavenger hunt means lots of attempts and retries. With a paid vision API every judged photo costs money; locally, it's free no matter how many times you play.
- It runs without internet. The model runs on my own laptop with no cloud API behind it.
-
I can swap the model. If a better small vision model comes out, it's a one-line change (
MODEL=...).
Where a closed model would win: a bigger cloud vision model would judge tricky missions more accurately. For a game where the referee being occasionally wrong is part of the fun, a free, private local model is the better trade.
Prize Categories
Best Use of Gemma

Top comments (0)