This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Field Dex is a pixel-art creature-collecting game where the creatures are real. You photograph a bird, a tree, a beetle or a flower on your normal iPhone camera, and when you get home you open the day's photos like a pack of trading cards.
My laptop identifies every photo with a local vision model and rolls a rarity, from Common to Legendary. It writes the card's flavour text, pays out XP, and dresses your pixel avatar in gear you unlock. Each day, finding something new also unlocks the next chapter of a story starring you, which ends on a cliffhanger. The only way to read what happens next is to go back outside.
It's for anyone who has caught themselves scrolling when they meant to go for a walk, and for people who like collecting things. That includes me.
Why a game?
I didn't want to build something that politely suggests going outside. Doomscrolling works because of three things: you don't know what the next swipe will show, there is always one more to see, and nothing ever finishes. So I tried to copy the mechanics but not the screen:
| Doomscrolling | Field Dex |
|---|---|
| Unpredictable next item | A rarity roll when you flip each card |
| Endless feed | A Dex with ??? slots for every creature you haven't found yet |
| Reward is on the screen | Reward is only earned outside, then opened at home |
| No "done" state | A new Wanted poster every day and a story that never quite ends |
The screen is the shortest part of the day. The photo walk is the game, and the pack opening is the reward at the end of it.
The game, in detail
The Wanted poster. Every day the game posts a bounty such as "a bee on a yellow flower". It's described by what you can see, never by species name, so the vision model can actually check whether your photo matches. A match is worth +50 XP.
Rarity is a mix of luck and effort. A harder subject (birds beat plants), a first-ever species, a sharp photo and a shot taken within about 75 minutes of sunrise or sunset all raise your odds. The final roll is seeded from the photo's hash, so re-uploading can't reroll it. From my simulation of 10,000 rolls, a good photo of a new bird comes out Rare 57% of the time, Epic 31% and Legendary 12%. A repeat plant is Common about three times out of four, so a Legendary actually feels like an event.
score = KIND_BASE[kind] # birds are harder than plants
score += NEW_SPECIES_BONUS if is_new else 0 # first time you've seen it
score += (quality - 3) * 0.5 # sharp photo beats a blurry one
score += 1.0 if golden_hour else 0 # dawn and dusk are luckier
score += (rng.random() ** 2) * 4.5 # the surprise, skewed low
It can't be gamed easily. It rejects photos of screens, photos with no living thing, photos more than a week old, and anything you've uploaded before. The rarity and XP rules are plain Python code, not model output, so the AI can't be flattered or prompted into handing out a Legendary.
Nothing is spoiled early. Your XP bar and Dex count only change when you flip a card, not when the laptop finishes processing it. The surprise is always saved for the flip.
The pixel world. Your avatar is drawn from text grids on a canvas, with no image files. It gains a camera, binoculars, a leaf crown, a bug net or a cape as you hit milestones. Each card's artwork is your own photo, shrunk and colour-reduced into pixel art, and tapping it shows the original. Sound effects are synthesised live with WebAudio, so there are no audio files either, and nothing loads from the internet.
Soft pressure on purpose. There are no streaks to lose and nothing is ever taken away. I wanted to pull people outside, not punish them for staying in.
Demo
Code
hardikgaikwad
/
FieldDex
A local 'touch grass' gamified application to get you off your screen.
Field Dex
A pixel-art creature collecting game played in the real world. You photograph living things outside, then come home and open them like a card pack. A local open-weight vision model identifies each find. A rarity roll, XP, levels, gear for your pixel avatar, and a story that continues one chapter at a time are what bring you back the next day.
Everything runs on your own laptop through Ollama. No accounts, no cloud, no API keys. Your photos never leave your devices.
Today Β· Reveal Β· Dex Β· Story. Screenshots use sample photos from Wikimedia Commons.
How a day works
- Outside: use the normal iPhone Camera. No app, no laptop, no signal needed.
- Home: the laptop is on your iPhone's Personal Hotspot. Open Field Dex in Safari, tap Open today's pack, and pick today's photos.
-
On the laptop:
-
qwen2.5vl:3bidentifies each photo (about 6 sβ¦
-
The README has setup steps, a map of the code and the screenshots. Everything runs with pip install -r requirements.txt and two ollama pull commands.
How I Built It
The constraint that shaped everything
My first idea was a "side quest generator": a local agent reads your profile and free time and hands you a short outdoor mission. I built it and it worked. Then I noticed the obvious problem: my laptop can't come outside with me, and it gets its internet from my iPhone's hotspot. A quest app that only works at a desk is useless.
So I flipped the design instead of fighting it. The phone collects outside; the laptop thinks at home. The iPhone camera needs no app and no signal. Back home, Safari talks to the laptop over the hotspot. That delay between taking the photo and seeing the card turned out to be the best part of the design, because it's exactly the gap between pulling the lever and seeing what you won.
OUTSIDE HOME (same hotspot)
ββββββββββββββββ upload βββββββββββββββββββββββββββββββββ
β iPhone cameraβ ββββββββββΆ β FastAPI + SQLite β
β (no app, no β β ββ Pillow: HEIC β JPEG, pixel art
β signal) β ββββββββββ β ββ qwen2.5vl:3b β what is it?
ββββββββββββββββ cards β ββ game.py β rarity, XP, unlocks
Safari UI β ββ qwen2.5:3b β card text, story
βββββββββββββββββββββββββββββββββ
all via Ollama, on the laptop
The open models
-
qwen2.5vl:3b(open-weight vision-language model) looks at each photo and returns structured JSON: what the subject is, its visible features, a category, a confidence, whether it's a photo of a screen, and a photo-quality score. -
qwen2.5:3b(open-weight text model) writes each card's flavour text and stats, the daily Wanted poster, and the story chapters. - Ollama runs both, locally. Its structured-output mode takes a JSON schema, which forces the answer into a shape my code can trust.
The whole thing runs on a laptop with an RTX 3050 with 4 GB of VRAM. Identifying a photo takes roughly 6β9 seconds, and a batch of six photos, with their cards and the story chapter, finished in about 90 seconds. Setting the context window to 4,096 tokens kept the text model entirely on the GPU. Ollama's default context pushed a third of it onto the CPU.
The idea that made small models usable
Three-billion-parameter models are good at some things and bad at others, so the pipeline gives each part to whoever is good at it:
- The model describes and names things, and writes. It's good at that.
- Python decides rarity, XP, levels, unlocks and rejections. These are deterministic, testable and can't be talked into cheating.
Two real examples of that split paying off:
- The model names well and categorises badly. On my test photos it called the peacock "Indian Peafowl" correctly, then filed it under "animal". A butterfly became a "plant". So the name now decides the category through a small vocabulary map, and the model's own category is only a fallback.
- Field order matters. I originally asked "is this a living thing?" first, and the model confidently said a tree beside a building was not. Reordering the schema so it describes first, then classifies, then judges fixed that case. It still labelled a hibiscus flower as "none" until the name-based category rule above caught it. Both fixes together got all six of my test photos right, including rejecting a street scene with no living thing in it.
A pipeline detail that matters on 4 GB
Ollama can only keep one model loaded on a card this small. Processing photo-by-photo would swap vision and text models back and forth every time. So a batch runs all vision calls first, then all text calls, and the models swap once per batch.
Other things worth saying
- One background worker processes the queue while the page polls for progress, so uploading returns instantly and the "developing film" screen shows live progress.
- A real bug I hit: uploading a duplicate photo left a SQLite transaction open and locked the worker out of the database. I fixed it with autocommit and wrote a regression test. The project has 11 tests covering the game rules, the vision-category logic and that bug.
- The UI has no build step. It's plain HTML, CSS and JavaScript, with bundled pixel fonts (Press Start 2P and Pixelify Sans, both under the SIL Open Font License).
Where the small models struggle (honestly)
- Flavour text sometimes invents facts. One card claimed a peafowl is "a master of disguise".
- Names are sometimes vague. A photo of a big tree came back as just "Tree".
- The stats on each card (charm, stealth, grit) are made up by the model for fun, not real biology.
- Story chapters stay in the right world and use the day's finds, but the prose can be clumsy. One chapter said "Mynah, the Mynah, sang a song".
- Telling lookalike species apart is beyond a 3B model, so I treat the names as a fun guess, not a field guide.
Taking It Outside
[TODO before publishing: replace this block with your real photo walk.]
Where did you go? How many photos? What was your best and worst find? Did a Wanted poster make you walk somewhere you wouldn't have? How did it feel to open the pack afterwards compared with scrolling? Add 2β3 real screenshots or photos.
Why Does Open Innovation Matter?
I could have built this around a hosted vision API in an afternoon. I'm glad I didn't, for four reasons that are specific to this app:
-
My photos are my life. A folder of "what I saw on my walks" shows where I go and when, which is exactly what you shouldn't hand to a server you don't control. Field Dex never sends a photo anywhere. Nothing in the app calls out to the internet: no CDN, no analytics, no fonts fetched from a server, and inference is a call to
localhost. - The cost is zero, which changes the design. A game that rewards you for taking more photos only works if each one is free. With a per-image API price, I'd have been quietly discouraging the very behaviour the game exists to encourage. I also ran dozens of prompt and pipeline experiments in one evening without watching a bill or a rate limit.
- Open weights let me constrain the model. Because I run the stack myself, I could force structured JSON output, control the context window to fit a 4 GB card, run the vision and text models in a deliberate order, and keep the rules in code rather than trusting a prompt. That's how a 3B model became reliable enough to ship.
-
I can change it. The models are two lines in
profile.yaml. Any vision model Ollama can run can be dropped in, and when a better small model ships next month, the game gets better without me touching the code. (I only testedqwen2.5vl:3bandqwen2.5:3bend to end, so I can't promise how others compare.)
The honest trade-off: a hosted frontier model would identify species more accurately and write better prose. For a game about the act of going outside, I think "mostly right, private, free and mine" beats "perfectly right and rented".
Built on open-weight models (Qwen2.5-VL and Qwen2.5) running through Ollama, with the session saved through DevRelay.




Top comments (0)