This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Voices of the Wild is a field guide that talks back. Photograph something outdoors (a tree, a crow, a parked car, the moon) and an open model works out what it is. If it knows, that thing answers in its own voice, with captions, and joins your collection. Then the app sends you for a walk.
"Well, hello down there. I'm Old Oak. Two hundred years in one spot. Touch some grass for me? I can't reach."
The screen is only on for the photo. The rest of the time you're listening and walking.
How the game works
- A daily trail of 5 to 7 quests, the same for everyone, unlocked one at a time: Meet a vehicle → Walk 500 steps → Touch grass → Walk 1,000 steps → …
- Real steps. The phone's motion sensor counts them, so the only way to finish a walk is to walk.
- XP for every find and every walk, a bonus for finding a different kind of thing, streaks, 12 badges, 5 ranks and an anonymous community leaderboard.
- A miss isn't a dead end. The app tells you what it half-recognised ("Something among the vehicles stirred… get closer"), and you can skip that quest for 25 XP. Walks can't be skipped.
- Next day starts tomorrow's trail early if you want to keep going.
The cast
81 characters in 13 kinds, from trees and birds to vehicles and street objects. Each kind has a guardian who steps in when the app isn't sure. A few of them:
- Parker, the parked car: "Twenty-three hours a day, I just sit here. You walked over? Show-off."
- Gossip the Crow: "I remember faces, so be nice to me."
- Sprinkles, the ice cream truck: "Walk to me. Never run into the road."
- Bronze Bert, the statue: "A hundred years in one pose. And my nose itches."
- Grumpy Gourd, the pumpkin: "It's October, so everyone wants to cut a face in me. Rude."
Three writing rules: every goodbye points you somewhere else outside, nobody ever tells you to touch, pick, chase or climb, and the words stay easy enough for a ten-year-old.
Who it's for: families on an evening walk, people new to a neighbourhood, and anyone who passes the same tree every day without looking up.
Demo
Live: https://voices-of-the-wild.netlify.app (no account, no install; best in Chrome on Android)
On a laptop, tap "No camera? Try one" to send sample photos through the real matcher, and add ?debug to the URL to see the match scores.
The video walks through a full trail: a miss, two finds, two walks, the leaderboard, Next day and a skip. The photos are Creative Commons images from Wikimedia Commons and the steps are simulated; the matching, voices, XP and timings are real.
Code
Mohibullah16
/
voices-of-the-wild
Photograph something outdoors and it talks back. On-device EmbeddingGemma 2, step-counted quest walks, 91 voiced characters, offline-first PWA.
Voices of the Wild
Point your phone at something outside. An open model works out what it is, and that thing talks back.
An oak, a crow, a parked car, a puddle, the moon: 81 characters and 13 guardians, 295 captioned voice lines.
Then it sends you on a walk. Phones use a small cloud listener running the same open model; laptops can run it entirely on-device
▶ Try it live: voices-of-the-wild.netlify.app
Best on an Android phone in Chrome. No account, no install. About 37 MB of voices downloads in the background on the first visit.
Entry for the DEV Hacktoberfest Open-Source AI Challenge, Week 1: Touch Grass. Demo video (1:55, sound on): docs/demo.mp4
The idea
Most apps want your eyes on the screen. This one wants the screen off. You only take the phone out for the photo; the rest of the time it is a voice…
-
app/: the PWA (Vite, TypeScript, lit-html, Transformers.js), with unit tests and Playwright end-to-end checks. -
listener/: a small server running the same model, deployed on Modal. -
pipeline/: build-time scripts for embeddings, calibration and voices. -
data/roster/roster.json: every character, description and line.
How I Built It
photo ─► EmbeddingGemma 2 ─► 768 numbers ─► compare with 717 character descriptions
│
character speaks ◄─ sure? ─► guardian asks for a closer look
EmbeddingGemma 2 (open weights, Apache 2.0) does the recognising. At build time it embeds 717 short visual descriptions of the 81 characters: "a dark car silhouette at night with its headlights glowing", "a close-up of a car wheel with an alloy rim". At play time it embeds only your photo, and the app finds the closest description. The same q4 weights run in two places: a small CPU server on Modal for phones, or in the browser on WebGPU for laptops.
Matching works in two steps: first the kind of thing, then the character within that kind. If the top character isn't clearly ahead, its guardian speaks instead. If nothing is close, "Nobody here wants to talk."
Why an embedding model and not a vision LLM? Because a character should say the same thing every time you meet it, which is what makes it worth collecting. An embedding model can only pick from characters that exist, so it never invents one. Every line was written, checked against the safety rules and recorded once. And its confidence is a number I can test.
Tested on 244 real photos (Wikimedia Commons, checked by eye), with thresholds tuned to be forgiving, because "nobody wants to talk" in front of a real tree is worse than a guardian asking for a closer look:
| Cross-validated | |
|---|---|
| Right character named | 77.9% |
| Guardian asks for a closer look | 12.7% |
| Wrong character named | 7.8% |
| "Nobody here" on a real outdoor photo | 1.6% |
| Indoor photos given a character | 0 of 10 |
Voices are made once, at build time. The 28 gold Elders are performed with ElevenLabs Eleven v4, using audio tags like [caws] and [hums a jingle]. The open-source Kokoro-82M voices everyone else. The app ships 295 captioned MP3s and never calls a voice API.
| Stack | |
|---|---|
| Recognition | EmbeddingGemma 2 (q4 ONNX) with Transformers.js |
| Voices | ElevenLabs (Elders), Kokoro-82M (everyone else), build time only |
| App | PWA: Vite, TypeScript, lit-html, Workbox, IndexedDB, hosted on Netlify |
| Server | Modal: the cloud listener and the anonymous rankings |
| Quality | Vitest, Playwright on a phone viewport, WCAG 2.2 AA (captions for every line) |
Limitations. The test photos are cleaner than a phone snap at dusk, so treat the numbers as an upper bound. Look-alikes still fool it now and then (a silver birch called a weeping willow). Phones need a connection for the cloud listener. And the characters can't hold a conversation: every line is pre-written.
Why Does Open Innovation Matter?
I could put the model wherever it works. The same open weights run in Node at build time, in a browser tab, and on a server that costs nothing when idle. When the model didn't fit in a phone's browser, I moved it to my own server instead of changing vendors.
I decide what's kept. The listener turns a photo into 768 numbers and forgets it. No third party sees photos of people's streets.
I could measure it. Open weights give raw scores, which made the thresholds, the guardian fallback and the published accuracy numbers possible.
It's cheap enough for a free game. A child might take dozens of photos on one walk. That's fine with a server that scales to zero, and not with a per-call bill.
ElevenLabs is the one closed part, and it's used only at build time to perform lines that were already written. The app never calls it.
My Agent Session
I built this with Claude Code. I ran several agents in parallel: one wrote and tested character descriptions, one re-ran the photo calibration, and one built game features test-first. A Playwright script then played the app on a phone-sized screen to check each feature. A CLAUDE.md file held the rules it had to follow: no photo leaves the device except to my own listener, ElevenLabs only at build time, and no line may tell anyone to touch wildlife or step into traffic.
Prize Categories
Best Use of Gemma. EmbeddingGemma 2 is the whole recogniser: 717 descriptions at build time and every photo at play time, on a server and in the browser. Its raw scores make the guardian fallback possible, which keeps wrong names to 7.8%.
Best Use of ElevenLabs. The voices are the game: a character is a voice you collect. The 28 gold Elders are performed with Eleven v4 audio tags, recorded once and shipped as captioned MP3s, so they play instantly and no API quota can break the game.
Go outside. Find a lawn, or the big tree outside your door. Old Oak would love some news.
What's one thing on your street you'd want to hear talk?
Credits: voices by **ElevenLabs* and Kokoro-82M. EmbeddingGemma 2 by Google DeepMind. Transformers.js and ONNX Runtime. Hosted on Netlify and Modal. Test and sample photos from Wikimedia Commons contributors, credited in ATTRIBUTION.md. Fonts: EB Garamond and Atkinson Hyperlegible Next. Icons: Phosphor.*
AI disclosure: built with Claude Code; character lines and this post were drafted with AI help and edited by me.






Top comments (1)
Demn, this is some good enginerring.