This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
WildQuest AI turns a short walk into a nature-discovery game. Pick a session of 3, 5, or 7 missions (roughly 5-20 minutes), and each mission asks you to find something real outside โ a leaf, a flower, bark, the sky. You open your camera, photograph what you found, and an on-device AI model checks whether it's a match. There are no feeds, no streaks, and no accounts โ just a short list of reasons to look up from your screen. When the session ends, the app tells you to put the phone away.
It's for anyone who wants a tiny, concrete excuse to step outside: someone on a work break, a parent keeping a kid occupied on a walk, or a remote worker who needs ten real minutes of air between meetings.
Demo
Live app: https://touch-grass-sepia-theta.vercel.app/
It's installable as a PWA โ add it to your home screen, and after the first visit (which downloads the ~90MB model) it keeps working offline.
Code
GitHub: https://github.com/Gautam5514/Touch-Grass
๐ฟ WildQuest AI
Explore more. Scroll less. Your next adventure is outside.
WildQuest AI turns a short walk into a nature-discovery game. Get a handful of small missions โ find a green leaf, a flower, a tree trunk, the sky โ photograph what you find, and an open-weight vision model running entirely on your own device checks each photo. When you finish, the app tells you to put the phone away.
Built for the DEV Hacktoberfest Open-Source AI Challenge ยท Week 1: Touch Grass.
๐ธ Screens
Screenshots use real photos with a stubbed score from the mock-AI test build, so they're deterministic. Real model behavior is covered by
tests/e2e/real.spec.tsanddocs/EVALUATION.md.
๐ Contents
How I Built It
The verification core runs entirely on-device, with open-source AI doing all the work โ no paid vision API, no inference server, and no uploaded photos.
- Model: Xenova/clip-vit-base-patch32, a quantized ONNX port of OpenAI's open CLIP ViT-B/32 (~90MB), run through Transformers.js and ONNX Runtime Web (WASM) inside a Web Worker so the UI never blocks.
- Matching: each captured photo is scored against text prompts for the target, its aliases, visual look-alikes, and the other missions' targets. A decision rule (the target must rank first, score at least 0.45, and lead the next-best match by at least 0.15) returns matched, uncertain, mismatch, or error.
- Stack: Next.js 16, React 19, TypeScript, Tailwind CSS 4, Dexie (IndexedDB) for local history, a service worker for the PWA/offline cache, and Vitest + Playwright for unit and end-to-end tests, including real-model end-to-end runs.
- Because the model ships with the app and runs locally, the only outbound network call in the whole experience is the one-time model download from Hugging Face. Everything else, including every photo you take, never leaves your device.
Why Does Open Innovation Matter?
An open-weight model is what makes the core promise of this app possible: photos that never leave your phone. A closed, API-based vision model would mean sending every "is this a leaf?" photo to a server โ which breaks the privacy story, adds per-request cost and latency, and makes offline use impossible. Because CLIP's weights are open and small enough to quantize down to a browser-friendly ONNX build, I could ship the entire verification pipeline inside the static app bundle, test it deterministically offline, and let anyone self-host or fork the project without touching a provider's API key or billing dashboard. Open weights turned "an AI feature" into "a free, static, forkable PWA."
Prize Categories
Entering for the overall challenge prize.








Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.