This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
What if AI could encourage us to spend less time looking at screens and more time exploring the world around us?
That question inspired me to build Goblin Mode: Nature Scavenger Bingo, an AI-powered outdoor nature scavenger hunt designed to turn an ordinary walk through a park, garden, or campus into an interactive discovery experience.
Instead of endlessly scrolling through an app, users receive nature-based missions, explore their surroundings, photograph real specimens, and use AI to verify their discoveries.
The project addresses two important challenges: keeping users focused on the real world instead of their screens, and identifying natural objects accurately from imperfect photographs.
Goblin Mode combines an illustrated field-journal interface with computer vision, taxonomic knowledge retrieval, location-aware quest generation, and offline functionality.
Its central character is Grimble, an eccentric goblin naturalist who guides users through their discoveries with voice narration and a field-notebook-inspired interface—complete with stitched leather borders, torn deckled parchment pages, brass corner brackets, washi masking tape, and handcrafted watercolor-and-ink SVG botanical sketches.
How It Works
1. Discover Nature Through Interactive Missions
The application presents a 3×3 nature scavenger bingo board filled with missions inspired by leaves, insects, flowers, and other natural objects.
When users select a mission, the interface enters Focus Mode and displays one active challenge. This encourages users to put their phones away while exploring and bring them back only when they need to photograph their discovery.
When a user clicks "Generate New Board" to refresh their missions, the 3×3 grid cleanly unmounts and displays a dual concentric emerald-and-amber loading spiral with an enforced >= 850ms minimum transition window—preventing jarring sub-second visual flashes while giving clear feedback that a fresh board is being synthesized.
2. AI-Powered Nature Verification
Users can photograph a specimen and submit it for verification. Before the image reaches the vision model, the backend enhances the image, evaluates its visual details, and creates a dual-scale inspection image containing both the full photograph and a magnified region.
The verification pipeline combines image analysis, taxonomic knowledge, and multimodal AI to distinguish genuine discoveries from incorrect lookalikes or featureless images:
-
Stage 1 — Computer Vision Pre-Processing (
image_enhancer.py): Raw camera photos are analyzed with a Laplacian edge filter (ImageFilter.FIND_EDGES) to measure structural edge density (ActiveEdgeRatio) and color variance (ColorStd), flagging blank walls or tiny dust specks before AI inference. The image is then converted into YCbCr color space to balance luminance (Ychannel) contrast at a45%blend without distorting natural green chlorophyll hues, sharpened viaUnsharpMask(radius=1.6, percent=145%, threshold=3)to recover blurry leaf veins and insect legs, and composited into a single784×448pxSide-by-Side Dual-Scale Naturalist Inspection Plate (Left:448×448pxFull Frame | Right:332×448px2×Magnified Center Detail Crop). -
Stage 2 — 4-Gate Taxonomic Referee (
verify.py): Evaluates Subject Integrity (Gate 1), Mandatory Taxonomic Anatomy (Gate 2), Hard-Negative Impostors (Gate 3), and Outdoor Camera Blur Tolerance (Gate 4).
For example, the system is designed to distinguish a Peepal leaf (Ficus religiosa)—which requires a heart-shaped cordate base, straight pinnate-reticulate yellow venation, and an elongated tail-like caudate drip-tip ($\ge 20\%$ of blade length)—from visually similar indoor Money Plant (Epipremnum aureum) vines or Betel (Piper betle) leaves, and strictly reject a tiny dark speck on a floor as evidence of a Weaver Ant (Oecophylla smaragdina) unless 3-part body segmentation, a narrow petiole waist, and 6 jointed legs are visible.
Server-side validation checks the model's response before awarding progress—forcing passed = false and 0 XP if is_impostor_or_fake == true, confidence < 0.85, or the CV speck detector triggered—helping reduce incorrect identifications to achieve 90–95%+ taxonomic precision.
3. Dynamic Quest Generation
Goblin Mode can generate nature missions using location context rather than relying exclusively on a fixed list of challenges.
Its Universal Two-Tier Taxonomic RAG system (botanical_rag.py) combines a curated biodiversity index (Tier 1, using strict whole-word boundary regexes \b...\b) with dynamically generated species-verification criteria (Tier 2). Whenever a player attempts a randomly generated location-based quest that isn't in the static catalog (such as Touch-Me-Not Mimosa pudica, Crimson Callistemon Bottlebrush, Seven-Spotted Coccinella Ladybug, or Plumeria rubra), our backend queries the open-weight 120B openai/gpt-oss-120b model (reasoning_effort="low") on the fly to synthesize its scientific name, mandatory diagnostic traits, and lookalike impostors to reject. Cached results in _DYNAMIC_RAG_CACHE are reused in 0ms, while deterministic domain-class fallback rules handle cases where the AI service is unavailable.
4. Offline-First Exploration
Outdoor connectivity is not always reliable, so the application includes client-side HTML5 Canvas image normalization (64px min to 1280px max bounding box, #FFFFFF alpha compositing, 0.88 JPEG compression under 300 KB) and IndexedDB persistence (idb-keyval) for local progress and queued specimens.
A 3-tier procedural quest-generation system cascades from Local Ollama (gemma2:2b) $\to$ Cloud Groq (openai/gpt-oss-120b) $\to$ a client-side Combinatoric Matrix (combinatoricMatrix.ts) that multiplies 15 sensory adjectives, 15 nature nouns, and 10 location contexts to produce up to 2,250 combinations for zero-signal offline exploration.
Account-specific local storage (goblin_bingo_board_v4_<uid> vs. goblin_bingo_board_v4_guest) also keeps individual users' boards and XP progress strictly isolated on the same device.
5. Voice, Atmosphere, and Accessibility
Grimble provides expressive voice narration through ElevenLabs (eleven_multilingual_v2) when configured, backed by a deterministic SHA-256 disk audio cache (backend/audio_cache/) that drops repeat voice latency from ~1,400ms to <15ms and supported by a dedicated Centralized Dual-Channel Audio Controller (audioManager.ts) that separates SFX and Voice channels so UI clicks never overlap or interrupt Grimble's field commentary.
The interface also adapts its visual atmosphere to the time of day in Indian Standard Time (UTC +05:30) on a fast 10-second polling heartbeat, with distinct Morning/Day (07:00–17:00 IST) sunlit pollen motes, Evening/Sunset (17:00–19:00 IST) amber-violet twilight washes, and Night (19:00–07:00 IST) moonlit washes with a crescent moon, twinkling stars, and drifting bioluminescent fireflies that are strictly hidden outside nocturnal hours.
Together, these features aim to make nature exploration feel like an interactive field adventure rather than another screen-based game.
Demo
Live Web App (Please try on phone for best experience):
Live Backend Health Telemetry: https://goblin-nature-bingo.onrender.com/api/health
Demo video: https://drive.google.com/file/d/1DqGSOqb6mjMUv41udrZGpEPFkdf78fXF/view?usp=sharing
Code
GitHub Repository: https://github.com/SoumyaGhosh2006/Goblin_Nature_Bingo
🌿 Goblin Nature Bingo
Offline-First, Location-Aware Outdoor Field Journal & AI Botanical Scavenger Hunt
📖 Overview
Goblin Nature Bingo is an interactive, offline-first outdoor scavenger hunt web application guided by Grimble, a voice-synthesized goblin naturalist who gets players off their screens and physically outside exploring parks, gardens, campuses, and neighborhood trails.
Instead of endlessly scrolling through a screen, users receive nature-based missions on a 3×3 weathered leather-and-parchment field binder, step outside to explore their surroundings, photograph real botanical and zoological specimens, and use a multi-stage Computer Vision + Universal Taxonomic RAG + Open-Weight Multimodal AI referee to verify their discoveries.
Core Capabilities
-
Tactile Naturalist Field Journal & Screen-Free Focus Mode: A
3×3deckled-paper Bingo grid featuring stitched leather borders, brass corner brackets, washi masking tape, and bespoke ink-and-watercolor SVG botanical illustrations. Selecting any quest tile enters Focus Mode, collapsing the UI into a single sensory outdoor mission…
Explore the implementation, report issues, and contribute to making outdoor AI experiences more reliable and accessible.
How I Built It
Frontend: React 19, TypeScript, Vite, Tailwind CSS, Framer Motion, HTML5 Canvas, and IndexedDB using idb-keyval.
Backend: Python 3.11, FastAPI, Pillow, httpx, and Pydantic.
AI and Computer Vision: Qwen 3.8 27B and GPT-OSS 120B through Groq, with local Ollama integrations for Gemma 2 and Moondream.
Authentication and Persistence: Firebase Authentication, Cloud Firestore, and account-scoped IndexedDB storage.
Voice and Monitoring: ElevenLabs text-to-speech and Sentry SDK integration.
Deployment: Vercel for the frontend and Render for the FastAPI backend.
The architecture separates image enhancement, taxonomic retrieval, visual verification, quest generation, persistence, and audio management into distinct components. This makes individual systems easier to maintain and gives the application multiple fallback paths when a service is unavailable.
Our Engineering Journey: Real Problems We Faced & How We Tackled Them
Building an AI app that works reliably outdoors required solving several non-obvious systems and computer-vision problems that emerged during real-world testing:
1. The "Everything Is Blurry" Bug & RGB Auto-Contrast Distortion
-
What Went Wrong: When we first deployed the backend to Render and tested real phone photos—including a real damp mossy wall in a bathroom and outdoor leaves—the verifier kept rejecting valid photos as "blurry or obscured." Tracing the pipeline revealed two culprits: first, our cloud container was attempting to reach a local Ollama vision instance (
127.0.0.1:11434) first and falling back to a blind placeholder description when Ollama wasn't running on Render; second, when we added standard RGBImageOps.autocontrastto sharpen blurry camera shots, it stretched the Red, Green, and Blue histograms independently. On a green leaf with very little red-channel variance, independent RGB stretching distorted natural chlorophyll greens into dark, unnatural blotches! -
How We Solved It: We routed cloud vision directly to multimodal
qwen/qwen3.8-27band re-engineeredimage_enhancer.pyto convert images intoYCbCrcolor space, applying a45%-blendedautocontrastexclusively on the Luminance (Y) channel before merging back with the untouchedCbandCrchrominance channels. Paired withUnsharpMask(radius=1.6, percent=145%, threshold=3), this recovered crisp leaf venation, serrated edges, and mossy bryophyte textures from slightly blurry phone shots while preserving 100% of true botanical RGB hues.
2. When the Model Became "Too Chill": Stopping Floor Dots & Vine Impostors
- What Went Wrong: After making the prompt tolerant of slight camera blur, the model became overly lenient. When we tested it by submitting a photo of tiny black and white dots on a floor for a "Weaver Ant Trail" quest, it approved them as ants! When we submitted a random heart-shaped indoor Money Plant vine leaf for a "Heart-Shaped Peepal Leaf" quest, it approved that too—only rejecting photos when the frame was a completely blank wall.
-
How We Solved It:
-
Laplacian Edge & Speck Telemetry: We added pre-inference computer vision telemetry using
ImageFilter.FIND_EDGESandImageStat. Any image withActiveEdgeRatio < 0.010andColorStd < 18.0(orEdgeMean < 2.0) is automatically flagged withis_featureless_or_speck = True, mathematically catching floor dots and blank surfaces. -
4-Gate Taxonomic RAG & Whole-Word Regexes: We built
botanical_rag.pywith explicit Mandatory Diagnostic Traits and Hard-Negative Impostors to Reject, plus a server-side post-validation guard (confidence >= 0.85andis_impostor_or_fake == False). During testing, we also caught a classic regex bug: substring matching caused the keyword"ant"in our arthropod classifier to falsely match"Touch-Me-Not Plant"(pl-ant) and"Fragrant Plumeria"(fragr-ant)! Switching to strict word-boundary regular expressions (\b{keyword}s?\b) fixed the collision immediately. In benchmark runs, floor dots and wrong vine leaves were cleanly rejected atconfidence = 0.00–0.05.
-
Laplacian Edge & Speck Telemetry: We added pre-inference computer vision telemetry using
3. Overcoming the 7,000 ITPM Token Wall & Supporting Infinite Location Quests
-
What Went Wrong: Because players can generate quests dynamically anywhere in the world, we couldn't rely on a static hardcoded species list. However, when we initially sent two high-resolution base64 images (
1280pxFull Frame +768px2×Detail Crop) and ran dynamic RAG synthesis on the same vision model (qwen/qwen3.8-27b), a single verification consumed 5,194 input tokens—hitting Groq's7,000 ITPMfree-tier rate limit on the very second photo! We also discovered that legacy.envconfigurations referencingllama-3.2-11b-vision-previewfailed withHTTP 400 model_decommissioned. -
How We Solved It:
-
5.5× Vision Token Reduction: Instead of sending two separate high-res images,
image_enhancer.pycomposites the448×448Enhanced Full View and the332×4482×Magnified Center Zoom side-by-side into a single784×448pxNaturalist Inspection Plate (~450 vision tokens). Combined with/no_thinkandmax_tokens: 220, total input tokens dropped from 5,194 to ~900 tokens per verification (an 82% reduction), enabling 6–7 consecutive photo verifications per minute with zero rate-limit errors. -
Decoupled 120B Dynamic RAG Synthesizer: We routed Tier-2 Dynamic Taxonomic RAG synthesis and board generation to Groq's 120B
openai/gpt-oss-120b(reasoning_effort="low"), which uses a separate rate-limit bucket and synthesizes full scientific rubrics for any unseen location quest in~0.3s(cached in_DYNAMIC_RAG_CACHEfor0msrepeat lookups). -
Self-Healing Model Config: Added automatic runtime sanitization in
config.pythat transparently upgrades decommissioned.envmodel strings toqwen/qwen3.8-27b.
-
5.5× Vision Token Reduction: Instead of sending two separate high-res images,
4. Eliminating Audio Layering, Clock Drift & Cross-Account State Leaks
-
Centralized Audio Preemption: Rapidly clicking specimen tiles, lighting toggles, and Grimble's voice button originally caused 2–3 audio instances to play simultaneously. We replaced scattered
.play()calls with a singleton dual-channel controller (audioManager.ts) that runspause()+currentTime = 0with a60msdebounce on theSFXchannel while protecting theVoicechannel from minor UI interruptions. -
Account-Scoped Storage Hydration: To prevent Player A's offline board and XP from leaking when Player B logs in on the same device, we bound IndexedDB keys to Firebase Auth's
onAuthStateChangedUID (goblin_bingo_board_v4_<uid>) and purged in-memory state on logout.
Why Open Innovation Matters
Open-weight models and accessible developer tools make it possible to experiment with AI beyond conventional chatbot interfaces.
Goblin Mode explores how these technologies can support outdoor learning, biodiversity awareness, and hands-on discovery. Its combination of cloud models (qwen/qwen3.8-27b and openai/gpt-oss-120b), local inference options (gemma2:2b and moondream via Ollama), and deterministic offline fallbacks (2,250-combination procedural matrix) is designed to reduce dependence on a single proprietary vendor, survive model deprecations seamlessly, and run at $0 cost for students and educators.
Local inference through Ollama can also support privacy-conscious exploration when the relevant models are installed and configured on the user's device, keeping backyard and campus photographs completely on local hardware.
The broader goal is to make AI a companion to real-world experiences, not a replacement for them.
My Agent Session
Prize Categories
-
Best Use of Render: Deploys and orchestrates our FastAPI AI verification, Pillow Computer Vision enhancer, and Dynamic Taxonomic RAG backend (
goblin-nature-bingo-api) via a declarativerender.yamlBlueprint andProcfile. -
Best Use of Gemma: Integrates Google's open-weight Gemma 2 (
gemma2:2b) model via Ollama (backend/services/ollama_client.py) for local zero-cost procedural quest generation and botanical verification. -
Best Use of ElevenLabs: Powers Grimble the Goblin's voice narration (
backend/services/elevenlabs_svc.py,backend/routes/voice.py) with deterministic SHA-256 disk caching and a dual-channel audio preemption manager. -
Best Use of Sentry Agent Tracing: Uses
sentry-sdk>=2.0.0(backend/main.py,backend/routes/verify.py) with FastAPI transaction tracing (traces_sample_rate=1.0) to monitor multimodal vision latency, rate-limit retries, and RAG pipeline performance.
What I Learned
Building Goblin Mode gave me an opportunity to explore the challenges of combining computer vision, retrieval-augmented generation, offline-first web development, and multimodal AI in one application.
I learned that a useful AI experience requires more than a model call. Image quality, color-space math (YCbCr vs. RGB), false-positive impostor gates, token budget optimization (5,194 $\to$ ~900 tokens), SHA-256 caching, local IndexedDB persistence, per-UID account isolation, and multi-tier failure recovery all matter when software is used outside controlled conditions.
The project also reinforced an important design principle: technology should encourage meaningful experiences rather than demand constant attention.
Step outside. Find something unexpected. Complete your next nature quest.

Top comments (0)