DEV Community

Cover image for Goblin Mode: Nature Scavenger Bingo — An Offline-First AI Nature Hunt
Soumya ghosh
Soumya ghosh

Posted on

Goblin Mode: Nature Scavenger Bingo — An Offline-First AI Nature Hunt

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

Grimble's Field Study board with Peepal Leaf, Neem Leaflet, Banyan Root, Wild Hibiscus and Garden Lizard tiles

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

What if AI could encourage us to spend less time looking at screens and more time exploring the world around us?

That question inspired me to build Goblin Mode: Nature Scavenger Bingo, an AI-powered outdoor nature scavenger hunt designed to turn an ordinary walk through a park, garden, or campus into an interactive discovery experience.

Instead of endlessly scrolling through an app, users receive nature-based missions, explore their surroundings, photograph real specimens, and use AI to verify their discoveries.

The project addresses two important challenges: keeping users focused on the real world instead of their screens, and identifying natural objects accurately from imperfect photographs.

Goblin Mode combines an illustrated field-journal interface with computer vision, taxonomic knowledge retrieval, location-aware quest generation, and offline functionality.

Its central character is Grimble, an eccentric goblin naturalist who guides users through their discoveries with voice narration and a field-notebook-inspired interface—complete with stitched leather borders, torn deckled parchment pages, brass corner brackets, washi masking tape, and handcrafted watercolor-and-ink SVG botanical sketches.

How It Works

1. Discover Nature Through Interactive Missions

The application presents a 3×3 nature scavenger bingo board filled with missions inspired by leaves, insects, flowers, and other natural objects.

When users select a mission, the interface enters Focus Mode and displays one active challenge. This encourages users to put their phones away while exploring and bring them back only when they need to photograph their discovery.

When a user clicks "Generate New Board" to refresh their missions, the 3×3 grid cleanly unmounts and displays a dual concentric emerald-and-amber loading spiral with an enforced >= 850ms minimum transition window—preventing jarring sub-second visual flashes while giving clear feedback that a fresh board is being synthesized.

2. AI-Powered Nature Verification

Users can photograph a specimen and submit it for verification. Before the image reaches the vision model, the backend enhances the image, evaluates its visual details, and creates a dual-scale inspection image containing both the full photograph and a magnified region.

The verification pipeline combines image analysis, taxonomic knowledge, and multimodal AI to distinguish genuine discoveries from incorrect lookalikes or featureless images:

  • Stage 1 — Computer Vision Pre-Processing (image_enhancer.py): Raw camera photos are analyzed with a Laplacian edge filter (ImageFilter.FIND_EDGES) to measure structural edge density (ActiveEdgeRatio) and color variance (ColorStd), flagging blank walls or tiny dust specks before AI inference. The image is then converted into YCbCr color space to balance luminance (Y channel) contrast at a 45% blend without distorting natural green chlorophyll hues, sharpened via UnsharpMask(radius=1.6, percent=145%, threshold=3) to recover blurry leaf veins and insect legs, and composited into a single 784×448px Side-by-Side Dual-Scale Naturalist Inspection Plate (Left: 448×448px Full Frame | Right: 332×448px 2× Magnified Center Detail Crop).
  • Stage 2 — 4-Gate Taxonomic Referee (verify.py): Evaluates Subject Integrity (Gate 1), Mandatory Taxonomic Anatomy (Gate 2), Hard-Negative Impostors (Gate 3), and Outdoor Camera Blur Tolerance (Gate 4).

For example, the system is designed to distinguish a Peepal leaf (Ficus religiosa)—which requires a heart-shaped cordate base, straight pinnate-reticulate yellow venation, and an elongated tail-like caudate drip-tip ($\ge 20\%$ of blade length)—from visually similar indoor Money Plant (Epipremnum aureum) vines or Betel (Piper betle) leaves, and strictly reject a tiny dark speck on a floor as evidence of a Weaver Ant (Oecophylla smaragdina) unless 3-part body segmentation, a narrow petiole waist, and 6 jointed legs are visible.

Server-side validation checks the model's response before awarding progress—forcing passed = false and 0 XP if is_impostor_or_fake == true, confidence < 0.85, or the CV speck detector triggered—helping reduce incorrect identifications to achieve 90–95%+ taxonomic precision.

3. Dynamic Quest Generation

Goblin Mode can generate nature missions using location context rather than relying exclusively on a fixed list of challenges.

Its Universal Two-Tier Taxonomic RAG system (botanical_rag.py) combines a curated biodiversity index (Tier 1, using strict whole-word boundary regexes \b...\b) with dynamically generated species-verification criteria (Tier 2). Whenever a player attempts a randomly generated location-based quest that isn't in the static catalog (such as Touch-Me-Not Mimosa pudica, Crimson Callistemon Bottlebrush, Seven-Spotted Coccinella Ladybug, or Plumeria rubra), our backend queries the open-weight 120B openai/gpt-oss-120b model (reasoning_effort="low") on the fly to synthesize its scientific name, mandatory diagnostic traits, and lookalike impostors to reject. Cached results in _DYNAMIC_RAG_CACHE are reused in 0ms, while deterministic domain-class fallback rules handle cases where the AI service is unavailable.

4. Offline-First Exploration

Outdoor connectivity is not always reliable, so the application includes client-side HTML5 Canvas image normalization (64px min to 1280px max bounding box, #FFFFFF alpha compositing, 0.88 JPEG compression under 300 KB) and IndexedDB persistence (idb-keyval) for local progress and queued specimens.

A 3-tier procedural quest-generation system cascades from Local Ollama (gemma2:2b) $\to$ Cloud Groq (openai/gpt-oss-120b) $\to$ a client-side Combinatoric Matrix (combinatoricMatrix.ts) that multiplies 15 sensory adjectives, 15 nature nouns, and 10 location contexts to produce up to 2,250 combinations for zero-signal offline exploration.

Account-specific local storage (goblin_bingo_board_v4_<uid> vs. goblin_bingo_board_v4_guest) also keeps individual users' boards and XP progress strictly isolated on the same device.

5. Voice, Atmosphere, and Accessibility

Grimble provides expressive voice narration through ElevenLabs (eleven_multilingual_v2) when configured, backed by a deterministic SHA-256 disk audio cache (backend/audio_cache/) that drops repeat voice latency from ~1,400ms to <15ms and supported by a dedicated Centralized Dual-Channel Audio Controller (audioManager.ts) that separates SFX and Voice channels so UI clicks never overlap or interrupt Grimble's field commentary.

The interface also adapts its visual atmosphere to the time of day in Indian Standard Time (UTC +05:30) on a fast 10-second polling heartbeat, with distinct Morning/Day (07:00–17:00 IST) sunlit pollen motes, Evening/Sunset (17:00–19:00 IST) amber-violet twilight washes, and Night (19:00–07:00 IST) moonlit washes with a crescent moon, twinkling stars, and drifting bioluminescent fireflies that are strictly hidden outside nocturnal hours.

Together, these features aim to make nature exploration feel like an interactive field adventure rather than another screen-based game.

Demo

Live Web App (Please try on phone for best experience):

Live Backend Health Telemetry: https://goblin-nature-bingo.onrender.com/api/health

Demo video: https://drive.google.com/file/d/1DqGSOqb6mjMUv41udrZGpEPFkdf78fXF/view?usp=sharing

Code

GitHub Repository: https://github.com/SoumyaGhosh2006/Goblin_Nature_Bingo

🌿 Goblin Nature Bingo

Offline-First, Location-Aware Outdoor Field Journal & AI Botanical Scavenger Hunt

Live Demo Backend API License: MIT


📖 Overview

Goblin Nature Bingo is an interactive, offline-first outdoor scavenger hunt web application guided by Grimble, a voice-synthesized goblin naturalist who gets players off their screens and physically outside exploring parks, gardens, campuses, and neighborhood trails.

Instead of endlessly scrolling through a screen, users receive nature-based missions on a 3×3 weathered leather-and-parchment field binder, step outside to explore their surroundings, photograph real botanical and zoological specimens, and use a multi-stage Computer Vision + Universal Taxonomic RAG + Open-Weight Multimodal AI referee to verify their discoveries.

Core Capabilities

  1. Tactile Naturalist Field Journal & Screen-Free Focus Mode: A 3×3 deckled-paper Bingo grid featuring stitched leather borders, brass corner brackets, washi masking tape, and bespoke ink-and-watercolor SVG botanical illustrations. Selecting any quest tile enters Focus Mode, collapsing the UI into a single sensory outdoor mission…

Explore the implementation, report issues, and contribute to making outdoor AI experiences more reliable and accessible.

How I Built It

Frontend: React 19, TypeScript, Vite, Tailwind CSS, Framer Motion, HTML5 Canvas, and IndexedDB using idb-keyval.

Backend: Python 3.11, FastAPI, Pillow, httpx, and Pydantic.

AI and Computer Vision: Qwen 3.8 27B and GPT-OSS 120B through Groq, with local Ollama integrations for Gemma 2 and Moondream.

Authentication and Persistence: Firebase Authentication, Cloud Firestore, and account-scoped IndexedDB storage.

Voice and Monitoring: ElevenLabs text-to-speech and Sentry SDK integration.

Deployment: Vercel for the frontend and Render for the FastAPI backend.

The architecture separates image enhancement, taxonomic retrieval, visual verification, quest generation, persistence, and audio management into distinct components. This makes individual systems easier to maintain and gives the application multiple fallback paths when a service is unavailable.


Our Engineering Journey: Real Problems We Faced & How We Tackled Them

Building an AI app that works reliably outdoors required solving several non-obvious systems and computer-vision problems that emerged during real-world testing:

1. The "Everything Is Blurry" Bug & RGB Auto-Contrast Distortion

  • What Went Wrong: When we first deployed the backend to Render and tested real phone photos—including a real damp mossy wall in a bathroom and outdoor leaves—the verifier kept rejecting valid photos as "blurry or obscured." Tracing the pipeline revealed two culprits: first, our cloud container was attempting to reach a local Ollama vision instance (127.0.0.1:11434) first and falling back to a blind placeholder description when Ollama wasn't running on Render; second, when we added standard RGB ImageOps.autocontrast to sharpen blurry camera shots, it stretched the Red, Green, and Blue histograms independently. On a green leaf with very little red-channel variance, independent RGB stretching distorted natural chlorophyll greens into dark, unnatural blotches!
  • How We Solved It: We routed cloud vision directly to multimodal qwen/qwen3.8-27b and re-engineered image_enhancer.py to convert images into YCbCr color space, applying a 45%-blended autocontrast exclusively on the Luminance (Y) channel before merging back with the untouched Cb and Cr chrominance channels. Paired with UnsharpMask(radius=1.6, percent=145%, threshold=3), this recovered crisp leaf venation, serrated edges, and mossy bryophyte textures from slightly blurry phone shots while preserving 100% of true botanical RGB hues.

2. When the Model Became "Too Chill": Stopping Floor Dots & Vine Impostors

  • What Went Wrong: After making the prompt tolerant of slight camera blur, the model became overly lenient. When we tested it by submitting a photo of tiny black and white dots on a floor for a "Weaver Ant Trail" quest, it approved them as ants! When we submitted a random heart-shaped indoor Money Plant vine leaf for a "Heart-Shaped Peepal Leaf" quest, it approved that too—only rejecting photos when the frame was a completely blank wall.
  • How We Solved It:
    1. Laplacian Edge & Speck Telemetry: We added pre-inference computer vision telemetry using ImageFilter.FIND_EDGES and ImageStat. Any image with ActiveEdgeRatio < 0.010 and ColorStd < 18.0 (or EdgeMean < 2.0) is automatically flagged with is_featureless_or_speck = True, mathematically catching floor dots and blank surfaces.
    2. 4-Gate Taxonomic RAG & Whole-Word Regexes: We built botanical_rag.py with explicit Mandatory Diagnostic Traits and Hard-Negative Impostors to Reject, plus a server-side post-validation guard (confidence >= 0.85 and is_impostor_or_fake == False). During testing, we also caught a classic regex bug: substring matching caused the keyword "ant" in our arthropod classifier to falsely match "Touch-Me-Not Plant" (pl-ant) and "Fragrant Plumeria" (fragr-ant)! Switching to strict word-boundary regular expressions (\b{keyword}s?\b) fixed the collision immediately. In benchmark runs, floor dots and wrong vine leaves were cleanly rejected at confidence = 0.00–0.05.

3. Overcoming the 7,000 ITPM Token Wall & Supporting Infinite Location Quests

  • What Went Wrong: Because players can generate quests dynamically anywhere in the world, we couldn't rely on a static hardcoded species list. However, when we initially sent two high-resolution base64 images (1280px Full Frame + 768px 2× Detail Crop) and ran dynamic RAG synthesis on the same vision model (qwen/qwen3.8-27b), a single verification consumed 5,194 input tokens—hitting Groq's 7,000 ITPM free-tier rate limit on the very second photo! We also discovered that legacy .env configurations referencing llama-3.2-11b-vision-preview failed with HTTP 400 model_decommissioned.
  • How We Solved It:
    • 5.5× Vision Token Reduction: Instead of sending two separate high-res images, image_enhancer.py composites the 448×448 Enhanced Full View and the 332×448 2× Magnified Center Zoom side-by-side into a single 784×448px Naturalist Inspection Plate (~450 vision tokens). Combined with /no_think and max_tokens: 220, total input tokens dropped from 5,194 to ~900 tokens per verification (an 82% reduction), enabling 6–7 consecutive photo verifications per minute with zero rate-limit errors.
    • Decoupled 120B Dynamic RAG Synthesizer: We routed Tier-2 Dynamic Taxonomic RAG synthesis and board generation to Groq's 120B openai/gpt-oss-120b (reasoning_effort="low"), which uses a separate rate-limit bucket and synthesizes full scientific rubrics for any unseen location quest in ~0.3s (cached in _DYNAMIC_RAG_CACHE for 0ms repeat lookups).
    • Self-Healing Model Config: Added automatic runtime sanitization in config.py that transparently upgrades decommissioned .env model strings to qwen/qwen3.8-27b.

4. Eliminating Audio Layering, Clock Drift & Cross-Account State Leaks

  • Centralized Audio Preemption: Rapidly clicking specimen tiles, lighting toggles, and Grimble's voice button originally caused 2–3 audio instances to play simultaneously. We replaced scattered .play() calls with a singleton dual-channel controller (audioManager.ts) that runs pause() + currentTime = 0 with a 60ms debounce on the SFX channel while protecting the Voice channel from minor UI interruptions.
  • Account-Scoped Storage Hydration: To prevent Player A's offline board and XP from leaking when Player B logs in on the same device, we bound IndexedDB keys to Firebase Auth's onAuthStateChanged UID (goblin_bingo_board_v4_<uid>) and purged in-memory state on logout.

Why Open Innovation Matters

Open-weight models and accessible developer tools make it possible to experiment with AI beyond conventional chatbot interfaces.

Goblin Mode explores how these technologies can support outdoor learning, biodiversity awareness, and hands-on discovery. Its combination of cloud models (qwen/qwen3.8-27b and openai/gpt-oss-120b), local inference options (gemma2:2b and moondream via Ollama), and deterministic offline fallbacks (2,250-combination procedural matrix) is designed to reduce dependence on a single proprietary vendor, survive model deprecations seamlessly, and run at $0 cost for students and educators.

Local inference through Ollama can also support privacy-conscious exploration when the relevant models are installed and configured on the user's device, keeping backyard and campus photographs completely on local hardware.

The broader goal is to make AI a companion to real-world experiences, not a replacement for them.

My Agent Session

Building Goblin Mode: Universal Taxonomic RAG & Dual-Scale Computer Vision Referee
You

Fix three issues in Goblin Mode: Nature Scavenger Bingo:
1. Atmosphere phase updates too slowly after device time changes (poll every 10-15s in Auto mode).
2. Redefine IST time boundaries to Day (07:00-17:00), Evening/Sunset (17:00-19:00), and Night (19:00-07:00) with fireflies strictly nocturnal.
3. Fix overlapping audio playback by implementing a centralized dual-channel audio manager (SFX and Voice) that resets currentTime = 0 and protects Grimble's voice commentary from minor UI clicks.

Agent

I will update RealTimeAtmosphere.tsx to poll every 10 seconds using exact UTC-to-IST (+05:30) minute boundary calculations (Day 07:00–17:00, Sunset 17:00–19:00, Night 19:00–07:00) and build a singleton Dual-Channel Audio Manager (audioManager.ts) separating SFX and Voice channels.

You

When someone clicks on 'Generate new board', hide the board and show an animated loading circle/spiral until the new board is ready, then reveal the board.

Agent

I will add an isGeneratingBoard state in App.tsx with a minimum 850ms transition threshold and render a dual concentric emerald/amber loading spiral inside BingoBoard.tsx while new quests are synthesized.

You

Now we need the image verification pipeline to achieve 90-95% taxonomic accuracy across both curated quests and any randomly generated location-based quests. It must not reject slightly blurry real photos (like damp mossy walls), but it must strictly reject fake impostors like black/white floor dots submitted as Weaver Ants or random Money Plant vine leaves submitted as a heart-shaped Peepal leaf. Use computer vision enhancers and a Taxonomic RAG system.

Agent

I will architect a 3-stage Computer Vision + Universal Two-Tier Taxonomic RAG + 4-Gate Multimodal Referee pipeline:
1. backend/services/image_enhancer.py: Computes Laplacian edge density (ImageFilter.FIND_EDGES) to flag featureless floor specks, applies 45%-blended YCbCr Y-channel auto-contrast (preserving true botanical RGB hues), sharpens venation via UnsharpMask(radius=1.6, percent=145, threshold=3), and builds a 784x448px Side-by-Side Dual-Scale Naturalist Inspection Plate (Left: 448x448 Full Frame; Right: 332x448 2x Magnified Center Zoom), cutting vision input tokens by 5.5x (5,194 -> ~900 tokens).
2. backend/services/botanical_rag.py: Implements Tier-1 whole-word regex matching (\b...\b) for 17+ core species and Tier-2 Live Dynamic Taxonomic RAG Synthesis via openai/gpt-oss-120b (with in-memory LRU caching) for 100% of unseen location-generated quests.
3. backend/routes/verify.py: Evaluates the inspection plate with qwen/qwen3.8-27b across 4 sequential gates and enforces a strict server-side post-validation override (confidence >= 0.85, is_impostor_or_fake == False, is_featureless_or_speck == False).

Prize Categories

  • Best Use of Render: Deploys and orchestrates our FastAPI AI verification, Pillow Computer Vision enhancer, and Dynamic Taxonomic RAG backend (goblin-nature-bingo-api) via a declarative render.yaml Blueprint and Procfile.
  • Best Use of Gemma: Integrates Google's open-weight Gemma 2 (gemma2:2b) model via Ollama (backend/services/ollama_client.py) for local zero-cost procedural quest generation and botanical verification.
  • Best Use of ElevenLabs: Powers Grimble the Goblin's voice narration (backend/services/elevenlabs_svc.py, backend/routes/voice.py) with deterministic SHA-256 disk caching and a dual-channel audio preemption manager.
  • Best Use of Sentry Agent Tracing: Uses sentry-sdk>=2.0.0 (backend/main.py, backend/routes/verify.py) with FastAPI transaction tracing (traces_sample_rate=1.0) to monitor multimodal vision latency, rate-limit retries, and RAG pipeline performance.

What I Learned

Building Goblin Mode gave me an opportunity to explore the challenges of combining computer vision, retrieval-augmented generation, offline-first web development, and multimodal AI in one application.

I learned that a useful AI experience requires more than a model call. Image quality, color-space math (YCbCr vs. RGB), false-positive impostor gates, token budget optimization (5,194 $\to$ ~900 tokens), SHA-256 caching, local IndexedDB persistence, per-UID account isolation, and multi-tier failure recovery all matter when software is used outside controlled conditions.

The project also reinforced an important design principle: technology should encourage meaningful experiences rather than demand constant attention.

Step outside. Find something unexpected. Complete your next nature quest.

Hacktoberfest #DEVChallenge #HF26Challenge #TouchGrass #AI #OpenSource #Gemma #TabPFN #ElevenLabs #Sentry #Render

Top comments (0)