This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
Touch Grass Bingo: Nature Exploration Powered by OpenCLIP & Render ๐ฟ
What I Built
Modern developers and students often spend 12+ hours a day behind glowing monitors, debugging code and attending virtual meetings. For the Week 1: Touch Grass challenge, the objective was simple: build an open-source AI project where the screen is the absolute shortest part of the experience, serving solely as an incentive to step outdoors.
Touch Grass Bingo is a mobile-first web app that generates a daily 3x3 nature bingo card seeded deterministically by the current calendar date. Each tile challenges the user to locate an outdoor element commonly encountered in nature, specifically localized for Sri Lanka, featuring both English and Sinhala terminology (e.g., coconut palms, fallen jackfruit leaves, wild millipedes, red flowers, and rough tree bark).
The Gameplay Loop
- Step Outside: Open the web app on any mobile browser to view today's 9-item randomized card.
- Explore & Photograph: Find an item in your garden, yard, or neighborhood park and take a picture using your mobile camera.
- Instant Zero-Shot Verification: The client compresses the photo and sends it to our self-hosted OpenCLIP vision model on Render. The model performs zero-shot image-text verification against target prompts and distractors.
- Pocket the Phone: If verified, the tile locks in green and your progress persists in local storage. Total screen time is roughly 5 seconds per item; the rest of the time is spent observing the natural world.
Demo
- ๐ Live Application: https://touch-grass-bingo.onrender.com (Best experienced on a mobile smartphone with camera access)
Code
The complete source code is public and licensed under the MIT License:
ChalanaDilshan
/
touch-grass-bingo
A mobile-first daily nature bingo game powered by OpenCLIP zero-shot vision. Step outside, snap photos, and touch grass. Built for Hacktoberfest 2026.
๐ฟ Touch Grass Bingo
A mobile-first daily nature bingo card powered by open-weight AI vision.
Step outside, photograph real-world Sri Lankan flora & fauna, and let AI verify your finds
What is it?
Touch Grass Bingo is a gamified nature walk app that generates a fresh 3ร3 bingo card every day filled with Sri Lankan outdoor nature items โ coconut trees, ant trails, spider webs, jackfruit leaves, and more.
Tap any tile โ snap a photo with your phone camera โ an open-weight vision AI (OpenCLIP ViT-B-32) verifies your photo in real time using zero-shot classification (no training, no fine-tuning needed).
Complete 3 in a row to score a BINGO ๐, and fill the full card to become a Garden Master ๐.
Features
Feature
Detail
๐ฟ Daily Nature Card
Fresh 9-item card every day, seeded by date (all players share the same card)
๐ค AI Verification
OpenCLIP ViT-B-32
Model Attribution & Open Weights
-
Model Architecture: OpenCLIP ViT-B/32 (
laion2b_s34b_b79kweights) - Model Source: mlfoundations/open_clip
- Model License: MIT License
How I Built It
The engineering philosophy behind Touch Grass Bingo was built around severe resource constraints: zero GPU cost, low bandwidth overhead on mobile networks, and lightweight container memory footprints.
[Mobile Phone Camera]
โ
โผ (HTML5 Canvas compression to 512px max dimension)
[Render Docker Container (FastAPI)]
โ
โผ (Single Uvicorn worker on PyTorch CPU)
[OpenCLIP Zero-Shot Verification]
โ
โโโบ Encode Image Embedding
โโโบ Encode Target Prompt & Distractor Embeddings
โโโบ Softmax Cosine Similarity Scoring
โ
โผ
[Client UI] โโโบ Tile turns green, stored in browser localStorage
1. The AI Core: Zero-Shot Classification with Distractor Scoring
Rather than collecting thousands of nature images and training a brittle custom image classification pipeline, the backend utilizes OpenCLIP (ViT-B-32). Because CLIP models are pre-trained on billions of image-text pairs, they generalize exceptionally well to diverse natural entities in an open-world setting.
To eliminate false positives (e.g., mistaking a red plastic bucket or red vehicle for a real red flower), each entry in items.json defines a Target Prompt, a set of Negative Distractor Prompts, and a confidence threshold:
{
"id": "1",
"name_en": "Red Flower",
"name_si": "เถปเถญเท เถธเถฝเถเท",
"prompt": "a close up photo of a real red flower growing in nature",
"distractors": [
"a red piece of cloth",
"a red car",
"a red plastic bucket",
"a person wearing red"
],
"threshold": 0.55
}
When an image arrives, the backend calculates cosine similarities between the image embedding and all text tokens (target + distractors). An item is only verified if the target prompt achieves the highest softmax probability score and crosses the required confidence threshold.
2. Frontend & Bandwidth Optimization
-
Native Rear-Camera Invocation: Built using pure semantic HTML5 with
<input type="file" accept="image/*" capture="environment">. This triggers the phone's native rear camera interface without cumbersome third-party camera libraries. -
Client-Side Canvas Downscaling: Native phone camera photos typically weigh between 5 MB and 12 MB. Uploading raw images would overwhelm mobile data connections and spike backend RAM. Using an in-memory HTML
<canvas>, the client downscales images to a maximum dimension of 512px at 80% JPEG quality before dispatch. This reduces upload payloads to under 60 KB without sacrificing CLIP's 224x224 input resolution. -
Privacy & State Storage: Tile progress is saved directly in browser
localStorage. No uploaded photos are stored on disk or database; images are inspected in-memory and discarded immediately upon response.
3. Backend & Render Container Architecture
- Built with FastAPI on Python 3.10.
- Dockerized with a lean, CPU-only PyTorch layer (
--index-url https://download.pytorch.org/whl/cpu) to avoid multi-gigabyte CUDA bloat. - Configured with a single Uvicorn worker (
--workers 1) to guarantee deterministic memory usage within Render's container constraints.
Real-World Outdoor Testing & Honesty Report
A core bonus of the Touch Grass challenge is actually taking the project into the field and reporting candidly on how it performed. I took my smartphone outside in my neighborhood in Sri Lanka to run the bingo card through actual outdoor lighting conditions.
What Worked Well
- High-Texture Subjects: Rough tree bark, stone surfaces, and rock moss consistently scored above 82% confidence on the first try. CLIP's vision transformer excels at high-frequency surface textures.
- Distractor Rejection: Testing the "Red Flower" prompt against a bright red plastic flower pot successfully failed validation; the distractor "a red plastic bucket" captured the highest probability share, accurately rejecting the false item.
- Inference Speed: Average verification latency on the Render container hovered between 1.1s and 1.8s per request, keeping outdoor screen interruptions minimal.
Failure Modes & Limitations
- High Tropical Glare: Midday direct sunlight washed out contrast on broad leaves (such as banana and jackfruit leaves), occasionally dropping target confidence below threshold until shaded manually by hand.
- Scale and Distance Issues: Tiny insects (such as ant trails) photographed from standing eye level failed to trigger positive matches because the ants occupied less than 2% of the frame. Close-up framing was necessary for verification.
- Network Dependency: Because the model runs in a cloud container on Render rather than on-device, an active mobile data connection is mandatory.
Why Does Open Innovation Matter?
This project highlights three core tenets of why open innovation and open-weight models are indispensable:
- User Privacy & Sovereign Compute: Outdoor exploration apps naturally photograph surroundings, personal property, and neighborhood environments. By hosting an open-weight OpenCLIP model on our own private Render instance, user imagery is never sent to proprietary APIs where it could be logged or used for model retraining.
- Zero Marginal Cost for Iteration: Adding or modifying bingo items is as simple as adding entries to a JSON file. There are no per-token API charges, surprise rate limits, or proprietary platform restrictions.
- Inspectable & Tunable Behavior: When classification failed during outdoor testing, having direct access to prompt embeddings allowed immediate debugging and calibration of negative distractors, a transparency impossible with opaque, closed vision endpoints.
Prize Categories
- Primary Challenge: Open-Source AI Challenge, Week 1: Touch Grass
- Partner Prize: Best Use of Render (FastAPI backend and OpenCLIP inference runtime deployed and hosted on Render via Docker)


Top comments (0)