This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Wellness apps usually trap you in a new cycle of screen time. I wanted to build something that forces you to use your phone simply as a viewfinder for the physical world ๐๏ธ.
Meet Terra-Gotchi ๐พโan open-source, fully offline digital pet that survives strictly by "eating" real-world natural textures ๐. It asks for a craving (e.g., "I want a rough piece of tree bark" ๐ชต or "Find me a green leaf" ๐), and you have to physically go outside, find it, and take a picture ๐ท. A local multimodal AI acts as a judge โ๏ธ, verifying your photo before the pet accepts the meal ๐. It is designed for developers, students, and anyone looking to gamify their time away from their desks ๐ถโโ๏ธ.
Demo
Code
๐ฑ Terra-Gotchi ๐
The digital pet that only eats real grass.
A 100% offline, multimodal AI companion that makes you go outside.
๐ Table of Contents
- โจ What is Terra-Gotchi?
- ๐ฎ How It Works
- ๐ Key Features
- ๐งฐ Tech Stack
- ๐๏ธ Architecture
- โ๏ธ Installation & Setup
- ๐น๏ธ Usage
- ๐ ๏ธ Troubleshooting
- ๐บ๏ธ Roadmap
- ๐ค Contributing
- ๐ฉโ๐ป Author
โจ What is Terra-Gotchi?
Virtual pets have always asked the same thing of us: more screen time. Terra-Gotchi flips the script.
Terra-Gotchi is an offline, multimodal AI digital pet built for the Hacktoberfest 2026 "Touch Grass" challenge. You can't feed it by tapping a button. Instead, your pet craves something from the real world:
๐ฃ๏ธ "I'm so hungry... find me a green leaf!" ๐ฃ๏ธ "I want something rough, like tree bark!"
To feed it, you have to put the phone down, walk outside, find the object, and photograph it. Aโฆ
How I Built It
Building secure, local-first systems is a major interest of mine, so I built this specifically to prove that powerful multimodal AI doesn't need cloud infrastructure โ๏ธ๐ซ. This entire project runs 100% offline on local hardware:
๐ง The Brain (Open-Weight LLM): Google's lightweight gemma:2b running locally via Ollama generates dynamic, creative cravings and reacts joyfully when fed.
๐๏ธ The Eyes (Vision Verification): The llava open-source multimodal model visually inspects the user's camera feed to verify if the photo actually contains the requested natural item.
โ๏ธ The Engine & Memory: A Python/Flask backend routes the AI logic, paired with an embedded SQLite database ๐พ to store the pet's state and interaction history locally.
๐ฅ๏ธ The Interface: A responsive HTML/Vanilla JS frontend utilizes the WebRTC API for live camera capture ๐ฅ and the Web Speech API ๐ฃ๏ธ so the pet can speak its cravings out loudโmeaning you don't even have to look at the screen while hunting for its food.
Why Does Open Innovation Matter?
The "Touch Grass" theme is about disconnecting ๐. If I built this using a closed API, the app would instantly break the moment you walked into the woods where there is no Wi-Fi ๐ฒ.
By running gemma:2b and llava locally with Ollama ๐ฆ, Terra-Gotchi becomes a true outdoor companion ๐๏ธ. It costs absolutely nothing to run ๐ธ, keeps all user camera data strictly on their own device ๐, and proves that we can run complex, multi-agent AI systems in the middle of a forest without relying on a centralized cloud server.
Prize Categories
๐ Best Use of Gemma: For utilizing the open-weight gemma:2b model locally to power the agent's core personality, cravings, and conversational reactions.


Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.