This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Reality Remix is an open-source AI app that turns a photo of your surroundings into a short real-world challenge.
Take a photo of what's around you, and Reality Remix uses a local vision model to understand the environment and create a physical-world challenge based only on what it can actually see.
Instead of giving you another screen-based activity, it gives you a reason to look up, move around, explore, and interact with the world around you.
The idea is simple:
Take a photo → Get a challenge → Put the phone down → Touch grass.
The project is designed for anyone who wants a small reason to step away from their screen and interact with their physical surroundings.
Demo
The demo shows the complete flow:
Upload a photo → AI analyzes the surroundings → Receive a real-world challenge → Ask follow-up questions through the challenge chat.
Code
The complete source code is available on GitHub:
https://github.com/chaitanya2850/Reality-Remix
The repository also includes the setup instructions for running the local AI model with Docker and Ollama.
How I Built It
Reality Remix is built with:
- React for the frontend
- Spring Boot for the backend
- Ollama for local AI inference
- Gemma 3 4B for vision and challenge generation
- Docker to run Ollama
The frontend sends an image to the Spring Boot backend.
The backend sends the image to Gemma through Ollama. Gemma analyzes the visible environment and generates a challenge based only on what it can see.
The generated challenge is then returned to the React frontend.
I also built a persistent chat for each challenge. Users can ask questions about the current challenge and continue the conversation without losing previous messages.
The core AI inference happens locally rather than relying on a proprietary cloud AI API.
Why Does Open Innovation Matter?
Open innovation is important to Reality Remix because the project is built around local AI.
Using an open-weight model and an open-source inference runtime means the user's image can be processed on their own machine instead of being sent to a proprietary AI service.
That matters for Reality Remix because the input is a photo of someone's physical surroundings. Keeping that image under the user's control is a much better fit for the idea.
It also makes experimentation easier. The model, prompts, and AI behavior can be inspected and changed, and different models can be tested through Ollama without redesigning the application.
A closed AI API could provide similar vision capabilities, but the open approach gives the project more control over how the AI runs, how the model is used, and where the user's data goes.
Most importantly, the open components aren't just an implementation detail. Local AI is what makes the core experience possible without depending on a proprietary AI service.
Prize Categories
Best Use of Gemma
Reality Remix uses Gemma 3 4B as the core vision model that analyzes the user's surroundings and turns what it sees into real-world challenges.
Gemma runs locally through Ollama, making the model an integral part of the application's core experience.
Top comments (0)