OmniScan AR: Turning the Real World into an Interactive AI Learning Experience 🌍📱
This is my submission for the Hacktoberfest Open-Source AI Challenge — Week 1: Touch Grass.
What I Built
OmniScan AR (EvoSnap) is an AI-powered, camera-first application that encourages people to step away from their screens and explore the world around them.
Instead of searching the internet every time we encounter something interesting, what if we could simply point our camera at it and start learning?
A plastic bottle, a leaf, an electronic device, or an everyday object can become the starting point for an interactive learning experience.
With OmniScan AR, users can:
- 🔍 Explore the world around them: Identify everyday objects using AI-powered visual analysis.
- 🌱 Understand environmental impact: Learn about materials, recycling, sustainability, and the lifecycle of the things we use.
- ♻️ Discover the hidden story of objects: Explore how materials degrade over time and understand the environmental consequences of waste.
- 🔮 Visualize the future: Generate illustrative visualizations of how an object might look after 10, 25, 50, 100, or even 500 years of environmental exposure.
- 💬 Ask questions and learn: Interact with an AI knowledge companion to explore an object's properties, uses, and environmental significance.
- 📷 Experience live object scanning: Use a camera-based interface with bounding boxes and tracking to make exploration feel more interactive.
The idea behind OmniScan AR is simple: technology should encourage us to experience the world, not just consume content about it.
It is designed for students, curious explorers, educators, and anyone who wants to understand the everyday world through a more interactive and environmentally conscious lens.
Demo
🚀 GitHub Repository: https://github.com/UMANG-SH941/omniscan-ar
The repository contains the application source code, backend services, research artifacts, and material-classification model.
I'm sharing the repository as the primary project link. A deployed demo or walkthrough video is attached.
Code
Explore the complete project on GitHub:
The project is organized into a few major components:
- Frontend: React Native, Expo, TypeScript, and Expo Router.
- Backend: FastAPI and Python.
- Vision and language: A provider-agnostic layer for visual understanding and conversational explanations.
- Object detection: AI-generated bounding boxes combined with client-side tracking and smoothing.
- Material classification: A MobileNetV3-Small model trained using TrashNet and exported to ONNX.
- Future visualizations: An image-generation pipeline with configurable providers.
- Offline support: A local scan vault that keeps previously scanned objects accessible on supported native devices.
The architecture is modular, so individual AI components can be replaced or improved without rebuilding the entire application.
How I Built It
I wanted to build something that combines practical AI engineering with a tangible, real-world purpose.
1. Camera-first exploration
The application captures objects from the user's surroundings and sends images through the visual-analysis pipeline. The resulting information is presented through an interactive mobile interface.
For live detection, the backend currently uses a vision-language model to identify visible objects and estimate their bounding boxes. The frontend then matches detections between frames using Intersection over Union (IoU) and exponential smoothing to reduce visual flickering.
2. An open-source material classifier
One of the important components is a MobileNetV3-Small material-classification model, fine-tuned using the TrashNet dataset and exported to ONNX for local inference.
The checked-in evaluation reports 88.93% validation accuracy on the project's validation split. This is an initial result rather than a claim of equivalent accuracy on real-world camera images.
The classifier provides a complementary material prediction alongside the vision-language analysis, connecting practical machine learning with waste awareness.
3. A modular AI backend
The backend uses FastAPI and separates visual analysis, knowledge generation, detection, storage, and image generation into independent services.
The LLM integration supports configurable providers, including Gemini, OpenAI, and locally hosted Ollama models. Image generation also supports different providers, including self-hosted Stable Diffusion WebUI.
This makes it possible to experiment with different models and deployment strategies rather than tying the entire application to a single AI service.
An important distinction: the application is built around an extensible AI architecture, but not every current component runs locally or uses open weights. The material classifier is an open-source ML component, while the actual inference path depends on the configured provider.
4. Visualizing long-term environmental impact
The Future Vision feature combines material-aware prompts with image generation to create illustrative visualizations of environmental degradation over time.
These images are intended to communicate environmental possibilities, not scientifically validated predictions of an individual object's exact future.
That distinction matters: AI can make abstract environmental timescales easier to understand, but generated imagery should not be confused with scientific evidence.
Why Does Open Innovation Matter?
Open innovation makes projects like OmniScan AR possible for independent developers and student builders.
I didn't need to train a vision model from scratch or build every infrastructure component myself. I could combine existing open-source frameworks, a pretrained MobileNetV3 model, ONNX Runtime, and modular AI services to build a complete application around a new idea.
It also gives the project room to evolve.
- Reproducibility: Other developers can inspect the implementation and understand how the components fit together.
- Experimentation: Different open-weight vision-language models and local inference backends can be tested without redesigning the application.
- Accessibility: Open-source tools reduce the initial cost of experimentation and make advanced ML workflows more approachable.
- Community collaboration: Contributors can improve detection, optimize inference, expand environmental knowledge sources, or make the application more accessible.
- User control: Local inference and self-hosted services offer potential paths toward better privacy, lower recurring costs, and reduced dependence on external providers.
Open innovation is not just about using free tools. It is about making the underlying technology inspectable, adaptable, and improvable by a wider community.
My Agent Session
I used AI-assisted development to help explore implementation choices, structure services, and iterate on the application.
Agent session: I'll add a DevRelay session link here once I have a shareable session URL.
Prize Categories
I'll select the partner categories that match the actual technologies used in the submitted build and the challenge's eligibility requirements.
What's Next?
OmniScan AR is an ongoing project, and there is plenty of room to improve it.
Some of the next steps I'd like to explore include:
- Replacing remote live-frame detection with an efficient on-device object detector.
- Integrating more open-weight vision-language models through local inference.
- Improving material classification with real-world camera images and stronger evaluation.
- Grounding environmental estimates in reliable, cited datasets and sources.
- Making the application easier to use outdoors, with better offline workflows and privacy controls.
The larger vision is to make learning more experiential: walk outside, discover something ordinary, point your camera at it, and leave with a better understanding of the world.
Let's build AI that helps us touch grass — and understand the planet we're standing on. 🌍
Top comments (0)