This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My parents often use household devices like the AC, microwave, TV, and dishwasher, but they sometimes forget what particular buttons do.
Usually, they would have to ask me:
βWhat does this button do?β
And because these devices have so many buttons and modes, remembering them isn't always easy.
So I built Explain This Thing β a small personalized AI assistant that lets them upload a photo of a remote or control panel or used a saved photo, tap the button they're confused about, and get a simple explanation in English or Hindi.
The goal wasn't to build an AI that understands every appliance in the world.
I wanted to build something that understands the few devices my parents actually use.
The app currently focuses on four devices:
- βοΈ Bedroom AC Remote
- π½οΈ Dishwasher
- π² Panasonic Microwave
- πΊ Samsung TV Remote
They can select their device, use a saved photo or upload a new one, tap a button, and get an explanation of what it does and what they should press.
The part that mattered most
I gave the app to my parents and asked them to try it themselves.
Their feedback was simple:
It helped them understand the buttons without needing my or anyone else's help.
They also told me that they often forget what certain buttons do, so being able to quickly look it up themselves was useful.
That was exactly the problem I wanted to solve.
The project isn't about replacing a manual. It's about giving someone a quick way to regain that little bit of independence when they forget something.
Demo
π₯ Demo video: https://drive.google.com/file/d/1Fneh7OW8dZGmtfM7jQ_0ph0bTnEW-gEt/view?usp=sharing
The demo shows:
- Selecting one of the saved devices
- Using a saved device photo
- Tapping a button
- Getting a simple explanation
- Switching between English and Hindi
- Using device-specific knowledge
Code
GitHub: https://github.com/muskan-110/Explain_This_Thing_HacktoberFest26
The project is structured as a React frontend and FastAPI backend, with the device-specific knowledge stored separately for each appliance.
How I Built It
The frontend is built with:
- React
- Vite
- TypeScript
The backend uses:
- Python
- FastAPI
- RapidOCR
The AI runs locally using:
- Ollama
- Gemma 3 4B for vision and explanation
- nomic-embed-text for embeddings
How a request works
The basic pipeline looks like this:
Photo of device
β
βΌ
User taps a button
β
βΌ
Crop button area
β
βΌ
OCR
β
βΌ
Identify button label
β
βΌ
Search device knowledge
β
βΌ
Gemma 3 4B
β
βΌ
Simple explanation
β
ββββββββ΄βββββββ
βΌ βΌ
English ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯
If OCR can't read a label, for example on an icon-only button, Gemma's vision takes over and reads the label or icon instead.
The app also keeps personalized knowledge for each device.
For example:
data/
βββ appliances/
βββ ac_remote/
βββ dishwasher/
βββ microwave/
βββ tv_remote/
Each device has its own notes and information about its buttons.
This means the AI doesn't have to rely entirely on general knowledge. It can retrieve information specifically written for that particular device.
Why Does Open Innovation Matter?
I deliberately built this around an open-weight model running locally rather than sending my parents' device photos to a closed AI API.
The core AI runs through:
Ollama β Gemma 3 4B
This matters for this project because the input can contain photos of things inside someone's home.
With local inference:
Photo
β
My computer
β
Ollama
β
Gemma 3 4B
β
Answer
instead of:
Photo
β
External AI API
β
Remote server
β
Answer
That gives me more control over where the data is processed.
It also means I don't need to pay for an API request every time my parents want to look up a button.
Most importantly, using an open-weight model means I can experiment with the model, change the surrounding pipeline, replace the model, and build the experience around my users rather than being locked into one closed API.
For this particular project, local AI isn't just a technical choice. It fits the problem.
Final note
I started this project thinking about AI models and computer vision.
I ended it thinking more about independence.
Sometimes the useful thing you can build with AI isn't something spectacular.
Sometimes it's just something that means your parents don't have to call you the next time they forget what a button does.
And that's exactly why I built Explain This Thing. β€οΈ
Prize Categories
Best Use of Gemma: the core of the project is Gemma 3 4B running locally through Ollama. It reads icon-only buttons when OCR can't read a label, writes the explanations in simple English and Hindi, and returns structured output the app can display reliably.
Top comments (0)