DEV Community

Muskan Khatoon
Muskan Khatoon

Posted on

Explain This Thing: An AI That Helps My Parents Understand Their Devices

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My parents often use household devices like the AC, microwave, TV, and dishwasher, but they sometimes forget what particular buttons do.

Usually, they would have to ask me:

β€œWhat does this button do?”

And because these devices have so many buttons and modes, remembering them isn't always easy.

So I built Explain This Thing β€” a small personalized AI assistant that lets them upload a photo of a remote or control panel or used a saved photo, tap the button they're confused about, and get a simple explanation in English or Hindi.

The goal wasn't to build an AI that understands every appliance in the world.

I wanted to build something that understands the few devices my parents actually use.

The app currently focuses on four devices:

  • ❄️ Bedroom AC Remote
  • 🍽️ Dishwasher
  • 🍲 Panasonic Microwave
  • πŸ“Ί Samsung TV Remote

They can select their device, use a saved photo or upload a new one, tap a button, and get an explanation of what it does and what they should press.

The part that mattered most

I gave the app to my parents and asked them to try it themselves.

Their feedback was simple:

It helped them understand the buttons without needing my or anyone else's help.

They also told me that they often forget what certain buttons do, so being able to quickly look it up themselves was useful.

That was exactly the problem I wanted to solve.

The project isn't about replacing a manual. It's about giving someone a quick way to regain that little bit of independence when they forget something.

Demo

πŸŽ₯ Demo video: https://drive.google.com/file/d/1Fneh7OW8dZGmtfM7jQ_0ph0bTnEW-gEt/view?usp=sharing

The demo shows:

  1. Selecting one of the saved devices
  2. Using a saved device photo
  3. Tapping a button
  4. Getting a simple explanation
  5. Switching between English and Hindi
  6. Using device-specific knowledge

Code

GitHub: https://github.com/muskan-110/Explain_This_Thing_HacktoberFest26

The project is structured as a React frontend and FastAPI backend, with the device-specific knowledge stored separately for each appliance.

How I Built It

The frontend is built with:

  • React
  • Vite
  • TypeScript

The backend uses:

  • Python
  • FastAPI
  • RapidOCR

The AI runs locally using:

  • Ollama
  • Gemma 3 4B for vision and explanation
  • nomic-embed-text for embeddings

How a request works

The basic pipeline looks like this:

          Photo of device
                 β”‚
                 β–Ό
        User taps a button
                 β”‚
                 β–Ό
          Crop button area
                 β”‚
                 β–Ό
              OCR
                 β”‚
                 β–Ό
       Identify button label
                 β”‚
                 β–Ό
       Search device knowledge
                 β”‚
                 β–Ό
           Gemma 3 4B
                 β”‚
                 β–Ό
       Simple explanation
                 β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
          β–Ό             β–Ό
       English        ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯€
Enter fullscreen mode Exit fullscreen mode

If OCR can't read a label, for example on an icon-only button, Gemma's vision takes over and reads the label or icon instead.

The app also keeps personalized knowledge for each device.

For example:

data/
└── appliances/
    β”œβ”€β”€ ac_remote/
    β”œβ”€β”€ dishwasher/
    β”œβ”€β”€ microwave/
    └── tv_remote/
Enter fullscreen mode Exit fullscreen mode

Each device has its own notes and information about its buttons.

This means the AI doesn't have to rely entirely on general knowledge. It can retrieve information specifically written for that particular device.

Why Does Open Innovation Matter?

I deliberately built this around an open-weight model running locally rather than sending my parents' device photos to a closed AI API.

The core AI runs through:

Ollama β†’ Gemma 3 4B
Enter fullscreen mode Exit fullscreen mode

This matters for this project because the input can contain photos of things inside someone's home.

With local inference:

Photo
  ↓
My computer
  ↓
Ollama
  ↓
Gemma 3 4B
  ↓
Answer
Enter fullscreen mode Exit fullscreen mode

instead of:

Photo
  ↓
External AI API
  ↓
Remote server
  ↓
Answer
Enter fullscreen mode Exit fullscreen mode

That gives me more control over where the data is processed.

It also means I don't need to pay for an API request every time my parents want to look up a button.

Most importantly, using an open-weight model means I can experiment with the model, change the surrounding pipeline, replace the model, and build the experience around my users rather than being locked into one closed API.

For this particular project, local AI isn't just a technical choice. It fits the problem.

Final note

I started this project thinking about AI models and computer vision.

I ended it thinking more about independence.

Sometimes the useful thing you can build with AI isn't something spectacular.

Sometimes it's just something that means your parents don't have to call you the next time they forget what a button does.

And that's exactly why I built Explain This Thing. ❀️

Prize Categories

Best Use of Gemma: the core of the project is Gemma 3 4B running locally through Ollama. It reads icon-only buttons when OCR can't read a label, writes the explanations in simple English and Hindi, and returns structured output the app can display reliably.

Top comments (0)