DEV Community

Aishani
Aishani

Posted on AI-assisted

TinkerSight: An AI Friend for When You're Stuck With a Device

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built TinkerSight, a vision-first AI technical guide for my friend who lives alone and sometimes struggles with unfamiliar appliances and everyday devices.

The problem is simple: when you're standing in front of an unfamiliar AC remote, washing machine, or device, the manual often isn't written for the one question you actually have.

TinkerSight lets someone show a device through a photo, explain what they want to do, and get simple, step-by-step guidance.

Its core idea is:

See → Ground → Simplify → Guide

It doesn't just try to recognize what's in the image. When possible, TinkerSight retrieves relevant official manufacturer documentation and uses it to ground the guidance. If the image is unclear or the situation could be dangerous, it asks for clarification or stops instead of confidently guessing.

I built TinkerSight because I wanted to create something that felt less like searching through a manual and more like having someone standing beside you when you need help.

Demo

Here is a short demo showing three situations:

  • Using Turbo Cool on a Blue Star AC remote, with the answer grounded in official manufacturer documentation.
  • Getting guidance for a washing-machine control.
  • A safety scenario involving a burning smell, where TinkerSight refuses to provide hazardous repair instructions and tells the user to stop.

Demo video:

Watch the TinkerSight Demo video:
https://1drv.ms/v/c/bbab0edd6634b417/IQDjln7aXVbrR5HvUco5p-YwAXIP2aJejHEnnWeCJaIsJGc?e=KX55aw

Code

The complete source code is available on GitHub:

https://github.com/aisha453/TinkerSight

The repository includes the React frontend, FastAPI backend, local AI integration, manufacturer-document grounding, and safety handling.

How I Built It

TinkerSight is built around Qwen3-VL 2B, an open-weight vision-language model running locally through Ollama.

The architecture is:

React + Vite → FastAPI → Ollama → Qwen3-VL 2B

The user uploads an image and describes what they want to do. The vision model first identifies the device, visible controls, brand information, and relevant context.

For supported manufacturers, TinkerSight then retrieves information from selected official manufacturer webpages and uses that information to ground the response.

I also added a deterministic safety layer for obvious hazardous situations such as exposed electrical wiring, gas leaks, smoke, sparks, burning smells, and internal appliance repairs.

One important design principle was: don't pretend to know something that isn't visible or verified.

For example, TinkerSight does not claim an exact appliance model unless it can actually read the model information.

Why Does Open Innovation Matter?

Open innovation made this project possible in a way that would have been difficult to achieve with a closed API alone.

Because TinkerSight uses an open-weight vision-language model through local inference, the AI can run on the user's own machine without sending every image to a proprietary vision API.

It also means the model layer can be replaced or improved without rebuilding the entire application.

More importantly, using open AI made it possible for me to experiment with the behavior of the system itself: how it interprets images, how it handles uncertainty, how manufacturer information is introduced, and when it should refuse to provide instructions.

For a tool dealing with someone's home, appliances, and potentially sensitive images, having the ability to run the AI locally is especially meaningful.

Prize Categories

No partner prize category claimed.

TinkerSight currently uses open-source/open-weight AI and local inference, but I did not add a partner technology solely to qualify for a prize category.

Top comments (0)