This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
The Last Manual
The Last Manual is an AI-powered visual troubleshooting assistant built for people who are stuck with a physical device, machine, appliance, or object and don't know how to fix it.
The idea came from a simple problem: when something breaks, people often don't know what the part is called, don't know which manual to search for, and don't know whether the next step they're about to take is safe.
Instead of forcing someone to search through complicated manuals or watch multiple videos, The Last Manual lets them take a picture of the problem.
The system analyzes the image using Gemma, identifies what it can see, reasons through the possible problem, and generates a structured troubleshooting manual with:
- What was detected
- What might be wrong
- Safety warnings
- Tools/materials that may be needed
- Step-by-step instructions
- When to stop and seek professional help
The instructions can then be converted into natural spoken guidance using ElevenLabs, so the user can keep their hands on the object instead of constantly looking at a screen.
I built it for the friend who always says:
"Bro... this thing isn't working. What do I do?"
The goal is to turn that moment from "I have no idea" into "Okay, let's fix this step by step."
Demo
Live Demo:
https://the-last-manual-frontend.onrender.com/
Backend API:
https://the-last-manual-backend.onrender.com/
GitHub:
https://github.com/vanihba0-lgtm/The-last-manual
Code
The complete source code is available on GitHub:
https://github.com/vanihba0-lgtm/The-last-manual
The project contains:
- Next.js frontend
- FastAPI backend
- Gemma vision/troubleshooting pipeline
- ElevenLabs voice output
- Image capture/upload
- Safety validation layer
- Structured AI responses
- Session-based troubleshooting flow
How I Built It
The core of The Last Manual is built around Gemma, an open-weight AI model.
Instead of using a closed vision API as the main reasoning engine, I built the troubleshooting pipeline around Gemma's ability to understand visual input and generate structured reasoning.
The basic flow is:
Camera → Image Capture → Gemma → Troubleshooting Analysis → Safety Validation → Manual → ElevenLabs Voice
The frontend is built with Next.js and the backend uses FastAPI/Python.
The backend sends the captured image to Gemma through an OpenAI-compatible interface and asks it to behave like a real troubleshooting expert rather than simply describing the image.
The response is structured into practical sections so that the user gets an actual manual rather than a generic chatbot response.
I also added a safety layer so potentially dangerous situations can be flagged before presenting instructions.
For accessibility and hands-free use, the generated instructions can be converted into spoken guidance using ElevenLabs.
Why Does Open Innovation Matter?
Open innovation matters because a troubleshooting assistant should not have to depend entirely on a black-box system that users cannot understand, modify, or control.
Using an open-weight model like Gemma gave me the ability to build the AI experience around the model instead of simply wrapping a closed chatbot API.
It also makes the project more adaptable.
The troubleshooting system can potentially be:
- customized for specific types of equipment
- fine-tuned for specialized domains
- run in more private environments
- adapted for offline or low-connectivity scenarios
- inspected and improved by other developers
Most importantly, the open model is not just an extra feature in this project.
It is part of what makes The Last Manual possible.
I wanted to build something where AI isn't just answering questions — it is helping someone understand the physical world in front of them.
My Agent Session
Prize Categories
- Build for a Friend
- Open-source AI / Gemma
Top comments (1)
it was really fun joining in hacktober fest 2026 and i am sure i'll learn many thing in th whole hacktober feast!