This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Developers often face errors that are much easier to show than explain.
You get a confusing terminal error, a stack trace, a cloud-console warning, or an IDE problem โ and then spend time copying text, explaining context, and figuring out what went wrong.
So I built Explain This Screenshot โ a privacy-first AI developer assistant that lets you simply take a screenshot and ask your AI friend to figure it out.
๐ธ Screenshot โ ๐ง Understand โ ๐ ๏ธ Fix
Upload a screenshot of a technical problem and the application analyzes it to provide:
- ๐ What I see โ understands the screenshot and identifies relevant information
- โ What's wrong โ identifies the likely problem
- ๐ง Why it happened โ investigates the possible root cause
- ๐ ๏ธ How to fix it โ provides practical solutions
- ๐ป Try this โ generates commands or code when appropriate
- ๐งโ๐ซ Beginner explanation โ explains the problem in simple language
- โ ๏ธ Warnings โ highlights potentially risky actions
- ๐ฏ Confidence โ indicates how confident the system is in its diagnosis
The project is built for developers, students, and anyone who has ever stared at an error message and thought: "What does this even mean?"
The "friend" I'm building for is essentially the developer who needs help debugging without having to perfectly explain the problem first.
Demo
๐ฅ Video Demo:
currently in development phrase
๐ Live Demo:
ShreyashChaugule-github
/
Plugin
This project is build for hacktoberfest 2026 Dev Challenge
Explain This Screenshot
Tagline
Understand the error. Fix the problem.
Explain This Screenshot is a privacy-first AI developer assistant that allows a user to upload a screenshot of a technical problem and receive a clear explanation and actionable solution. It turns that screenshot into an understandable diagnosis and practical fix.
Why Open Innovation?
Screenshots may contain source code, API keys, internal infrastructure, customer information, internal dashboards, logs, and private development environments.
Local inference provides:
- Privacy
- Offline capability
- Model choice
- Model swapping
- Experimentation
- No per-request cloud API cost
Features
- Fast Mode: Get a quick structured explanation of the error from a single vision model call.
- Deep Analysis: A multi-agent orchestrated pipeline to verify root causes, simplify explanations, and verify the generated solution.
- Privacy-first: No database, no user accounts, no cloud storage, and no permanent screenshot storage.
- Model Swappable: Bring your own open-weight vision model via Ollama.
- Offlineโฆ
The core demo flow is:
Upload Screenshot
โ
Fast Explanation
โ
Deep Analysis
โ
5 Specialized AI Agents
โ
Verified Solution
Code
๐ป GitHub Repository:
ShreyashChaugule-github
/
Plugin
This project is build for hacktoberfest 2026 Dev Challenge
Explain This Screenshot
Tagline
Understand the error. Fix the problem.
Explain This Screenshot is a privacy-first AI developer assistant that allows a user to upload a screenshot of a technical problem and receive a clear explanation and actionable solution. It turns that screenshot into an understandable diagnosis and practical fix.
Why Open Innovation?
Screenshots may contain source code, API keys, internal infrastructure, customer information, internal dashboards, logs, and private development environments.
Local inference provides:
- Privacy
- Offline capability
- Model choice
- Model swapping
- Experimentation
- No per-request cloud API cost
Features
- Fast Mode: Get a quick structured explanation of the error from a single vision model call.
- Deep Analysis: A multi-agent orchestrated pipeline to verify root causes, simplify explanations, and verify the generated solution.
- Privacy-first: No database, no user accounts, no cloud storage, and no permanent screenshot storage.
- Model Swappable: Bring your own open-weight vision model via Ollama.
- Offlineโฆ
The project is open source and includes the frontend, backend, AI provider integration, agent orchestration, prompts, tests, and documentation.
How I Built It
The most important design decision was to not build another application that simply sends everything to a closed AI API.
I wanted the AI to be:
- Open-weight
- Local
- Replaceable
- Privacy-conscious
- Useful without requiring an account
- Usable without a database
๐ง Open-Weight Vision AI
The project uses an open-weight vision-language model through Ollama.
The model can understand screenshots containing things like:
- Terminal errors
- Source code
- Stack traces
- IDE warnings
- Docker output
- Cloud console errors
- Configuration problems
- API responses
The model is configurable, so the application isn't permanently tied to one model.
๐ฅ๏ธ Local Inference
The architecture looks like this:
โโโโโโโโโโโโโโโโโโโโโโโโ
โ React Frontend โ
โ โ
โ Upload Screenshot โ
โโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โ
โโโโโโโโโโโโโโโโโโโโโโโโ
โ Node.js + Express โ
โ โ
โ Agent Orchestrator โ
โโโโโโโโโโโโฌโโโโโโโโโโโโ
โ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Ollama โ
โ โ
โ Open-weight Vision Model โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The core inference can therefore happen on the user's own machine.
๐ค The Multi-Agent Pipeline
For the Deep Analysis mode, I split the problem into five specialized AI agents.
Screenshot
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Screenshot Analyzer โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Error Investigator โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Solution Engineer โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Beginner Explainer โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโ
โ Solution Verifier โ
โโโโโโโโโโโโฌโโโโโโโโโโโ
โ
โผ
Verified Fix
1. Screenshot Analyzer
First, the system determines what is actually visible in the screenshot.
It identifies things such as the error, programming language, code, terminal output, and relevant context.
2. Error Investigator
The next agent investigates the likely root cause.
It separates what was directly observed from what is inferred or uncertain.
3. Solution Engineer
This agent generates practical fixes, commands, code, and alternatives.
The application never automatically executes AI-generated commands.
4. Beginner Explainer
Technical debugging can be intimidating, especially for students and newer developers.
This agent converts the diagnosis into a simple explanation of what happened and why.
5. Solution Verifier
Finally, another agent reviews the proposed solution.
It checks whether the solution actually addresses the observed problem, identifies unsupported assumptions, and flags potentially dangerous actions.
This gives the final response an additional verification step rather than blindly displaying the first AI-generated answer.
โก Fast Mode vs ๐ง Deep Analysis
I also wanted the application to be useful for both quick questions and deeper debugging.
Fast Mode
Screenshot
โ
Vision Model
โ
Quick Explanation
Useful when you just want to understand an error quickly.
Deep Analysis
Screenshot
โ
5 Specialized Agents
โ
Verified Solution
The UI shows the progress of each stage:
โ Screenshot analyzed
โ Root cause investigated
โ Solution generated
โ Explanation simplified
โ Solution verified
๐ Privacy by Design
Screenshots aren't always harmless.
A developer's screenshot might contain:
- Source code
- API keys
- Internal URLs
- Customer information
- Cloud infrastructure
- Logs
- Configuration
- Private development environments
That's why I designed the project around local inference.
There is:
- โ No authentication
- โ No database
- โ No permanent screenshot storage
- โ No required closed AI API
- โ Local AI inference
- โ Configurable open-weight model
- โ Session-only follow-up context
The goal isn't to claim that local AI makes data automatically "100% secure." Instead, it gives developers the option to keep their screenshots and inference within their own environment.
๐ก๏ธ Treating Screenshots as Untrusted Input
Another important part of the implementation is AI safety.
A screenshot might contain text that looks like an instruction or command.
The system explicitly treats everything inside the screenshot as untrusted data to analyze, not instructions for the AI to follow.
The application also does not:
- Automatically execute commands
- Automatically modify the user's system
- Execute code from screenshots
- Treat instructions inside screenshots as system instructions
This was especially important because the project is designed to analyze developer environments where screenshots can contain commands and potentially sensitive information.
๐งฐ Tech Stack
Frontend
- React
- TypeScript
- Vite
- Tailwind CSS
- shadcn/ui
- Lucide React
Backend
- Node.js
- TypeScript
- Express
AI
- Ollama
- Open-weight Vision-Language Model
- Custom multi-agent orchestration
Architecture
The five agents are implemented as specialized modules inside the Node.js application rather than separate microservices.
This keeps the project lightweight and easy to run locally.
Why Does Open Innovation Matter?
This is probably the most important part of the project for me.
A closed AI API could certainly analyze a screenshot.
But using open-weight AI and local inference makes a different architecture possible.
Instead of:
Screenshot
โ
Your Application
โ
Closed AI API
โ
External Cloud
the project can work like:
Screenshot
โ
Your Application
โ
Ollama
โ
Open-weight Model
โ
Your Computer
That changes what developers can experiment with.
๐ Model Freedom
The application can be configured to use different compatible models.
The AI provider isn't deeply embedded throughout the application.
๐ Local Processing
Developers can run inference locally instead of automatically sending screenshots to an external AI provider.
This is particularly useful for screenshots containing source code, logs, infrastructure information, or other sensitive development context.
๐งช Experimentation
Because the model and agent prompts are accessible, developers can modify the system itself.
Want to change how root-cause analysis works?
Modify the investigator.
Want a different explanation style?
Modify the beginner explainer.
Want another verification step?
Add another agent.
๐ฐ Different Cost Model
There is no per-request charge from a hosted AI API for local inference.
There are still hardware, electricity, and model-running costs, of course.
๐ Offline Potential
Once the application dependencies and model have been downloaded, the core analysis workflow can operate without an internet connection.
๐ก What Open AI Made Possible Here
The biggest thing open innovation enabled wasn't simply "using a free model."
It allowed me to make the AI layer part of the application architecture.
The model can be swapped.
The prompts can be inspected.
The agents can be modified.
The inference can run locally.
The application doesn't need a database or user account.
And developers can take the project, change it, and build something completely different from it.
That's the part of open AI that I wanted to explore with this project.
My Agent Session
Currently used my local model this for short time but will be working on Gemma 4 model and build for friend
Prize Categories
Best Use of ElevenLabs
Best Use of Tinker
Best Use of Backboard
Best Use of Render
๐ What's Next?
There are several directions I'd like to explore:
- Support more open-weight vision models
- Add more specialized debugging agents
- Improve structured output reliability
- Add more language/framework-specific diagnosis
- Improve local model performance
- Add optional screenshot redaction
- Add richer developer-tool integrations
- Support more complex multi-screenshot debugging workflows
โค๏ธ Final Thought
Developers don't always need another chatbot.
Sometimes they just need to show someone the problem and hear:
"I see what's happening. Here's why. Here's what you can try."
That's what I wanted Explain This Screenshot to be โ a small, local, open AI-powered debugging friend.
๐ธ Show it the problem.
๐ง Let it understand.
๐ ๏ธ Get a fix.
Built for the developers who have ever taken a screenshot and said: "Can someone tell me what's wrong here?"
Top comments (0)