This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
I built Family Tech Support Assistant for my parent, to make everyday phone and computer troubleshooting easier to follow.
The problem I wanted to address is the gap between technical instructions and what someone can comfortably do on their screen. An unfamiliar setting or error message can be difficult to describe, and a long list of instructions can make the situation more confusing.
The application accepts a plain-language question and an optional screenshot, then provides one troubleshooting step at a time. Its AI integration uses Gemma 3 4B running locally through Ollama.
Before giving the first instruction, it shows its understanding of the problem. The user can confirm that understanding or correct it before continuing.
The interface supports English, Hindi and Gujarati. Selecting a language changes the application’s labels, buttons and guidance, and instructs Gemma to respond in that language. Translated interface text and model-language accuracy are evaluated separately.
The application also includes:
- Screenshot editing: crop the relevant area, zoom for inspection and cover private details with opaque black redaction.
- Explicit step outcomes: record whether an instruction worked, failed or needs a clearer explanation.
- Failure notes: describe what happened so the next response has more context.
- Optional local history: save conversation text and outcomes in SQLite, then search or reopen previous solutions.
- Deletion controls: remove saved conversations from the application.
- Error recovery: preserve the conversation and recorded feedback when a model request fails.
The current version runs on a Windows PC. For a phone problem, the user can transfer a screenshot to the PC and follow the suggested instructions on their phone. Direct access from a phone browser is a future improvement.
Demo
Watch the two-minute feature walkthrough
The silent walkthrough shows:
- Connection checking and English, Hindi and Gujarati interfaces.
- Device selection and describing a problem.
- Screenshot upload, crop, zoom, redaction, undo and attachment controls.
- Reviewing and correcting the assistant’s understanding.
- One-step guidance and requests for simpler explanations.
- Failed-step notes, error handling, retry and success confirmation.
- Searching saved solutions, persistence across a server restart, deletion and temporary chats.
- Image validation, the ten-reply limit, responsive layout and the application architecture.
The walkthrough uses scripted test-model responses. It demonstrates the actual interface and backend workflow, including screenshot editing and SQLite storage; it does not demonstrate real Gemma inference or establish model accuracy.
Parent usability feedback and measured real-model performance have not yet been documented.
Code
View the source repository on GitHub
The project includes application source, Windows startup instructions, automated tests and design documentation.
A separate 34-page technical guide explains the system architecture, all 14 UML diagram types, class responsibilities, API flow, SQLite storage, testing and privacy decisions.
The application code is MIT licensed. Gemma model weights are downloaded separately and remain subject to Gemma’s license terms.
How I Built It
I kept the stack small so the project would be approachable to understand and maintain.
| Component | Technology | Responsibility |
|---|---|---|
| Interface | HTML, CSS and JavaScript | Forms, screenshot editing, language selection and history |
| Backend | Python standard-library HTTP server | Local API and frontend delivery |
| AI runtime | Ollama | Local model execution |
| Model | Gemma 3 4B | Problem and screenshot understanding, troubleshooting guidance |
| Image processing | Pillow | Image validation, resizing and metadata removal |
| Persistent storage | SQLite | Optional saved conversation text and outcomes |
| Backend testing | Python unittest | Workflow, validation and failure handling |
| Windows startup | Batch script and virtual environment | Dependency installation and application launch |
Initial setup requires internet access to download dependencies and the model. Normal inference uses the local Ollama service without a paid model API key.
A class-based, layered design
I used a layered monolith with focused classes:
- ApiController handles API input and coordinates requests.
- SupportService manages the troubleshooting workflow.
- OllamaClient handles communication with the model runtime.
- ImageProcessor validates and rebuilds screenshot data.
- SessionRepository manages active conversations in memory.
- HistoryRepository stores selected conversation text in SQLite.
The frontend separates language handling, screenshot editing, answer display and saved-history interactions into their own classes.
The service layer holds workflow decisions, the repository pattern isolates storage, and the adapter pattern isolates communication with Ollama. Constructor dependency injection allows automated tests to replace the real model client with a predictable fake.
These boundaries let me change and test individual parts while keeping deployment to one local application.
Review before instructions
Screenshot interpretation can be wrong. Showing the assistant’s understanding gives the user a chance to correct an assumption before acting on its advice.
The review includes the problem, the available evidence and what remains uncertain. The approved understanding then guides the first troubleshooting response.
Record failures explicitly
“Failed” and “unclear” mean different things. The application records these outcomes separately, along with the user’s notes.
Feedback is committed before requesting another model response. If that request fails, the recorded attempt remains available for retry.
The backend rejects an exactly repeated failed action. Recognising the same action expressed in different words remains a limitation.
Make history optional and visible
Active sessions live in memory. When saving is enabled, conversation text and outcomes are archived in a local SQLite database.
Saved history survives application restarts and can be searched or deleted. Reopened archives are read-only records; they do not automatically resume a live troubleshooting session.
Screenshots are excluded from the archive. The database is not encrypted by the application, and deletion does not promise forensic erasure.
Edit screenshots before submission
Screenshot editing happens in the browser before submission. Redaction changes the image pixels instead of adding a removable overlay.
The backend validates and rebuilds accepted images, removing metadata. The application binds to localhost, and model requests use the configured local Ollama service.
Users still need to identify and cover the private areas before submitting an image.
Testing and current limitations
The recorded development verification includes:
- 55 backend tests, covering persistence, opt-out, deletion, failed-step records, Unicode search and other application behaviour.
- Deterministic browser workflow tests.
- Screenshot crop and redaction pixel checks.
- Selected automated accessibility scans.
Backend verification used Python 3.12 and two supported Pillow versions. These checks verify application behaviour; they do not establish real Gemma accuracy or performance on the target Windows PC.
The target machine has 16 GB RAM and an NVIDIA RTX 3050 with 4 GB VRAM. Response latency and resource usage have not yet been documented, and I do not assume the model fits entirely in GPU memory.
Real-model evaluation still needs to assess screenshot interpretation, Hindi and Gujarati response quality, repeated advice and the clarity of troubleshooting instructions. Testing with my parent is also needed to determine whether the interaction is understandable without assistance.
AI-assisted development
I used Codex to assist with implementation, debugging, tests and documentation.
That assistance made verification important. Predictable tests check application state and error handling, while real-model testing and a parent usability session are needed to assess usefulness.
Pillow and the bundled Noto fonts retain their respective licenses, with credits included in the project.
Why Does Open Innovation Matter?
The open components make the application adaptable to my family’s language and workflow.
Running Gemma locally allows screenshot analysis without sending each image to a hosted model API. I can inspect the prompts, change the interaction, examine the image-processing pipeline and replace the model adapter as the project develops.
That matters because screenshots can contain personal information. Local inference, editable screenshots and optional local history give the user practical control over how their information is handled.
Local operation also avoids needing a paid model API key for normal use. The trade-off is that the computer must have enough resources to run the model, and performance needs to be measured on that hardware.
The project is also a learning opportunity. Its architecture and technical guide explain how requests move through the application, how conversations change state and how failures are handled.
Gemma is an open-weight model with its own terms. The application’s MIT license is separate; the components do not all have identical licensing.
A key lesson from building this project is that troubleshooting needs more structure than a question-and-answer box. Users need to correct misunderstandings, explain failed attempts and recover earlier advice.
The next improvement I would prioritise is paired access from a phone browser while Gemma continues running on the Windows PC. That would reduce the need to transfer screenshots manually. Feedback from my parent will help determine whether this or another improvement should come first.
Prize Categories
Best Use of Gemma
Gemma 3 4B is the application’s central AI component. The Ollama integration uses it to interpret questions and screenshots, produce an understanding for user review, and generate one-step troubleshooting guidance in the selected language.
Top comments (1)
crazy brother