DEV Community

Cover image for Screen Memory | Search Engine For My Screenshots
Mir Shah
Mir Shah Subscriber

Posted on

Screen Memory | Search Engine For My Screenshots

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Screen Memory is a Windows app that finds any screenshot when you describe it in plain words.
No file names. No endless scrolling. No cloud.

I had 560+ screenshots named like Screenshot 2026-07-18 204819.
I knew I had taken the one I needed. I just could not find it.
So I built something that remembers what is inside every screenshot.

It all runs on my laptop, which has no dedicated GPU. Your screenshots never leave your machine.

Who it is for

It is for me, and for my friend, who also takes screenshots of everything and then can never find the one they need. We both had the same problem: a folder full of screenshots with names like Screenshot 2026-07-18 204819, and no way to search what is inside them.

How it works

Step 1 Take a screenshot like you always do.
Step 2 A small popup asks: "Describe it so you can search it later?"
Step 3 Gemma 4 looks at the image and writes a title, a description, the text it can see, and tags.
Step 4 Ask a question later, like "That ad about connecting to a Raspberry Pi remotely?"
Step 5 An AI agent goes looking. It searches, reads the best match, and shows you the screenshot with a reason why.

A real search from my own folder

I asked that ad about connecting to a Raspberry Pi remotely
It found A sponsored HackerNoon post about Tailscale, inside a screenshot of my X feed.

The first search found nothing, and the agent fixed it on its own. This is what it did:

The saved description says "sponsored post", not "ad", so the first search missed.
The agent tried again with different words, found the screenshot, and read its details to be sure.

I never typed a file name. I described what I remembered, and the answer came from text Gemma had written when I took the screenshot.

Demo

Code

GitHub logo abbasmir12 / screenmemory

Search Engine For Your Screenshots

Screen Memory

Find any screenshot by describing it in plain words. Gemma 4 reads each screenshot once and saves what it sees as text. Later, a small search agent looks through that text and brings the screenshot back.

Demo video: https://youtu.be/oH65CRKHw5Q

I had 560 screenshots named like Screenshot 2026-07-18 204819 I knew I had taken the one I needed. I just could not find it.

What it does

Type a question like "that ad about connecting to a Raspberry Pi remotely" and get the screenshot back, with a reason why it matched. The AI does not scan your folder. It uses tools to search, read, and double-check, then answers.

When you take a screenshot, a small popup asks "Describe it so you can search it later?". If you ignore it, nothing is sent anywhere. With a local model, your screenshots, your database, and your searches stay on your PC.

…

Run it yourself

Step 1 Install Python, then run pip install pillow watchdog.
Step 2 Install Ollama and pull a Gemma 4 model.
Step 3 Copy .env.example to .env, then run python main.py check to test your setup.
Step 4 Run python main.py, click Index existing screenshots, then search.

How I Built It

The stack

Model Gemma 4 (open weights), running locally with Ollama.
Agent Tool calling over an OpenAI-compatible API.
Storage SQLite, saving the real file path of every screenshot.
App Python and tkinter.
New screenshots The watchdog library watches the Screenshots folder.

I built it over one weekend, with an AI assistant helping me write and test the code.

An agent that uses tools

The model never sees my whole database.
It gets four tools and decides what to call, like a person searching.

search_screenshots Searches descriptions, visible text, and tags.
list_recent Lists the newest screenshots in a date range.
get_screenshot_details Reads the full text of one screenshot.
look_at_screenshot Opens the real image and asks a question about it.

Here is what the agent did when I asked for the Ollama command:

search_screenshots(keywords=["ollama", "windows", "download", "command"])
   -> 2 result(s)
get_screenshot_details(id="963f6cf1")
Enter fullscreen mode Exit fullscreen mode

It searched, found two candidates, read the right one, and gave me the exact command. You can watch these steps live in the app.

Ask first, always

Screenshots can hold private things like passwords and chats, so the app never describes one on its own.
A small popup appears in the corner with Describe, Skip, or Always describe.
If you ignore it, nothing happens. When the work is done, a second popup confirms it.

How It Works Behind the Scenes

When you search, the AI does not look through your screenshots. It looks at each one once, when you take it. After that, it only searches text.

Remember once

When you say yes to the popup, Gemma 4 looks at the screenshot and writes down what it sees: a title, a short description, the important text on screen, and a few tags. This is the heavy AI work, and it happens only one time per screenshot.

A simple database

All of that is saved in SQLite, which is just one file on my computer. There is no server to install or run. Each screenshot is one row with its text and its path, which is where the real image lives on disk.

The image itself is never stored in the database. Only its path is.

SQLite also builds a word index over that text, like the index at the back of a book. It already knows which screenshots contain each word, so it never has to read them all.

Ask in plain words

The agent is not staring at your folder. It is Gemma 4 with four small tools, like a librarian who checks the card catalogue instead of reading every book. It searches the text with search_screenshots, reads one result in full with get_screenshot_details, and gives you an answer. If the stored text is not enough, it can open one real image with look_at_screenshot.

It answers with short ids, and the app looks up each path in the database. So the agent can only point to screenshots that really exist.

Why it feels fast

The slow part, understanding the image, is done in advance. A search only sends a few short pieces of text to the model, never images. Looking up words in SQLite is almost instant.

Understand once. Search many times.

Why Does Open Innovation Matter?

A screenshot folder is some of the most private data on a computer.
With an open model, I did not have to send it to anyone.

What open models made possible

Private by default The model, the database, and the search all run on my own machine.
Free to run There is no subscription and no per-image cost.
Swap the model in three lines I tested with a hosted provider first, then switched to local Ollama without changing any code.
No lock-in If a model changes or disappears, I change one line. The prompts are plain text and the tools are small Python functions, so I can change how the agent behaves.

These are the three lines that control everything:

LLM_BASE_URL=http://localhost:11434/v1
LLM_API_KEY=ollama
LLM_MODEL=gemma4:e4b
Enter fullscreen mode Exit fullscreen mode

What I would do next

Search by meaning So "cheap flights" finds "budget airfare".
Start with Windows So it is always watching, with nothing to launch.
One-click install Package it as a single .exe so friends can use it without Python.

Prize Categories

Best Use of Gemma: Gemma 4 is the core of the project. It runs locally with Ollama, reads every screenshot with vision, and powers the search agent with tool calling.

Top comments (1)

Collapse
 
mirshah12 profile image
Mir Shah •

i built something meaningful in this short time, hope you'll give it a try! šŸ’–