This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built Recall, a local-first personal memory system for a friend who has a habit of saving everything: screenshots, useful webpages, PDFs, notes, images, and random files that might be useful someday.
The problem was not storing those things. It was finding them again.
A normal keyword search works when you remember the exact words. But memory usually does not work that way. You remember “that screenshot about the GitHub DNS problem”, “the article about FAISS”, or “the image with the terminal error”. You remember the idea, not necessarily the exact filename or sentence.
Recall turns those scattered files into searchable memories.
You can ingest text, webpages, images, screenshots, PDFs, and other files, then search the collection using natural language. Images are processed with OCR and image captioning so that the information inside them can become searchable too.
For example, instead of remembering a filename, you can search:
“the screenshot where GitHub could not resolve the host”
and retrieve the relevant memory.
The important part is that this is local-first. The data stays on the machine, there are no accounts or cloud APIs required, and once a memory has been processed, searching does not need a generative AI model running in the background.
Demo
Code
The repository is open source and contains the backend, frontend, ingestion pipeline, search system, tests, and technical documentation.
How I Built It
The core idea was to use AI where it provides the most value, rather than making an LLM responsible for every query.
Recall processes content during ingestion:
content
↓
metadata + content extraction
↓
OCR / image captioning
↓
searchable text
↓
embeddings
↓
SQLite + FTS5 + FAISS
For image understanding, Recall uses BLIP-base for image captioning and RapidOCR for OCR. For semantic retrieval, the default embedding model is Google's EmbeddingGemma 300M, with automatic fallbacks to all-MiniLM-L6-v2 and finally deterministic hash vectors when model availability is limited.
The backend is built with FastAPI and Python. Memories are stored in SQLite, with SQLite FTS5 providing BM25 lexical search and FAISS providing vector retrieval.
The final search is hybrid:
BM25 search ─────┐
├── Reciprocal Rank Fusion ──→ ranking
Vector search ───┘
+ metadata
+ recency
+ exact-match signals
This solves a practical problem with relying on embeddings alone.
Semantic search is good at understanding that “DNS failure” and “could not resolve host” may refer to the same thing, while lexical search is much better at exact strings such as error messages, filenames, or identifiers.
AI therefore does the expensive work once at ingestion, and retrieval stays deterministic and fast.
The frontend is built with React and Vite, with the production build served by the FastAPI application.
I also designed the system around a few constraints that mattered for this project:
- The database is the source of truth; the FAISS index can always be rebuilt.
- Ingestion is idempotent, using content hashes to avoid processing the same file repeatedly.
- AI failures should not make the entire ingestion pipeline fail.
- Search should continue working even when the optional AI models are unavailable.
- The application should be usable on a normal laptop rather than requiring a dedicated GPU.
Why Does Open Innovation Matter?
For this project, open innovation is not just about avoiding an API bill.
Memory is personal.
A system that is indexing screenshots, documents, saved pages, and private notes should not require sending that collection to a remote service just to make it searchable.
Open models made it possible to build the important parts of Recall locally: image captioning, OCR, and semantic embeddings can all happen on the user's machine.
That changes the trade-off completely.
Instead of:
“Upload my personal data to a service so an AI can understand it.”
the model becomes:
“Let the AI understand my data locally, store the useful representations, and search them without an online AI dependency.”
It also makes the system easier to experiment with. Models can be swapped, compared, or removed without redesigning the entire application around one proprietary API.
For a personal memory system, that control is as important as the model quality itself.
Prize Categories
Google Gemma
Recall uses Google's EmbeddingGemma 300M as its primary semantic embedding model.
EmbeddingGemma converts memories into vector representations that allow Recall to retrieve content based on meaning rather than exact keywords. This is particularly important for personal memories because users often remember the context of something without remembering the exact words contained in it.
For example, a user can search for:
“the screenshot where GitHub couldn't resolve the host”
even if the original screenshot contains an error such as:
“fatal: unable to access... Could not resolve host: github.com”
EmbeddingGemma provides the semantic layer that makes this kind of retrieval possible.
The model runs locally, so the user's personal memories do not need to be sent to a remote embedding API. Recall stores the generated embeddings locally and uses FAISS for vector retrieval.
I chose EmbeddingGemma because it provides a strong semantic representation while remaining practical for a local-first application. It also supports longer input than the lightweight embedding models I initially considered, which is useful when memories contain documents or substantial extracted text.

Top comments (0)