This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
"Thaththa", in Sinhala language, is my father. In his cupboard there is a box of things he has kept for about 30 years: certificates,awards, medals, newspaper articles, old photographs and coins. Every one of them has a story and almost none of those stories are written down anywhere. Thaththa is the only one who can tell them.
So I built Thaththa's Memory Box, a personal AI agent that runs on a laptop at home. You photograph a keepsake and the agent fills in a catalogue card for it: what it is, every word written on it and careful guesses about when, where, who and why. Then it asks questions and Thaththa answers them. His answers become part of the record and the agent reads every later keepsake knowing what he said.
The moment the app exists for
One of the first things we photographed was an old photo of Thaththa, sitting at a desk in uniform. The agent described what it could see, noting the emblem visible on his shirt and that the photo appears to be an older,possibly faded print. Then, instead of inventing his life story, it asked:
Does this photo show Thaththa at a particular time in his life?
Is the emblem on his shirt related to his work or a group he belonged to?
Thaththa answered:"Yes. This photo was taken during my work at Samanthurei, Sri Lanka".
The details turned blue, with "from Thaththa" next to them. The agent's earlier guesses moved into the card's history. And every keepsake photographed after that one was read knowing what Thaththa had said.
The rule: it is not allowed to invent a memory
An AI that confidently writes "your 1998 school sports meet" when the medal says no such thing is not building a memory collection. It is writing fiction about someone's life and handing it to their grandchildren. So the whole design rests on one rule and the rule is enforced in code, not in the prompt.
Every detail on a card carries where it came from, shown in three inks:
- Blue italic: the family told it. Nothing the model does later can overwrite it.
-
Typewriter: it is written on the object. The model has to transcribe every word it can see, exactly. A date or place only counts as "written on it" if the code finds it in that transcription. If Gemma says "I read 1998" but the transcription of the medal is
INTER-SCHOOL ATHLETICS 400 M, then 1998 stays a guess. - Grey with a question mark: a guess. The story may only use it with "appears to" and the card turns the biggest gaps into questions for the family.
def grounded(value: str, visible_text: str) -> bool:
"""Is `value` actually written on the object? Years must all appear in the transcription;
anything else must appear in it as a whole phrase."""
The model proposes. Plain Python decides what counts as known. When the family corrects something, the old value is never deleted; it moves into the history with who changed it and when. And if the agent can't work out what someone meant, their words are still saved on the card exactly as they said them. Nothing a person tells it is thrown away.
Demo Video
Code
Thaththa's Memory Box
Your personal AI agent for family memories, running at home on Gemma 4.
▶ Demo video: https://youtu.be/TQnalzTmMiE
Thaththa has a box of things he has kept for about 30 years: certificates, awards, medals, newspaper articles old photographs and coins. Only he knows the stories behind them.
Thaththa's Memory Box is a personal AI agent that runs on a laptop at home. Photograph a keepsake, and the agent:
- reads it with Gemma 4: what it is, every word written on it, and careful guesses about when, where, who and why;
- asks the family the questions that would fill the biggest gaps;
- remembers what they answer, and reads every later keepsake knowing it.
Nothing leaves the house: no account, no API key, no cloud. The whole archive is one folder you can copy to a USB stick.
The rule it is built on:
…Python, FastAPI and htmx, with Gemma 4 served locally by Ollama and 28 tests. The interesting part is src/memorybox/memory.py: the provenance rules, in plain Python, that sit between the model and the card.
How I Built It
I built it over the challenge weekend with Claude as a pair programmer, then tested it on Thaththa's real memory box.
Gemma 4 does three jobs, all through Ollama on Azure Vivobook, 16GB RAM
- Looking (vision). The photo, plus everything the family has said so far, goes in. Out come the object's kind, an exact transcription of any writing, a guess with a stated reason for date, place, people and occasion and one to three questions for the family.
-
Writing (text). Two to four sentences for the card, given every detail labelled
[family],[read],[seen]or[guess], with rules about how each may be used. - Understanding (text). Thaththa's free-text answer becomes structured updates ("date → [year]") plus anything worth remembering beyond this object. Those become memory notes and every later photo is read against them.
Three things that mattered when running a small model locally:
-
Ollama's native API with a JSON Schema in
format. Ollama constrains decoding to the schema, so a small model's reply always parses. Asking for JSON in the prompt is fine for a big hosted model; on a laptop, enforcing it is the difference between working and not. -
Setting
num_ctxexplicitly. Ollama's default context window silently drops the start of a long prompt and the start is where the instructions are. As the memory notes grow, the agent passes only the notes that share words with the object being read. -
Thinking off (
think: false). On a laptop the reasoning is most of the wait. Each keepsake took about 3 minutes on my laptop.
One keepsake at a time. A background worker reads the box in order, so the family can add twenty photos and keep browsing while it works. You can add the back of a photo, or the next page of a letter and every side is
read together. After a restart it picks up where it left off.
Why open innovation matters for this
These are the most private things a family owns. Thaththa's certificates, his awards, photos of him at work.I would not upload his box to a company to be described and nobody should have to. Here the photos go from a phone to a laptop in the same house and the model that reads them is a file on that laptop. Turn off the Wi-Fi and everything still works.
A memory archive has to outlive the service it was made with. An app built on an API key stops working when the key, the price or the company changes. This archive is a folder: a readable JSON file and the photos.
"Take a copy" downloads it as one zip, with a memory book that opens in any browser. The model is open weights, so in ten years it can be run again, or swapped for a better one by editing one line.
I could put the rules around the model, not inside it. Because the whole stack is open, the "never invent a memory" rule is ordinary, testable code that sits between the model and the card. With a closed API I could only have asked nicely in the prompt.
It costs nothing per photo, so the whole box gets done. Not just the ten best things, but the coins and the newspaper cuttings too.
Where a closed model was better Hosted models are faster and generally better with faded handwriting and small print; Gemma took about 3 minutes per keepsake on my laptop and on Thaththa's uniform photo it asked what the emblem was instead of recognising it. I chose local anyway: his photos and certificates stay in the house and for a memory archive a good question beats a confident guess, because Thaththa's answer is the real record.
Handing it over
I sat down with Thaththa at home with the laptop and his box. The first thing he chose was an old photo of himself, sitting at a desk in uniform.
The agent described what it could see, the emblem on his shirt and the faded print and then, instead of guessing his story, it asked: "Does this photo show Thaththa at a particular time in his life?"
He answered it: yes and he told it when the photo was taken. His words appeared on the card in blue, marked "from Thaththa" and the agent's guesses moved into the history. Every keepsake after that was read knowing what he had said.
That was the moment I knew the idea worked. The agent didn't tell him about his own life. It asked and he was the one who remembered.
What didn't work: while testing, a new keepsake showed the picture from an earlier one. Gemma had read the right photo, but the browser was reusing a cached image because keepsake numbers repeat when you start a new box. I changed the caching so the browser always checks for the latest photo.
What's next: the box isn't finished and that's the point. There are still certificates, medals, newspaper cuttings and coins to go through, one evening at a time, with Thaththa answering the questions only he can. When we're done, "Take a copy" will turn it into a memory book the whole family can keep, on a USB stick, with no app and no internet needed.
Thirty years of keepsakes sat in a box because nobody had time to ask him about them. Now something does ask and it waits for his answer.
Prize Categories
-
Best Use of Gemma. Gemma 4 (E4B, with E2B as a fallback) runs locally through Ollama and does all three
jobs: it reads keepsake photos with vision, transcribing printed text and handwriting exactly; it writes
provenance-constrained stories; and it turns the family's free-text answers into structured updates. It uses
schema-constrained output with thinking off and switching to the 26B model is one line in
config/models.toml.



Top comments (2)
A great project. Thanks for making this memory box software
Thank you! ❤️ Hope this helps!