This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Thread is a local-first life archive built from recordings.
I built it for a close friend who wants to preserve memories from different periods of her life as material for an autobiography.
That gave me an interesting problem to solve. Recording those memories is easy. But as the archive grows, finding a particular story again β who was there, where it happened, when it happened, what other memories connect to it β becomes much harder.
AI seems like an obvious way to organize that material. But it creates another problem: if we use AI to preserve someone's memories, at what point does the AI's interpretation start replacing the memory itself?
That became the central constraint behind Thread:
AI can organize a memory, but it should never become the memory.
Thread takes recorded or imported conversations and turns them into a navigable archive of one person's life: stories, people, places, dates, and the connections between them.

Each mark is a story, drawn as the shape of the voice and placed in the year it happened. The years with no recordings stay visible.
But the original recording always remains the source of truth.
A recording can hold no story at all, and it is still kept exactly as it was received.
Thread can identify that a story happened in 2007, for example, but that claim does not have to be trusted just because a model generated it. The interface shows where the information came from, and you can jump back to the exact moment in the recording and hear the person saying it.

While a memory plays, each name lights up at the second it is said, and threads reach every other moment where that person appears.
I think of the archive as four different layers:
- Recording β the preserved evidence.
- Story β an interpretation of something told in that recording.
- People and places β entities connected back to supporting evidence.
- Life β a representation built from those stories, never a replacement for them.

The recording is the evidence. The stories are spans laid over it, never a replacement for it.
The result is not an AI-generated autobiography.
It is a tool that helps my friend organize and revisit her own memories while keeping her words β not the model's reconstruction of them β at the center.
AI organizes the memory. The voice remains the evidence.
Demo
π₯ Video demo: https://www.youtube.com/watch?v=Q1Uah6K4Xck
π Read-only live demo: https://thread.marianacastro.dev/
The video shows the complete flow using a real recording: importing audio, processing it locally, discovering stories, inspecting provenance, jumping from an extracted fact back to the exact moment that supports it, and finally seeing the new memories appear in the person's Life view.
The hosted demo is intentionally different from a normal local installation.
A local Thread installation can import and process recordings using the local AI pipeline. The public deployment is a read-only snapshot of a preprocessed archive, so visitors can explore the product without the server needing access to the local models or accepting personal recordings.
Code
π» Source: https://github.com/maricastroc/thread
Thread is open source, and the repository includes the application, setup instructions, demo archive, fixtures, and documentation for running the full pipeline locally.
How I Built It
Thread is a Next.js and TypeScript application backed by SQLite, but the interesting part for me was deciding exactly how much authority to give the AI.
The processing pipeline looks roughly like this:
record
β
preserve original audio
β
transcribe + timestamp
β
discover stories
β
extract people, places and dates
β
verify provenance
β
index
β
explore
β
return to the original voice
Listening: Whisper + Silero VAD
Recordings are transcribed locally with whisper.cpp using Whisper large-v3-turbo.
Word-level timing matters because Thread needs more than a transcript: it needs to be able to take someone from an interpreted fact back to the moment in the recording that supports it.
I also use Silero VAD to detect speech and avoid treating long periods of silence as content.
Understanding: Gemma
The interpretation layer uses Gemma 4 E4B through Ollama.
Gemma receives the transcript and performs a deliberately constrained job:
- find individual stories;
- suggest titles;
- identify people, places, and dates;
- point to the transcript segments that support those claims.
It produces structured output rather than becoming a general-purpose chatbot inside the archive.
And importantly, the model does not get the final word.
After Gemma proposes an annotation, Thread verifies its provenance deterministically against the transcript before storing it. If the cited words cannot be found, the claim is dropped.
The interface then distinguishes different levels of provenance, including:
- said β explicitly present in the recording;
- from the words β derived directly from what was said;
- [inferred] β interpretation rather than an explicit statement.

Every fact points back to its words and its second in the recording: "1978" comes from "78", said at 0:04.
That distinction is important because a plausible hallucination is especially dangerous in an archive of someone's life.
Finding: EmbeddingGemma + SQLite
For retrieval, Thread combines EmbeddingGemma, also running through Ollama, with SQLite FTS5.
But search follows the same rule as the rest of the project.
It does not generate an authoritative answer about the person's life.
It finds the relevant part of the archive and takes you back to the evidence.

Search never answers with generated text: it finds the moment and plays the voice from there.
Keeping interpretation replaceable
The AI layer sits behind a StoryInterpreter interface.
That means the recordings and transcripts are not tied permanently to one model. The interpretation layer can be changed and the archive reprocessed while the underlying evidence remains intact.
That separation became one of the most important architectural decisions in the project:
the model is replaceable; the memory is not.
Everything required for the normal processing pipeline runs locally:
- Next.js
- TypeScript
- SQLite
- ffmpeg
- whisper.cpp
- Silero VAD
- Gemma through Ollama
- EmbeddingGemma through Ollama
No recording needs to leave the machine for Thread to build the archive.
Why Does Open Innovation Matter?
For Thread, using open models was not just a technology preference. It changed what I could make the product promise.
These recordings can contain family stories, names, relationships, places, personal events, and voices. They are exactly the kind of data I do not want the architecture to assume can be sent to an external service.
Because the models can run locally, Thread can process those recordings without requiring them to leave the person's computer.
But openness also matters for a second reason: interpretation changes.
A story extracted by today's model should not become permanently fused with the archive simply because that was the model available when the recording was imported.
Thread preserves the recording and transcript separately from the interpretation layer. A different open model can be plugged in later, and the archive can be reprocessed without changing its underlying evidence.
That makes it possible to treat AI output as what it actually is: an interpretation that can be inspected, challenged, improved, or replaced.
A closed API could perform many of the individual tasks in this pipeline. But building around open models made it possible to combine local processing, replaceable interpretation, inspectable provenance, and long-term ownership of the archive as properties of the system rather than promises made by an external provider.
Open innovation also made experimentation practical. I could constrain the model, test how its interpretations behaved across different recordings, create fixtures in Portuguese, English, and Spanish, and build a larger synthetic archive containing dozens of stories to see whether the interface and retrieval model still made sense beyond the original demo.
For a project about preserving someone's memories, I think that distinction matters.
The intelligence helping organize the archive can change.
The person's voice should remain.
Prize Categories
- Best Use of Gemma β Gemma is the interpretation layer at the core of Thread: it discovers stories and extracts structured people, places, and dates from transcripts, while Thread independently verifies the provenance of its claims before they enter the archive.
Top comments (0)