<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mariana Castro</title>
    <description>The latest articles on DEV Community by Mariana Castro (@maricastroc).</description>
    <link>https://dev.to/maricastroc</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158711%2Fdfa74b39-fde6-4229-9b3d-91b2762d3cfe.jpeg</url>
      <title>DEV Community: Mariana Castro</title>
      <link>https://dev.to/maricastroc</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/maricastroc"/>
    <language>en</language>
    <item>
      <title>Can You Find It? An open model hides one real thing in a place you know</title>
      <dc:creator>Mariana Castro</dc:creator>
      <pubDate>Thu, 08 Oct 2026 13:44:29 +0000</pubDate>
      <link>https://dev.to/maricastroc/can-you-find-it-an-open-model-hides-one-real-thing-in-a-place-you-know-3fg3</link>
      <guid>https://dev.to/maricastroc/can-you-find-it-an-open-model-hides-one-real-thing-in-a-place-you-know-3fg3</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;I go to the same café every week. Same counter, same stool, same order. Last night a 4-billion-parameter model running on my laptop looked at a photo of that counter, picked one thing in it, and gave me one line:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;I found something hiding in plain sight.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It took me two hints to find it: a small framed award on the tiled wall, the kind of thing you stop seeing the second time you walk in. My field note, typed while I was still a bit annoyed with myself: &lt;em&gt;"A café I go to every week. Finally! I've sat at that counter dozens of times and never noticed it!"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's the whole game.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can You Find It?&lt;/strong&gt; is I Spy, but the AI does the spying, and it only uses what is really in front of you.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You take a photo of where you are, or pick one of a place you pass every day: your street, the way to the bakery, the view from a window.&lt;/li&gt;
&lt;li&gt;Gemma 4, running on your own computer, secretly chooses &lt;strong&gt;one real object in that photo&lt;/strong&gt;, with a box around it.&lt;/li&gt;
&lt;li&gt;You get a single clue line. Not a description: an invitation to look at the place through a lens (&lt;em&gt;"It only does its job after dark."&lt;/em&gt;, &lt;em&gt;"Someone wanted this place to remember something."&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;You look at the place with your own eyes, now or the next time you're there. Five hints if you need them, from meaning to appearance to direction, a pixelated glimpse, and finally the part of your photo where it is.&lt;/li&gt;
&lt;li&gt;You take a close-up and the model checks it: &lt;strong&gt;FOUND IT&lt;/strong&gt;, &lt;strong&gt;ALMOST&lt;/strong&gt; (right kind of thing, not the one I saw) or &lt;strong&gt;NOT QUITE&lt;/strong&gt;. At the end you see both side by side: &lt;em&gt;what I saw / what you saw.&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every "go outside" app I could think of has the same problem: the app is the thing you look at. Here the screen does two small jobs, the clue and the check, and the place does the rest.&lt;/p&gt;

&lt;p&gt;The model can see the street. It can't walk down it. That part is yours.&lt;/p&gt;

&lt;p&gt;It's for anyone who walks the same streets every day and has stopped seeing them, and for anyone who'd like a reason to look up from their phone that doesn't come from their phone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/P09wG8R31oY" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A real round at my café, recorded on my phone. The model's wait, the hint wait and my typing are sped up; nothing else is edited. I played this one from a photo at home, so the "search" happens in the photo. In the street, that part happens with your eyes and the phone in your pocket.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There's no hosted demo on purpose: the model runs on your own computer, and that's the point (more on that below). Not every round is that good. In another photo, an old one of mine from a square in France, it picked an empty bike dock, one of five identical ones. Technically there, not much of a hunt. More on that below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/maricastroc" rel="noopener noreferrer"&gt;
        maricastroc
      &lt;/a&gt; / &lt;a href="https://github.com/maricastroc/can-you-find-it" rel="noopener noreferrer"&gt;
        can-you-find-it
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Can You Find It?&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;I Spy, but the AI hides something real in a place you know — and checks that you found it.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Take a photo of where you are, or pick one of a place you pass every day: your
street, the way to the bakery, the view from a window. A local, open model
(Gemma 4) looks at it, secretly chooses one real thing in it, and gives you a
single line:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;I FOUND SOMETHING.&lt;/strong&gt;
&lt;em&gt;Someone wanted this place to remember something.&lt;/em&gt;
Find it with your own eyes, now or next time you're there.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;You look at the place, not the screen. When you spot it, you snap a close-up —
right away, or a five-second photo on your way past, sent whenever you like
The model compares it with what it saw: FOUND IT, ALMOST or NOT QUITE. At the
end you see both side by…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/maricastroc/can-you-find-it" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Next.js 16, React 19, TypeScript, Node, &lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; and Gemma 4 E4B. &lt;code&gt;npm run doctor&lt;/code&gt; checks Ollama, the model and your memory, and prints a QR code for your phone.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;Everything runs on &lt;strong&gt;Gemma 4 E4B&lt;/strong&gt; (open weights, Apache 2.0) through &lt;strong&gt;Ollama&lt;/strong&gt;, on my own laptop: local inference, no cloud API anywhere. The phone is just a browser on the same Wi-Fi; a small Next.js server on the laptop holds the rounds and talks to the model. The whole design is shaped around what a small open vision model can and can't do, so I started by measuring that.&lt;/p&gt;

&lt;h3&gt;
  
  
  First, a spike: can a small open model even do this?
&lt;/h3&gt;

&lt;p&gt;Before building any UI I wanted to know if the core trick works: can a model look at a wide photo of a real place and point at one specific, real, findable thing, without making it up?&lt;/p&gt;

&lt;p&gt;I took 20 photos of public places (parks, squares, streets, gardens, a playground, a forest trail), asked the model for targets with boxes, &lt;strong&gt;cut every box out as a real crop&lt;/strong&gt;, and labelled all 184 candidates by hand. No trusting the model's opinion of itself.&lt;/p&gt;

&lt;p&gt;What I learned:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemma 4 E4B localises well.&lt;/strong&gt; Gemma 4 natively returns boxes as &lt;code&gt;box_2d&lt;/code&gt; on a 0–1000 grid. With the final prompt, 88% of boxes held the thing it described, and 0% were invented.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bigger was worse.&lt;/strong&gt; Gemma 4 12B, on the same task, put only 41% of its boxes on the thing it named, and took twice as long. The game runs on E4B.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Asking for "interesting" things made it invent them.&lt;/strong&gt; Moss, drains and bollards that weren't there. Asking for fewer, shorter targets ("only what you can see with certainty") fixed it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Yes/no verification is useless.&lt;/strong&gt; Asked "is there a red mailbox in this crop?", a small model says yes. So every candidate is verified with &lt;strong&gt;multiple choice&lt;/strong&gt; on its crop, against the other things it proposed and a few decoys. It even had to answer with the option text, not a letter: it would describe the crest correctly and then pick the wrong letter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It writes bad clues.&lt;/strong&gt; Clues written by the model scored 1.03 out of 2 with me; when it zoomed in and wrote from the crop, 0.64, and it invented details. So the clue lines are &lt;strong&gt;written by people&lt;/strong&gt;, and rules pick the one that's true for the target. That got to about 1.55.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In the end, about 1 in 2 photos gives a genuinely good round, 1 in 5 gets an honest &lt;em&gt;"I couldn't find anything I'd trust here"&lt;/em&gt; (a forest trail has no plaques), and the rest are playable but mundane.&lt;/p&gt;

&lt;h3&gt;
  
  
  The pipeline
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;photo ──► Gemma proposes 2 short targets with boxes
      ──► cheap filters: no areas, people, animals, vehicles, huge boxes, look-alikes
      ──► multiple-choice check on the real crop (+ "is a person at it?")
      ──► human-written clue line, chosen by rules       ← shown right away
      ──► hints written in the background while you start looking
      ──► your close-up ──► FOUND IT / ALMOST / NOT QUITE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Making it fast enough on a laptop
&lt;/h3&gt;

&lt;p&gt;The first version took a median of 47 seconds to show a clue. Two targets instead of three, shorter labels, a smaller crop for verification, and showing the clue before the hints are written brought it to 26 seconds. A tiny warm-up pass when you open the camera wakes the model while you frame the photo.&lt;/p&gt;

&lt;p&gt;Then I found the real problem: memory. The model needs about 9.5 GB, my laptop has 16, and with a browser, an editor and a few chat apps open, macOS starts swapping. The same call that takes 11–22 seconds took 27–45 seconds, and once &lt;strong&gt;sixteen minutes&lt;/strong&gt;. &lt;code&gt;npm run doctor&lt;/code&gt; now warns you when the computer is swapping, which is the most useful line of code in the project.&lt;/p&gt;

&lt;p&gt;I also tried sending a smaller photo (1536 px instead of 1920). It's 25% fewer image tokens, but it lost exactly the best small targets, like a memorial plaque, so I kept 1920.&lt;/p&gt;

&lt;h3&gt;
  
  
  The bug that made me rewrite the check
&lt;/h3&gt;

&lt;p&gt;To test the losing screen, I photographed my laptop screen instead of the target, a black pedestal fan. The game said &lt;strong&gt;FOUND IT&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The log showed why. The model described my laptop photo as "Black pedestal fan with metal grille", the target's own words. My prompt told it what the target was &lt;em&gt;before&lt;/em&gt; it looked at my photo, and a small model, told what to expect, sees it.&lt;/p&gt;

&lt;p&gt;The check now happens in three steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The model says what the player's photo shows &lt;strong&gt;without being told the target&lt;/strong&gt; (the laptop became "a small, brown, round object", not a fan).&lt;/li&gt;
&lt;li&gt;A text-only question decides whether that description could be the target, among the round's other proposals and a few decoys.&lt;/li&gt;
&lt;li&gt;Only then are the two images compared: same object, same kind, or neither.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On my 24 test pairs it scores the same as before (10/10 real finds, 6/7 look-alikes, 7/7 unrelated), and it rejects the laptop. A wrong photo now gets its NOT QUITE in about 3 seconds.&lt;/p&gt;

&lt;h3&gt;
  
  
  Changing the game halfway through
&lt;/h3&gt;

&lt;p&gt;The first version only worked one way: stand somewhere, take the photo, wait, hunt, all with the phone out and a laptop on the same network. I couldn't test that in the middle of the street, and not everyone can, or should, walk around staring at a phone. So now the photo can come from your library too, hunts stay open for days, and the close-up can be a five-second photo on your way past, checked whenever you like. You can still play it all on the spot.&lt;/p&gt;

&lt;p&gt;That turned the slow model and the laptop at home into non-issues: you set up a hunt at home, look on your normal route, and check at home again. The screen part happens indoors. The street part doesn't need a screen.&lt;/p&gt;

&lt;h3&gt;
  
  
  Testing a game that runs on a model
&lt;/h3&gt;

&lt;p&gt;276 tests, including the engine replayed against &lt;strong&gt;recorded real Gemma replies&lt;/strong&gt;, so the pipeline is tested on what the model actually says, deterministically, without a GPU. Accessibility checks with axe on every screen.&lt;/p&gt;

&lt;h3&gt;
  
  
  What doesn't work yet
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Look-alikes.&lt;/strong&gt; A row of identical lamps or bike docks: the model's own count of similar objects is unreliable, so sometimes it picks one of five. ALMOST softens it; it's still the main open problem.&lt;/li&gt;
&lt;li&gt;About 1 in 2 photos is a great round. Point it at places with &lt;em&gt;things&lt;/em&gt;: signs, plaques, lamps, carvings.&lt;/li&gt;
&lt;li&gt;Your computer has to be on, and on the same network as your phone when you set up or check a hunt.&lt;/li&gt;
&lt;li&gt;English only, for now.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your photos never leave your network.&lt;/strong&gt; The phone resizes the photo (dropping GPS data), sends it to your own computer, and the model runs there. No API, no account, no company seeing your street or your café. For a game built on photos of where you live and walk, a closed API would mean sending exactly that to someone else's server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It costs nothing to play.&lt;/strong&gt; No tokens, no rate limits. A round is a handful of model calls on hardware I already own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I could look inside and fix it.&lt;/strong&gt; The two biggest improvements came from seeing exactly what the model saw and said: the crops that proved it localises well, and the log line where it called my laptop a fan. With an open model I can test another size (12B: worse), change the image resolution, and re-run my whole test set on my laptop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It works without the internet&lt;/strong&gt;, as long as the phone can reach the computer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Best Use of Gemma.&lt;/strong&gt; Gemma 4 E4B does all the seeing: it proposes targets with native &lt;code&gt;box_2d&lt;/code&gt; boxes, verifies each one on its own crop, describes the player's close-up without knowing the answer, and compares the two images. All locally, on a laptop.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
      <category>gemma</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Thread: AI Can Organize a Memory. It Shouldn't Become the Memory.</title>
      <dc:creator>Mariana Castro</dc:creator>
      <pubDate>Sun, 04 Oct 2026 18:29:28 +0000</pubDate>
      <link>https://dev.to/maricastroc/thread-ai-can-organize-a-memory-it-shouldnt-become-the-memory-h5h</link>
      <guid>https://dev.to/maricastroc/thread-ai-can-organize-a-memory-it-shouldnt-become-the-memory-h5h</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Thread is a local-first life archive built from recordings.&lt;/p&gt;

&lt;p&gt;I built it for a close friend who wants to preserve memories from different periods of her life as material for an autobiography.&lt;/p&gt;

&lt;p&gt;That gave me an interesting problem to solve. Recording those memories is easy. But as the archive grows, finding a particular story again — who was there, where it happened, when it happened, what other memories connect to it — becomes much harder.&lt;/p&gt;

&lt;p&gt;AI seems like an obvious way to organize that material. But it creates another problem: if we use AI to preserve someone's memories, at what point does the AI's interpretation start replacing the memory itself?&lt;/p&gt;

&lt;p&gt;That became the central constraint behind Thread:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI can organize a memory, but it should never become the memory.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Thread takes recorded or imported conversations and turns them into a navigable archive of one person's life: stories, people, places, dates, and the connections between them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffblg9kxm2w6bhyrhzxhe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffblg9kxm2w6bhyrhzxhe.png" alt="Kiara's Life in Thread: each story is drawn as the shape of its voice, placed in the year it happened" width="800" height="482"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Each mark is a story, drawn as the shape of the voice and placed in the year it happened. The years with no recordings stay visible.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;But the original recording always remains the source of truth.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd2zkxraf6qr2edg75xxt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd2zkxraf6qr2edg75xxt.png" alt="A recording where no story was found: kept exactly as received, with its transcript searchable" width="799" height="530"&gt;&lt;/a&gt;&lt;em&gt;A recording can hold no story at all, and it is still kept exactly as it was received.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Thread can identify that a story happened in 2007, for example, but that claim does not have to be trusted just because a model generated it. The interface shows where the information came from, and you can jump back to the exact moment in the recording and hear the person saying it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuq7v0jo5l4vm4sywqpbo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuq7v0jo5l4vm4sywqpbo.png" alt="A memory playing on the Life: when Graça is named, threads reach the other stories she appears in" width="800" height="549"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;While a memory plays, each name lights up at the second it is said, and threads reach every other moment where that person appears.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I think of the archive as four different layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Recording&lt;/strong&gt; — the preserved evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Story&lt;/strong&gt; — an interpretation of something told in that recording.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;People and places&lt;/strong&gt; — entities connected back to supporting evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Life&lt;/strong&gt; — a representation built from those stories, never a replacement for them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy94w4cjs6a9e6swmnc93.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy94w4cjs6a9e6swmnc93.png" alt="A recording page: the whole waveform, with each story found in it marked as a numbered span" width="799" height="673"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The recording is the evidence. The stories are spans laid over it, never a replacement for it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The result is not an AI-generated autobiography.&lt;/p&gt;

&lt;p&gt;It is a tool that helps my friend organize and revisit her own memories while keeping her words — not the model's reconstruction of them — at the center.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI organizes the memory. The voice remains the evidence.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;🎥 &lt;strong&gt;Video demo:&lt;/strong&gt; &lt;a href="https://www.youtube.com/watch?v=Q1Uah6K4Xck" rel="noopener noreferrer"&gt;https://www.youtube.com/watch?v=Q1Uah6K4Xck&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;🌐 &lt;strong&gt;Read-only live demo:&lt;/strong&gt; &lt;a href="https://thread.marianacastro.dev/" rel="noopener noreferrer"&gt;https://thread.marianacastro.dev/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The video shows the complete flow using a real recording: importing audio, processing it locally, discovering stories, inspecting provenance, jumping from an extracted fact back to the exact moment that supports it, and finally seeing the new memories appear in the person's Life view.&lt;/p&gt;

&lt;p&gt;The hosted demo is intentionally different from a normal local installation.&lt;/p&gt;

&lt;p&gt;A local Thread installation can import and process recordings using the local AI pipeline. The public deployment is a read-only snapshot of a preprocessed archive, so visitors can explore the product without the server needing access to the local models or accepting personal recordings.&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;💻 &lt;strong&gt;Source:&lt;/strong&gt; &lt;a href="https://github.com/maricastroc/thread" rel="noopener noreferrer"&gt;https://github.com/maricastroc/thread&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Thread is open source, and the repository includes the application, setup instructions, demo archive, fixtures, and documentation for running the full pipeline locally.&lt;/p&gt;
&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;Thread is a Next.js and TypeScript application backed by SQLite, but the interesting part for me was deciding exactly how much authority to give the AI.&lt;/p&gt;

&lt;p&gt;The processing pipeline looks roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;record
   ↓
preserve original audio
   ↓
transcribe + timestamp
   ↓
discover stories
   ↓
extract people, places and dates
   ↓
verify provenance
   ↓
index
   ↓
explore
   ↓
return to the original voice
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Listening: Whisper + Silero VAD
&lt;/h3&gt;

&lt;p&gt;Recordings are transcribed locally with &lt;strong&gt;whisper.cpp using Whisper large-v3-turbo&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Word-level timing matters because Thread needs more than a transcript: it needs to be able to take someone from an interpreted fact back to the moment in the recording that supports it.&lt;/p&gt;

&lt;p&gt;I also use &lt;strong&gt;Silero VAD&lt;/strong&gt; to detect speech and avoid treating long periods of silence as content.&lt;/p&gt;

&lt;h3&gt;
  
  
  Understanding: Gemma
&lt;/h3&gt;

&lt;p&gt;The interpretation layer uses &lt;strong&gt;Gemma 4 E4B through Ollama&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Gemma receives the transcript and performs a deliberately constrained job:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;find individual stories;&lt;/li&gt;
&lt;li&gt;suggest titles;&lt;/li&gt;
&lt;li&gt;identify people, places, and dates;&lt;/li&gt;
&lt;li&gt;point to the transcript segments that support those claims.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It produces structured output rather than becoming a general-purpose chatbot inside the archive.&lt;/p&gt;

&lt;p&gt;And importantly, the model does not get the final word.&lt;/p&gt;

&lt;p&gt;After Gemma proposes an annotation, Thread verifies its provenance deterministically against the transcript before storing it. If the cited words cannot be found, the claim is dropped.&lt;/p&gt;

&lt;p&gt;The interface then distinguishes different levels of provenance, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;said&lt;/strong&gt; — explicitly present in the recording;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;from the words&lt;/strong&gt; — derived directly from what was said;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;[inferred]&lt;/strong&gt; — interpretation rather than an explicit statement.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgvvzmn9rnmgu7lm1dy4j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgvvzmn9rnmgu7lm1dy4j.png" alt="A story's transcript: each fact is marked as said, or derived from the words, with the second it was said" width="800" height="410"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Every fact points back to its words and its second in the recording: "1978" comes from "78", said at 0:04.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That distinction is important because a plausible hallucination is especially dangerous in an archive of someone's life.&lt;/p&gt;

&lt;h3&gt;
  
  
  Finding: EmbeddingGemma + SQLite
&lt;/h3&gt;

&lt;p&gt;For retrieval, Thread combines &lt;strong&gt;EmbeddingGemma&lt;/strong&gt;, also running through Ollama, with SQLite FTS5.&lt;/p&gt;

&lt;p&gt;But search follows the same rule as the rest of the project.&lt;/p&gt;

&lt;p&gt;It does not generate an authoritative answer about the person's life.&lt;/p&gt;

&lt;p&gt;It finds the relevant part of the archive and takes you back to the evidence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcpobeuls8mxuketi2q2k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcpobeuls8mxuketi2q2k.png" alt="Search results: the moment where Kiara talks about the festa de São João, with a button to listen from that second" width="800" height="499"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Search never answers with generated text: it finds the moment and plays the voice from there.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Keeping interpretation replaceable
&lt;/h3&gt;

&lt;p&gt;The AI layer sits behind a &lt;code&gt;StoryInterpreter&lt;/code&gt; interface.&lt;/p&gt;

&lt;p&gt;That means the recordings and transcripts are not tied permanently to one model. The interpretation layer can be changed and the archive reprocessed while the underlying evidence remains intact.&lt;/p&gt;

&lt;p&gt;That separation became one of the most important architectural decisions in the project:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the model is replaceable; the memory is not.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everything required for the normal processing pipeline runs locally:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Next.js&lt;/li&gt;
&lt;li&gt;TypeScript&lt;/li&gt;
&lt;li&gt;SQLite&lt;/li&gt;
&lt;li&gt;ffmpeg&lt;/li&gt;
&lt;li&gt;whisper.cpp&lt;/li&gt;
&lt;li&gt;Silero VAD&lt;/li&gt;
&lt;li&gt;Gemma through Ollama&lt;/li&gt;
&lt;li&gt;EmbeddingGemma through Ollama&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No recording needs to leave the machine for Thread to build the archive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;For Thread, using open models was not just a technology preference. It changed what I could make the product promise.&lt;/p&gt;

&lt;p&gt;These recordings can contain family stories, names, relationships, places, personal events, and voices. They are exactly the kind of data I do not want the architecture to assume can be sent to an external service.&lt;/p&gt;

&lt;p&gt;Because the models can run locally, Thread can process those recordings without requiring them to leave the person's computer.&lt;/p&gt;

&lt;p&gt;But openness also matters for a second reason: &lt;strong&gt;interpretation changes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A story extracted by today's model should not become permanently fused with the archive simply because that was the model available when the recording was imported.&lt;/p&gt;

&lt;p&gt;Thread preserves the recording and transcript separately from the interpretation layer. A different open model can be plugged in later, and the archive can be reprocessed without changing its underlying evidence.&lt;/p&gt;

&lt;p&gt;That makes it possible to treat AI output as what it actually is: an interpretation that can be inspected, challenged, improved, or replaced.&lt;/p&gt;

&lt;p&gt;A closed API could perform many of the individual tasks in this pipeline. But building around open models made it possible to combine &lt;strong&gt;local processing, replaceable interpretation, inspectable provenance, and long-term ownership of the archive&lt;/strong&gt; as properties of the system rather than promises made by an external provider.&lt;/p&gt;

&lt;p&gt;Open innovation also made experimentation practical. I could constrain the model, test how its interpretations behaved across different recordings, create fixtures in Portuguese, English, and Spanish, and build a larger synthetic archive containing dozens of stories to see whether the interface and retrieval model still made sense beyond the original demo.&lt;/p&gt;

&lt;p&gt;For a project about preserving someone's memories, I think that distinction matters.&lt;/p&gt;

&lt;p&gt;The intelligence helping organize the archive can change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The person's voice should remain.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma&lt;/strong&gt; — Gemma is the interpretation layer at the core of Thread: it discovers stories and extracts structured people, places, and dates from transcripts, while Thread independently verifies the provenance of its claims before they enter the archive.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
