DEV Community

Cover image for Can You Find It? An open model hides one real thing in a place you know
Mariana Castro
Mariana Castro

Posted on AI-assisted

Can You Find It? An open model hides one real thing in a place you know

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

I go to the same café every week. Same counter, same stool, same order. Last night a 4-billion-parameter model running on my laptop looked at a photo of that counter, picked one thing in it, and gave me one line:

I found something hiding in plain sight.

It took me two hints to find it: a small framed award on the tiled wall, the kind of thing you stop seeing the second time you walk in. My field note, typed while I was still a bit annoyed with myself: "A café I go to every week. Finally! I've sat at that counter dozens of times and never noticed it!"

That's the whole game.

Can You Find It? is I Spy, but the AI does the spying, and it only uses what is really in front of you.

  1. You take a photo of where you are, or pick one of a place you pass every day: your street, the way to the bakery, the view from a window.
  2. Gemma 4, running on your own computer, secretly chooses one real object in that photo, with a box around it.
  3. You get a single clue line. Not a description: an invitation to look at the place through a lens ("It only does its job after dark.", "Someone wanted this place to remember something.").
  4. You look at the place with your own eyes, now or the next time you're there. Five hints if you need them, from meaning to appearance to direction, a pixelated glimpse, and finally the part of your photo where it is.
  5. You take a close-up and the model checks it: FOUND IT, ALMOST (right kind of thing, not the one I saw) or NOT QUITE. At the end you see both side by side: what I saw / what you saw.

Every "go outside" app I could think of has the same problem: the app is the thing you look at. Here the screen does two small jobs, the clue and the check, and the place does the rest.

The model can see the street. It can't walk down it. That part is yours.

It's for anyone who walks the same streets every day and has stopped seeing them, and for anyone who'd like a reason to look up from their phone that doesn't come from their phone.

Demo

A real round at my café, recorded on my phone. The model's wait, the hint wait and my typing are sped up; nothing else is edited. I played this one from a photo at home, so the "search" happens in the photo. In the street, that part happens with your eyes and the phone in your pocket.

There's no hosted demo on purpose: the model runs on your own computer, and that's the point (more on that below). Not every round is that good. In another photo, an old one of mine from a square in France, it picked an empty bike dock, one of five identical ones. Technically there, not much of a hunt. More on that below.

Code

Can You Find It?

I Spy, but the AI hides something real in a place you know — and checks that you found it.

Take a photo of where you are, or pick one of a place you pass every day: your street, the way to the bakery, the view from a window. A local, open model (Gemma 4) looks at it, secretly chooses one real thing in it, and gives you a single line:

I FOUND SOMETHING. Someone wanted this place to remember something. Find it with your own eyes, now or next time you're there.

You look at the place, not the screen. When you spot it, you snap a close-up — right away, or a five-second photo on your way past, sent whenever you like The model compares it with what it saw: FOUND IT, ALMOST or NOT QUITE. At the end you see both side by…

Next.js 16, React 19, TypeScript, Node, Ollama and Gemma 4 E4B. npm run doctor checks Ollama, the model and your memory, and prints a QR code for your phone.

How I Built It

Everything runs on Gemma 4 E4B (open weights, Apache 2.0) through Ollama, on my own laptop: local inference, no cloud API anywhere. The phone is just a browser on the same Wi-Fi; a small Next.js server on the laptop holds the rounds and talks to the model. The whole design is shaped around what a small open vision model can and can't do, so I started by measuring that.

First, a spike: can a small open model even do this?

Before building any UI I wanted to know if the core trick works: can a model look at a wide photo of a real place and point at one specific, real, findable thing, without making it up?

I took 20 photos of public places (parks, squares, streets, gardens, a playground, a forest trail), asked the model for targets with boxes, cut every box out as a real crop, and labelled all 184 candidates by hand. No trusting the model's opinion of itself.

What I learned:

  • Gemma 4 E4B localises well. Gemma 4 natively returns boxes as box_2d on a 0–1000 grid. With the final prompt, 88% of boxes held the thing it described, and 0% were invented.
  • Bigger was worse. Gemma 4 12B, on the same task, put only 41% of its boxes on the thing it named, and took twice as long. The game runs on E4B.
  • Asking for "interesting" things made it invent them. Moss, drains and bollards that weren't there. Asking for fewer, shorter targets ("only what you can see with certainty") fixed it.
  • Yes/no verification is useless. Asked "is there a red mailbox in this crop?", a small model says yes. So every candidate is verified with multiple choice on its crop, against the other things it proposed and a few decoys. It even had to answer with the option text, not a letter: it would describe the crest correctly and then pick the wrong letter.
  • It writes bad clues. Clues written by the model scored 1.03 out of 2 with me; when it zoomed in and wrote from the crop, 0.64, and it invented details. So the clue lines are written by people, and rules pick the one that's true for the target. That got to about 1.55.

In the end, about 1 in 2 photos gives a genuinely good round, 1 in 5 gets an honest "I couldn't find anything I'd trust here" (a forest trail has no plaques), and the rest are playable but mundane.

The pipeline

photo ──► Gemma proposes 2 short targets with boxes
      ──► cheap filters: no areas, people, animals, vehicles, huge boxes, look-alikes
      ──► multiple-choice check on the real crop (+ "is a person at it?")
      ──► human-written clue line, chosen by rules       ← shown right away
      ──► hints written in the background while you start looking
      ──► your close-up ──► FOUND IT / ALMOST / NOT QUITE
Enter fullscreen mode Exit fullscreen mode

Making it fast enough on a laptop

The first version took a median of 47 seconds to show a clue. Two targets instead of three, shorter labels, a smaller crop for verification, and showing the clue before the hints are written brought it to 26 seconds. A tiny warm-up pass when you open the camera wakes the model while you frame the photo.

Then I found the real problem: memory. The model needs about 9.5 GB, my laptop has 16, and with a browser, an editor and a few chat apps open, macOS starts swapping. The same call that takes 11–22 seconds took 27–45 seconds, and once sixteen minutes. npm run doctor now warns you when the computer is swapping, which is the most useful line of code in the project.

I also tried sending a smaller photo (1536 px instead of 1920). It's 25% fewer image tokens, but it lost exactly the best small targets, like a memorial plaque, so I kept 1920.

The bug that made me rewrite the check

To test the losing screen, I photographed my laptop screen instead of the target, a black pedestal fan. The game said FOUND IT.

The log showed why. The model described my laptop photo as "Black pedestal fan with metal grille", the target's own words. My prompt told it what the target was before it looked at my photo, and a small model, told what to expect, sees it.

The check now happens in three steps:

  1. The model says what the player's photo shows without being told the target (the laptop became "a small, brown, round object", not a fan).
  2. A text-only question decides whether that description could be the target, among the round's other proposals and a few decoys.
  3. Only then are the two images compared: same object, same kind, or neither.

On my 24 test pairs it scores the same as before (10/10 real finds, 6/7 look-alikes, 7/7 unrelated), and it rejects the laptop. A wrong photo now gets its NOT QUITE in about 3 seconds.

Changing the game halfway through

The first version only worked one way: stand somewhere, take the photo, wait, hunt, all with the phone out and a laptop on the same network. I couldn't test that in the middle of the street, and not everyone can, or should, walk around staring at a phone. So now the photo can come from your library too, hunts stay open for days, and the close-up can be a five-second photo on your way past, checked whenever you like. You can still play it all on the spot.

That turned the slow model and the laptop at home into non-issues: you set up a hunt at home, look on your normal route, and check at home again. The screen part happens indoors. The street part doesn't need a screen.

Testing a game that runs on a model

276 tests, including the engine replayed against recorded real Gemma replies, so the pipeline is tested on what the model actually says, deterministically, without a GPU. Accessibility checks with axe on every screen.

What doesn't work yet

  • Look-alikes. A row of identical lamps or bike docks: the model's own count of similar objects is unreliable, so sometimes it picks one of five. ALMOST softens it; it's still the main open problem.
  • About 1 in 2 photos is a great round. Point it at places with things: signs, plaques, lamps, carvings.
  • Your computer has to be on, and on the same network as your phone when you set up or check a hunt.
  • English only, for now.

Why Does Open Innovation Matter?

  • Your photos never leave your network. The phone resizes the photo (dropping GPS data), sends it to your own computer, and the model runs there. No API, no account, no company seeing your street or your café. For a game built on photos of where you live and walk, a closed API would mean sending exactly that to someone else's server.
  • It costs nothing to play. No tokens, no rate limits. A round is a handful of model calls on hardware I already own.
  • I could look inside and fix it. The two biggest improvements came from seeing exactly what the model saw and said: the crops that proved it localises well, and the log line where it called my laptop a fan. With an open model I can test another size (12B: worse), change the image resolution, and re-run my whole test set on my laptop.
  • It works without the internet, as long as the phone can reach the computer.

My Agent Session

Prize Categories

Best Use of Gemma. Gemma 4 E4B does all the seeing: it proposes targets with native box_2d boxes, verifies each one on its own crop, describes the player's close-up without knowing the answer, and compares the two images. All locally, on a laptop.

Top comments (0)