DEV Community

Matheus Ferreira
Matheus Ferreira

Posted on

Duelo de Fotos: a photo duel for two, judged by a local Gemma that is sometimes confidently wrong

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge: Week 1.

What I Built

My girlfriend and I like going out, by day or by night. We also both own a phone, which means a walk can quietly turn into two people looking at two screens in a nicer location.

So I built Duelo de Fotos ("photo duel"), a game for the two of us that cannot be played from the couch.

  1. Before leaving, we type our two names and say whether the walk is by day or by night. A Gemma model running on my computer picks five secret themes for each of us: "something rusty", "something that blinks", "something that looks like it has a face".
  2. Each of us opens only our own card, takes a picture of it with the phone, and closes it. From here on the phone is a camera and nothing else.
  3. Outside, we each hunt for one photo per theme. Interpretation is allowed.
  4. Back home, the photos go in and Gemma becomes the judge. It looks at each photo on its own, gives it a score from 0 to 10 and one sentence explaining why. The page adds it up and says who won.
  5. Then we argue with the judge, because it is a 4B model and it is sometimes confidently wrong. I built the game around that instead of hiding it, and I will come back to why.

The challenge says the best builds "should make the screen the shortest part of the experience". That sentence shaped everything. The app has two short moments, one before the walk and one after. In between there is no app to open, nothing to fill in, and no reason to look at the phone except to take a picture of a rusty gate.

The scoreboard after a test duel

The scoreboard from a test duel. These are drawings I generated to test the judge, not real photos.

One thing I want to say up front: the game is finished and working, but we have not played it on a real walk yet. Everything below about how the judge behaves comes from test runs at my desk. The first real duel is the next thing on the list.

Demo

The app runs locally, so there is no hosted link. The video shows it running on my own computer.

The interface is in Brazilian Portuguese, because the game is for the two of us. What you are seeing, in order:

  • "Sortear os temas" (draw the themes): the form with the two names and the day or night choice.
  • "Carta secreta" (secret card): one player's five themes. Each theme starts with "Algo", which means "something".
  • "Chamar o juiz" (call the judge): the upload page, one photo per theme.
  • The scoreboard: the two totals, and under each photo the theme in bold and the judge's reason in grey.

Code

Duelo de Fotos

A photo duel for two people, judged by a Gemma model running on your own computer.

Before you go out, Gemma hands each player five secret themes ("something rusty", "something that blinks"). Outside, your phone is only a camera. Back home, you drop the photos in, Gemma scores each one from 0 to 10 with a one-line reason, and the scoreboard tells you who won.

The screen is the short part: about a minute before the walk and a minute after. Everything in between happens outside.

Built for the DEV Hacktoberfest 2026 challenge, week 1: "Touch Grass".

The scoreboard after a test duel

The photos in the screenshots are test drawings, not real pictures. The app's interface is in Brazilian Portuguese.

How a duel goes

  1. Draw the themes. Type the two names, say whether the walk is by day or by night, and optionally what kind of place it is. Gemma picks…

MIT licensed. To run it you need Java 25 and Ollama:

ollama pull gemma3:4b
./mvnw spring-boot:run
Enter fullscreen mode Exit fullscreen mode

Then open http://127.0.0.1:8080.

How I Built It

Browser  ->  Spring Boot  ->  Spring AI ChatClient  ->  Ollama  ->  Gemma 3 (4B)
Enter fullscreen mode Exit fullscreen mode

Java 25, Spring Boot 4.1, Spring AI 2.0, Thymeleaf pages, and gemma3:4b through Ollama. Duels are stored as a folder on disk with the photos and one JSON file. There is no database.

The first idea was wrong

I did not start with a game. I started with a walk diary: you come home, drop in your photos, and Gemma writes a caption for each one and lays them out like polaroids on a corkboard. I even had a star map of the sky above the place and time of the last photo.

It worked, and it was dry. It organized a walk that had already happened. Nothing in it gave anyone a reason to leave the house, and the theme of the week is exactly that. I threw it away the same evening, star map included.

The duel fixes this at the root: the app has nothing to judge until you go out and come back.

Letting a small model invent the themes did not work

My first version of the duel asked Gemma to write ten themes from scratch. I tried four different prompts, and each one failed in its own way:

  • Too poetic. "Shared laughter in silence", "Time in every instant". Nobody can photograph those. One theme, "People in search of moments", asked for photos of strangers, which the prompt had explicitly forbidden.
  • Too specific. After I asked for concrete things: "Blue street-cleaning truck", "Large polished metal chandeliers". Good luck finding those on a Tuesday night.
  • Stilted. After I asked for "something" plus a visible quality and listed the kinds of quality (colour, shape, size, material), the model copied my category words straight into the output: "Something size grandiose imposing", "Something position horizontal calm".
  • Poetic again. After I removed the list of categories: "Something that resonates mysteriously".

Last week, building Papel Claro, I learned that a list of examples is a list of suggestions to a 4B model. This week I learned the same is true of a list of categories.

So I stopped asking the model to write. The themes now come from a hand-written list of 75, each tagged as day, night or any time. For each duel the code draws 24 candidates and asks Gemma to choose the ten that fit the walk. Any line it returns that is not in the list is ignored, and if it picks fewer than ten the code fills the rest. In four test runs it returned ten valid themes every time, and the themes read like something a person would say: "something with a double shadow", "something that grew where it should not have".

The honest cost: Gemma now chooses instead of creating, and the "kind of place" you type barely changes the result. Day versus night makes a real difference, and that comes from the list.

One photo per call, and a strict answer format

Each photo goes to the model alone, with its theme. Last week the same model read a close-up well and a full page badly, so I never ask it to compare two images.

String resposta = chat.prompt()
    .user(u -> u.text(pedido).media(tipo, new ByteArrayResource(foto)))
    .call()
    .content();
Optional<Veredito> veredito = vereditoDe(resposta, json);
Enter fullscreen mode Exit fullscreen mode

The model must answer with a JSON object holding a score and a reason. A small model sometimes returns JSON without the fields you asked for, so an answer missing either one is a failure, not an empty verdict:

if (veredito == null || veredito.nota() == null || veredito.nota() < 0 || veredito.nota() > 10
        || veredito.motivo() == null || veredito.motivo().isBlank()) {
    return Optional.empty();
}
Enter fullscreen mode Exit fullscreen mode

The app tries once more and then shows "the judge could not judge this photo", worth zero points.

What the judge gets wrong

I tested the judge with simple drawings. One of them was a red ball on a grey floor.

  • For the theme "something vivid blue", the judge gave it 3. Correct. Its reason: "The photo shows the flag of Japan." It is not the flag of Japan.
  • For the theme "something made of smooth wood", the judge gave the same red ball 8, because it saw "a block of smooth wood".

It is also generous: seven of the nine scores in that test landed on 7 or 8.

In most projects this would be the section where I explain how I worked around the model. Here I decided not to. A referee who is occasionally, confidently wrong is a familiar thing in every sport, and I expect arguing about the call to be half the fun. The reason is printed under every photo so there is always something specific to argue with. What I did instead was make the failure visible and cheap: a low-stakes game is the right place for a 4B judge. A tool that decided anything important with these verdicts would not be.

Things I only found by running it

  • The scoreboard crashed the first time a full duel finished. Thymeleaf could not call a non-public method on my record, and the unit tests had never rendered that page. There is now a test that renders every page.
  • Spring AI retries were already off. I carried that over from last week: with the default, a closed Ollama freezes the page for minutes. In Spring AI 2.0 the setting counts retries, so 1 still makes two calls and 0 makes one.
  • Logs record only the exception class. The message of a failed model call can quote what was sent to the model.

On my machine (an AMD RX 9060 XT and 16 GB of RAM) choosing the themes takes about 2 seconds, and judging nine small test drawings took 18 seconds. Full-size phone photos will take longer, and I have not measured by how much.

Honest limits

  • It has not been played on a real walk yet. The judge was tested with drawings, not with photos taken outside, and night photos are the case I trust least.
  • The judge is generous and sometimes wrong, as described above.
  • HEIC photos are not accepted, so an iPhone has to save JPEG.
  • There is no list of past duels. You return to a duel through its address.
  • The interface is in Portuguese only.

Why Does Open Innovation Matter?

The photos are of us. A game like this collects exactly the pictures I would not upload anywhere: the two of us, the streets around where we live, at the times we are usually out. With an open-weight model on my own computer, the photos go from the browser to a folder on the same machine. The server listens on 127.0.0.1 only. Nothing in the project can send a photo anywhere, because nothing in it talks to anything but the local Ollama.

A silly game should cost nothing to play. One duel is one call to choose themes and ten vision calls to judge. With a paid API I would start counting, and a game you ration is not a game. With a local model the tenth duel of the month costs the same as the first.

You can read what the judge was told. The judge's instructions and the list of themes are plain text in the repository. When a verdict is absurd, I can see exactly what the model was asked. And the list is the part I expect people to change: add your own themes, your own inside jokes, one line each.

The judge can be replaced. The model is one line in application.properties. If you have the hardware for a larger Gemma and want a stricter referee, change the name and play again.

None of the heavy pieces are mine. Gemma, Ollama, Spring AI and Spring Boot are all open. My part was finding a use for them that ends with two people outside.

My Agent Session

I built this with Claude Code, in one long session, and I would rather describe it plainly.

  • The code was written by Claude. I chose the stack and the constraints, and I ran and looked at every step.
  • The direction was mine, mostly by saying no. The walk diary was my idea and I was the one who called it dry. The star map was also my idea, and I dropped it a few minutes after seeing it work. I asked for other ideas, got five, and picked the duel. When the first design looked tacky, I asked for a cleaner one.
  • This article was drafted with Claude too. The idea, the decisions and the person it is for are mine.

Prize Categories

Best Use of Gemma. A single gemma3:4b, running locally through Ollama, does both jobs in the game. As text model it chooses the themes that fit the walk. As vision model it looks at every photo and writes the verdict. Its size is what makes the project possible, since it runs on an ordinary home computer and the photos never leave it. Its size is also what makes the project fun: the game is built around a judge that can be wrong, instead of pretending it never is.

Top comments (0)