DEV Community

Cover image for fieldcards: a local Gemma turns your photos into a printed scavenger hunt
Kunal
Kunal

Posted on

fieldcards: a local Gemma turns your photos into a printed scavenger hunt

Hacktoberfest: Maintainer Spotlight

This is a submission for the Hacktoberfest Open-Source AI Challenge: Week 1 (theme: Touch Grass).

What I Built

fieldcards is a command-line tool that turns the photos on your laptop into a printed scavenger hunt.

You point it at a folder of photos from a park, a garden or a trail. Gemma 3, running locally through Ollama, looks at each one, works out the most findable thing in it, and writes a clue that doesn't name it. Each clue is printed as an ASCII card. You print the sheet, leave your phone at home, and go find the real things.

The screen is the shortest part: about two minutes to make the sheet, then the rest of the afternoon outside.

+-[ #02 bird ]-------------------------+
|                                      |
|    Find the one that often stares    |
|     with a pointed beak and dark     |
|               plumage.               |
|                                      |
+--------------------------------------+
- Shiny black feathers
- Strong, curved beak
- Alert, upright posture
Where: Near trees and open water areas.
[ ] found it   time: ______
Enter fullscreen mode Exit fullscreen mode

That card is real gemma3:4b output for a photo of a crow on a branch.

Demo

It's one command:

ollama pull gemma3:4b
npm install -g github:KunalSiyag/fieldcards
fieldcards ./photos --title "Saturday walk"
Enter fullscreen mode Exit fullscreen mode

In a terminal, every card loads with the same effect as the hero of altsvg, my earlier project: the card frame fills with shimmering noise and a scan line while Gemma reads the photo, and when the answer arrives, the noise dissolves into the card. Here is a real run on three photos. The recording shortens each wait to about two seconds; the real times are in the next section.

fieldcards in a terminal: noise and a scan line while Gemma reads each photo, then each card dissolves into place

Every run writes three files:

  • index.html: the print-ready sheet, with the answer key on its own page.
  • cards.txt: the same cards as plain text. --cols 32 fits a 58 mm receipt printer, and --stdout | lp prints them straight away.
  • deck.json: everything Gemma said. Fix any mistakes, then reprint with --from without running the model again.

The printed sheet: clue cards in ASCII frames, things to look for, and tick boxes

--mode guide turns the same photos into a pocket field guide, with names and descriptions instead of riddles.

A real photo of mine

The first photo in that run is the only photo of my own I had on hand while writing this: the bar counter of a brewery I visited recently, taken from the floor above. It isn't an outdoor scene, but it shows the whole pipeline on a real, busy photo.

An oval bar counter seen from above, under a large ring chandelier

Gemma called it "round, reflective surfaces", with the clue "Find something that mirrors the sky and reflects the light around." That's fair for such a busy scene, but it isn't what I'd want on a card. It's exactly the case deck.json is for: change the subject to "Oval bar counter", run fieldcards --from fieldcards-out/deck.json, and the sheet reprints without calling the model again.

Code

GitHub logo KunalSiyag / fieldcards

Turn outdoor photos into printable scavenger-hunt cards with a local Gemma model. Offline, private, screen-light.

fieldcards

Turn your outdoor photos into a printable scavenger hunt, with an open-weight model running on your own computer.

Each photo goes to Gemma 3, running locally through Ollama. Gemma works out what the photo shows and writes a clue for finding it that doesn't name it. Each clue is printed in a +-[ #02 bird ]--+ frame, in the style of the altsvg alt-text placeholders. In a terminal, every card loads like the altsvg hero: shimmering noise and a scan line while Gemma reads the photo, then the noise dissolves into the card.

fieldcards in a terminal: noise and a scan line while Gemma reads each photo, then each card dissolves into place You print the sheet, leave your phone at home, and go find the real things.

+-[ #05 other ]----------------+
|                              |
|    Find a dark, weathered    |
|    giant resting beside a    |
|   rushing, silver stream.    |
|                              |
+------------------------------+
- Green moss covering its
  surface
- Dark, grey wood
- A smooth,
…

How I Built It

  1. Describe. Each photo goes to POST /api/chat on the local Ollama server, together with a JSON schema in format. Gemma has to answer with subject, category, description, clue, look_for, where and confidence, so a chatty reply can't break the layout.
  2. Check. The answer is validated and trimmed. If a clue contains its own answer, the answer is blanked out, and low-confidence answers print as "best guess" on the card.
  3. Print. Each card is one grid of characters: the frame and the clue. The same grid is drawn as SVG for the printed page, as plain text for cards.txt, and animated in the terminal, so all three always match. Each run of characters is placed on an exact column, so cards line up in any monospace font, on screen and on paper. The style comes from altsvg's alt-text placeholders, and altsvg wraps and measures the text.
  4. Animate. The loader and the dissolve are drawn on stderr with plain ANSI codes. It only runs in an interactive terminal, so pipes, CI and --stdout stay clean, and --no-anim turns it off.

Photos that were already described are matched by content hash and skipped on reruns, so you can add photos to the folder and run it again.

Can a 4B model draw?

I wanted each card to show a picture of its subject, the way the altsvg hero draws a mountain scene for "Mountains at sunrise". I tried four ways:

  • A fixed library of hand-drawn ASCII pictures that Gemma chose from. They looked good, but they're the opposite of generated: the clue "something that mirrors the sky" got a house. I removed it.
  • Gemma draws the subject itself as named shapes on a 100 by 60 canvas ("head: circle, beak: line, branch: rect"), which fieldcards renders as ASCII line art. The drawing always matches what Gemma described, because Gemma made it. But a 4B model draws crudely. The mirror came out as a circle in a frame; this is its crow:
                 -----\
              //        \\
            /--\\         \
          | |    \        |
          | |    |       \||
           \\\  /         /\
            |\\         //||
            |  \\-- --//  ||
            |             ||
            ---------------|
Enter fullscreen mode Exit fullscreen mode
  • A bigger model, gemma3:12b. On my 4 GB GPU it ran mostly on the CPU at about 115 seconds per drawing, and the drawings were no better.
  • Asking for SVG line icons, in the style of open-source icon sets. Worse again: a tick for the crow, a dash for the dandelion.

So the drawings are opt-in: --art drawing adds Gemma's doodle above each clue, for about 25 seconds per photo. The default card is the clue, which is fast and always says the right thing. I also tried printing the photo itself in characters; at card size, a busy photo turns into texture, so that's the opt-in --art photo.

Tuning a 4B model

The prompt took three tries, and two of them taught me something:

  • Example names leak. I added "Dandelion seed head" to the prompt as an example of a good name. On the next run Gemma labelled a waterfall photo "Dandelion seed head". Small models copy examples, so the final prompt has none.
  • Strict wording makes it timid. When I told it to use "high" confidence only when certain, everything came back "low", and the names got vague ("Dark feathered observer"). The original, plainer prompt worked best. I kept it and added one rule so the clues don't all start with "Seek".

What it gets right and wrong

On my test photos Gemma reliably recognised common things: a dandelion clock, a maple leaf, a waterfall, a crow. Its clues don't give the answer away. It was weaker on specifics: the crow came back as "a dark, watchful bird", a soft pine pollen cone was described as "woody", and the brewery became "round, reflective surfaces". Its self-rated confidence is only a rough signal. That's why every answer lands in deck.json for you to check before printing, and why every sheet says "look, don't pick".

Why Open Matters Here

  • No signal needed. The model runs on the laptop, so you can make cards at a cabin, a campsite or anywhere offline.
  • Your photos stay yours. Photos carry faces, houses in the background and GPS in their metadata. Nothing is uploaded; the only request goes to localhost.
  • It costs nothing to run. There's no API key and no per-photo bill, so a teacher can make a sheet for every class.
  • It's swappable. --model and --draw-model take any model Ollama can run, and the JSON schemas keep the output the same shape. That's how I could test a 12B model in one flag.

On my laptop (RTX 3050, 4 GB), the first photo takes about two minutes because it also loads the model. Each photo after that takes 7 to 14 seconds, using about 2.9 GB of GPU memory.

What's Next

I haven't taken a printed hunt outside yet. That's the next step, and it's the test the whole design depends on. When I do, I want to find out:

  • whether six cards is the right size for a 30 to 60 minute walk;
  • whether the clues are fun to solve or too vague, especially the cards marked "best guess";
  • whether people want a picture on the card, or whether the riddle is the fun part.

I'll add an update to this post with photos of the sheet in use.

Prize Categories

Best Use of Gemma. Gemma 3 4B is the core of the project. It reads every photo locally and writes each card's clue and notes through a JSON schema, and with --art drawing it draws each picture too.

Credits

  • Gemma 3 by Google, run locally with Ollama.
  • altsvg is my own earlier open-source project (started in March 2026, published on npm). Its ASCII style was added this week for fieldcards.
  • jpeg-js and pngjs decode photos for the optional --art photo mode.
  • Test photos from Wikimedia Commons, including the crow by Alexis Lours (CC BY 4.0) and the waterfall by Dietmar Rabich (CC BY-SA 4.0). The brewery photo is mine.
  • Built with help from Claude Code.

Top comments (0)