DEV Community

Yash Adake
Yash Adake

Posted on AI-assisted

Trailnotes: a field journal my own GPU writes from my walk photos

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

Trailnotes turns the photos from a walk into a field journal: a timeline, a map of where I stopped, and a short note for each photo, written by an open-weight vision model running on my own laptop GPU. Time and GPS come from the photos' own metadata. The photos and the notes never leave my machine.

I love travelling, especially trekking, and I've always wanted to document those trips in a well-formatted way. That's where the idea came from: one place to keep my memories with photos, short descriptions, timestamps and locations.

I built it for this challenge, then tested it the only honest way: I went for a walk. On 8 October I spent 27 minutes walking from my street to the Pavana riverbank and then through a local market in Pune, taking 18 photos on my phone.

Demo

Code

Trailnotes

Turn a folder of walk photos into a field journal. A vision model running on your own computer looks at each photo, and Trailnotes adds the time and GPS from the photo's own metadata. You get a map, a timeline, walk stats and a journal. Your photos and the model's descriptions never leave the machine. One exception: opening the page loads the map library and map tiles from the internet, which tells those servers roughly which area you are looking at.

Built for the Hacktoberfest 2026 "Touch Grass" challenge: open-weight AI that gets you out of the house and then helps you remember what you saw.

Why local matters here

Your walk photos carry your exact GPS position, your routine and often your home. Sending those to a hosted vision API means handing that over. With an open-weight model on your own GPU, the photos never go anywhere, it…

How I Built It

  • Built with Claude Code as my coding assistant: I directed the design, ran every real-model test on my own laptop, took the walk and checked the results.

  • Python, with Pillow reading each photo's EXIF data (time, GPS, orientation).

  • Model: qwen2.5vl:7b, an open-weight vision model, served by Ollama on an RTX 5060 laptop GPU. A 19-photo run took 113.5 seconds, about 6 seconds a photo.

  • Structured output: the model returns JSON (title, living things with a confidence, terrain, description). The schema caps list and string length, and an unusable answer is retried once. I added that after one test run produced a runaway 9,400-character list for a single photo.

  • Untrusted output: everything the model writes is HTML-escaped, and a photo it can't describe is shown as failed, never silently dropped.

  • Honest stats: distance only joins photos taken within 3 hours of each other. My first version added up a 217 km "walk" from old phone photos taken months apart in two cities. Real photos caught that, not my unit tests.

  • Privacy zone: my walk started close to home, so the first map showed exactly where I live. I added --privacy-zone LAT,LON,METRES: photos inside the radius keep their note but lose their location everywhere (no pin, no coordinates in the page or the JSON, not counted in the distance). Four of my 18 photos fall inside it, so the journal shows 14 pins and records at least 0.71 km. The thumbnails are re-encoded without EXIF, so they carry no GPS either.

  • The map: OpenStreetMap's own tile servers showed "Access blocked" when I opened the journal as a local file, and another free provider now wants an API key. I switched to OpenTopoMap, which uses the same open OpenStreetMap data, and added a --tiles flag to choose another.

  • 43 tests, and for the important ones I deliberately broke the code to make sure they fail.

What it got right and wrong

I checked the entries against my photos.

Right:

  • The riverbank: "a calm river stretches through a lush, tree-lined bank."
  • A coconut stall, a sugarcane-juice cart, a glass shop, a sunglasses stand, a rack of belts and stacked tyres, all described plainly and correctly.
  • It read the shop signs correctly: "S.K. GLASS", "Shree Vision LED Screen", "Imported Glasses", even a small "Rs. 9" price on the sunglasses stand.
  • It named a flowering bush as Lantana camara, and it was right.
  • In my first run it spotted a small child's bicycle in an alley photo and marked it "low" confidence. In the next run, on the same photo, it didn't mention the bicycle at all. Same model, same photo, different answer.

Wrong:

  • It listed things that aren't alive as living things: a building, cut coconuts, cut sugarcane stalks, and (in the first run) that bicycle, even though the prompt tells it not to. It saw the right objects and put them in the wrong list.
  • One label came out as "dieser plant", and in the second run just "dieser": the model slipped into German mid-answer.
  • It called a Marathi shop sign "Hindi": same script, different language.
  • A street tree held up by a metal support became "a healthy forest environment" on a "forest floor".
  • It called the flower market "indoor" while describing a dirt floor.

The journal says on every page that identifications are guesses by a small model, and that "high" confidence is the model's opinion, not a measurement. After this walk I agree with it.

Why Does Open Innovation Matter?

  • Privacy is the product. A walk's photos carry exact locations and routines. A hosted vision API would receive all of it. Here nothing leaves the laptop, and the privacy zone exists because I could see and change exactly where location data went.
  • Choice. I swapped models with one flag and compared them on the same photos. gemma3:4b reported a squirrel at "high" confidence on a gate covered in signs; qwen2.5vl:7b didn't. A closed API gives you one model and no comparison.
  • Offline and free per photo. Once the model is downloaded, describing a photo needs no connection and costs nothing. Only the map tiles come from the internet, and the page says so.

Top comments (0)