DEV Community

Pranesh Nikhar
Pranesh Nikhar

Posted on AI-assisted

WanderLog: an offline field journal that answers "where did I see that kingfisher last time?"

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission ๐ŸŒฟ

I have a habit of seeing things outside and then losing them to my own memory. The peacock that called from the fog. The kingfisher that dove twice off the jetty and came up with a fish. The lizard big enough to hiss like a bicycle pump.

So I built WanderLog โ€” a local-first field journal that turns my coding agent into a naturalist's notebook. It's a Model Context Protocol (MCP) server. I go outside; when I get back I record a thirty-second voice memo of what I saw, and one command turns it into structured sightings and a walk. Everything else โ€” remembering, finding, connecting โ€” happens on my own machine, with open-source AI, with no network.

Stack: TypeScript ยท MCP (Model Context Protocol) ยท SQLite ยท ONNX Runtime ยท Transformers.js ยท open-weight embeddings (all-MiniLM-L6-v2) ยท local gemma3:4b via Ollama ยท ElevenLabs Scribe (optional) ยท Sentry gen_ai tracing
Code: github.com/praneshnikhar/wanderlog

What I Built

WanderLog is a field journal for the person who notices things outside and forgets them: the birder without a checklist, the hiker who wants to remember which trail had fog on the lake, the person who sees a peacock on the commute and wants to know whether it's the same one from last month.

It's built around one rule: the screen should be the shortest part of the experience. There's no dashboard to maintain, no logbook to transcribe, no "curated profile of your nature observations." There's a walk, a thirty-second voice memo, and a question you can answer a month later.

On the trail, you just talk to your phone like you're telling a friend what you saw. At home, one command:

$ npm run log-voice -- ~/walk.m4a
transcribing walk.m4a with ElevenLabs...
transcript (scribe_v2, eng):
  "Saw a white-breasted kingfisher at the lake jetty. It dove twice and got a fish.
   Also spotted a rain lily on the path side, fresh after last night's rain.
   Walked the Central Park loop, three point two kilometers in fifty minutes. Fog on the lake."
logged walk #1
a walk on Central Park loop, 3.2 km, 50 minutes, fog on the lake. Walked on 2026-10-08.
logged sighting #1
bird white-breasted kingfisher, at lake jetty, dove twice and got a fish. Spotted on 2026-10-08.
logged sighting #2
plant rain lily, at path side, fresh after last night's rain. Spotted on 2026-10-08.
Enter fullscreen mode Exit fullscreen mode

No typing, no forms. The memo became a walk, two sightings linked to that walk, and three searchable journal entries โ€” and the spoken "three point two kilometers" became a real number.

Then the journal answers questions about itself, from the same local model:

$ WANDERLOG_OLLAMA_MODEL=gemma3:4b npm run demo

  "where did I see the kingfisher last time?"

  1. [2026-10-03] sighting #10 โ€” white-breasted kingfisher at Maota Lake, Amber
     score=0.465 semantic=0.500 keyword=0.186

  Answer from local gemma3:4b:

  According to entry [1] (2026-10-03), the white-breasted kingfisher was last
  sighted at Maota Lake, Amber. The note specifically states "bird
  white-breasted kingfisher, at Maota Lake, Amber".
Enter fullscreen mode Exit fullscreen mode

It doesn't search for the word "kingfisher" โ€” it searches for what the entry means. "Where did I see it last time" and "bird with the blue flash near the lake at golden hour" both land on the same entries, because the journal is understood by the same open-weight model that understands you. And the answer above was written by a 4B model on the same laptop: it cited the entry number, parsed the date, and got the place right.

Demo

The full loop on a real walk โ€” memo recorded on the trail, logged at home, question answered โ€” is being field-tested on Friday, October 9. The video and an honest field report will land right here after the walk:

Everything in the terminal sessions above is real output from the actual project, not mock-ups: the voice run is a real transcription of a real audio file, and the search output comes from the seeded journal.

Code

WanderLog

A local-first field journal for people who spend time outside.

WanderLog is a Model Context Protocol (MCP) server that turns your coding agent into a naturalist's notebook. Log the birds, plants, and trails you see on a walk; later ask in plain language "where did I see the kingfisher last time?" โ€” and get an answer computed entirely on your machine.

  • Open-source AI at the core. Semantic search runs on an open-weight sentence-transformer (Xenova/all-MiniLM-L6-v2, Apache-2.0) executing in-process via Transformers.js and ONNX Runtime. Everything that understands your journal โ€” Q&A and voice-note parsing โ€” runs on Gemma (gemma3:4b) through Ollama, on your machine. No API keys, no cloud.
  • No internet required. After the first run, the embedding model is cached and every query is answered with zero network calls. Your sightings, notes and locations never leave the SQLite file on your disk.
  • Voice-first. Aโ€ฆ

MIT licensed, one npm test away from reproducible: a hermetic smoke test drives every tool over real MCP stdio, plus unit tests for the voice pipeline that stub both the ElevenLabs and Ollama calls.

How I Built It

flowchart LR
    subgraph trail [On the trail]
        A[Voice memo<br/>no typing]
    end
    subgraph home [At home, on-device]
        B[Transcribe<br/>ElevenLabs Scribe ยท optional]
        C[Parse<br/>gemma3:4b via Ollama]
        D[(SQLite journal<br/>sightings ยท walks ยท vectors)]
        E[Semantic search + answers<br/>MiniLM embeddings ยท gemma3:4b]
    end
    A --> B --> C --> D --> E
    A -.->|--transcript: no network at all| C
  • Storage โ€” one SQLite file (~/.wanderlog/journal.db): sightings, walks, and a small embeddings table that stores each entry's vector as a blob. At field-journal scale (hundreds to low thousands of entries) a vector database would be ceremony, not engineering โ€” cosine similarity over the table runs in milliseconds.
  • Embeddings โ€” @huggingface/transformers loads an open-weight ONNX sentence-transformer (all-MiniLM-L6-v2, Apache-2.0) and runs it in-process with ONNX Runtime. First run downloads ~23 MB once; after that, every query is answered with zero network calls. If the model can't load, search degrades gracefully to keyword overlap โ€” the journal still works at the edge of the trail.
  • Voice โ€” the new piece. A memo is transcribed (ElevenLabs Scribe by default, optional), then a local open-weight model parses the transcript into JSON โ€” walk, sightings, places, notes โ€” validated against a zod schema. If the model invents a category the schema doesn't know, the entry is kept and bucketed as other rather than dropped. Pass --transcript and the entire pipeline runs offline; no bytes leave the machine. One honest wrinkle: Scribe runs zero-retention by default (enable_logging=false), which is an enterprise feature on ElevenLabs' side โ€” on smaller plans the CLI says so up front and offers WANDERLOG_STT_LOGGING=1 (explicit retention) or the fully local path instead of silently downgrading your privacy.
  • Q&A โ€” answer_question retrieves the top evidence locally, then asks your local Ollama model (I tested llama3.2:3b and Google's open-weight gemma3:4b) to answer strictly from those entries, citing entry numbers. No Ollama? It returns the ranked evidence, which for a field journal is usually answer enough.
  • Ranking โ€” hybrid score: embedding similarity (75%) merged with keyword overlap (25%), with a freshness tiebreak so a "when did I last see X" question prefers the recent entry over an equally-relevant older one.
  • Tracing โ€” every journal operation is instrumented with Sentry's gen_ai semantic attributes: tool calls are gen_ai.execute_tool spans, transcription is gen_ai.transcription, parsing and Q&A are gen_ai.chat with token counts, and search and embeddings are spans too. Set SENTRY_DSN and you can watch the agent think.

Measured on my MacBook (no GPU fan spin-up, no cloud): cold model load 244 ms, one embedding 3 ms, full semantic search over a 16-entry journal 9 ms. The only thing slower than that is me, on the trail.

Why Does Open Innovation Matter?

The closed version of this product would be an app with a cloud account. And the cloud version of a field journal is exactly wrong, for four reasons:

  1. Location data stays private. Every sighting has a place attached โ€” sometimes precise coordinates. eBird and iNaturalist are amazing, but they ask me to send every observation (and where I was standing) to their servers. WanderLog keeps the notebook in a file I own. Not a server I don't control โ€” a file I can cat, back up, or delete with my own hands.
  2. The model is swappable because it's open. WANDERLOG_EMBED_MODEL and WANDERLOG_OLLAMA_MODEL are environment variables. Don't like MiniLM? Point it at another open-weight embedder and the journal starts understanding you differently โ€” no vendor change, no data migration, no "new API version." That's what open weights buy you: your data is interpreted by models you choose, not by a company's pricing page.
  3. It costs nothing and runs anywhere. A laptop, a Raspberry Pi, a 2020 MacBook โ€” the whole inference stack is ONNX Runtime plus a 23 MB model, and the optional Q&A and parsing layer runs on gemma3:4b, a 4B open-weight model that fits comfortably in RAM. There is no per-token bill, no rate limit, no "free tier."
  4. You can change how the agent behaves. It's an MCP server; the journal-keeper prompt, the tools, and the scoring are all source in a repo you can fork. Fine-tuning your agent is editing the prompt โ€” the whole behavior contract is inspectable.

The one cloud call in the whole system โ€” transcription โ€” is opt-in in both directions: you can use it, skip it (--transcript), or swap it, and everything downstream of the transcript is open weights on your disk. The open approach won precisely where it matters: everywhere the network ends. The trail is the last place on earth that has birds but no bandwidth, and WanderLog was designed for exactly that boundary โ€” the model is already on the disk, the data is already on the disk, and the two of them never need to meet a server.

The tools

Tool What it does
log_sighting Record a species, category, place, weather, field notes
log_walk Record an outing: trail, distance, duration, notes
log_voice_note Voice memo (or raw transcript) โ†’ transcription + structured entries
search_journal Hybrid semantic + keyword search, computed on-device
answer_question Local-Gemma answer with citations; evidence fallback
journal_overview Totals, category counts, last outing

Plus MCP resources (journal://stats, journal://day/{date}, journal://category/{category}, journal://walk/{id}), so any client can browse the notebook like files.

My Agent Session

The build โ€” including verifying the voice pipeline with a real memo and catching the zero-retention tier limit the hard way โ€” is saved on DEV as an agent session:

WanderLog: verifying the voice pipeline end-to-end and catching a zero-retention 403
You

My Mac lagged and I had to force restart โ€” I think you were in the middle of verifying the voice extension?

Agent

Picking up where we left off: re-running the typecheck and the full test suite (voice unit tests + MCP smoke test), then the real end-to-end run with gemma3:4b.

Agent

All green. Next: the on-device path through scripts/log-voice.ts with gemma3:4b, writing to a temp DB so the real journal is untouched.

You

Here's an ElevenLabs API key so you can test a real voice memo. (Key omitted from this transcript for security.)

Agent

Generating a real voice memo with macOS text-to-speech, then running the full audio โ†’ transcript โ†’ structured entries pipeline.

Agent

That's a real finding, not a fluke: the code asks for zero-retention (enable_logging=false) by default, and ElevenLabs gates that to enterprise/trial plans. Confirming the rest of the pipeline works when logging is explicitly allowed, and that the user is told the tradeoff instead of it happening silently:

Agent

Now fixing the failure mode: a 403 from the zero-retention gate must become an actionable error โ€” never a raw HTTP dump, and never a silent downgrade of the privacy default.

Agent

Adding a unit test that simulates the ZRM 403 and asserts the helpful message.

Agent

Re-running everything, plus a live check against the real API to see exactly what a user now gets.

You

Commit it, and don't worry about rotating the key.

Agent
You

Pasted the DEV submission template for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.

Agent

Pulled the challenge rules, confirmed the project's first commit (Oct 6) is inside the challenge window, then restructured the existing DEV draft to the template: voice-first story, gen_ai tracing, honest zero-retention wrinkle, and the three prize categories it genuinely qualifies for.

Prize Categories

WanderLog genuinely uses three of this week's partner technologies, so I'm entering all three:

  • Best Use of Gemma โ€” gemma3:4b via Ollama is the intelligence of the project: it parses spoken notes into structured field entries and answers journal questions with citations, strictly from retrieved evidence. It runs on the same laptop that holds the journal.
  • Best Use of ElevenLabs โ€” Scribe is the transcription leg of the voice flow, and the only cloud step. Zero-retention by default, optional by design (--transcript skips it entirely), and honest when the plan can't honor zero-retention instead of quietly logging your audio.
  • Best Use of Sentry Agent Tracing โ€” every operation the agent performs emits gen_ai.* traces to Sentry when SENTRY_DSN is set: tool calls, transcription, chat with token counts, embeddings, and search. It's how I watched a 4B model on a laptop behave like an agent.

What's next

I'm taking WanderLog to Central Park, Jaipur this week for a real field test โ€” logging actual sightings from a voice memo, and asking the journal questions at the end. This post will be updated with the field report and the video.

Built in three days, entirely inside the Hacktoberfest Week 1 window, MIT licensed, with a hermetic test suite that exercises every tool over real MCP stdio.

Open source made this project possible โ€” the model is open, the runtime is open, and the notebook is yours. Go outside. Take notes. Ask it questions later.


Built for Hacktoberfest 2026 Open-Source AI Challenge: Week 1 โ€” Touch Grass. Built on open-source AI: Transformers.js + open-weight embeddings, Gemma 3 via Ollama, MCP, SQLite โ€” all open, all local, all yours.

Top comments (0)