DEV Community

Cover image for Pocket Field Guide: an AI that wants you to put your phone away
Arsh Goyal
Arsh Goyal

Posted on

Pocket Field Guide: an AI that wants you to put your phone away

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

Most nature apps are built like every other app: they want your attention. You point the camera at a bird and get an answer, and then you're scrolling a species page while the bird flies off.

Pocket Field Guide is built the other way round. It's a browser app with two models running entirely on your device, and every screen is designed to send you back outside.

  • Go outside. Tell it you have 30 minutes in a park. Gemma writes four things to find ("Touch the rough surface of a mossy rock", "Listen for a bird singing from a rooftop"). Then the app switches to a plain screen with a timer and a checklist, and tells you to put the phone away.
  • Identify. When you spot something, take one photo. SigLIP names it from a built-in field guide of about 150 common species. Gemma then writes a short field note that doesn't tell you facts to read: it tells you what to check in person to confirm it, and one related thing to go find next.
  • Journal. Finds go into an on-device journal, with an optional location that never leaves the browser. You can export it as Markdown.

There's no account, no API key and no server. After the first download it works in airplane mode, which is the point: the places worth walking usually have the worst signal.

It's for anyone who wants a reason to go for a walk and a little help noticing things on the way. It's made for beginners.

Demo

Try it: pocket-field-guide.onrender.com. Use Chrome or Edge on a laptop. Open Setup, download the models once (~900 MB), and then it works offline.

New to it? The built-in guide explains every feature, how to check that offline mode works, and troubleshooting.

Code

GitHub logo Eternity0207 / pocket-field-guide

On-device AI field guide that gets you outside. Gemma 3 + SigLIP running in the browser, works offline.

Pocket Field Guide 🌿

An AI field guide that runs entirely in your browser and is built to get you off the screen.

  • Go outside: pick 15, 30 or 60 minutes and where you are. Gemma writes a short list of things to look for, then the app switches to a quiet walking screen with a timer and a checklist.
  • Identify: take or drop in a photo of a plant, bird, bug, mushroom or leaf. SigLIP matches it against a built-in field guide of about 150 species, and Gemma writes a short field note: what to check in person to confirm it, plus something related to find nearby.
  • Journal: finds are saved on your device with an optional location, and you can export them as Markdown.

No account, no API key, no server. After the first download it works with no internet at all.

New to the app? The…

How I Built It

The whole app is static files: Vite and vanilla JavaScript, with no framework. All the AI runs in a Web Worker through Transformers.js and ONNX Runtime Web.

Two open models, each doing what it's good at

Job Model Where it runs
"What is this?" SigLIP base (Google, Apache-2.0), 8-bit vision tower, ~95 MB CPU (WASM, all threads)
Field notes, walk ideas Gemma 3 1B instruct, 4-bit, ~785 MB GPU (WebGPU)
Fallback when there's no GPU Gemma 3 270M, fp16 CPU

Identification is zero-shot. Every species in labels.js is described by three short phrasings ("a photo of a hoopoe, a type of bird.", "a close-up photo of a hoopoe."...). I embed those with SigLIP's text encoder at build time, average them per species, and ship the result as a ~495 KB file. In the browser, identifying a photo is one pass through the vision encoder plus a dot product. It takes about 1–2 seconds on my laptop's CPU.

I started with CLIP ViT-B/32 and it wasn't good enough: on the first 8 test photos it got 4 right as its top guess, and it called a hoopoe a starling. Swapping in SigLIP was a few lines of code (same library, same API) and it got 15 of 16 Wikipedia lead photos right. It's an open model, so changing it was easy to try.

Keeping a 1B model honest

The first version simply asked Gemma 1B to "write a field note about a European robin". It confidently wrote that the robin has a white throat patch. It doesn't. Small models are fluent and wrong, and in a nature app that's worse than useless.

So Gemma never recalls facts any more. The fix has four parts:

  1. Trusted traits. Every species in the field guide has one line of ID traits written by hand ("pinkish-orange body; black-and-white striped wings; fan-like crest tipped black..."). Gemma only gets to rephrase those into checks you can do in person.
  2. A worked example and a prefill. The prompt includes one complete example note, and the answer is started for the model with **Check it:**. That stopped the "Okay, here's a field note for October 26, 2023..." openings completely.
  3. A check on every answer. Each note is checked: both sections are there, the "Check it" line uses words from the trusted traits, and it didn't echo the prompt. If it fails, the app builds a plain note straight from the traits instead.
  4. Safety isn't generated. "Never eat a wild mushroom identified by an app" comes from a fixed list per group. I don't want a 1B model improvising mushroom advice.

The result is notes like this (Gemma 3 1B, on the GPU, about 10 seconds):

Check it: Observe thin leathery brackets in fan shapes, colored with vibrant bands. Tiny pores are visible beneath.
Look nearby: Search for dead wood with colourful bands, noting the subtle pores.
Careful: Never eat a wild mushroom identified by an app.

Making it work offline

Getting this to run offline took the most work, and I learned a few things along the way:

  • Big files and the Cache API don't mix well. Transformers.js caches models in the Cache API by default, and the 763 MB Gemma weights file kept failing to save. I wrote a small Cache-shaped store on the Origin Private File System instead. It streams to disk and writes atomically, so a half-finished download never looks like a valid model.
  • The 4-bit Gemma builds don't run on the CPU backend. ONNX Runtime Web has no WASM kernel for the quantized embedding op they use. Worse, a failed session load left the runtime unable to create any new session in that worker. So the app never "tries and sees": it plans the device up front, and if the GPU path fails it restarts the worker and comes back on the CPU with Gemma 270M.
  • Threads need headers. WASM only uses all CPU cores when the page is cross-origin isolated (COOP: same-origin, COEP: credentialless). On my laptop that took identification from about 2.7 s to about 1.7 s.
  • The ONNX Runtime WASM is served from the app's own origin, not a CDN, and a service worker caches the app shell. Once the models are downloaded, nothing needs the network.

Hosting

The app is hosted on Render as a free static site, set up by a render.yaml in the repo. The one setting that matters is the cross-origin isolation headers above: without them the CPU path runs on a single thread. Every push to main redeploys automatically.

Why Does Open Innovation Matter?

I couldn't have built this on a closed API, for four reasons.

It has to work with no signal. Trails, riverbanks and big parks are exactly where coverage drops. A cloud model would fail in the one place this app is meant for. Open weights let me ship the model to the trail.

Where you walk is private. A journal of your locations with timestamps is sensitive data. With local inference there's no server to leak it, subpoena it or train on it. The privacy promise is simple to check: open DevTools and watch the network tab stay empty.

It costs nothing to run, so it can afford to be generous. No per-photo fee means no reason to ration identifications or put a paywall in front of a walk. Hosting is a free static site.

I could look inside and change things. When CLIP wasn't accurate enough, I swapped in SigLIP in an afternoon. When I wanted identification to be instant, I split the model and precomputed half of it at build time. That's only possible because I have the weights, not just an endpoint. The same applies to anyone using it: the field guide is a plain file, so a birder in Kerala or a naturalist in Oregon can add their local species, run npm run embed, and the model knows them right away, with no fine-tuning.

The closed version of this app would be easier to build. It would also stop working on the trail, cost money for every photo, and send everyone's walking routes to someone else's server.

Taking It Outside

I didn't want to build another app that keeps you on your phone. I wanted one that is happy when you close it.

So my rule while building was simple: if a feature makes you look at the screen longer, it doesn't go in. No feed, no streaks, no notifications. Just a short list, a timer, and a nudge to go and look closer.

The best result for this app isn't more time in the app. It's more time outside.

Prize Categories

  • Best Use of Gemma: Gemma 3 (1B on WebGPU, 270M on CPU) runs locally in the browser and writes every field note and walk plan, grounded in trusted traits and checked before display.
  • Best Use of Render: the app's front end is hosted on Render as a static site, with a Blueprint that sets the cross-origin isolation headers multithreaded in-browser inference needs.

Top comments (0)