This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
TrailMind is an offline-first Progressive Web App (PWA) that lets you photograph any bird, plant, mushroom, or wildlife on the trail and get an instant AI-powered identification — with zero internet connection, zero cloud accounts, and zero monthly fees.
Point your camera at a mystery mushroom. Get the species name, key identifying features, whether it's edible or deadly, and where it typically lives. All of that runs locally on your laptop or phone, powered by open-weight vision models through Ollama.
The goal was simple: make the AI part of the experience last less than 30 seconds, so you spend the rest of your hike actually looking at things.
Who is it for? Hikers, naturalists, birders, foragers, parents with curious kids, anyone who has ever pointed at something and said "what IS that?" It's for anyone who wants to learn from nature in real time — without pulling out a phone, searching an app store, creating an account, or burning mobile data on a remote trail.
Demo
🌐 Run it yourself locally (30 seconds):
git clone https://github.com/devkumar2313

/trailmind
cd trailmind
node server.js
# → Open http://localhost:3000
Then set up Ollama once:
# Install: https://ollama.com
ollama pull llama3.2-vision
OLLAMA_ORIGINS=* ollama serve
That's it. Take a photo outside, upload it, and get your identification.
The entire AI runs on your machine. No API key. No signup. No data sent anywhere.
Code
TrailMind 🌿 — Offline Nature AI Companion
Identify birds, plants, mushrooms & wildlife with open-source AI — no internet required on the trail.
What It Does
TrailMind is a Progressive Web App (PWA) that uses locally-running open-weight AI vision models (via Ollama) to identify nature in your photos. Take a picture of a bird, plant, mushroom, or any wildlife — TrailMind gives you:
- 🦅 Species name (common + scientific)
- 🔬 Key identifying features
- 📚 Interesting ecological facts
- 🥾 Trail notes (habitat, season, edibility/safety)
The entire AI runs on your device. Your photos never leave your laptop or phone.
Why Open Source AI?
| Closed API | TrailMind (Open-source) |
|---|---|
| Sends photos to corporate servers | Photos never leave your device |
| Requires internet on the trail | Works offline, mid-hike |
| Monthly subscription costs | Runs free, forever |
| Black-box model | Swap models, fine-tune, inspect |
| Rate-limited | Unlimited queries |
Tech Stack
| Component | Technology |
|---|---|
| AI Inference | Ollama — local open-weight |
Project structure:
trailmind/
├── public/
│ ├── index.html # The entire PWA — vanilla HTML/CSS/JS, zero npm dependencies
│ ├── sw.js # Service Worker for offline caching
│ └── manifest.json # PWA install manifest
├── server.js # Minimal Node.js static file server (stdlib only)
├── package.json
├── LICENSE # MIT
└── README.md
The frontend is intentionally dependency-free. No React, no bundler, no node_modules to wrestle with. It's a single HTML file that runs in any browser, installs as a PWA, and talks to a local Ollama server.
How I Built It
The Open-Source AI Stack
Ollama is the hero here. It's an open-source runtime that makes running large language models (including vision models) as simple as ollama pull llama3.2-vision. It exposes a local REST API on localhost:11434 — which means a plain browser fetch() call is all you need to talk to it.
The app supports four open-weight vision models:
| Model | Size on disk | Best for |
|---|---|---|
llama3.2-vision (Meta, Apache 2.0) |
~8 GB | Best accuracy, richest descriptions |
llava (LLaVA Team, Apache 2.0) |
~4 GB | Great speed/accuracy balance |
moondream (Vikhyat Korrapati, Apache 2.0) |
~1.7 GB | Works on older/low-RAM hardware |
llava-phi3 (Microsoft, MIT) |
~2.9 GB | Excellent on CPU-only machines |
All four are Apache 2.0 or MIT licensed. You can inspect them, fine-tune them, distill them, or swap them entirely — something a closed GPT-4V call can never give you.
The Architecture
┌─────────────────────────────────────┐
│ TrailMind PWA │
│ (Browser — fully cached offline) │
│ │
│ 📸 Camera/Upload → Base64 encode │
│ ↓ │
│ 🔍 POST /api/chat (localhost) │
└──────────────┬──────────────────────┘
│ (stays on your machine)
┌──────────────▼──────────────────────┐
│ Ollama Server │
│ (Running locally on your device) │
│ │
│ llama3.2-vision / llava / etc. │
│ Open-weight, Apache 2.0 / MIT │
└─────────────────────────────────────┘
The browser takes a photo → converts it to Base64 → sends it to localhost:11434/api/chat with a structured naturalist prompt. The model replies with species name, confidence, key identifying features, ecological facts, and trail notes. The entire round trip stays on your hardware.
The Offline Part
A Service Worker caches the entire app shell after the first load. Once cached, TrailMind loads and renders with zero network — even in airplane mode. The only thing that needs to be running is the local Ollama server, which doesn't need internet either once the model is downloaded.
History is stored in localStorage. Your identifications persist across sessions. No server, no database, no sync.
The Prompt Engineering
Getting a vision LLM to behave like a field naturalist took some iteration. The final prompt structures the output with Markdown headings that the app renders nicely:
You are a field naturalist expert. Analyze this photo taken outdoors.
**Species / Subject**: [Common name and scientific name]
**Category**: [Bird / Plant / Mushroom / Insect / Mammal / Other]
**Confidence**: [High / Medium / Low] — explain briefly why
**Key Identifying Features**:
- 3-5 observable features
**Interesting Facts**:
- 2-3 ecological facts
**Trail Notes**:
- Season, habitat, safety notes
This structure means the output is useful in 5 seconds — you can read the species name and confidence at a glance and pocket your phone.
Why Does Open Innovation Matter?
Let me give you the concrete version:
1. It works where you are.
I live in India. The best hikes here have no cell signal for hours. A closed-API app that requires an internet call for every identification is useless in those conditions. TrailMind works on a dirt trail with no signal because the model runs on your device. Open-weight models that can be pulled locally made this possible. A commercial API could not.
2. Your photos don't become training data.
When you photograph a bird and upload it to a closed service, that image may be retained, logged, or used to improve their model. When you use TrailMind, the image goes from your camera → your RAM → your Ollama process → your screen. It never leaves your machine. That's not a policy promise — it's an architectural guarantee, made possible by open-source AI.
3. The cost is zero — forever.
Closed vision APIs charge per API call. At scale — say, a school field trip program using this for 30 kids identifying 20 things each — that's hundreds of API calls, real money, and a dependency on a company's pricing decisions. TrailMind has no per-query cost after the model is downloaded. Open-source AI made it viable for education and conservation without a funding model.
4. You can change it.
Want to swap llama3.2-vision for a fine-tuned model trained specifically on regional bird species? Do it. Want to add a RAG layer that pulls from iNaturalist records? Build it. The model isn't a black box you're calling through an API — it's a weight file you own. Open-weight AI means the project can evolve in ways a closed API never could.
The screen should be the shortest part of the experience. Open innovation is what made that possible.

Top comments (0)