This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
I love to explore nature, and while I'm familiar with the minerals, fossils, and stones around my area, I wanted something that could help people that may not know as much, and more importantly, something that could be used completely offline so that it would still work even if you are far into a trail.
The idea is that a photo alone isn't normally enough to identify a rock, and even using flagship close models like gemini can give mixed results. There are several tests you can perform at the site or at home such as scratching it, weighing it in their hand, or putting vinegar on it to see if there is a reaction. This app takes a photo for a first guess, then walks you through those same physical tests one at a time:
- Can your fingernail scratch it? A coin? A nail?
- Rub it on an unglazed tile. What color is the powder?
- Does a magnet stick to it?
- Put a drop of vinegar on it. Does it fizz?
- Is it a shell with two halves? Is each half symmetrical?
Each answer updates the odds, and the app always asks whichever test will narrow things down the most next. Most identifications take 5 to 8 tests. Between questions the phone is in your pocket and the rock is in your hand, which is the part I like best: the screen is the shortest part of the experience.
The model knows 94 things: 53 minerals, 22 rocks and 19 fossils. It also knows where you're standing. It looks up the bedrock under you from open geologic map data, so in limestone country it leans toward calcite and crinoids, on granite toward feldspar and mica. It even knows the age of the rock. A trilobite can't turn up in rock that's 100 million years too young.
When you're done, you get a specimen label (type, hardness, streak, where to look for more) and can save the find with its GPS location. Finds export as GeoJSON, so you can map everything you've found if desired.
It's for anyone who picks up rocks on a hike, at a creek or on a beach and wonders what they're holding. That covers kids, rockhounds, and people like me with pockets full of crinoids.
Demo
Try it: [https://blue-math-567b.gavin-rose.workers.dev/]. Open it in Safari or Chrome on your phone and add it to your home screen. The first launch downloads the model (about 63 MB, Wi-Fi is recommended). After that, it works wherever you go! While a connection is available, it will use a larger hosted open weight model to give better results more quickly, and will fall back to the lightweight model when necessary or if you toggle the other one off.
I took it outside
I went out on a few walks near home, where the bedrock is shale and sandstone, and here's how it went.
Crinoids, three different ways. Loose stem discs in my palm: crinoid stems, 92%. A stem lying along a rock face: 87%. A slab of fossil hash: 88%. Each one took about five tests.
It got things wrong too. A small white chip came back as calcite, and it wasn't. One find only reached 19% for "conglomerate", which is really a shrug dressed up as an answer. And when I photographed nothing to test the app, it confidently named a rock anyway.
It crashed. On my iPhone the app downloaded the model, then Safari killed the page. In the field I had to time my taps to get in before it went down.
The questions felt off. The fossil-shape question ("stacked discs? coiled shell?") kept coming up before anything suggested a fossil, so for most rocks I was tapping "None of these".
Every one of those turned into a fix, covered below. Taking it outside was the most useful testing I did.
How I Built It
It is an installable web app (PWA) written in TypeScript with Vite. There's no app store and no account. The AI runs directly in the browser.
The photo model: SigLIP 2, on the phone
The photo step uses SigLIP 2, Google's open-weight image and text model, running fully on-device through Transformers.js and ONNX Runtime Web. On phones it runs as a 4-bit model (63 MB) on the CPU. On desktops with WebGPU it uses the fp16 version on the GPU.
I didn't start with SigLIP. I started with CLIP, then benchmarked five open backbones on the same images and the same held-out split, training the same small classifier on top of each:
| Backbone | Top-1 | Top-3 |
|---|---|---|
| CLIP ViT-B/32 (where I started) | 43.5% | 61.9% |
| CLIP ViT-B/16 | 48.6% | 66.3% |
| DINOv2 small | 47.7% | 62.4% |
| DINOv3 small | 51.2% | 67.0% |
| SigLIP 2 base | 56.0% | 72.9% |
SigLIP 2 won clearly, at the same 63 MB on the phone. Swapping it in took about an hour. That's only possible because every one of these models is open and runs in the same runtime.
Training it on rocks
Off the shelf, SigLIP 2 manages 29% top-1 across all 94 classes. I trained a small classification layer on top of it, using about 5,000 openly licensed images:
- Wikimedia Commons: Creative Commons photos of minerals and rocks, with attribution saved for every file.
- GBIF: CC0 and CC-BY photos of fossil specimens from natural history museums. These were the key to fossils, including plenty of the loose crinoid discs I actually find.
- My own field photos: always trusted and never filtered.
Web data is messy. A search for "granite" returns mountains, statues and kitchen counters. So the training script uses the model itself as a filter: it drops any image that looks more like a landscape, building or diagram than a rock specimen. Museum photos had the opposite problem: a tiny fossil next to a big ruler and a color card. For those, the script tries a grid of crops and keeps the one that looks most like a fossil and least like a ruler.
The trained model gets 62.3% top-1 and 78.6% top-3 on held-out images, up from 29% and 45% for the untrained model. Top-3 matters most, because the photo is only a starting point and the physical tests do the rest.
The tests: odds that update with every answer
Each of the 94 entries has its properties written down: hardness range, streak color, luster, how it breaks, whether it's magnetic or fizzes, and for fossils, shape and age. The app keeps a probability for every candidate and updates it with each answer. Answers are treated as noisy, because people misjudge luster and vinegar is a weak acid.
To pick the next question, it calculates which one is expected to shrink the uncertainty most (information gain). A few rules keep it sensible. It always opens with a broad "What did you find?", and fossil-only questions wait until the evidence actually leans toward a fossil.
Geology from where you're standing
With your GPS position (GPS works without signal), Streak looks up the bedrock in Macrostrat, an open geologic map database. It turns rock types into odds: a granite map unit boosts quartz, feldspar and mica, and "crinoidal limestone" in the description boosts crinoids directly.
It never rules anything out. Creeks, glaciers and landscaping gravel carry rocks for miles, so local geology only boosts candidates, and there's a "Found it somewhere else?" switch. Fossils also get an age check: outside their time range they're down-weighted, unless you're standing on river gravel, which can hold fossils of any age. Before a trip you can save the geology for about 20 km around you, so it works offline too.
Optional online mode: Llama 4 Scout
With signal, Streak can also send the photo to Llama 4 Scout, Meta's open-weight vision model, running on Cloudflare Workers AI. The request goes through the app's own Cloudflare Worker, so the phone only ever talks to the app's domain. It sends the photo and the name of the bedrock, never your coordinates. The larger model's answer is blended with the on-device one, and if anything fails, it quietly falls back to offline. A toggle on the home screen turns it off.
Fixing what the field test found
- The crash: iOS was killing the page because loading the 176 MB fp16 model needed several copies in memory at once. Phones now get the 4-bit model, and there's a crash guard: if the page dies while loading the model, the next launch automatically drops to the lightest setting instead of crashing in a loop.
- "That's a rock" about nothing: before identifying, the app now checks whether the photo shows a rock, mineral or fossil at all, by comparing it against descriptions of both specimens and non-specimens (an empty hand, the sky, a room). If not, it says so. Any result under 35% confidence now shows as Unknown with its closest matches, instead of a confident wrong name.
- Fossil questions too early: fixed by the question ordering described above.
- Rare minerals coming up too often: a rarity prior. Halite dissolves in rain, so it almost never lies on the ground outside deserts, and the app shouldn't jump to it.
Why Does Open Innovation Matter?
For this app, open models weren't just a nicer choice. Several features only work because of them.
It works with no signal. Good rocks are in road cuts, creek beds and canyons, where you usually have no bars. A closed vision API is useless there. Because SigLIP 2 is open, its weights sit on the phone and run in the browser. The geology lookup is open data too, cached ahead of time.
I could host the weights myself, and that saved the project. On the first real test, the model wouldn't download. My network was blocking huggingface.co entirely. With a closed API that would have been the end of it. Because the weights are open, I bundled them into the app and served them from my own Cloudflare site, split into 20 MB pieces the browser stitches back together. Now the app works on any network that can load it.
I could swap models and measure. I benchmarked five open backbones on my own data and switched to the winner for an 11-point gain, without asking anyone's permission or changing any API contracts.
I could train it on my own rocks. Museum and web photos don't look like a phone photo of a dusty pebble in your palm. With open weights I train my own layer on top of the model and keep adding my own field photos. Every crinoid I photograph makes the next ID better.
Your location stays on your phone. Offline mode sends nothing anywhere. Online mode sends a photo and a bedrock name to a model I chose and deployed, not your coordinates to a company whose retention policy I'd have to trust. Where you find good rocks is something collectors like to keep to themselves.
It costs nothing to run. On-device inference is free. Hosting is on Cloudflare's free tier, and online mode fits in Workers AI's free daily allowance, roughly 120 to 160 photo IDs a day.
Everything underneath is open. Open models (SigLIP 2, Llama 4 Scout), open runtimes (Transformers.js, ONNX Runtime), open data (Macrostrat, GBIF, Wikimedia Commons). It's also honest about its limits: it shows you the odds, the closest alternatives, and when it doesn't know.
My Agent Session
I built the app with the help of Claude, where I asked the agent to do the user interface and wiring so I could focus on the logic and the AI under the hood.



Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more