This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
"Go touch grass" is good advice that rarely works, because a walk with no destination loses to the couch every time. What gets people out of the house is a reason: a dog that needs walking, a kid who needs tiring out, a thing to go and see.
Sidequest gives a walk a reason. You tell it three things:
- where (your location, or one of four famous parks if you are trying it from a desk),
- how long (20 minutes, 40, or an hour),
- who is coming (just you, you and a kid, or you and a dog).
It reads OpenStreetMap around you, picks the most interesting real things within reach (a 500-year-old tomb, a statue, a bench with an inscription, a tree someone bothered to name), and routes the shortest loop through them that fits your time. Then Gemma, running on your device, writes you one small thing to do at each stop.
It is for anyone who wants to walk more and keeps not doing it, and for the parent at 4 pm on a Saturday who needs the next 40 minutes to have a plot. The quest is written at home on wifi and saved on the phone, so it still works in the middle of a park with one bar of signal.
Demo
🔗 vanshajpoonia.github.io/sidequest
No sign-up. Pick Lodhi Garden, Golden Gate Park, Hampstead Heath or Central Park, choose a length, and press Make my quest. Or press Start where I am and get a walk around your own block.
The first quest downloads Gemma once (about 860 MB, so use wifi). After that, on my laptop in Chrome with WebGPU, the three-stop quest above took 21 seconds to write, retries included. Every stop has a Walk there ↗ link for directions and a "Why this mission?" panel, which turns out to be the most important part of the app:
That is a real one from the quest above. Gemma's first idea for a radio mast was to climb it. The checker said no, told Gemma why, and the second draft kept you on the ground. The panel shows you all of it.
Code
VanshajPoonia
/
sidequest
Turns the real map around you into a short walking quest. OpenStreetMap stops, missions written by Gemma 3 running in your browser.
Sidequest 🧭
Turns the real map around you into a short walking quest.
Pick a start and a length (20 minutes, 40, an hour). Sidequest reads OpenStreetMap around you, picks the most interesting real things within reach (a 500-year-old tomb, a blue plaque, the 1894 Liberty Tree, a whispering bench), routes the shortest loop through them that fits your time, and then Gemma, running in your browser, writes you one small thing to do at each stop.
Live: https://vanshajpoonia.github.io/sidequest/
Stop 2 · 🏛️ Viaduct Bridge. Cross the Viaduct Bridge and examine the bridge’s construction – is it solid stone or a cleverly constructed timber?
The model writes. The code navigates.
Given only a place's name, a 1B model fills the gaps with confident inventions (the table below counts them). So Gemma is never asked where to go: choosing stops and adding up distances is arithmetic, and arithmetic belongs in…
Static files, no build step, no backend, no API keys. MIT.
How I Built It
The model writes. The code navigates.
The first decision was what not to ask Gemma. Picking stops and adding up distances is a job ordinary code does perfectly, so Gemma never chooses where you go. src/plan.js scores OpenStreetMap features by what their tags prove (has a name? an inscription? a Wikipedia link? a date?), keeps one of each kind where it can, and brute-forces the shortest start → stops → start loop, walked at 75 m a minute with a 1.3× allowance for paths not being straight, plus three minutes at each stop. If a loop does not fit your time, you are not offered it.
That leaves Gemma one job, the one a language model is actually good at: turning a list of facts about a place into a sentence a person wants to act on. "Race your kid to the English oak and see if the two of you can reach all the way around its trunk."
The hard part was getting it to do only that.
Measuring it
I took the six best stops around each of the four parks (24 in all, a mix of tombs, statues, plaques, trees, viewpoints and benches), cycled through solo, kid and dog walkers, and wrote a checker that flags a mission if it:
- mentions a year or a proper noun that is not in the stop's OpenStreetMap tags,
- asks the walker to find out a date the map does not have ("determine the original construction date" turned out to be Gemma's favourite filler),
- asks for anything unsafe (climbing, wading, picking, feeding, going inside),
- describes the place instead of giving an instruction, or talks to a kid on a dog walk,
- runs past 35 words.
Then I ran Gemma 3 1B (4-bit, the exact weights the browser loads) over all 24 stops in four configurations, with greedy decoding so the numbers come out the same on every run.
| Prompt | Invented a name or date | Asked for a date the map lacks | Not a mission, or wrong walker | Clean |
|---|---|---|---|---|
| The place's name only (rules + examples) | 15 | 0 | 2 | 8 / 24 |
| OSM tags + rules, no examples | 0 | 0 | 12 | 12 / 24 |
| OSM tags + rules + 4 examples as chat turns | 4 | 1 | 3 | 16 / 24 |
| + checker, retry with feedback, template (shipped) | 0 | 0 | 0 | 24 / 24 |
Nothing unsafe and nothing over-long turned up in any configuration, so those columns are left out.
Attempt 1: give it the name and let it use what it knows. This is the obvious prompt, and 15 of 24 missions invented something. Gemma knows of Lodhi Garden and Golden Gate Park, but not well enough to send someone anywhere real, so it fills the gaps with confident detail:
Lion (a statue in Golden Gate Park): "Search the Lion statue in Golden Gate Park for a discarded leash."
Richard Tucker (a bust near Central Park), for a parent and kid: "Search the Red Tomb for your kid's favorite toy, a red rubber ball, and see if it's visible from the entrance."
There is no leash, there is no ball, and there is no Red Tomb in New York. More on the Red Tomb shortly.
Attempt 2: give it the tags and a set of clear rules. The rules worked: no invented names at all. But only half the answers were missions. For a 1B model, rules written in prose are not instructions so much as text to continue, and it often just echoed the input back. Asked for a mission at the Athpula Bridge, it replied "Athpula Bridge". Twice it replied, in full, "Place:".
Attempt 3: show it, do not tell it. I kept the rules and added four worked examples, each as a real turn in the conversation (user: the facts, assistant: the mission) rather than as a paragraph of examples in the prompt. Small models copy the shape of what they have already "said" far more reliably than they follow instructions. Clean missions went from 12 to 16 out of 24.
They also copied things they were never supposed to copy. One of my examples was a made-up place called the Red Tomb, and it started turning up everywhere:
Karl Marx's original grave, Highgate, London: "Count the stones surrounding the Red Tomb and note the faint scent of lavender."
Attempt 4: do not trust it, check it, and tell it what was wrong. What ships runs every draft through the checker. If it fails, the draft goes back to Gemma with the specific reason (It mentions "Red", which is not in the facts. Write the mission again, using only the facts.). After three failures, a plain template takes over: "Find the view, then name the farthest thing you can see." A boring mission beats an invented one.
Result: 24 of 24 clean. Sixteen passed on the first try, five after being told what was wrong, and three fell back to a template.
Check the checker
That Karl Marx line was in my shipped results on the first full run, marked clean. The checker had compared words with a substring search, and the stop's inscription says the remains were buRIED, REmoved and re-inteRRED. As far as the checker could tell, "Red" was right there in the facts.
It now matches whole words, that case is a regression test, and the numbers in the table are from the fixed version.
It was not the only one. Taking the screenshots for this post, I watched the checker reject "Listen carefully to the faint hum emanating from the St. Columba Radio Mast" for "describing the place instead of giving you something to do". The full stop in "St." looked like the end of a sentence. Worse, the draft before it, "Ascend the tower of the St. Columba Radio Mast", had been rejected only by that same accident, because "ascend" was not on the unsafe list. Both are fixed and both are tests now.
I found every one of these bugs the same way: by reading the raw output to choose quotes for this post. I now think that is the most underrated debugging technique in AI work.
What the checker cannot catch
It catches names, dates, unsafe verbs and the wrong audience. It cannot catch an invented ordinary object: at a temple whose tags say nothing about lions, Gemma asked for "the most detailed carving of the lion's head", and that shipped. Some sentences that pass are just odd ("Listen carefully to the plaque"). And Gemma has no sense of occasion. At the spot in Delhi where Indira Gandhi was assassinated, its first draft asked a parent and child to "estimate the exact time of the assassination". The date rule rejected it, but by luck, not judgement.
So the app does not hide any of this. Every stop's "Why this mission?" shows exactly what Gemma was given, links to the OpenStreetMap object, and lists every draft that was rejected and why. If a mission is odd, you can see where it came from.
Choosing the model
- Gemma 3 1B instruction-tuned, 4-bit ONNX, through transformers.js on WebGPU, falling back to WASM. About 860 MB, cached after the first visit.
- Gemma 3 270M is about a quarter of the size, and in my tests it mostly repeated the prompt back. Too small for this job.
- Gemma 4 E2B is the obvious next step up, but its ONNX build is around 3.2 GB, which is too much to ask of a phone browser for a walking app. I did not ship what I could not run on a phone.
Privacy, by construction
"Start where I am" snaps your position to a roughly 500 m grid before asking Overpass for map data, pads the search radius to cover the snap, and does the real distance filtering on the device. The map server learns which neighbourhood you are in, not which house. The AI learns nothing, because the AI is a file on your phone.
Why Does Open Innovation Matter?
Sidequest is two open things working together, and both halves matter.
Open data is what makes the stops real. Every stop is an OpenStreetMap object that a volunteer mapped, often down to the text on a bench plaque. When a mission is wrong because the map is wrong, the "Why this mission?" link takes you to the exact object, and anyone can fix it. That fix then improves every app that uses OSM, not just mine. A closed maps API gives you a rating and a photo carousel, and if it is wrong, you file feedback and wait.
An open model is what makes the walk private and offline. Where you walk, when, and with whom is about as personal as data gets. With Gemma running in the tab, the place facts go into a model on your own device and a sentence comes out. There is no request to an AI service to log, and nothing breaks when the signal does, which in a park is often.
And I could take it apart. Every number in the table above came from running the same weights in Node that the browser downloads, against frozen map snapshots, with greedy decoding. Anyone can run npm test and get the same table. I could compare prompt strategies, swap in a smaller Gemma with one environment variable, and find out exactly where the model needed help. With a hosted model the thing I measured could change next week behind an endpoint, and every stop would cost money forever. Here it costs nothing.
My Agent Session
I built this with Claude Code as a pair programmer. I didn't save the session to DevRelay, so here is the next best thing: the whole evaluation is reproducible.
git clone https://github.com/VanshajPoonia/sidequest
cd sidequest && npm install && npm test
That runs 31 planner and checker checks, then Gemma 3 1B over the 24 stops in all four configurations, and prints the table above along with every failed draft. MODEL=onnx-community/gemma-3-270m-it-ONNX npm test runs the smaller model for comparison.
Prize Categories
Best Use of Gemma: Gemma 3 1B runs entirely in the browser and writes every mission, constrained to OpenStreetMap facts by a measured prompt, a checker and a retry loop.





Top comments (1)
tr.ee/dev-to