What I Built
The screen shows one line:
I am stone, flat, and written on. One of the dates on me is 1956.
Then it goes away. You put the phone in your pocket, walk to a stele in the walled city of Ahmedabad, and find it. The phone confirms you are looking at the right thing and tells you what the inscription actually says.
Sphinx turns the map of your own neighbourhood into a riddle walk. You give it a town. It finds real things mapped there — statues, blue plaques, drinking fountains, war memorials, stepwells — and writes a two-line riddle for each, describing only what you can check with your own eyes when you arrive. You walk the walk with no signal, find each thing, and unlock a real fact about a place you have probably passed a hundred times without looking.
Then there is the twist. Every stop ends with an option that is not "I've found it": "There is nothing here." If the fountain has been removed, or the plaque vandalised, or the statue is not where the map says, the child taps that — and Sphinx drafts a precise, polite OpenStreetMap note for a grown-up to send.
That is the educational payload, and it is the reason I built this. The lesson is not "look at the pretty statue". It is the map is made by people, and people can be wrong, and you can fix it. A child who reports one wrong drinking fountain at ten has learned something most adults never do.
I chose children as the audience deliberately, and it changed almost every decision. There is no score, no streak, no leaderboard, no notification and no confetti. The reward for finding something is a fact. After twenty seconds of reading, the screen is allowed to dim itself almost to black, because a child walking around at dusk staring at a glowing phone is not touching grass. The product metric is not time-in-app; it is time-to-pocket.
Demo
Watch the 4-minute walkthrough →
The demo is a real hunt, built from live OpenStreetMap data — npx tsx scripts/hunt.ts "Ahmedabad, Gujarat, India" returned 17 candidates and 7 playable stops, and that hunt is committed to the repository as data/riddles/ahmedabad.json. The riddles in the film are not sample text.
Three things make it a demo rather than a pitch:
- It works with no API key, no account and no network. The hunt is baked into the repo and seeded into IndexedDB, read by the same player, behind the same consent gate. There is no separate demo path — a demo that shares no code with the product proves nothing.
- Measured in the browser, not asserted: 8 requests on load, 0 cross-origin. 0 requests of any kind during the entire play loop.
-
Every number in this post is produced by this repository, and
npm run check:claimsfails the build if any of them drifts.
What works where. The GitHub Pages build is the static export: the baked Bristol, Ahmedabad and Madison hunts play in full, offline, including the photo check's degraded path. Building a live hunt needs the /api/hunt proxy to Overpass, which does not exist in a static export — run npm run dev locally for that. The deployed demo is deliberately the part that needs no infrastructure.
The play loop
The twist
The child is told: "You found something the map got wrong." Not "wrong answer". Not "try again". They have done something genuinely valuable, and every word on that screen exists to protect the feeling. The draft is shown in full before anything is sent, a grown-up presses send, and OpenStreetMap's own policy — "notes should be a human-to-human communication" — is respected by construction: there is no queue, no batch and no auto-post path anywhere in the codebase.
Code
github.com/0dvsh0/sphinx · npm install && npm run dev
No test framework, no ORM, no analytics SDK, no UI kit. The whole app is Next.js, React, and the transformers.js runtime that loads the on-device vision model. Five verification suites run as assertions in plain scripts:
npm run check # typecheck → safety → contrast → vision → notes → build → privacy → claims
| Suite | What it proves |
|---|---|
check:safety |
125 assertions. Riddles never instruct movement, never name the target, never invent a fact. Consent gates behave. The stranger script is intact. |
check:privacy |
No third-party call site exists in the source, the service worker refuses other origins, no analytics package is installed. |
check:claims |
35 assertions auditing this post against the dataset, the evaluation JSON and the source files. |
audit:contrast |
Riddle text measures 15.87:1, holding 8.73:1 through a simulated outdoor glare model. |
check:vision |
The on-device vision tower and the precomputed text index genuinely share a vector space. |
check:notes |
No auto-post path exists anywhere. Nothing was ever sent to real OpenStreetMap. |
The one I am proudest of is that check:safety reads the UI source files. A rule that lives in src/lib but is never rendered is not a safety feature, so 26 of the assertions open the components and confirm the control is actually on screen — that the review screen renders {stop.riddle} and {stop.fact}, contains no slice( or substring(, and that no child-facing component links to the parent panel. That last check scans every component rather than a hardcoded list, because the first version named four files and silently ignored the fifth.
How I Built It
Eight increments, each ending runnable, because building a thing for children in one shot is how you ship a demo instead of a product. The full reasoning, including every bug, is in PROJECT_CONTEXT.md and docs/roadmap.md.
The dataset nearly poisoned everything, and I did not see it
The fine-tune needed data, so I swept fourteen cities across four continents — Bristol, Freiburg, Utrecht, Ljubljana, Wellington, Portland, Madison, Edinburgh, Kraków, Ghent, Valparaíso, Aix-en-Provence, Porto and Bilbao — and pulled 568 real OSM objects plus 29 hand-written house-voice seeds. (Fifteen were attempted; Tallinn failed its Overpass request on the build that produced this dataset and is not in it.) Many cities because vary the prompts, not the completions: Bristol's tagging habits are not Portland's, and a model trained on one city's conventions will confidently misdescribe another's.
The first build produced 94% broken rows — 415 of 443:
prompt: category = tourism:artwork
target: RIDDLE: Someone cast me out of bronze... I have been standing here since 1733.
The prompt says nothing about bronze, an artist, or 1733. Training on that teaches a model to invent facts — the single failure this project exists to prevent — and it is invisible in the output, because the output looks fine. The cause was a lookup miss: composeHunt re-orders the candidate pool and my id map was built from the re-ordered array.
Three guards exist now, because a check that only warns is a check that gets ignored. Skip on lookup miss. Skip when any number in the target is absent from the prompt, because numbers are what a child checks by eye. And a final gate that aborts the build and prints the offenders.
That third gate immediately caught me breaking my own rule in the hand-written seeds — a monument riddle that said "I went up in 1901" with no year in its own tags. I had written the guard and violated it in the same file. Hand-written data is not safer than generated data. All the seeds now pass the same 125 assertions as generated text.
The fine-tune
Qwen/Qwen3-8B, LoRA rank 32, lr 2e-4, 2 epochs, batch 16, 477 training rows, evaluated against 120 held-out objects. Held out by OSM object, never by row — a row-level split puts the same memorial on both sides and measures memorisation instead of generalisation.
| Loss | 2.71 → 0.75 |
| Cost to train | ~$0.02 |
| Throughput | 2.22 clock cycles/step |
| Generation | 48 output tokens, $0.000029 per riddle |
| Median latency | 2.4 s at the shipped setting (0.4), 3.3 s at the 0.7 default |
The throughput number is the result of submitting forward_backward and optim_step back-to-back and then awaiting both. Awaiting between them costs 3 clock cycles per step instead of 1 — the same result at triple the bill. It is documented at the call site so a later reader does not "tidy it up".
The fine-tune failed its safety gates, and that is the interesting part
On the first evaluation the model produced invented dates:
prompt: category = tourism:artwork | material = metal | artwork_type = statue
model: I am metal, flat, and written on. One of the dates on me is 1923.
It had learned "one of the dates on me is YYYY" from the 140 inscription examples, then filled the year slot even when the prompt contained no date at all. A slot-filling hallucination — and precisely the failure a child must never meet.
I did not loosen the gates. I changed what was being measured.
Every generation now passes the same lint generated riddles already pass, plus a groundedness check against its own prompt — the check the lint cannot make, because the lint does not know the tags. Anything failing falls back to the template riddle, exactly as a template frame falls back today. So the table reports raw / shipped / baseline, and the winner is decided on shipped: deciding it on raw would mean declaring a model safe on generations that never reach a child.
Sampling temperature turned out to be the lever, and I found it by sampling the saved checkpoint, so the sweep cost nothing to train:
| Temperature | Raw groundedness | Rescued | Specificity | Variety |
|---|---|---|---|---|
| 0.7 (default) | 98.8% | 3.3% | 84.2% | 91.7% |
| 0.4 — shipped | 99.6% | 1.7% | 84.2% | 91.7% |
| 0.0 (greedy) | 99.6% | 1.7% | 84.2% | 90.0% |
Below 0.7 the rescue rate halves. Greedy is slightly more repetitive — determinism narrows what the model can vary — so 0.4 beats 0.0.
What the model actually won, stated honestly
It won variety, and it lost on cost. Baseline specificity is 83.3% against the model's 84.2%; variety is 87.5% against 91.7%. Against that: the template costs $0 and takes 0 ms, the model costs $0.000029 and 2.4 s per riddle, and it still needs rescuing on 1.7% of the held-out riddles.
That is a real win and it is a narrow one. The defensible claim is:
A fine-tuned writer produces more varied riddles from thin OSM tags, at a measurable and small cost, provided every generation passes the gate.
Not "the model beat the baseline". Both the gates and the rescue rate are in the table, because without them the claim would be the second one, and the second one would be false.
One caveat I want on the record: temperature was chosen on the same held-out set it is reported on, so those numbers are optimistic. The un-swept 0.7 row is the unbiased one, and it is the row to quote if anyone presses.
The photo check, and the 146 MB problem
CLIP ViT-B/32 wants 85 MB of vision tower plus 61 MB of text tower — 146 MB downloaded to a child's phone. Not shippable. But our visual vocabulary is closed: twenty signatures, all derived from OSM categories, all known before the phone is opened. So the text embeddings are computed once at build time into an 84 KB JSON and the phone downloads only MobileCLIP-S0's 11.3 MB vision tower.
Scoring returns three states, never a bare number. ✅ "Correct." ❓ "I can't tell — try again, or ask the grown-up." ❌ "That's not it. Keep walking." There is deliberately no path that confidently tells a child they are wrong when they are holding the right thing, and the check is opt-in with a one-tap skip. The honest position, which the parent panel also states: a CLIP zero-shot score is a plausibility check, not proof of identity.
Safety is a gate, not a setting
A consent gate that can only be verified by clicking through it is one nobody verifies, so evaluateGate is a pure function and a test proves it. Three decisions do the real work:
- Unknown age is treated as under-12. Defaulting to allowed for a missing value is how gates get bypassed in practice, one missing value at a time.
- Stale consent reopens the gate. The record stores the stop count that was reviewed, so a hunt that changed after review needs looking at again. Without this, nothing stops a hunt growing after an adult read it, and the review screen is decorative.
- It is enforced where the child is. Consent is granted at build time, but the player re-evaluates on the way in. A gate enforced only where a grown-up happens to be looking is not a gate.
The stranger-safety script lives in data, not JSX, so a refactor cannot truncate the most important sentence in the app — and a check asserts it contains no exception clause:
Only play with someone whose grown-up you already know. No matter what a stranger tells you, get back to your own grown-up and tell them.
A rule with an "unless" in it gets read the wrong way by a nine-year-old under pressure.
Why Does Open Innovation Matter
The challenge asks for four claims. Here they are, each with the measurement behind it and — where it does not hold — the part that does not hold.
1. It runs offline, with no internet
True for the whole game, and measured. Once a hunt is saved it lives in IndexedDB behind a service worker that only ever caches our own origin. Measured: 0 cross-origin requests, and 0 requests at all during play. A hunt built at home works in a park with the phone in aeroplane mode.
This is a safety requirement before it is a feature. A phone that dies mid-hunt in an unfamiliar place is a real-world failure, and no bar of signal should decide whether a child gets to finish.
Where it does not hold: building a new hunt needs a connection, because it fetches live OSM data, and that route is not in the static Pages build at all. The tuned model path also needs network at generation time — though after generation the phone needs nothing.
2. Data stays off a server you don't control
True for everything about the child. Photos are drawn to a canvas and scored on-device; they are never uploaded and never written to storage. Location, riddles, answers and the consent record stay in the browser. No account, no device identifier, no analytics, no telemetry, no CDN — because a CDN is a third party learning who is playing and from where.
The photo check is what makes that structural rather than aspirational: there is no server-side copy to leak, subpoena or breach. Zero network requests were measured while the camera was open.
Where it does not hold, and I want to be exact about it: the training path sends OSM tags to a third-party API. 425 of 477 training rows contain a real person's name and 118 contain an inscription — "Equestrian Statue of William III", "Sir Stanley Hooker" — and those went to Thinking Machines. All of it is public ODbL data, so this is not a privacy breach. It is also plainly not "nothing leaves a server you don't control", and I would rather say so here than let the headline imply it.
The only two requests the app itself makes, both from the server and both to public volunteer infrastructure, are a place name to Nominatim and a bounding box to Overpass. The browser never talks to a third party at all.
3. It lets you fine-tune and swap models
Demonstrated end to end. The model id goes through config and is never hardcoded, because the model, the Overpass endpoint and the Nominatim URL all churn upstream. Swapping the writer means changing one value.
More interestingly, the model is not load-bearing for safety. The app's safety story depends on the gate, the lint and the category allowlist — all of which run whether a riddle came from a template or from an 8B model. I demonstrated this the hard way: the model failed two safety gates, the system caught it, and the fallback shipped. A swap cannot make the app unsafe, because nothing about the swap can bypass the checks.
4. It costs nothing to run
- The app: $0. Static hosting, no database, no API keys, free volunteer infrastructure.
- Building a hunt: $0.
- A hunt with the tuned writer: $0.000029 per riddle, plus ~$0.02 to retrain from scratch.
Free if you want it free, and about a third of a tenth of a penny per riddle if you want the tuned model. Nothing here needs a paid service to be useful, which matters for a school or a community group that cannot buy one.
The part I would actually argue
Open innovation is not "I used an open model". It is who gets to run this, and who gets to change it.
Every hardcoded dependency is behind config: the riddle writer, the vision tower, the OSM endpoints. A school in a country with different map data can point it at a different Overpass and keep everything else. The hunt files ship in the repository under ODbL, so anyone can fork it, audit every riddle a child will read, and change the category allowlist without asking me.
And the safety decisions are assertions in a test suite, not good intentions in a README. Anyone can run npm run check and see that a riddle cannot instruct a child into a road, that consent cannot be bypassed with a missing value, and that nothing auto-posts to OpenStreetMap.
A safety claim you can check is worth more than a safety claim you have to trust. That is the part of this project I would most like other people to take.
Prize Categories
Best Use of Tinker — a LoRA fine-tune of Qwen/Qwen3-8B on 477 rows derived from 568 real OpenStreetMap objects across fourteen cities, with a held-out split by OSM object, a reproducible dataset builder and evaluation harness in the repository, and a measured comparison against the baseline on metrics chosen for this use case rather than a generic benchmark. The honest result is included: a narrow win on variety, a regression on cost and latency, and two safety gates that failed on the first run and are now enforced by the shipping gate.
Touch Grass (theme) — the app's stated goal is to get a child out of the house. No streaks, no notifications, no infinite scroll, and the screen is designed to go dark after twenty seconds. The puzzle content is real geography; the engagement content is none.
Built with curiosity for Hacktoberfest 2026 — to help kids look up, touch grass, and understand the open map of Earth.










Top comments (0)