DEV Community

Cover image for Achievement unlocked: left my room, thanks to my walking game
daniel
daniel

Posted on

Achievement unlocked: left my room, thanks to my walking game

Gemma 4 Challenge: Build With Gemma 4 Submission

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

Built with my guy @hsmuyu ._.

What I Built

I love exploring. Trying new food, finding streets I've never walked, going out just to enjoy the day. But most of my days happen in one room, staring at a screen.

Part of it is being a student. Most fun things cost money: arcades, skiing, trips. And living on campus, even simple hobbies like gardening aren't really possible. So most of the time, the only option left is walking out the door.

But honestly, I barely walked anywhere for fun. I only go out when I have to: grab food, to buy groceries, to get through the week. Same route there & back. I'd lived here for months and still only knew the way to the canteen and the supermarket :')

When I tried to explain this to my teammate, it came down to games. I can spend hours wandering a game map for no real reason, because games reward you. You take a quest, you go somewhere, you get something back: a sound, a badge, a number going up. Every small step feels like an achievement. My walks had none of that. I'd go out, get what I needed, and come back. Nothing found, and nothing to show or brag about.

So we built footfolio. It turns the real world around you into a map, wherever you are: your campus, your hometown, a city you're visiting. And it rewards you for walking it.

The fog clears only where you actually walk, and the real cafes, parks or shops underneath appear as you uncover them. Gemma, running on your phone, gives you quests: hints about real places still hidden in the fog. You read one, put the phone away, and go look. If you're stuck for a while, a small light shows up to help, but only after you've tried.

When you find the place, the app notices. The quest completes, the place joins your map, and you earn your way toward badges: Explorer for the places you've uncovered, Local legend for a famous spot to eat, Landmark for a famous place to see.

It's small, but it's the exact same feeling that keeps me playing a game: I did something, and I have something to show for it.

Over time, the cleared map becomes a record of where your feet have been. That's the name: a folio of your footsteps. And the fog that's left tells you where to go discover next :D

Demo

Live: https://footfolio.onrender.com/

footfolio is made for phone. Open the link, allow location, and it starts wherever you are. Then go for a walk. Keep the screen on while you walk, since phones pause GPS when the screen locks.

Not outside right now? Add ?debug to the link and tap the map to walk. You start wherever you are, ?debug=10 walks ten times faster.

Gemma is optional. Turn it on from the quest card or Settings. It downloads once (about 800 MB, or 240 MB on phones without WebGPU), so Wi-Fi is a good idea. Without it, quests still work with plain clues.

Code

footfolio logo: a pixel-art sneaker

footfolio

A walking game: your map starts covered in fog that clears where you walk, and Gemma, running on your phone, writes riddles about real places still hidden in it.

Architecture

Everything personal (your path, fog, finds, badges) stays in the browser. Render hosts the app and one small service that fetches places around a point rounded to ~550 m.

footfolio/
├── index.html
├── src/                      the app (Vite + TypeScript, runs on the phone)
│   ├── main.ts               wires every module together
│   ├── gps.ts, debugWalk.ts  where you are: real GPS, or tap-to-walk
│   ├── fog.ts, fogLayer.ts   ~11 m fog cells, cleared as you walk, drawn on a canvas
│   ├── region.ts             the 2 km circle of places that follows you
│   ├──
…

How I Built It

footfolio is a phone web app: Vite and plain TypeScript, Leaflet for the map, and OpenStreetMap for the tiles and the places. Almost everything runs on the phone. The fog is a grid of ~11 m squares drawn on a canvas; every good GPS reading clears the squares around you, and what you've cleared is saved in the browser.

Render: hosting the app, and feeding Gemma real places

The app is a Render Static Site. That matters more than it sounds: phones only share GPS with HTTPS pages, so it needs real hosting to work at all. The build command runs the test suite before building, so a broken build never goes live.

The places come from OpenStreetMap through the Overpass API, and that's where I needed a backend. Browsers can't call Overpass reliably: the public servers are slow, rate-limited, and want a header browsers aren't allowed to set. So there's one Render Web Service, server/places.ts, about 85 lines with one job: GET /places?lat&lng.

The heart of it is a few lines:

// server/places.ts (trimmed)
const c = roundCentre({ lat, lng });         // round to ~550 m again: never an exact position
const entry = await region(c.lat, c.lng);    // cached for a week; one Overpass call per region
send(200, entry.gz, { 'content-encoding': 'gzip', 'cache-control': 'public, max-age=86400' });
// ...
console.warn('upstream failed:', String(e)); // the error only: no coordinates are ever logged
Enter fullscreen mode Exit fullscreen mode

The app sends a position already rounded to ~550 m, and the server rounds it again. It asks Overpass for the places within 2 km, filters them down to the ones worth finding, gzips the answer (a dense city goes from about 1 MB to 100 KB) and caches it for a week. No database, no accounts. When the official Overpass server went down on the day I built this, I added two mirrors it falls back to, and if they all fail, it serves the last copy it has.

That one service is why footfolio works wherever you are. The app loads a 2 km circle around you, and when you've walked 1.5 km away it quietly loads the next one.

Every riddle Gemma writes is about a place that came through Render. Render doesn't run the model, on purpose (more on that below), but it's what gives the model something real to write about.

Gemma: on the phone, writing riddles

The rule I set early: the code decides, Gemma writes. The code picks the hidden place, how far it is and when to give more help, and your GPS decides when you've found it. Gemma only writes the clues. A small model can't be trusted with game rules or map data, but it's good at turning "a cafe, 300 m away, near the park you found yesterday" into a riddle.

It runs in the browser, two ways:

  • With WebGPU: Gemma 3 1B (4-bit ONNX, ~800 MB) through Transformers.js.
  • Without WebGPU (older iPhones, many Android phones, in-app browsers): Gemma 3 270M (4-bit GGUF, ~240 MB) on the CPU through wllama, which is llama.cpp compiled to WebAssembly.

The second path exists because of a real moment: I turned Gemma on and my phone said "needs WebGPU", while my teammate's worked fine. Plenty of phones are like mine, and telling those users "no" felt wrong, so they get a smaller model on the processor instead. It's slower, so their clues are written ahead in the background.

I picked the models by testing them side by side. The 270M model on its own mostly copied my examples back or wrote as if it were the player. The 1B wrote 16 usable clues out of 16, at about a second each on a laptop GPU. Nothing downloads until you switch Gemma on, and the switch shows the size first.

Making a small model sound human (no fine-tuning)

My first riddles were robotic: "Head north-east to a cafe concealed by the fog." Correct, but it read like a GPS, and honestly, most people can't use "north-east" on a walk. I didn't fine-tune anything. Four things in src/questText.ts did the work.

1. Give it a character and a tiny job. Gemma doesn't play an assistant here; it plays the fog itself:

const RULES = 'You are the fog in a walking game. Speak to the player (you) in one short, playful sentence of at most 12 words. '
  + 'Say what the hidden place is like, never its name. Do not give a direction, a distance or a number. Do not invent other places, people or stories.';
Enter fullscreen mode Exit fullscreen mode

The direction moved out of the words entirely. The wisp on the map shows where to go, so Gemma only has to describe what a place feels like: something a small model is actually good at.

2. Show, don't tell. Small models follow examples far better than rules. So every prompt opens with two hand-written examples, sent as real back-and-forth chat turns, picked for the stage of the quest:

return [
  { role: 'user', content: `${RULES}\n\n${ask(a[0])}` },  // "Hidden place: Swee Choon Dim Sum (restaurant). Give the first riddle."
  { role: 'assistant', content: a[1] },                   // "Somewhere ahead, little baskets of dumplings are steaming all day."
  { role: 'user', content: ask(b[0]) },
  { role: 'assistant', content: b[1] },
  { role: 'user', content: ask(f) },                      // the real quest
];
Enter fullscreen mode Exit fullscreen mode

Sending them as separate turns mattered. When I put both examples in one message, both models just kept continuing the pattern instead of answering.

3. Feed it facts, one stage at a time. A quest has three clues. The first is a riddle. The second comes when the wisp appears, and it gets one more fact: the nearest place you've already found, so the clue can start from somewhere you know ("Past the library, follow the smell of steamed buns."). The third comes when you're close: "it's right around you, one last detail."

4. Never trust the answer. Everything Gemma writes goes through clean() before you see it. It keeps the first line, strips emoji, markdown and labels like "Clue:", and cuts to whole sentences. Then it throws the answer away if it:

  • uses a number we didn't give it
  • uses a compass word
  • talks as the player ("I'm ready…")
  • copies one of the examples
  • gives the place away

That last check is the subtle one. Banning the whole name would be too strict ("Bedok Reservoir Park" would ban the word park), so only the distinctive words are banned:

// "Bedok Reservoir Park" → bans "bedok" and "reservoir", not "park" (the kind already says that)
export const giveaways = (f) => {
  const said = `${f.what} ${f.near ?? ''}`.toLowerCase();
  return f.hidden.toLowerCase().split(/[^\p{L}\p{N}]+/u).filter((w) => w.length >= 4 && !said.includes(w));
};
Enter fullscreen mode Exit fullscreen mode

Every rule in that list came from a real bad answer. Gemma gets up to five tries (with some randomness, so each try is different), and every rejected answer is logged so I could keep tuning the checks. A plain line shows instantly while Gemma writes, and its clue replaces it a second or two later. If all five fail, the plain line simply stays. The game never waits on the model.

Why Does Open Innovation Matter?

For footfolio, open isn't a nice extra. It's the reason the idea can exist at all.

A walking game is, by nature, a stream of where you are. Building the riddles on a closed API would mean sending each player's position, and the places around them, to someone else's server on every quest. With Gemma, that never happens:

  • Your location stays private. Gemma runs in the browser on your phone, so your path, your fog and your finds never leave it. The only thing that goes out is a point rounded to ~550 m, to the places service we run.
  • No third party. No AI company sees where anyone walks, what they found, or what they were looking for.
  • Nothing to pay per quest. No API key to hide, no charge per riddle. That matters for who this is for: walking is the free option, and an app that costs money every time it speaks wouldn't fit that. The only running cost is one small places service on Render.
  • No lock-in. The model is one line of config. I swapped the 270M for the 1B when the riddles weren't good enough, then brought a 270M build back alongside it for phones without WebGPU. With a closed API, I'd get whichever models the vendor decides to offer.
  • It keeps going when the signal drops. Once the page is open, GPS, the fog, the quests and Gemma all run on the phone, and the places for your area are already saved. Lose signal halfway down a back street and the game carries on; only new map tiles wait for the connection.

Open wasn't the easy path, though. Google's own web-ready Gemma 3 models are gated behind a Hugging Face login, and a public website has nowhere to keep that key, so I used the openly published ONNX and GGUF builds instead. A small model also needs guidance a big API wouldn't: the narrator prompt, the examples, the checks on every answer. And the download, 240 to 800 MB, is a real ask, which is why it's opt-in and shows its size first.

A hosted model would have written better riddles with no download. But it would have meant a server watching everyone's walk, and for a game that is literally a map of where you go, that's the wrong trade.

The data is open too. Every place in footfolio comes from OpenStreetMap, mapped by volunteers, and when the official Overpass server went down mid-build, the open ecosystem had mirrors to fall back on.

My Agent Session

This is the session where I worked on the Gemma part. First checking if it could run on a phone at all, then getting it to write clues without saying the place's name :)

Gemma on the phone: from 'is that possible?' to riddles that hide the place
You

i plan to have quest generated when the user is at one place
make it top upper right corner, thinking of webgpu with gemma is that possible?

Agent

Yes, it's possible, and it suits the project: Gemma would run on the phone, so location still never leaves the device. There are a few catches.

How it would work
- Google's MediaPipe web package (@mediapipe/tasks-genai) runs Gemma in the browser on WebGPU. Google publishes "-Web" builds of Gemma 3 (270M up to 27B), Gemma 3n E2B/E4B and Gemma 4 E2B/E4B. For a phone, start with the smallest: Gemma 3 270M or 1B.
- WebGPU now works in Safari on iOS 26 and in Chrome on Android. It doesn't work in Firefox on Android yet, and older phones lack it. Those users would keep today's code-made quests as the fallback.

The catches
1. Download size. The model is roughly hundreds of MB for the small ones and several GB for 3n/4. It can't load on first open over mobile data. Make it an opt-in download, cache it, and host it on Hugging Face. GitHub Pages won't serve files over 100 MB.
2. Memory. iOS kills tabs that use too much memory, so the bigger models are risky on phones.
3. Speed. Writing a short quest takes a few seconds. Start generating the next quest while the current one is still going, so the text is ready when you arrive.
4. Made-up facts. A small model will invent things like "the famous mural". Keep the current split:
- Code: picks the target, checks completion by GPS (the existing reach/find rule) and handles anything that can be checked.
- Gemma: only writes the wording, from the place name, its type and nearby unfound places.
- Output: cap its length, show it with textContent, and use a template if the output is bad.

Top-right corner

The top bar (#top) spans the full width. The chip is on the left with max-width: 100%, so a long label or large text would run into a top-right card. The stats card also opens full width underneath. Two decisions follow:
- I'd cap the chip's width so the top-right card has room.
- In Phase 6 the quest took over the chip ("one goal at a time"). If the quest moves to the top-right, the chip can go back to showing the daily goal. That change is fine as long as you mean it.

Where it hooks in: finishing a quest already starts the next one after QUEST_NEXT_MS (in discovery.ts). That's where generation would start, with the place you're at as input.

Next step I'd suggest: before building any UI, make a small ?llm test page that loads Gemma 3 1B and times one quest prompt on your actual phone. If it's too slow or crashes, we learn that before building the rest.

PLAN.md says the AI quest design lives in ../walking-maxxing, but I couldn't find that folder next to this one.

Sources:
- MediaPipe LLM Inference guide for Web
- WebGPU in iOS 26
- WebGPU browser support

You

yes do it a serverless one must use gemma others you decide, output your decision to a md and do ir

Agent

I'll read the quest, discovery, HUD and wiring code first.

Before deciding on a runtime, I need to know which Gemma web builds a browser can download without a Hugging Face login. I'll check the model listings.

The Gemma 3 builds for MediaPipe are gated, meaning the browser would need a Hugging Face login. The Gemma 4 E2B web build and the ONNX Gemma 3 builds are open. Next I'll check their file sizes.

Decisions made. Next I'll create a branch (the current one is named "no AI") and install Transformers.js 4.3.0. I'm pinning 4.3.0 because 4.3.1 is only 3 days old.

Transformers.js 4.3.0 reports total download progress and loads its ONNX runtime from jsDelivr. Next I'll write the pure quest-text module: the prompt, the output check and the fallback line.

Next is the Gemma loader.

Before building UI on it, I'll check Gemma actually runs and writes usable lines. I'll start the dev server and drive real Chrome on this laptop.

WebGPU works on this laptop (Intel GPU, half-precision supported). Vite reloaded the page because it found the new package, so I'll run the spike again.

The 270M model can't do this task: it mostly echoes the prompt ("The player is out for a walk. Next: …"). I'll try the 1B model.

That command misfired, so I'll read the file to see what actually changed.

The edit applied correctly. Next I'll tidy one line, widen the "about finding" words, and rerun with 1B and then 270M.

Agent

Gemma now writes the quests, entirely on the device with no server, and the quest has moved to the top-right corner. I wrote the decisions to GEMMA.md. It runs end to end on this laptop, but nobody has tried it on a phone yet, and that's the main risk.

What you'll see
- A round flag button sits top-right, with a ring that fills as you get closer. The chip on the left goes back to today's goal.
- When a quest starts, which happens when you arrive at a place, a card opens under the flag. It shows one line of quest text plus the exact distance, and says "Written by Gemma on this phone" when Gemma wrote it.
- You turn Gemma on with a new switch in the stats card. It shows the download size first ("800 MB download, runs on this phone"), then download progress. Once on, it stays on and loads from the browser's cache on later visits.
- Until Gemma has loaded, or when its answer fails the checks, the card shows a plain line instead. Browsers without WebGPU (Firefox on Android, older phones) get the plain line, and the switch says why it can't turn on.

Decisions that change what you might expect
- Transformers.js, not Google's MediaPipe. The Gemma 3 web models for MediaPipe need a Hugging Face login to download, which a static site can't provide. The only open MediaPipe build (Gemma 4) is 2 GB.
- Gemma 3 1B (about 800 MB), not 270M. I tested both on this laptop. The 270M model mostly copied the prompt or wrote things like "I'm ready to explore!". The 1B model's lines were usable, and each took about 1.2 s.
- Gemma only writes the wording. Code still picks the target and decides when the quest is done. Gemma is told the direction and kind of place (never the hidden place's name), and answers that invent numbers, name the hidden place or drop the direction are thrown away.

Tests
- The unit tests, typecheck and build pass.
- The full browser suite passes, including a run that downloads the real model. That run got "Head west to a restaurant veiled in the fog." for a target that really was west, and the line stayed after a reload.
- One older check ("fog beside trail after pan") fails in about 2 of 3 runs on this Windows machine. It fails the same way on the code from before my changes, and passes on a retry.

What's left
- Try it on a phone. Phone GPUs are slower, and 800 MB may be too much memory for older iPhones. If it is, switching to the smaller model is one line in src/config.ts (GEMMA_MODEL), at the cost of more plain lines.
- The plain line shows first, and Gemma's replaces it a second or two later.
- Switching Gemma off frees memory but leaves the 800 MB in the browser's cache. There's no "delete download" button yet.

The work is on a new branch, feat/gemma-quest-text, because the current one was named "no AI". It also carries your uncommitted Vite version bump. Nothing is committed yet. The dev server is still running at localhost:3000/?debug if you want to try it: tap the chip, turn on "Gemma writes quests", then tap the map to walk.

Agent

The only shell still running is the Vite dev server (npm run dev), which serves the app at http://localhost:3000/. Open ?debug there to try the Gemma quests.

The test runs and their watchers have all finished. The SSH server from earlier is a Windows service, not a shell, and it keeps running in the background.

You

short ans, which part is generated by gemma

Agent

Gemma writes only the bold first line in the quest card, for example "Head west to a restaurant veiled in the fog." When Gemma wrote it, the card also says "Written by Gemma on this phone".

Everything else comes from code: picking the target, the distance line, the flag and its ring, the map marker, and deciding when the quest is done.

You

<p id="questBy" hidden="">Written by Gemma on this phone</p> i search in devtools, why is it hidden?

Agent

It's hidden because the line on screen isn't Gemma's. It's the fallback line. hud.ts sets questBy.hidden = !quest.text, and the quest only gets a text once Gemma has written a line that passes the checks.

It's usually one of these:
1. Gemma is off. It's off by default. Tap the chip and turn on "Gemma writes quests".
2. Gemma is still downloading or loading. Wait for the switch to say "Runs on this phone, offline".
3. Gemma's answer was rejected by the checks (for example, it got the direction wrong). The fallback line then stays for that quest, and there's no retry until the next quest.

To check in the console:
js
fogMap.quest().text // undefined = fallback line
document.getElementById('gemmaNote').textContent // Gemma's state

If it says "Runs on this phone, offline" and text is still undefined, it's case 3.

You

TODO
Ask gemma to generate interesting hint to the location, hide the POI flag icon
Without disclosing where the place is
Only one POI at a time
When user can’t find it for too long, show the POI flag icon
Moving too far away from the point, alert user
Suppress some places if icons are too crowded at one place

Agent

Do you want me to build all of these? Here's how I'd do each one; say "go" or change any point.

  1. Gemma hint, flag hidden. For a reach quest, Gemma gets the place's name and type and writes a riddle-style hint (for example "Somewhere north-east, people queue for something cold and sweet"). The check rejects any hint that contains the name. The map marker starts hidden. The card keeps the direction and the distance, so it plays as a hotter/colder game.
  2. One POI at a time. Reach quests already have a single target, so I'd only make sure that holds. If you meant something else (for example, showing only one found-place pin at a time), tell me.
  3. Can't find it for too long. The flag marker appears after 5 minutes, or sooner if you've walked past and the distance is growing. The card would say "Here's the spot".
  4. Moving too far away. If you get 50% farther than where you started (minimum +100 m), you get a one-time toast and a vibration ("You're heading away from the quest"). It re-arms once you come back closer.
  5. Crowded icons. When place pins overlap on screen, show one and hide the others, recomputed after each zoom or pan. No clustering library. Tapping a visible pin still shows its own name.

These will be config constants so you can tune them: 5 min, 50%/+100 m, and the overlap distance in pixels. One cost: item 1 sends the place's name to Gemma, but Gemma runs on the phone, so nothing leaves the device.

You

yes do all

Agent

I'll start with the riskiest part: whether Gemma 1B can write a hint without giving away the name. Then I'll build the rest. First, the files I haven't read yet.

Here's the new hint prompt. Gemma now gets the hidden place's name and type, and the check rejects any hint that uses a distinctive word from the name.

Agent

All six items are built and working, including on the real Gemma model. Like before, it's all on this laptop and hasn't been tried on a phone.

  1. Gemma hint, flag hidden. A new reach quest has no marker on the map. Gemma gets the place's name and type and writes a hint without giving the name away. For "PolyMate Praxis Makerspace" it wrote "Follow the scent of soldering and code as you head north-west to a place filled with tools and inventive minds." Any hint that contains a distinctive word of the name is thrown out. Gemma gets 5 tries per quest, then the plain line is used instead.
  2. One POI at a time. This was already the case (one quest, one target, at most one marker), so nothing changed.
  3. Can't find it. The spot gets marked on the map after 5 minutes, or sooner if you got at least 60 m closer and then drifted 60 m back out. You get the message "The quest spot is on the map". The check runs on every GPS update and once a minute, so it also works if you stand still.
  4. Moving too far away. Once you're 1.5 times the starting distance away (and at least 100 m further), you get one message with the way back ("It's about 400 m north") and a vibration. It won't repeat until you've come back within the starting distance.
  5. Crowded icons. Place pins closer than 22 px on screen show only the newest one, and zooming in brings the rest back.

Things you might trip over:
- Hints can be vague or a bit wrong. They come from the model's general knowledge of the name, so Sheng Siong (a supermarket) became "a bustling market". The 5-minute reveal covers this.
- Quests saved before this update come back already marked, because they never had a hint.
- Rejected Gemma answers now show in the browser console as Gemma line rejected: …, so you can see what the checks throw out.

Tests: the unit tests, typecheck and build pass. The full browser suite passes, including a run with the real model, and covers the hidden spot, the alert appearing only once, both ways the spot gets marked, the marker staying put as you get closer, and the crowded pins. The only failures were the old Phase 2 timing check that also fails on the original code.

All the thresholds are constants in src/config.ts, and GEMMA.md has a new section on these changes. Still uncommitted on feat/gemma-quest-text. The dev server is still running at localhost:3000/?debug if you want to try it.

Prize Categories

Best Use of Gemma: Gemma runs locally on the player's phone, in the browser, with no server involved. Gemma 3 1B runs on WebGPU through Transformers.js, and Gemma 3 270M runs on the CPU through wllama (llama.cpp in WebAssembly) for phones without WebGPU. It writes the riddle clues about real places hidden in the fog, guided by a narrator prompt and hand-written examples. The code picks the places and checks every answer before it's shown.

Best Use of Render: The app is a Render Static Site, and a small Render Web Service fetches, filters and caches the OpenStreetMap places for wherever you are. Every riddle Gemma writes is about a place Render delivered. The model stays on the phone by design, so nobody's position is sent to an AI server.

Future Features

If we had more time, we'd add more achievements and make your progress shareable.

  • More achievements. Badges are just the first level. A few I'd love to add:
    • Full clear: uncover every place in an area
    • Night owl: find a place after dark
    • Traveller: explore in a new city
  • Shareable progress. Your map, your badges and how much you've explored, on a profile your friends can see. Part of what makes games fun is showing your progress and trying to beat your friends, and I want footfolio to have that too. You could even compare with a friend in another city, without ever walking the same street.

Sharing would always be your choice. Your walks stay on your phone unless you decide to show them.

Top comments (1)

Collapse
 
stanleynguyen profile image
Stanley Nguyen •

🔥