DEV Community

Penta Dank
Penta Dank

Posted on AI-assisted

LORE WALK: I built an app that wants you to put it away

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

I run Next Realm Interactive. My friend and creative director, Bully (@BullyStepHarder), is the guy on our team who holds the tablet. I see the blueprint; he translates it into something people can actually see. We call him the Lorekeeper, half as a joke. But most of his lore gets made indoors, under a ceiling, staring at a screen.

So I built him something that argues with that, and gets him curious about what open AI can do in his hands.

LORE WALK is a phone app whose main instruction is put me away. You tap Start a walk, and the screen tells you exactly one thing: find something worth remembering. Then you go. When something stops you, like a tree with a split trunk, a door painted a color nobody would choose, or a rock that looks like it's been waiting, you pull the phone out for about ten seconds and take one photo.

A small open-weight model running on the phone reads the photo. A second small model takes what it saw and writes a title and a few sentences of lore, treating the place as a location in some old world. Mythic, but tied to what's actually there. That's a Lore Shard. Then the phone goes back in your pocket.

When you end the walk, you get the Lore Codex: every shard in the order you found it, with an optional little route line if you turned on location. You can export it as Markdown or as one tall image to share.

The screen should be the shortest part of the walk. That was the whole design brief.

Demo

Live: https://lore-walk-nextrealm.dankpenta.workers.dev (open it on your phone; it installs as a PWA)

Lore Walk home screen

A Lore Shard being written

The Lore Codex

Screenshots from a full run on my machine: photo in, caption, tags, lore, Codex.

Code

GitHub logo 5thlegend / lore-walk

LORE WALK by Bully · Next Realm Interactive. Turn a walk into world lore with open models running on your phone.

LORE WALK — by Penta Dank · Next Realm Interactive

Walk the world. Write its lore.

LORE WALK is a tiny phone app that wants you to put it away. Start a walk, pocket your phone, and only pull it out when something stops you — a crooked tree, a doorway, a mural, a rock that looks like it has a name. Snap it. An open-weight AI model running entirely on your device reads the place and writes a short piece of world lore about it. That's a Lore Shard. End the walk and you get a Lore Codex: your shards in order, with an optional route line, exportable as Markdown or an image.

Live: https://lore-walk-nextrealm.dankpenta.workers.dev

Built by Penta Dank (General Dank) of Next Realm Interactive for his friend and creative director Bully (@BullyStepHarder), a.k.a. The Lorekeeper, to pull his creativity outside and into open…

MIT licensed. It's plain HTML, CSS, and one JS module. No build step and no backend.

How I Built It

The whole app is static files. All the intelligence runs in the browser through Transformers.js v3, which uses WebGPU when the phone supports it and falls back to WASM on the CPU when it doesn't.

The pipeline for each shard:

  1. Downscale on device. The photo goes through a canvas and comes out at 512px for the models and 480px as a JPEG thumbnail. The original never gets uploaded anywhere, because there's nowhere to upload it to.
  2. Read the place. Xenova/vit-gpt2-image-captioning (8-bit quantized) gives a plain caption, something like "a large tree in a grassy field."
  3. Name the thing. Xenova/clip-vit-base-patch32 does zero-shot classification against a short list of things worth stopping for (tree, doorway, mural, rock, bridge, stairs, gate…). That gives the lore writer a noun to hold onto.
  4. Write the lore. HuggingFaceTB/SmolLM2-360M-Instruct gets the caption and tags with a tight prompt: a 2–5 word place name, then three sentences of lore, in a fixed TITLE: / LORE: format that I parse and trim.
  5. Never break. If a model fails to load, the phone runs out of memory, or the small model rambles past the format, a seeded template generator builds the shard from the same caption and tags. On a trail, the demo can't be the thing that breaks.

Every shard records which models actually produced it, and the Codex shows that. The ⚙ panel lets you swap models: SmolLM2-135M for older phones, Qwen2.5-0.5B-Instruct if you've got the headroom, or template-only for zero download.

For offline use, a service worker caches the app shell and the library, and Transformers.js keeps the model weights in the browser's Cache API. Load it once on Wi-Fi at home and the walk itself needs no signal. Shards live in localStorage as small thumbnails.

The look comes from our studio: violet-black, galaxy purple, burning gold, fire orange, and white used sparingly. Big condensed type, like the spine of a tome. There's a maple leaf hidden in there too.

Why Does Open Innovation Matter?

Because of where this app is meant to be used: outside, where the signal is bad.

  • Offline on a trail. A closed API needs a round trip to a datacenter for every photo. Here the model is a file sitting on my phone. Signal bar or no signal bar, it writes.
  • The data stays on the device. I'm asking people to photograph their neighborhood, their walk home, the tree outside their kid's school. Those photos shouldn't be training data for anyone. With local inference, privacy isn't a promise in a policy. The photos physically never leave.
  • Swap the model. Open weights mean the app isn't married to one vendor. When a better small vision model ships next month, changing it is one line.
  • Zero cost. No API key, no per-request bill, no free tier that runs out. It costs nothing to run whether ten people use it or ten thousand.
  • Where open beat closed. A giant closed model would write prettier sentences. But it would also need a connection, a key, a budget, and trust, and it would make the phone a pipe to somewhere else. The small open models are a little rougher, and honestly that roughness reads like found folklore. For an app whose job is to get you off the screen, small, local, and free beats big and remote.

My Agent Session

This was built with an AI agent working alongside me, start to finish in one evening. Here's the real session: model choices, the first run where the small model ignored my output format, the CORS wall on the live site, the fallback that kept it from breaking, and the final run where all three open models wrote "The Luminous Tree."

LORE WALK: building an on-device open-model app that wants you to put it away
You

Hacktoberfest Open-Source AI Challenge, Week 1: Touch Grass. Build an app for my friend and creative director Bully, the 'Lorekeeper', that gets his creativity outside and into open AI. Use open-weight models running on-device, make it look like Next Realm Interactive, and deploy it to Cloudflare.

Agent

Plan: a phone PWA whose main instruction is 'put me away'. Start a walk, snap one photo when something stops you, and on-device models turn it into a Lore Shard. End the walk to get a Lore Codex. Static files only, no backend, so photos never leave the phone.

Agent

Model stack, all in the browser via Transformers.js v3 (WebGPU if available, else WASM): vit-gpt2 captioning to read the place, CLIP zero-shot to name the thing, SmolLM2-360M-Instruct to write the lore, and a seeded template generator so a shard always gets made.

Agent

GPU detection gets a 2.5s timeout so a hanging adapter request can't freeze startup. Model downloads begin as soon as the walk starts.

Agent

The lore prompt asks for a strict TITLE: / LORE: format: a 2-5 word place name, then 3 sentences.

Agent

Headless end-to-end test harness: a 390x844 phone viewport, a real tree photo (test.jpg), a full walk, and screenshots of home, walk and codex.

Agent

First local run with the real models. All three models loaded on CPU (no GPU adapter in the headless box). The caption and CLIP outputs were good, but SmolLM2 drifted off the format and the parser couldn't use it, so the template fallback wrote the shard.

Agent

Fallback worked as designed: the shard still came out and was honestly labeled 'template (seeded by vision output)'. The model had written a markdown 'Lore:' header instead of the requested format. The shipped parser tolerates TITLE variants and recovers the LORE text:

Agent

Deployed to Cloudflare as a Workers static-assets site. The whole config is one line pointing at ./dist.

Agent

Ran the same end-to-end test against the live URL.

Agent

CORS problem: in the live headless run, every Hugging Face model file was blocked, so all three models failed. The app still didn't break. It fell back to the template and produced a shard in 0.4s.

Agent

Made the offline story concrete: a service worker caches the app shell and the library, and Transformers.js keeps model weights in the Cache API. Shards go into localStorage as small thumbnails, and if storage fills up the app degrades to keeping only the most recent walks.

Agent

Committed v1: app, icons, OG image, manifest, service worker, README, MIT license, DEV post draft, and test logs.

You

Credit me (Penta Dank) as the maker. It's made for Bully.

Agent

Updated the byline across the page title, OG tags, header, manifest, README and post.

Agent

Added screenshots to the post and hosted them from the deployed site so DEV can embed them.

You

Do a Next Realm signature design pass.

Agent

Design pass: a cinematic violet-black world, galaxy purple, burning gold, fire orange, white used sparingly, condensed tome-spine type, tome-styled shards, new icons and OG image, and polished exports.

Agent

Responsive QA script. It forces template-only mode so it runs fast, makes two shards, ends the walk, and checks for horizontal overflow at 390x844, 1280x800 and 320x640.

Agent

Reran the full real-model test after the design pass. This time SmolLM2-360M's output got through the parser, so all three open models produced the shard (124.4s on CPU). The output is wild, rambling found folklore.

Agent

Committed the design pass with refreshed screenshots for mobile, desktop and small phones.

Agent

Final DEV post is honest about status: the full flow was tested end to end on a real tree photo, a shard takes about 1.5-2 minutes on a laptop CPU, phone WebGPU speed hasn't been measured, and the first real walk is Bully's.

Taking It Outside

Honest status: this one was built against the clock tonight, and the walk hasn't happened yet. I ran the full flow end to end on my machine with a real tree photo. The caption model saw "a tree with a large, colorful… tree trunk," CLIP tagged it as a tree, and SmolLM2 wrote back a few lines about "The Tree Oak" standing tall with its trunk reaching for the sky. That's a little clumsy, a little mythic, and exactly the kind of found folklore I was after.

On a laptop CPU, each shard took about a minute and a half. On a phone with WebGPU it should be faster, but I haven't measured that, so I won't claim it.

The first real walk is Bully's. That was always the point: I build the engine, the Lorekeeper writes the world. When he takes it out, I'll post the Codex in the comments, including whatever the models get wrong.

Top comments (2)

Collapse
 
launchgatecheck profile image
Launch Gate •

For the first phone walk, I'd test the interruption boundary as well as the happy path: start a shard offline, lock the screen during captioning, then return and take a second photo before the first finishes. Each Codex entry should keep its own photo/caption/tags/model provenance, without a late result landing on the newer photo. I'd also try a missing cached caption model: does template fallback clearly say the visual caption wasn't produced, rather than turning a default description into supposedly observed scenery? The honest laptop-only status is useful; phone timing and recovery can stay separate until that walk happens. Article-based suggestions, not a run of the app.

Collapse
 
pentadank profile image
Penta Dank •

I like that!!! I need to get deeper on some areas fs! Thanks