This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built this for my grandmother.
Because of her age, she can't remember a lot of things these days, and keeping track of her daily routine, such as when to take which medicine, has become hard for her. I wanted to give her something calm and simple that shows her day and what is left to do.
Little World is a daily routine tracker shown as a small 3D island. You describe the routine in plain English. Gemma, an open model running inside the browser, turns the description into a schedule. Each task appears on the rim of the island at its time of day, and the sun and moon move across the sky with the real clock, so the current time and the next task are visible at a glance.
This is the example routine I used during development. It is made up, not hers:
Thyroid tablet, 50 mcg, when I wake up, on an empty stomach.
Blood pressure tablet, Amlodipine 5 mg, after breakfast.
Metformin 500 mg before lunch and before dinner.
Eye drops in both eyes at 4 PM.
Calcium tablet after dinner, not with the thyroid tablet.
A 20 minute walk at 6 PM.
When a task is due, its marker lights up and a reminder appears. Pressing Done marks it complete. The Ask tab answers questions such as "What's next?" or "Did I take my eye drops?" using only the routine that was entered.
The routine never leaves the device. There is no account, no server and no analytics, and after the first load the app works offline. For something as personal as a medicine list, that was a firm requirement.
Little World is not a medical tool, and the app says so. It only repeats what it was told. When I asked "Can I take calcium with my thyroid tablet?", it answered that it could not see that in its notes and that I should check with a doctor or pharmacist, which is the behavior I wanted.
Showing it to my grandmother
I showed Little World to my grandmother. Her reaction was "oh wow, this is so good". Then she said "you made something for me", and she was genuinely pleased.
Designed for an older user
After that, I focused on making it easier to use for someone her age:
- "What do I do now?" A single button opens a full-screen view with one task in very large type, one large "I've done it" button with an undo option, and the next few tasks below. It can be set as the default screen.
- Spoken reminders. Reminders can be read aloud with the browser's built-in speech synthesis, so no external voice service is involved. If a task is still not done, the reminder repeats after 20 and 40 minutes.
-
Setup by a family member. Gemma needs a computer with a capable GPU, but the person using the routine does not. A family member can enter the routine on a computer, select "Send to another device", and share the QR code or link. The routine is encoded in the part of the link after the
#, which browsers do not send to a server. Opening the link on a phone or tablet shows the routine and asks for confirmation before using it. - Readability. Larger text throughout, buttons at least 48 pixels tall, labels that stay readable against the daytime sky, and task states written as words ("Now", "Waiting", "Done") rather than shown only by color. A reduced-motion setting turns on automatically when the device requests it.
Demo
Live app: https://starknightt.github.io/little-world/
- "Try the demo island" works on any device, including phones. It loads a routine that Gemma already parsed, so nothing is downloaded. From there, "What do I do now?" opens the large view and "Send to another device" shows the QR code.
- "Describe my own routine" requires desktop Chrome or another browser with WebGPU. Gemma 2 2B is downloaded once (about 1.4 GB) and stored by the browser. After that it loads in a few seconds and works offline.
The clip below shows the real model with nothing mocked: I enter a routine, Gemma parses it, I review the preview and save it, then ask two questions. The GIF runs at 2.5x speed; the 27-second video runs at 1.25x.
The same routine at night, and on a phone:
Code
StarKnightt
/
little-world
A daily routine tracker on a small 3D island. Gemma runs in the browser through WebGPU, and nothing leaves the device.
Little World
A daily routine tracker shown as a small 3D island. Describe a routine in plain English ("Thyroid tablet when I wake up, Metformin before lunch and dinner, eye drops at 4 PM"). Gemma, running entirely in the browser, turns it into a validated schedule. Each task is placed on the rim of the island at its time of day, and the sun and moon follow the real clock. Mark a task as done when it is complete. Ask "What's next?" and Gemma answers using only the routine.
Nothing leaves your device. No account, no server, no analytics. After the first load it works offline.
Live: https://starknightt.github.io/little-world/ (open ?demo for the instant demo, no download)
Not medical advice. Little World only repeats what you told it. It does not know about medicines, doses or interactions, and it is told to send those…
Built with Vite, TypeScript, Three.js, WebLLM and zod, under the MIT license. All code in the repository was written during the challenge weekend, starting October 4.
How I Built It
Stack
-
Gemma 2 2B (
gemma-2-2b-it-q4f16_1-MLC), Google's open-weight model, quantized to 4 bits. - WebLLM from MLC, which runs the model on the GPU through WebGPU. It runs in a Web Worker so the 3D scene stays responsive while the model is working.
- Three.js for the island, with a custom sky shader, bloom and HTML labels that avoid overlapping each other.
- zod to validate every routine before it is saved, and vite-plugin-pwa for offline support.
Pipeline
routine text
→ Gemma 2 2B, constrained by a grammar
→ JSON, one object per task: quoted source text, name, dose, time anchor, food rule, icon
→ quote check (plain code)
→ zod validation (one retry if it fails)
→ time anchors converted to clock times
→ editable preview
→ saved routine
Key decisions
1. A grammar instead of asking for JSON. WebLLM can constrain decoding to a grammar. I wrote a small EBNF grammar for the exact routine format with fixed separators, so the model chooses only the content, never the structure. With it, every output in my test set was valid JSON.
2. The model chooses a meaning, and code computes the time. My first version asked Gemma for an hour and a minute, and it often returned 0:00 for "after breakfast". Now the model chooses from a fixed list of anchors (on_waking, after_breakfast, before_lunch, bedtime and a few others), and code converts them to times such as 8:30, 12:30 and 22:00. The model returns a clock time only when the text contains one, such as "at 4 PM". The food rule is derived from the anchor as well, so "before lunch" cannot become "after food".
3. Quote first, then verify. For each task the model must also copy the exact source text into a said field. A short function then checks that quote with regular expressions. If the quote names exactly one time or food rule and the model chose something else, the quote takes precedence. A quote such as "before lunch and before dinner" is split into two tasks, and a dose containing a number that does not appear in the source text is removed. This check produced the largest accuracy improvement, in about 100 lines of code.
Before anything is saved, the user sees a preview in which every time can be edited and every task can be removed. A 2B model will sometimes be wrong, so the person always has the final say.
Questions follow the same principle. Code first computes the facts: what is done, what is waiting and for how long, and what is next. Gemma only phrases those facts as a short answer, and it is instructed not to give medical advice.
Results
I wrote 20 test routines with known correct answers. They cover medicines, insulin, an inhaler, eye drops, plant care, feeding a pet, explicit clock times and descriptive wording such as "the white tablet with breakfast, the pink tablet at bedtime". I scored Gemma on all of them in headless Chrome on my PC's RTX 4060.
| Gemma 2 2B | Valid JSON | Correct task count | Correct times | Correct food rule |
|---|---|---|---|---|
| Free text, no grammar | 14/20 (17/20 after one retry) | 13/20 | 38/54 | 16/21 |
| Grammar, model alone | 20/20 | n/a | 36/54 | 11/21 |
| Grammar + quote check (shipped) | 20/20 | 17/20 | 47/54 (87%) | 17/21 |
- Download: 1.4 GB, once, in 42 parts. On the live site, the first run including the download took about 7.7 minutes on my connection. After that, loading from the browser cache takes about 4 to 6 seconds.
- Speed: about 40 to 55 tokens per second for generation and 500 to 700 tokens per second for reading the prompt. Short routines take 4 to 6 seconds. The six-line example above is 688 tokens in and 330 out, about 8 seconds.
- Questions: answers take about 1.5 to 3 seconds.
- Offline: on the live site with a fresh browser profile, I downloaded the model, created a routine, disconnected the network and reloaded. The app loaded from its own cache (0.9 MB), Gemma loaded from the browser cache, and a new routine was created in about 12 seconds.
- First paint: under one second on the live site at desktop and phone sizes. The CSS is inlined into the page.
- Automated checks: a headless test opens the live site at desktop and phone sizes, covers the large view, the share link round trip and a browser without WebGPU, and fails on any console error. All 28 checks pass.
What did not work
- Gemma 3 1B was my first choice. At 537 MB it is a much smaller download. In the WebLLM build I used, its sliding-window attention configuration either failed to load or, once forced, produced unusable output. Without the grammar it produced no valid routines in 3 test cases. With the grammar it produced valid JSON for all 5 cases, but none of the 11 times were correct. I switched to Gemma 2 2B, at the cost of a download nearly three times larger.
- JSON schema mode looped in early tests. With a plain JSON schema, the small models sometimes generated whitespace until they reached the token limit. The custom grammar with fixed separators avoids this. In the final evaluation, schema mode was also valid for all 20 cases.
- Times were the hardest part. Asking for hours and minutes directly produced many wrong times, often midnight. The anchor list plus the quote check brought it to 47 of 54.
- The model still misses tasks. In the example routine above, Gemma leaves out the walk. I kept that in the demo deliberately, because it shows why the editable preview matters. During the offline test it also produced a dose of "500 mg" for a Vitamin D tablet whose dose I had not written, which is why the dose rule exists.
-
A bug in my own regular expressions. I generated part of the quote checker with a script that turned every
\b(word boundary) into a literal backspace character. Nothing matched, and accuracy dropped until I found it. - Testing a 1.4 GB model is slow. Each test with a fresh browser profile includes an eight-minute download. A force-closed Chrome once left the model cache half-written, and the app then waited indefinitely while checking the GPU. That check now times out after 3 seconds.
- Not every device can run the model. It needs WebGPU with 16-bit float support and a few GB of GPU memory, and I only verified it on my own GPU. Phones get the full app with the demo routine and can open a routine sent from a computer. When the model cannot run, the first screen says so and suggests that option.
Why Does Open Innovation Matter?
A daily routine, especially one with medicines in it, is private. It shows what a person is treated for, how often, and how consistently they keep up. My grandmother should not have to send that to a company's server to get a reminder or an answer to "Did I take the evening one?"
Because Gemma's weights are open, the whole model can run inside the page. The routine, the completed tasks and the questions stay in one browser on one device. There is no API key to leak, no usage bill, and no dependency on a service that could change its terms or shut down. Once loaded, it keeps working without a network connection.
Open weights also meant I could control how the model behaves, not only what I ask it. WebLLM let me constrain decoding with my own grammar, which raised JSON validity from 14 out of 20 to 20 out of 20. I could switch from Gemma 3 1B to Gemma 2 2B with a one-line change when the first did not work, and I could run the 20-case evaluation repeatedly on my own GPU at no cost. With a closed API, I would have been limited to adjusting the prompt.
The trade-off is a 1.4 GB first download and a WebGPU requirement. For a health routine I think that is the right trade, and smaller open models will reduce it further.
Prize Categories
- Gemma: Gemma 2 2B runs entirely in the browser through WebLLM and WebGPU and is the core of the app. It turns a plain-language description into the routine and answers questions using only that routine.














Top comments (2)
This is such a thoughtful project. The 3D work is seriously impressive on its own, but I really like how much attention you gave to the actual person using it too. The large “What do I do now?” view, spoken reminders, reduced motion support, and being able to set everything up on another computer and send the routine over are all really smart choices. I also love the way the island reflects the day with the flowers, sun, moon, cottage lights, etc. It makes something that could feel clinical or stressful feel much calmer and more personal. And keeping the routine completely local feels especially appropriate here. Really nicely done.
that's some interesting 3d stuff built on a simple idea man.
hats off