This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My friend's daughter has just started her first instrument: a toy xylophone with eight coloured bars. She reads a little, and practice works best when a grown-up sits beside her, points at the next bar, and waits. Plink is that patient partner.
Plink is a web app that runs on the family laptop. It listens through the microphone, learns the sound of her bars, lights up the bar to hit next and waits, with no buzzer, until she finds it. Between parts, two open models help: TabPFN learns from every note she plays and picks the part she'll get right about 80% of the time, and Gemma 4 turns any tune into a short lesson and writes a two-line note for her parent. Nothing she plays leaves the laptop, and it works with the Wi-Fi off.
A session goes like this:
- Plink's turn: Plink plays the part while the bars light up, a little slower than normal.
- Her turn: the next bar glows and Plink waits. A miss just makes the right bar wiggle.
- Help that fades: lots of help (only the right bar sounds, and Plink says its colour), some, or little (she plays from memory).
- Every note right first time is a star, and every 15 stars unlock a sticker.
- After every three parts of her song, Plink slips in a short part on a jump she finds hard. The last part is always one she's likely to get right, so she ends on a win.
The app is loud and colourful everywhere except during a part. Then the page calms down, the other bars dim, and only the target glows, in its real colour, because Plink's hints name the colours.
There's also an on-screen xylophone (mouse, touch, or the A S D F J K L ; keys) for when the real one isn't to hand. A grown-ups area behind a simple sum lets her parent add songs by playing them in, typing letters, singing them, or taking a photo of a song card.
Handing it over
Hosted Demo
Code
souptik4572
/
plink-xylophone-practice-app
Plink is a local-first web app built for children playing their first xylophone, and for the parent who sits beside her. It has a built-in on-screen xylophone, playable by mouse, touch or the home-row keys, so she can play and practise even without the real instrument.
Plink
A patient xylophone practice partner, built for one child and the parent beside her. Hacktoberfest 2026 Weekend Challenge, "Build for a Friend". Full design in spec.md.
Run
make setup # uv sync, pnpm install, ollama pull
make dev # backend on 127.0.0.1:8000, app on http://localhost:5173
make test
make eval # TabPFN drill-picker benchmark
The first time, make an account in the app (see Accounts).
Accounts
A grown-up signs up with an email and a password (8 characters or more), on the laptop and on the demo alike. Everything Plink keeps (her songs and their lessons, practice, settings and calibration) belongs to that account, and no account can see another's.
- Passwords are stored only as Argon2id hashes.
- The avatar at the top right opens the grown-up's Account page and Log out, both behind the grown-up sum, so she can't log herself out. The page holds your nameβ¦
make setup && make dev runs it locally: FastAPI and SQLite on the back end, React, TypeScript and Vite on the front, and Gemma through Ollama. There are 215 backend and 114 frontend tests, plus headless-Chrome runs that play whole sessions.
How I Built It
her xylophone ββ mic βββΆ browser: AudioWorklet β strike gate β which bar?
on-screen xylophone βββ (plain signal processing; audio never leaves the tab)
β only bar numbers and timings
βΌ
FastAPI ββ SQLite on the laptop (Postgres on the Render demo)
β every note she plays = one row in her practice log
βββ TabPFN: what to practise next, when to help, her trickiest jumps
βββ Gemma 4: lessons, praise, the parent note, song cards, Ask Plink
singing βββΆ Basic Pitch (in the browser) β notes β fitted to her eight bars
Hearing her bars is not a model's job
Detection is plain signal processing, on purpose. It has to answer the moment she strikes, the audio must stay in the browser, and it has to fail predictably.
Calibration has her strike each bar three times. Plink stores what each bar sounds like on a musical scale (quarter-semitone bands), so a calibration made on one laptop still works on another with a different sample rate. A sound only counts as a strike if it peaks at the mallet and then only decays, as a few clear partials. Claps, knocks, key clicks and voices don't behave like that, so Plink ignores them, and it re-measures the room's noise as she plays.
On synthetic wooden and metal bars it accepted all 96 test strikes and rejected every clap, knock, click and voice. In the browser, fed a recording through Chrome's fake microphone, the Mic check scored 40 of 40. [TODO: her real xylophone's Mic check score from Sunday.]
TabPFN: seven small decisions, from her first session
TabPFN is a tabular foundation model: you hand it a small table and ask for predictions, with no training loop. That fits a child's practice log well. Every note she plays is one row (the bar, the jump from the last one, the help level, how many times she's seen the partβ¦) plus whether she got it right first time.
From about 20 rows, which is her first session, TabPFN:
- chooses the next part at the level she'll get right about 80% of the time. It predicts every candidate part "as if she played it next" and picks the one closest to 0.8;
- times the help for each note: where it expects a miss, the glow comes sooner and Plink says the colour;
- advises more, less or the same help, but only for levels it has seen her play. It can't know what it hasn't seen, so otherwise it suggests trying that level once;
- rates each song for her and suggests one on the home screen;
- finds her trickiest jumps between bars;
- turns those jumps into practice: three short parts per jump, easiest first ("Stepping stones", "One big hop", "Hop and back"), slotted in after every three song parts at her sweet spot;
- ends the session on a win: when the next part probably won't fit in the time left, Plink says "Last one!" and plays the part TabPFN is surest she'll get right.
A fit and prediction takes 0.8 s warm on a laptop, so the next pick is computed in the background while she's celebrating the last one. On the 36 notes I played while testing, TabPFN edged logistic regression (ROC AUC 0.858 against 0.819). That's far too little data to call a win, and it isn't the point. The point is that TabPFN works on the tiny table you actually have in week one.
Gemma 4: six jobs, and code that owns the notes
Gemma 4 (gemma4:e4b, through Ollama) does the language and vision work:
- Lessons: it splits any tune into short named parts, with a tip for each.
- Praise and the parent note: the note is two plain lines on what to practise next, with TabPFN's trickiest jumps fed in.
- Song cards from a photo: beginner songbooks often print songs as coloured circles or letters. Gemma counts the notes first (with thinking on), then lists them (thinking off), and code maps each one to her bars. That two-pass approach read 14 of 14 notes on test cards printed with letters, numbers and colours.
- Ask Plink: a grown-up can ask "Is she ready for Jingle Bells?" Gemma calls read-only tools over her practice data and TabPFN, and answers with the numbers ("87% first try"). TabPFN decides; Gemma explains.
- The line Plink says before each jump practice part, naming the colours: "Go orange and turquoise back and forth, yay!"
The rule everywhere is that code owns the notes. Gemma names and explains, but every reply goes through a validator, and when one fails, Plink falls back to plain parts and plain text. A child never sees a broken lesson.
Two surprises along the way. Gemma 4 thinks by default: a lesson took 44 s with thinking and 10 s without, so thinking is only on for counting the notes on a song card. And with no name set, Gemma called her "Lily", in the praise and in the note. The prompts now forbid inventing names.
What didn't work: asking Gemma twice
The usual advice for structured output is a bounded retry-repair loop: when the JSON fails validation, send the error back and let the model fix it. Structured Output From LLMs: A Retry-Repair Loop Your Parser Never Sees Through by @devshakib lays the pattern out well. I added one (simplified from backend/app/coach.py):
REPAIR = ("That answer can't be used: {problem}. Reply again with only the corrected JSON, "
"nothing before or after it, keeping every rule in the first message.")
for attempt in range(config.GEMMA_REPAIRS + 1):
reply = _post(body, timeout)
try:
return validate(json.loads(reply["message"]["content"]))
except (ValueError, KeyError, TypeError) as e:
messages += [{**reply["message"], "role": "assistant"},
{"role": "user", "content": REPAIR.format(problem=e)}]
It barely helped. The repair fixed 2 of 10 broken lessons; the rest repeated the mistake or made a new one. The mistakes were arithmetic. I'd asked for each part's first and last note, and Gemma wrote things like "phrase 4 starts at note 2" when it had to start at note 21.
So I changed the question. Gemma now says how many notes each part holds ({"notes": 4, "nickname": "β¦", "tip": "β¦"}), and code places the parts:
| Lessons on the Gemini API (gemma-4-26b-a4b-it) | Valid first try | Repairs that worked | Valid |
|---|---|---|---|
| First and last note, with repair (10 runs) | 6 | 1 of 4 | 7 |
| + complaints that say how to fix it (20 runs) | 13 | 1 of 6 | 14 |
| Part sizes, placed by code (20 runs) | 11 | 1 of 1 | 12 |
Every lesson that came back was valid. The other 8 hit the API's 60 s timeout, and those songs keep plain parts until a grown-up asks again. On the laptop's e4b, 5 of 5 were valid.
The jump practice taught me the same thing. Asked to name six nearly identical practice parts, Gemma came back with junk half the time ("nickname-2", "=", "pattern-pattern"), and the repair fixed none of it. Now each kind of part has a fixed name, which is also easier for a child to learn, and Gemma only writes the spoken line, which it got right in 11 runs of 12. A line that names a colour the part doesn't use goes back for repair, because she finds her bars by colour.
The lesson: a retry loop can't fix a question the model finds hard. Change the question, and let code do the counting.
Render: the public demo
Her copy stays on the laptop. For everyone else there's a public demo on Render, defined in one render.yaml Blueprint: a web service that serves the app and the API from one Docker image, plus a free Render Postgres database for logins and each account's practice. Render generates the login-signing secret, and each new demo account is seeded with simulated practice (6 sessions, 373 notes), so TabPFN is live on the first visit. Render has no GPUs, so the demo runs Gemma 4 through Google's Gemini API (gemma-4-26b-a4b-it), which built a lesson in about 6 s.
Two things I learned getting there. PyTorch in a 1-CPU container saw all 10 of the host's CPUs, started 10 threads, and was throttled for 486 s until I set OMP_NUM_THREADS. And leaving uv's download cache in the image made it 3.47 GB instead of 2.06 GB.
Why Does Open Innovation Matter?
Because of who it's for. Building Local-First AI Apps: What Changes When the Data Stays on the Device by @sanskarin covers the general case. For a child's practice, it isn't a nice-to-have:
- Her audio never leaves the room. Detection runs in the browser, and only bar numbers and timings are saved. There is no recording of her to leak.
- It works with no internet. Gemma runs through Ollama on the laptop and TabPFN runs locally, so practice doesn't depend on a connection or a subscription.
-
It fits her. It's calibrated to her bars, Plink's words can be in her family's language, and a smaller or bigger Gemma is one environment variable (
GEMMA_MODEL) away. - It costs nothing per session, so she can play every day.
A closed API could write lessons too. But every song, every practice note and every question about a child would go to someone else's server, and practice would stop when the Wi-Fi does.
The honest limits: Gemma 4 and Basic Pitch are Apache 2.0, but TabPFN's weights are open under a non-commercial licence. That's fine for a gift, not for selling it. And the public demo trades some of this away for convenience: there, song-card photos, the numbers behind a parent note and Ask Plink questions go to Google's Gemini API. Her laptop copy sends nothing.
My Agent Session
I built Plink with help from Claude Code. This session is day two: logins (JWT access and refresh tokens), the calibration rework, moving the demo's Gemma to the Gemini API, and the Gemma and TabPFN upgrades above, measurements included.
Prize Categories
- Best Use of Gemma: six jobs across text, vision and tool calling (lessons, praise, the parent note, song cards from a photo, Ask Plink, and the lines for jump practice), with validators, a repair step and fallbacks, so a child never sees a bad answer.
- Best Use of TabPFN: seven live decisions from her own practice log, starting in her first session: what to practise next, when to help, which help level to try, how hard each song is for her, her trickiest jumps, practice on those jumps, and a last part she'll get right.
Plink's on-screen xylophone was inspired by my earlier project souptik4572/xylophone and rebuilt from scratch here. Thanks to Prior Labs (TabPFN), Google DeepMind (Gemma), Spotify (Basic Pitch) and Ollama. I built Plink with help from Claude Code.
Top comments (0)