I can’t read most of the filenames in this library. So the rule is that nothing gets filed unless two models with opposite bad habits happen to agree.
TL;DR
My partner is from Shanghai and loves karaoke. We own a karaoke machine with good speakers, good microphones and a terrible tablet bolted to the top. So over three days the house got its own karaoke server: songs live on the UNAS Pro, the box that holds all our media, the TV shows the lyrics, and anyone in the room scans a code and queues songs from their phone.
It started with about 770 songs in Korean, Mandarin and Cantonese. Then people added their own, under untidy names in characters and languages I can’t read.
So every morning at four, two AI models running in the house each say which language a new song is sung in. If they agree, it’s renamed and filed. If they both say Chinese but can’t agree which kind, a third model called Jev settles it. Anything else waits for a human. Forty filed so far, eight waiting, and one under a name I’ll own up to further down.
If you’re here for the workflow, it’s at the bottom.
The good speakers with the bad tablet
The machine is an iKarao Shell S2. The speakers are good, the two wireless mics are good, and the sound processing does that flattering thing karaoke machines do to a voice. The Android tablet on top of it is slow, locked down, and the only way to choose a song.
My partner grew up with KTV, which is karaoke done properly: a private room, a wall of screens, a catalogue that goes back decades, and Mandarin and Cantonese songs I will never be able to pronounce. What she wants from a karaoke night is K-pop, Mandopop and Cantopop. What the tablet offers is a slow search box.
The plan had been obvious for a while. Demote the machine to what it’s good at, speakers and mics, and put the brains somewhere else. There’s an open-source project called PiKaraoke that does exactly this: it runs on a server, shows a player page on any browser, and puts a QR code on the screen so guests can search and queue from their phones. My old gaming PC is already plugged into the TV. Plug the karaoke machine into the PC as a sound card and the job is done.
This one wasn’t a someday project in the usual sense. It wasn’t blocked on anything. It was a bit of fun, for someone else, and I expected it to take an evening.
I’ve written about this house before, most recently the task board the agent built without seeing it and the spa that only heats on spare sunshine. Same collaborator, same house, same principle: I supply direction, the agent supplies execution. This time the direction was “karaoke, for her, in her languages”, and that last clause turned out to be the whole project.
The brief said to install fonts
I wrote a brief first, as I do now. It said to build a small Linux container on the cluster, install PiKaraoke, keep the song library on the NAS, a UNAS Pro, and install a package of Chinese, Japanese and Korean fonts, because PiKaraoke burns the lyrics into the video on the server and without those fonts every lyric would come out as a row of empty boxes.
That last instruction was confident and wrong. The agent read PiKaraoke’s source before installing anything, and the server never draws a single character. It only re-encodes video, for things like changing the key. When a song has a separate lyrics file, the file is sent to the TV as it is, and the browser draws it, using two fonts that ship inside PiKaraoke itself.
So the font package would have installed cleanly, changed nothing, and left a note in the build record saying Chinese rendering was taken care of. It was left out, on purpose, with the reason written down.
It also got moved. I’d said to build it on the first node of the cluster, out of habit. The agent counted: that node already runs 40 containers, shares its graphics chip with the media servers, and hosts the front door for every web address in the house. Changing a song’s key means re-encoding video on the CPU while someone is singing. It went on the node with the fastest processor and room to spare.
Zero results, in exactly the languages we built it for
The server came up, the QR code worked, and an English search returned ten results.
Then the agent set up monitoring, which meant picking a search to run every hour as a health check, and it tried the searches this house would really make. PiKaraoke has a tick box, on by default, that adds the word “karaoke” to whatever you type.
- 周杰伦 晴天 (Jay Chou, in Simplified characters): 0 results.
- 周杰倫 晴天 (the same, in Traditional): 0.
- 陳奕迅 富士山下 (Eason Chan): 0.
- IU 좋은 날 : 0.
Every Chinese and Korean query, nothing. Untick the box and type “KTV” or “노래방” yourself: ten results each time.
The cause was one pair of quotation marks. PiKaraoke was wrapping the whole query, the appended English word included, in quotes, so YouTube was being asked for that exact phrase: Chinese characters followed by the English word “karaoke”. Almost no video on earth is titled that way.
Somebody had already noticed the quoting was off. There was a pull request open on the project, three weeks old, from another contributor, waiting on a discussion about approach. Nothing in its test plan mentioned Chinese or Korean. The agent applied that change to our copy, re-ran the same queries, got ten out of ten on every one, and drafted a comment for the pull request with the before and after numbers. I read it and posted it under my name.
The maintainer replied a day later that he was surprised it hadn’t been caught sooner, and merged it. It isn’t in a release yet, so our copy still carries the patch, and the upgrade notes say to check for it every time.
Lesson: Test the thing in the language it will be used in. An English health check on a karaoke server for Chinese speakers proves the server is up and nothing else.
There is no official Chinese karaoke on YouTube
Korea was easy. The two big Korean karaoke companies, TJ Media and KY, both publish their catalogues on YouTube. The brief said to verify each channel is real and official before using it, and a channel name costs nothing to fake, so the test was a link from the company’s own website. KY passed. TJ’s website, oddly, links only to a small corporate channel with product videos, so TJ is “official” on the strength of YouTube’s badge and 1.82 million subscribers, and the notes say so in those words.
The first Korean list had a hole in it. A channel search turned out to stop at about 500 results without saying so, and three songs my partner would notice were missing. The list is now built from the artist search plus targeted searches for specific songs, merged. It also nearly shipped a girl group’s song in the male key, because the channels publish several versions of everything and the tie-break picked the longest. 118 songs went in, starting with her favourite group.
China was not easy. The agent ran 24 typical KTV searches and 19 label searches, then went through twelve verified official channels one by one: Jay Chou’s, Mayday’s label, Eason Chan’s, the Hong Kong labels. Millions of subscribers between them and not one official karaoke upload. The channels that do have them say in their own descriptions that they’re unofficial.
I had to make a call there, and it was that official is preferred, not required. So the agent picked on quality instead, and it didn’t go by titles. It downloaded a 30-second sample from each candidate, looked at a frame to see which script the lyrics were in, and measured the audio to find out whether the original singer was still in there. Two channels came out on top, one for Mandarin with Simplified lyrics and pinyin, one for Cantonese. 650 of 652 songs downloaded. The two that didn’t are written down.
Then the names. I’d chosen Simplified characters, because the person who’ll be searching reads Simplified. Most Cantonese and Taiwanese sources write in Traditional, and PiKaraoke searches filenames, so a search for 周杰伦 would sail straight past 周杰倫. Every name gets converted, and every Chinese title gets its pinyin in brackets, Cantonese songs included, because the searcher thinks in Mandarin.
Then people started using it

The player page on the TV, with the control panel open. Changing key is the slider at the bottom.
It worked on the TV. Songs got sung. And songs got added, because when PiKaraoke doesn’t have what you want it will fetch it from YouTube for you, then and there.
Those land in the top folder of the library, under whatever the uploader called the video. Here is a real one:
純音樂 毛不易 ─《像我這樣的人》Wild West KTV 伴唱 Karaoke 伴奏 西野
Beside a library of tidy, searchable names, that’s a mess, and it’s a mess I can’t sort. I can’t tell a Mandarin song from a Cantonese one by looking. Often neither can anyone else: both are written in Chinese characters, and Traditional script is used for Taiwanese Mandarin and Hong Kong Cantonese alike.
The agent’s first answer was a script with rules. My question back was whether we could do something smarter. The house has n8n, a tool for wiring services together into workflows. It has a box running AI models locally. Could the two of them look at the new songs each night and decide?
One model was 95% sure of everything
The agent wouldn’t wire anything up until it had measured the models. It built a test: 80 titles made to look like phone downloads, each labelled with the right answer, with the Chinese script deliberately flipped on some of them so a model couldn’t cheat by treating Traditional characters as Cantonese. Three local models sat it.
English was solved by all three, 16 out of 16. Pulling the artist and song out of the clutter was solved too. The whole problem was Mandarin against Cantonese, and here the two bigger models were both mediocre in a useful way:
- gpt-oss:20b got every Mandarin song right, 21 of 21, and only 13 of 22 Cantonese ones. When unsure, it says Mandarin.
- qwen3:30b-a3b got 21 of 22 Cantonese songs and only 10 of 22 Mandarin. When unsure, it says Cantonese.
Both scored 84% overall. The smallest model actually scored highest, at 90%.
The obvious gate is confidence: ask each model how sure it is and only act on the sure ones. The prompt asked for exactly that. One of the two said 0.95 on all 13 of its wrong answers, and 0.95 on 66 of its 67 right ones. The other spread its wrong answers from 0.1 to a flat 1.0, with most of them at 0.85 or higher, which is where its right answers sat too. The small model filed a Korean song by IU as Mandarin pop, at 0.95.
So the rule became agreement. If the Mandarin-leaning model and the Cantonese-leaning model give the same answer, file the song. On the test that was 55 right out of 55, with no wrong moves , and 25 songs left for a person. Adding the highest-scoring model to the vote made things worse: every pairing that included it filed songs wrongly, because its mistakes weren’t the opposite of anyone’s.
Lesson: A model’s confidence is another thing it wrote, not a measurement of anything. If you need a check, find a second opinion that’s wrong in a different direction.
The four o’clock shift
It runs like this. At 4 am the karaoke server lists the songs in its top folder that are more than an hour old and aren’t queued or playing, because nobody wants a file moved from under them mid-chorus. It sends the titles to the n8n workflow. The workflow asks both models the same question about each title and gets back a folder and the names. If they agree, the song is filed.

The whole workflow. Two questions, one decision, and a side road for the songs that are Chinese but nobody agrees which kind.
Two things are deliberately not left to the models.
The names. The models were asked for Simplified characters and pinyin and returned Traditional characters and English translations instead; one title came back as “Fuji Mountain Below”. So the server builds every Chinese name itself with two ordinary libraries, one for script conversion and one for pinyin.
The spelling of artists. On the first run the models called the singer 毛不易 “Mao Buyi” twice and “Mao Yibei” twice, which meant a search for him found half his songs. The library already knew the answer: hundreds of files in it were named with the artist’s established English name. So now the library’s own spelling always wins, and a model only gets to name an artist nobody has filed before.
That fix had a bug worth telling. The first time the agent re-filed the four songs, the two “Mao Yibei” files were counted as evidence for their own bad name. Two votes each. The tie was judged “already right”. Files being corrected no longer get a vote.
Everything it decides, and why, is appended to a log file in the library itself, at my request, dry runs included. This morning’s entry, in part:
=== 2026-10-10 04:09:32 AEDT karaoke-tidy EXECUTE
REVIEW G-DRAGON - TOO BAD (feat. Anderson .Paak) (Official Video)
[A English / B K-pop] models disagree: English vs K-pop
MOVE 古巨基 - 情歌王
-> Cantopop/Leo Ku 古巨基 - 情歌王 (Qing Ge Wang)
[A Cantopop / B Cantopop] both models: Cantopop
summary: 9 considered, 1 moved, 8 left for review
The G-Dragon one has been in the review pile for five mornings and both models have a point. He’s a K-pop artist and that song is in English. The folders are by the language a song is sung in, and I haven’t decided what I want there.
One more piece went in a few days later. When both models said “Chinese” but split on which kind, the song had been going to Mandarin by default. On the test, that default was wrong for 10 of 20. Now those songs, and only those, get one yes-or-no question put to Jev, a small hosted decision model from TypeSafe: is this sung in Cantonese? It got 19 of the 20. That’s the one step where a song title leaves the house, and if the service doesn’t answer, the old default applies.
The rename that said yes and did nothing
I asked for a few artist spellings to be made official: “MayDay” to “Mayday”, “AMei” to “A-Mei”. Thirteen files.
The rename ran, reported success, and the library went from 782 songs to 778.
Four Mayday songs had vanished from the app. They were still on the NAS, each under a temporary name. The NAS doesn’t care about capital letters, so “MayDay” and “Mayday” are the same file as far as it’s concerned, and when asked to rename one to the other through the server’s existing connection it answered “done” and changed nothing. The second step of the rename then had nothing to finish.
The agent found it because it counts the library before and after every change. It finished the four renames over a fresh connection to the NAS, the count went back to 782, and the script learned to send any capitals-only rename that way.
And that fix had a hole in it. The agent built the new rename as a line of shell text with the filenames pasted in. A different session, asked to review the repository, flagged it the same day: a song title containing an apostrophe would have ended the quoted text early, and whatever came after it in the title would have run as a command, as the administrator, on the machine that hosts the server. Song titles are full of apostrophes.
Nothing had triggered it. The agent rewrote it so no filename is ever part of a command, refused names it couldn’t pass safely instead of trying to escape them, and tested it with a file named it's a test; touch $(echo pwned).txt. The file was renamed. Nothing called pwned appeared anywhere.
Lesson: “Success” is what the system said, not what it did. Count the thing before and after. And a filename somebody downloaded is input from a stranger, even when it’s a love song.
Smaller traps, for the record
- The lyrics-file test never ran. Every song so far has its lyrics baked into the video. The one real risk the agent identified, Cantonese-only characters missing from PiKaraoke’s built-in font, is still a prediction.
- PiKaraoke updates its YouTube downloader every time it starts , by default. That makes every restart an unreviewed upgrade. Ours is pinned, with a rehearsed procedure for bumping it.
- A daily download check cried wolf. YouTube answered one check with “sign in to confirm you’re not a bot” and the monitor sat red for a day with nothing wrong. It now tries three times, five minutes apart, before it complains.
- Moving a file loses its place in PiKaraoke’s library , which only recognises a move when the name is unchanged. Play history survives anyway, because it’s keyed on the YouTube id, which is why every filename keeps that id on the end. Checked on nine plays, not assumed.
- The admin password proves nothing by its login page , which redirects the same way for a right password and a wrong one. The proof was a request to an admin-only page with each.
- The after-song score is a random number. The scores looked wrong, so the agent read the code: it never listens to anything. It rolls a number, nudged upwards to be kind, and picks a matching round of applause. A gimmick, and we turned it off.
- A JSON reply that isn’t JSON. Some mornings one model’s answer for a title doesn’t parse. That song waits. The Leo Ku song above failed that way five mornings running and was filed on the sixth, with nothing changed.
What it looks like now
- 813 songs : 123 K-pop, 467 Mandopop, 206 Cantopop, 9 English, and 8 in the top folder waiting for a person.
- Twelve 4 am runs since it went live. 40 songs filed by the two models.
- Four Mandarin-or-Cantonese splits settled by the yes-or-no question, all four as Mandarin.
- One log file on the NAS with every decision and the two models’ answers beside it.
- One fix upstream , merged, that makes Chinese and Korean searches work for everybody who installs PiKaraoke after the next release.
- Two songs that still won’t download.
The honest part
The earlier pieces each ended on what got the thing unstuck. This one was never stuck. It was for fun, and the fun part, the server, was the easy three-quarters.
What I didn’t see coming is that I can’t check any of it. On every other project in this series I’m the one who knows whether the answer is right. I know when the spa is warm. I know what my dashboard should show. Here the library is in Korean, Mandarin and Cantonese, and I can read none of it. When a model says a song is Cantonese, I have no opinion.
My usual contribution was gone, so the work had to stand without it. The agent built a marked exam before it trusted a model. It refused the confidence scores once it had seen them. It picked a rule that needs no judge, only two parties who tend to disagree. And everything the rule can’t settle goes in a pile, in plain view, for the one person in the house who can read it. I didn’t design any of that. I asked whether we could do something smarter and then approved things I could verify only by their numbers.
Now what isn’t right. The folders are mostly trustworthy and the names are less so. Two mornings ago a song called 沉都, by a rapper called Kafe Hu, was filed under the artist “Chen Dou 沉都” with the title “高音质”, which is not a song. It’s the uploader’s note that the video is high quality audio. Both models agreed on the language, so it moved, and agreement on the language says nothing about which words were the title. It’s still there as I write this. The name-building has no second opinion, and it needs one.
The folders aren’t spotless either. Two copies of a Jordan Chan song are sitting in Cantopop because he’s a Hong Kong singer and both models said so. As far as I can find out, it’s a Mandarin song. The models that disagree by nature agreed, and were wrong together, on exactly the kind of artist the test said would be hard.
The eight songs in the review pile have mostly been there a week, re-argued every morning at four with the same result. A pile that nobody’s been told to look at is just a slower mess. And the two-model rule has been tested on 80 titles I had labelled by their source, some of which were probably labelled wrong.
Still. There’s a karaoke server in the house because someone I love wanted to sing in her own language, and every morning at four two AI models argue about Cantopop so she can find the song the next day.
Build the thing for fun. Then measure the part you can’t judge yourself.
Under the hood: if you want to build one
What it runs on
- PiKaraoke 1.23.0, in headless mode as a systemd service, in a Debian 13 container on Proxmox. Installed with uv as an unprivileged user. Needs ffmpeg, and deno for YouTube’s JavaScript challenge.
-
yt-dlp 2026.8.19, pinned, with PiKaraoke started using
--skip-ytdl-upgrade. - PiKaraoke pull request 941: the search quoting fix. Merged, not yet released as of 1.23.0.
- n8n 2.41.6, self-hosted: one workflow behind a webhook.
- Ollama 0.35.1 with gpt-oss:20b and qwen3:30b-a3b. About 4.8 s and 0.9 s a title.
- OpenCC 1.1.9 for Traditional to Simplified, and pypinyin 0.54.0. Both are Debian packages.
- Jev, a hosted decision model, for one yes-or-no question on Mandarin against Cantonese splits.
Decisions
- File on agreement between two models with opposite biases, not on confidence. 55 of 55 on the test. One model reported 0.95 on 13 of 13 wrong answers and 66 of 67 right ones.
- Don’t add a third voter because it scores higher. A model whose mistakes aren’t opposite to the others’ makes the vote worse.
- Classify by the language the song is sung in , and tell the model that script is not decisive.
- n8n classifies; the server decides paths. The server re-checks every decision against a fixed list of folders, cleans the name and refuses collisions.
- Build Chinese names in code, not in the model. OpenCC and pypinyin are deterministic. The models return Traditional script and translate titles.
- The library’s existing spelling of an artist beats the model’s.
- Never move a queued or playing song , and nothing under an hour old.
- Fail safe. If n8n or Ollama is down, nothing moves and the log says so. If the hosted model doesn’t answer, the default stands.
- Keep the YouTube id in every filename. Play history is keyed on it.
- Rejected: letting the hosted model break every tie. On non-Chinese disagreements it filed songs wrongly, including a promo video as a song.
Gotchas
- With the karaoke filter on, PiKaraoke 1.23.0 returns nothing for Chinese and Korean searches. Apply the fix above, or untick the filter and type
KTV,伴奏or노래방. - A YouTube channel search stops at around 500 results without telling you.
- Karaoke channels publish male-key, female-key and melody-removed versions of the same song. Pick by the tag in the title, not by duration.
- A capitals-only rename on a case-insensitive SMB share, through an existing mount, can return success and do nothing. Use a fresh session, and count files before and after.
- Never build a shell command from a filename. Pass names on standard input or as arguments, and refuse the ones you can’t pass safely.
- PiKaraoke draws subtitle files in the browser, from its own bundled fonts. Fonts installed on the server do nothing.
- Its login page redirects whether the password is right or wrong.
- A local model at temperature 0 will still sometimes return a reply that doesn’t parse. Treat that as “no decision”.
- pypinyin has to guess at characters with two readings. 都 came out as “dou” where “du” was meant.
The rule itself
Copied from the deployed workflow’s decision step:
const CHINESE = ['Mandopop', 'Cantopop'];
let folder = null, reason, split = false;
if (!a || !b) reason = 'a model reply was malformed';
else if (ga === gb && ga !== 'Other') { folder = ga; reason = `both models: ${ga}`; }
else if (CHINESE.includes(ga) && CHINESE.includes(gb)) { folder = 'Mandopop'; split = true; reason = `Chinese split ${ga} vs ${gb}: Mandopop by default`; }
else if (ga === gb) reason = `both models: ${ga} (cannot tell)`;
else reason = `models disagree: ${ga} vs ${gb}`;
And the part of the prompt that does the most work:
Use your knowledge of the artist and the song. Script is NOT decisive: Traditional Chinese characters
are used for both Taiwanese Mandarin songs and Cantonese songs, and K-pop is often written in English.
Both models get the same prompt, a JSON schema for the reply, and temperature 0.
Take it with you
The workflow, the 4 am script and the timer units, with the house taken out: homelab-recipes/karaoke-tidy.
- Substitute: the webhook address, the Ollama address, the library path, your folder names, and the name of your PiKaraoke service.
- Proven: twelve 4 am runs on a real library, the agreement rule, the name building, the log.
- Not proven: anything outside Korean, Mandarin, Cantonese and English, and the artist and title split, which has no second check.
- The shared workflow leaves Mandarin or Cantonese splits for review. Mine asks Jev, which needs its own account, so that branch isn’t included. The raw test answers are, with a script that recomputes every figure above.

Top comments (0)