DEV Community

Cover image for Spellcheck can't fix "chai GPT": building an LLM caption checker on $5 a month
Riley Sy
Riley Sy

Posted on

Spellcheck can't fix "chai GPT": building an LLM caption checker on $5 a month

Here's a line from YouTube's auto-captions on a lecture about attention in neural networks:

consider heart attention looking at the architecture heart attention is very similar to soft attention

The speaker said "hard attention" aka the counterpart of soft attention. The captions get it wrong seven times in one video. Every word is technically spelled correctly, so no traditional spellchecker will ever flag it.

The same video has "tange" for tanh, "RN" for RNN, "grin you ality" for granularity, and "atencion" for attention, in a video that's about attention.

I built Misheard to find these. You upload an SRT or VTT file and it flags the words that were probably misheard, suggests a fix for each, and lets you accept or reject them before you export. This post covers how it got there: the detector that didn't work, how I measured that, and what it took to put it online for strangers on a $5/month budget.

Why spellcheck isn't enough

Caption errors come in two kinds. A non-word (or non-English word in this case) like "atencion" isn't in any dictionary, so it's easy to catch. A real-word error like "heart" for "hard" is a word, it's just in the wrong place. You can only catch it by understanding the sentence.

To see which kind mattered, I listened to five videos all the way through and logged every caption error I heard: 105 in total. Over half were real-word errors (55), and only 41 were non-words. The rest were formatting slips. So a tool that only catches misspellings misses most of the problem.

Attempt one: clever local detectors

My first version ran entirely on my laptop, with no LLM. It stacked a few detectors:
• Out-of-vocabulary: flag words that are rare in wordfreq and not a known term.
• Phonetic match: flag words that sound like a domain term (Metaphone codes plus string similarity), so "gann" might be a misheard name.
• Split words: catch "grin you ality" by joining neighbours and checking the result.
Against my 105 logged errors, it caught nearly every non-word but only 9 of 55 real-word errors. So I tried to close the gap:
• Sentence embeddings (MiniLM) caught 0 of the 46 real-word misses.
• A masked language model (RoBERTa fill-mask, asking "is a similar-sounding word far more likely here?") caught about 5 more, at the cost of 17 to 22 extra flags. Only about 1 in 5 of those flags was right.

The phonetic detector also had a habit of "correcting" names that were right. Chollet became "should", Demis became "times", Yoshua became "wish".

The most useful thing I did in this phase was the listening pass itself. My first test set was built from what the tool had flagged, so it could never show me what the tool missed. Once I had an exhaustive list, I could measure two things honestly: recall (what share of real errors got flagged) and flag precision (what share of flags were real errors).

Attempt two: let an LLM read it

The idea that worked was the obvious one: have an LLM read the whole transcript, the way a person proofreading it would. I call this Read-through. The transcript goes out in chunks of about 400 words. Each chunk carries the local detectors' flags as hints, plus priming terms (names and jargon from the video title, for example). The model replies with the words it thinks were misheard and a fix for each.

Two prompt changes made it usable:
• Ask for a cause. Each verdict says whether the error is misheard, grammar or style, and I keep only misheard. Without this, the model rewrites the speaker's grammar or delivery.
• "Never list words you judged correct." Models love reporting "checked, fine". Banning that cut output tokens about 5x, and output tokens were most of the cost.

I didn't want to fool myself by tuning on the data I'd judge on. So before running anything, I wrote down the rule for which approach wins and set aside test sets I wouldn't look at while tuning. One was five new videos I listened through by hand. The other was eleven calls from Earnings-21, a public set of earnings-call recordings, captioned by a commercial speech-to-text service.

Held-out set Approach Real-world errors caught Flag precision
5 hand-checked videos Local detectors 6 of 79 0.34
5 hand-checked videos Read-through 22 of 79 0.78
11 earnings calls Local detectors 49 of 3,985 0.51
11 earnings calls Read-through 409 of 3,985 0.87

That's 3.7x and 8.3x the real-word catches, with better precision, at about $0.05 per hour of audio.

Then I compared six models through OpenRouter on a $15 experiment budget, again judged on fresh held-out sets. Qwen 3.6 Plus beat my previous default, Gemini 2.5 Flash: 713 vs 383 real-word catches on ten more earnings calls, at similar precision (0.85 vs 0.84) and the same $0.05/hour.

Recall is still far from perfect: about 3 real-word errors in 10 on hand-checked YouTube videos. That's why Misheard suggests fixes rather than applying them.

Putting it online for strangers on $5 a month

I don't know anyone who'd use this, so it had to work for strangers. I also didn't want a surprise bill. These are the choices that came out of that:
• A free tier with hard limits. Each visitor gets 10,000 words a day (about an hour of speech). All visitors share a daily budget of $0.25. The server's OpenRouter key has its own $5/month credit cap, so the worst case is a refusal message, not a bill.
• Bring your own key, kept in your browser. If you paste your own OpenRouter key, it lives only in your browser's localStorage and goes along with each request. The server never stores it, and a Forget button clears it.
• Nothing kept for long. Uploads are deleted 24 hours after your last activity, and there's a Delete button next to Export.
• One small box. It's a single always-on Fly.io machine with a volume. Correcting runs synchronously, and the server refuses a second run on the same transcript while one is going. A background job queue can wait until someone needs it.
• No third-party analytics. The server counts visits, uploads, runs and exports in a log file. That's how I'll know if this post sent anyone.

The last piece was a demo. Downloading caption files from YouTube is too much to ask of someone who clicked a link out of curiosity. So the home page offers a saved Example: CodeEmporium's "Attention in Neural Networks" (CC BY), already checked, with the video playing beside each Flag. That's where "heart attention" came from.

What it still gets wrong

The Example shows the misses too, because I didn't want to hide them:
• "neural machine translation nmt systems": the captions were right here. My local detector flagged "nmt" as sounding like "Nomad", because its term list leans toward AI company and model names.
• "Microsoft's attention gann": the speaker meant AttnGAN. The tool flagged the right word, but its suggestions ("Qwen", "gene") were wrong.
• "attention Gantz": probably "attention GANs". The LLM suggested "generation" with 95% confidence.

More generally, it catches roughly 3 in 10 real-word errors, so a clean result doesn't mean clean captions. On one test set, the LLM caught fewer of the plain misspellings than my local detectors did (9 vs 14 of 15). That's why the local flags still go along as hints. It's also English-only for now.

What I'd tell myself at the start

  1. Label everything before you measure anything. A test set built from your tool's own output can't show you what the tool misses.
  2. Write the rule for winning before you look at the results. I changed rules twice mid-experiment, and both times I posted the change before running anything it could affect.
  3. Prompt wording costs money. Asking for less output did more for cost than switching models.
  4. Cap your spending in code and at the provider. Hard limits are what let me put a paid API in front of strangers at all.

Try it

Open misheard.fly.dev and click the Example to see the Flags above in context, with the video beside them. You don't need an account or a key. If you have an SRT or VTT file of your own, upload it. The free tier covers about an hour of speech a day.

I'd really like to hear where it's wrong. That means misses, bad suggestions, and confusing parts of the review page. Leave a comment here or email misheard_app@outlook.com

Lastly if you've made it all the way here then thanks for reading!

Top comments (6)

Collapse
 
makeev profile image
Mikhail Makeev •

Real-word errors outnumbering the misspellings is the part that maps onto ticker tagging. Our check that an LLM-suggested ticker really belongs to an article accepted any word of the company's name longer than two letters. That's how Voyager Therapeutics ended up on a story about Viking Therapeutics, matched on "therapeutics". T1 Energy got onto a TotalEnergies story the same way. The fix asks for a word that isn't in the dictionary, and that check already existed in our code but only ran on crypto names. Since your term list leans toward AI names, have the earnings calls shown you company names that look like ordinary words?

Collapse
 
ruoning profile image
Riley Sy •

Yes, the earnings calls are actually full of them, and in both directions too.

Some company names are ordinary words (Spire, Mosaic, Travelers, Yeti), so a dictionary check passes them without question. More often, the ASR (which is from 2021 so it's admittedly outdated) turns an unusual name into an ordinary word: Culp became "cult" and "coal", Spire became "fire", and Danaher's Cytiva and Beckman came out as "sativa" and "Batman".

Those are the real-word errors a misspelling check can't catch, and that's why the call's company name goes in as a Priming term for added context. Your fix, requiring a match on a word that isn't in the dictionary is kinda like the same lesson from the other side. Only a distinctive word in a name is evidence.

Collapse
 
reidmarlow profile image
Reid Marlow •

Splitting the error cause into misheard versus grammar or style is the constraint that keeps transcript pipelines from turning into uninvited copy editors. When I ran LLM cleanup passes over Whisper transcripts without a cause tag, the model quietly rewrote spoken hesitation, dropped filler phrases, and drifted the word counts enough that timestamp alignment broke downstream. On the 400-word chunk boundary, overlapping the last two sentences as read-only context helps with real-word homophones that depend on a term introduced right before the cut. A chunk that starts cold on "heart attention" has to guess until the next mention, whereas carrying the prior sentence's "soft attention" across the boundary resolves the contrast on line one without adding to the output token budget.

Collapse
 
ruoning profile image
Riley Sy •

Yup splitting the error is definitely what limits the LLM's "corrections" from being it's own idealised paraphrasing!

The drift problem is also handled structurally: the model never returns text, only word-index spans plus a replacement. Each fix is spliced into its caption at the exact character offset, and nothing outside it changes, so a chatty model can't shift word counts or timings.

Regarding boundaries each chunk does carry about 40 words of read-only context on each side. I use a word count instead of sentences because YouTube auto-captions sometimes have no punctuation to find sentences by (the example video's transcript for example has none). Chunks are meant to close at a sentence end after ~400 words, but on unpunctuated captions they run to an 800-word cap.

Fun fact: the first mention of "heart attention" is about 100 words into its chunk, with "soft attention" in the same chunk, so the model had the contrast in view. Your cold-start case is real, though, and I haven't measured how often it happens.

Collapse
 
analista_83 profile image
Sammi De Blas •

Nice breakdown of the two error classes. For the real-word kind, one cheap layer between spellcheck and the LLM: run a domain lexicon with phonetic/Levenshtein matching over a sliding window (tanh vs tange, hard vs heart) so the model only adjudicates flagged spans instead of re-reading every line - that's what keeps the $5/month budget honest as uploads grow. And your accept/reject UI is quietly collecting labeled data: after a few hundred decisions you can fit a small classifier on (context, caption word, suggested fix) triples and use it to pre-rank the LLM's suggestions.

Collapse
 
ruoning profile image
Riley Sy •

Thanks! There actually is a cheap phonetic layer: one detector matches against a domain lexicon, and another catches a term spelled differently elsewhere in the same transcript. It deliberately skips common English words, though, because "heart" vs "hard" is a perfectly good word on both sides.

Unfortunately checking every ordinary word against the lexicon buried the real errors in false alarms. I actually started with the design you described, where the LLM only judged flagged spans. Letting it read everything, with the local flags as hints, it found several times more real-word errors at higher precision, for about 5–6¢ per hour of audio. So for this kind of error, reading everything turned out to be what makes it worth the money.

On the labeled data I do think that it's a good idea but right now uploads and decisions are deleted after 24 hours. So learning from reviews is something I'd revisit once people are actually using it, and doing it right means changing that retention promise transparently.