<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Riley Sy</title>
    <description>The latest articles on DEV Community by Riley Sy (@ruoning).</description>
    <link>https://dev.to/ruoning</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4169295%2F7081a482-a904-4394-92c7-584ad6a29f7f.jpg</url>
      <title>DEV Community: Riley Sy</title>
      <link>https://dev.to/ruoning</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ruoning"/>
    <language>en</language>
    <item>
      <title>Spellcheck can't fix "chai GPT": building an LLM caption checker on $5 a month</title>
      <dc:creator>Riley Sy</dc:creator>
      <pubDate>Thu, 08 Oct 2026 09:10:46 +0000</pubDate>
      <link>https://dev.to/ruoning/spellcheck-cant-fix-chai-gpt-building-an-llm-caption-checker-on-5-a-month-1c94</link>
      <guid>https://dev.to/ruoning/spellcheck-cant-fix-chai-gpt-building-an-llm-caption-checker-on-5-a-month-1c94</guid>
      <description>&lt;p&gt;Here's a line from YouTube's auto-captions on a lecture about attention in neural networks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;consider heart attention looking at the architecture heart attention is very similar to soft attention&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The speaker said "hard attention" aka the counterpart of soft attention. The captions get it wrong seven times in one video. Every word is technically spelled correctly, so no traditional spellchecker will ever flag it.&lt;/p&gt;

&lt;p&gt;The same video has "tange" for tanh, "RN" for RNN, "grin you ality" for granularity, and "atencion" for attention, in a video that's about attention.&lt;/p&gt;

&lt;p&gt;I built &lt;a href="https://misheard.fly.dev/" rel="noopener noreferrer"&gt;Misheard&lt;/a&gt; to find these. You upload an SRT or VTT file and it flags the words that were probably misheard, suggests a fix for each, and lets you accept or reject them before you export. This post covers how it got there: the detector that didn't work, how I measured that, and what it took to put it online for strangers on a $5/month budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why spellcheck isn't enough
&lt;/h2&gt;

&lt;p&gt;Caption errors come in two kinds. A non-word (or non-English word in this case) like "atencion" isn't in any dictionary, so it's easy to catch. A real-word error like "heart" for "hard" is a word, it's just in the wrong place. You can only catch it by understanding the sentence.&lt;/p&gt;

&lt;p&gt;To see which kind mattered, I listened to five videos all the way through and logged every caption error I heard: 105 in total. Over half were real-word errors (55), and only 41 were non-words. The rest were formatting slips. So a tool that only catches misspellings misses most of the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt one: clever local detectors
&lt;/h2&gt;

&lt;p&gt;My first version ran entirely on my laptop, with no LLM. It stacked a few detectors:&lt;br&gt;
• Out-of-vocabulary: flag words that are rare in wordfreq and not a known term.&lt;br&gt;
• Phonetic match: flag words that sound like a domain term (Metaphone codes plus string similarity), so "gann" might be a misheard name.&lt;br&gt;
• Split words: catch "grin you ality" by joining neighbours and checking the result.&lt;br&gt;
Against my 105 logged errors, it caught nearly every non-word but only 9 of 55 real-word errors. So I tried to close the gap:&lt;br&gt;
• Sentence embeddings (MiniLM) caught 0 of the 46 real-word misses.&lt;br&gt;
• A masked language model (RoBERTa fill-mask, asking "is a similar-sounding word far more likely here?") caught about 5 more, at the cost of 17 to 22 extra flags. Only about 1 in 5 of those flags was right.&lt;/p&gt;

&lt;p&gt;The phonetic detector also had a habit of "correcting" names that were right. Chollet became "should", Demis became "times", Yoshua became "wish".&lt;/p&gt;

&lt;p&gt;The most useful thing I did in this phase was the listening pass itself. My first test set was built from what the tool had flagged, so it could never show me what the tool missed. Once I had an exhaustive list, I could measure two things honestly: recall (what share of real errors got flagged) and flag precision (what share of flags were real errors).&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt two: let an LLM read it
&lt;/h2&gt;

&lt;p&gt;The idea that worked was the obvious one: have an LLM read the whole transcript, the way a person proofreading it would. I call this Read-through. The transcript goes out in chunks of about 400 words. Each chunk carries the local detectors' flags as hints, plus priming terms (names and jargon from the video title, for example). The model replies with the words it thinks were misheard and a fix for each.&lt;/p&gt;

&lt;p&gt;Two prompt changes made it usable:&lt;br&gt;
• Ask for a cause. Each verdict says whether the error is misheard, grammar or style, and I keep only misheard. Without this, the model rewrites the speaker's grammar or delivery.&lt;br&gt;
• "Never list words you judged correct." Models love reporting "checked, fine". Banning that cut output tokens about 5x, and output tokens were most of the cost.&lt;/p&gt;

&lt;p&gt;I didn't want to fool myself by tuning on the data I'd judge on. So before running anything, I wrote down the rule for which approach wins and set aside test sets I wouldn't look at while tuning. One was five new videos I listened through by hand. The other was eleven calls from &lt;a href="https://huggingface.co/datasets/Revai/earnings21" rel="noopener noreferrer"&gt;Earnings-21&lt;/a&gt;, a public set of earnings-call recordings, captioned by a commercial speech-to-text service.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Held-out set&lt;/th&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Real-world errors caught&lt;/th&gt;
&lt;th&gt;Flag precision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;5 hand-checked videos&lt;/td&gt;
&lt;td&gt;Local detectors&lt;/td&gt;
&lt;td&gt;6 of 79&lt;/td&gt;
&lt;td&gt;0.34&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5 hand-checked videos&lt;/td&gt;
&lt;td&gt;Read-through&lt;/td&gt;
&lt;td&gt;22 of 79&lt;/td&gt;
&lt;td&gt;0.78&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11 earnings calls&lt;/td&gt;
&lt;td&gt;Local detectors&lt;/td&gt;
&lt;td&gt;49 of 3,985&lt;/td&gt;
&lt;td&gt;0.51&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11 earnings calls&lt;/td&gt;
&lt;td&gt;Read-through&lt;/td&gt;
&lt;td&gt;409 of 3,985&lt;/td&gt;
&lt;td&gt;0.87&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's 3.7x and 8.3x the real-word catches, with better precision, at about $0.05 per hour of audio.&lt;/p&gt;

&lt;p&gt;Then I compared six models through OpenRouter on a $15 experiment budget, again judged on fresh held-out sets. Qwen 3.6 Plus beat my previous default, Gemini 2.5 Flash: 713 vs 383 real-word catches on ten more earnings calls, at similar precision (0.85 vs 0.84) and the same $0.05/hour.&lt;/p&gt;

&lt;p&gt;Recall is still far from perfect: about 3 real-word errors in 10 on hand-checked YouTube videos. That's why Misheard suggests fixes rather than applying them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting it online for strangers on $5 a month
&lt;/h2&gt;

&lt;p&gt;I don't know anyone who'd use this, so it had to work for strangers. I also didn't want a surprise bill. These are the choices that came out of that:&lt;br&gt;
• A free tier with hard limits. Each visitor gets 10,000 words a day (about an hour of speech). All visitors share a daily budget of $0.25. The server's OpenRouter key has its own $5/month credit cap, so the worst case is a refusal message, not a bill.&lt;br&gt;
• Bring your own key, kept in your browser. If you paste your own OpenRouter key, it lives only in your browser's localStorage and goes along with each request. The server never stores it, and a Forget button clears it.&lt;br&gt;
• Nothing kept for long. Uploads are deleted 24 hours after your last activity, and there's a Delete button next to Export.&lt;br&gt;
• One small box. It's a single always-on Fly.io machine with a volume. Correcting runs synchronously, and the server refuses a second run on the same transcript while one is going. A background job queue can wait until someone needs it.&lt;br&gt;
• No third-party analytics. The server counts visits, uploads, runs and exports in a log file. That's how I'll know if this post sent anyone.&lt;/p&gt;

&lt;p&gt;The last piece was a demo. Downloading caption files from YouTube is too much to ask of someone who clicked a link out of curiosity. So the home page offers a saved Example: CodeEmporium's &lt;a href="https://www.youtube.com/watch?v=W2rWgXJBZhU" rel="noopener noreferrer"&gt;"Attention in Neural Networks"&lt;/a&gt; (CC BY), already checked, with the video playing beside each Flag. That's where "heart attention" came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it still gets wrong
&lt;/h2&gt;

&lt;p&gt;The Example shows the misses too, because I didn't want to hide them:&lt;br&gt;
• "neural machine translation nmt systems": the captions were right here. My local detector flagged "nmt" as sounding like "Nomad", because its term list leans toward AI company and model names.&lt;br&gt;
• "Microsoft's attention gann": the speaker meant AttnGAN. The tool flagged the right word, but its suggestions ("Qwen", "gene") were wrong.&lt;br&gt;
• "attention Gantz": probably "attention GANs". The LLM suggested "generation" with 95% confidence.&lt;/p&gt;

&lt;p&gt;More generally, it catches roughly 3 in 10 real-word errors, so a clean result doesn't mean clean captions. On one test set, the LLM caught fewer of the plain misspellings than my local detectors did (9 vs 14 of 15). That's why the local flags still go along as hints. It's also English-only for now.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell myself at the start
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Label everything before you measure anything. A test set built from your tool's own output can't show you what the tool misses.&lt;/li&gt;
&lt;li&gt;Write the rule for winning before you look at the results. I changed rules twice mid-experiment, and both times I posted the change before running anything it could affect.&lt;/li&gt;
&lt;li&gt;Prompt wording costs money. Asking for less output did more for cost than switching models.&lt;/li&gt;
&lt;li&gt;Cap your spending in code and at the provider. Hard limits are what let me put a paid API in front of strangers at all.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Open misheard.fly.dev and click the Example to see the Flags above in context, with the video beside them. You don't need an account or a key. If you have an SRT or VTT file of your own, upload it. The free tier covers about an hour of speech a day.&lt;/p&gt;

&lt;p&gt;I'd really like to hear where it's wrong. That means misses, bad suggestions, and confusing parts of the review page. Leave a comment here or email &lt;a href="mailto:misheard_app@outlook.com"&gt;misheard_app@outlook.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Lastly if you've made it all the way here then thanks for reading!&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>ai</category>
      <category>python</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
