<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: deepika p</title>
    <description>The latest articles on DEV Community by deepika p (@deepika_p_df360028ce7e98a).</description>
    <link>https://dev.to/deepika_p_df360028ce7e98a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167256%2F3b8327ae-42fc-4489-b866-995da787bde7.jpg</url>
      <title>DEV Community: deepika p</title>
      <link>https://dev.to/deepika_p_df360028ce7e98a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/deepika_p_df360028ce7e98a"/>
    <language>en</language>
    <item>
      <title>Why Whisper "scores 8% WER" on dictation, and what actually fixed it</title>
      <dc:creator>deepika p</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:28:59 +0000</pubDate>
      <link>https://dev.to/deepika_p_df360028ce7e98a/why-whisper-scores-8-wer-on-dictation-and-what-actually-fixed-it-nh4</link>
      <guid>https://dev.to/deepika_p_df360028ce7e98a/why-whisper-scores-8-wer-on-dictation-and-what-actually-fixed-it-nh4</guid>
      <description>&lt;p&gt;I record dictation: notes to myself, commit messages, short specs, with a lot of product and library names in them. I ran faster-whisper large-v3-turbo over 101 of my own clips and got &lt;strong&gt;8.5% word error rate&lt;/strong&gt;. That sounded bad, so I read every error. Most of it wasn't what I expected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Setup, so you can judge the numbers:&lt;/strong&gt; one speaker (me), a quiet room, 101 clips, 1,057 reference words, &lt;code&gt;language="en"&lt;/code&gt;, beam 1, VAD on. The references are the text I meant to write, so an "um" or a spoken "scratch that" counts as an error if it survives into the transcript. Everything below comes from that one test set, and I'll say where it won't generalize.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Most of the "errors" weren't mishearings
&lt;/h3&gt;

&lt;p&gt;Of 90 error words, about 60 were not recognition mistakes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fillers ("um", "uh")&lt;/li&gt;
&lt;li&gt;spoken corrections ("Tuesday, no, Wednesday")&lt;/li&gt;
&lt;li&gt;voice commands ("new paragraph")&lt;/li&gt;
&lt;li&gt;number formatting ("three thirty p m" instead of "3:30 PM")&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A plain rules pass for fillers, voice commands and number formatting took WER from &lt;strong&gt;8.51% to 6.05%&lt;/strong&gt; with no model change. In an earlier run, a small LLM cleanup on top got it to about 2.5%, at roughly +450 ms. If you benchmark dictation, check what your reference text assumes.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. A glossary prompt helps, if you phrase it as a sentence
&lt;/h3&gt;

&lt;p&gt;Whisper accepts an initial prompt. Instead of a bare list of terms, I use a sentence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;faster_whisper&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;WhisperModel&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;WhisperModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;large-v3-turbo&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cuda&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;compute_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;float16&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;terms&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Kubernetes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PostgreSQL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Jetpack Compose&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Notes about &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;terms&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; and &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;terms&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;segments&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;info&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clip.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;# fix the language, don't auto-detect
&lt;/span&gt;&lt;span class="n"&gt;initial_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="n"&gt;vad_filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;# keep this on, see below
&lt;/span&gt;&lt;span class="n"&gt;beam_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things I measured that matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keep VAD on.&lt;/strong&gt; With VAD off and a prompt set, all 80 of my silence clips came back with hallucinated text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The prompt has a budget.&lt;/strong&gt; It's about 223 tokens, so it works for roughly a dozen to a hundred terms. It fails at around a thousand, with truncation and false insertions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On my 15 terms (19 spoken occurrences), the plain model got 12 right. The prompt alone got 16, and the full version below got all 19.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. A rewriter for the misses, and a bug I shipped
&lt;/h3&gt;

&lt;p&gt;A prompt doesn't catch everything ("Postgres sequel", or a name spelled three different ways). So after decoding I run a small phonetic rewriter: slide a window of 1 to 3 words over the transcript, compare each window to each dictionary term, and pick the best non-overlapping replacements.&lt;/p&gt;

&lt;p&gt;My first version scored a candidate as &lt;code&gt;similarity x window length in characters&lt;/code&gt;. That rewards longer windows, so "to Kubernetes" beat the exact word "Kubernetes", and the rewriter swallowed the "to". It happened on 4 of the 6 clips where a rewrite fired. The fix was one line: weight by the length of the &lt;strong&gt;matched term&lt;/strong&gt;, not the window.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode (101 clips)&lt;/th&gt;
&lt;th&gt;WER&lt;/th&gt;
&lt;th&gt;Jargon-clip WER (13 clips)&lt;/th&gt;
&lt;th&gt;Terms right (of 19)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;plain&lt;/td&gt;
&lt;td&gt;8.51%&lt;/td&gt;
&lt;td&gt;12.04%&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prompt + rewriter, with the bug&lt;/td&gt;
&lt;td&gt;7.95%&lt;/td&gt;
&lt;td&gt;8.33%&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prompt + rewriter, fixed&lt;/td&gt;
&lt;td&gt;7.57%&lt;/td&gt;
&lt;td&gt;5.56%&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;fixed, plus rules cleanup&lt;/td&gt;
&lt;td&gt;5.49%&lt;/td&gt;
&lt;td&gt;5.56%&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I replayed the fix against the other test sets I had: 485 real clips with no dictionary terms (zero outputs changed, zero false insertions), 450 synthetic term clips, and dictionaries padded with distractor terms. It was neutral or better everywhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. What didn't work, and what I'd distrust
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning on public data hurt.&lt;/strong&gt; LoRA on public sets cut error on a meeting-style benchmark from 12.5% to 5.9% in 5 hours of data, then flattened, and made my own dictation set worse (8.2% to 11.3% at 50 hours). Spoken-form labels teach the model to unlearn numerals and currency formatting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The "Notes about..." prompt has side effects.&lt;/strong&gt; It makes Whisper title-case list items and sometimes drop a small word ("a jacket, a charger" became "Jacket, Charger"). On some clips it also types a spoken "period" as the literal word. I haven't fixed either.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Noise breaks it first.&lt;/strong&gt; Mixing many-talker babble in at 10 dB added about 20 points of WER, and digits suffered most.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The jargon numbers are in-sample.&lt;/strong&gt; Those 15 terms are the ones I tuned on, with one speaker. The overall gain from the prompt and rewriter (8.51% to 7.57%) is within noise. The gain from the rules pass is not. I'd want 20+ speakers before claiming anything on new audio. I'm not publishing the clips, since they're my voice.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. If you want to try it
&lt;/h3&gt;

&lt;p&gt;I run this work behind a small batch transcription API, because I wanted it for my own tools. You can try it in the browser with no API key at [&lt;a href="https://www.google.com/url?q=http://readaloudai.org/transcribe%5D(https://readaloudai.org/transcribe?utm_source%3Ddevto%26utm_medium%3Dpost%26utm_campaign%3Ds1-stt-20261006)&amp;amp;source=gmail&amp;amp;ust=1791397468821000&amp;amp;sa=E" rel="noopener noreferrer"&gt;https://www.google.com/url?q=http://readaloudai.org/transcribe%5D(https://readaloudai.org/transcribe?utm_source%3Ddevto%26utm_medium%3Dpost%26utm_campaign%3Ds1-stt-20261006)&amp;amp;source=gmail&amp;amp;ust=1791397468821000&amp;amp;sa=E&lt;/a&gt;. With an API key it's $0.11 per hour of audio, billed by the second. A new account gets free credits worth about 54 minutes of transcription, and you don't need a card to start.&lt;/p&gt;

&lt;p&gt;Straight talk, taken from the docs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Batch only: no live streaming, no speaker labels.&lt;/li&gt;
&lt;li&gt;Accuracy is measured on English only.&lt;/li&gt;
&lt;li&gt;The first request after a quiet period can take 10 to 18 seconds while a GPU starts, then a short clip takes about a second.&lt;/li&gt;
&lt;li&gt;Whisper occasionally invents text over silence or noise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Disclosure: I build it. If it gets your terms wrong on a clip, I'd like to hear the example.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>whisper</category>
      <category>python</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
