I record dictation: notes to myself, commit messages, short specs, with a lot of product and library names in them. I ran faster-whisper large-v3-turbo over 101 of my own clips and got 8.5% word error rate. That sounded bad, so I read every error. Most of it wasn't what I expected.
Setup, so you can judge the numbers: one speaker (me), a quiet room, 101 clips, 1,057 reference words, language="en", beam 1, VAD on. The references are the text I meant to write, so an "um" or a spoken "scratch that" counts as an error if it survives into the transcript. Everything below comes from that one test set, and I'll say where it won't generalize.
1. Most of the "errors" weren't mishearings
Of 90 error words, about 60 were not recognition mistakes:
- fillers ("um", "uh")
- spoken corrections ("Tuesday, no, Wednesday")
- voice commands ("new paragraph")
- number formatting ("three thirty p m" instead of "3:30 PM")
A plain rules pass for fillers, voice commands and number formatting took WER from 8.51% to 6.05% with no model change. In an earlier run, a small LLM cleanup on top got it to about 2.5%, at roughly +450 ms. If you benchmark dictation, check what your reference text assumes.
2. A glossary prompt helps, if you phrase it as a sentence
Whisper accepts an initial prompt. Instead of a bare list of terms, I use a sentence:
from faster_whisper import WhisperModel
model = WhisperModel("large-v3-turbo", device="cuda", compute_type="float16")
terms = ["Kubernetes", "PostgreSQL", "Jetpack Compose"]
prompt = "Notes about " + ", ".join(terms[:-1]) + " and " + terms[-1] + "."
segments, info = model.transcribe(
"clip.wav",
language="en", # fix the language, don't auto-detect
initial_prompt=prompt,
vad_filter=True, # keep this on, see below
beam_size=1,
)
Two things I measured that matter:
- Keep VAD on. With VAD off and a prompt set, all 80 of my silence clips came back with hallucinated text.
- The prompt has a budget. It's about 223 tokens, so it works for roughly a dozen to a hundred terms. It fails at around a thousand, with truncation and false insertions.
On my 15 terms (19 spoken occurrences), the plain model got 12 right. The prompt alone got 16, and the full version below got all 19.
3. A rewriter for the misses, and a bug I shipped
A prompt doesn't catch everything ("Postgres sequel", or a name spelled three different ways). So after decoding I run a small phonetic rewriter: slide a window of 1 to 3 words over the transcript, compare each window to each dictionary term, and pick the best non-overlapping replacements.
My first version scored a candidate as similarity x window length in characters. That rewards longer windows, so "to Kubernetes" beat the exact word "Kubernetes", and the rewriter swallowed the "to". It happened on 4 of the 6 clips where a rewrite fired. The fix was one line: weight by the length of the matched term, not the window.
| Mode (101 clips) | WER | Jargon-clip WER (13 clips) | Terms right (of 19) |
|---|---|---|---|
| plain | 8.51% | 12.04% | 12 |
| prompt + rewriter, with the bug | 7.95% | 8.33% | 19 |
| prompt + rewriter, fixed | 7.57% | 5.56% | 19 |
| fixed, plus rules cleanup | 5.49% | 5.56% | 19 |
I replayed the fix against the other test sets I had: 485 real clips with no dictionary terms (zero outputs changed, zero false insertions), 450 synthetic term clips, and dictionaries padded with distractor terms. It was neutral or better everywhere.
4. What didn't work, and what I'd distrust
- Fine-tuning on public data hurt. LoRA on public sets cut error on a meeting-style benchmark from 12.5% to 5.9% in 5 hours of data, then flattened, and made my own dictation set worse (8.2% to 11.3% at 50 hours). Spoken-form labels teach the model to unlearn numerals and currency formatting.
- The "Notes about..." prompt has side effects. It makes Whisper title-case list items and sometimes drop a small word ("a jacket, a charger" became "Jacket, Charger"). On some clips it also types a spoken "period" as the literal word. I haven't fixed either.
- Noise breaks it first. Mixing many-talker babble in at 10 dB added about 20 points of WER, and digits suffered most.
- The jargon numbers are in-sample. Those 15 terms are the ones I tuned on, with one speaker. The overall gain from the prompt and rewriter (8.51% to 7.57%) is within noise. The gain from the rules pass is not. I'd want 20+ speakers before claiming anything on new audio. I'm not publishing the clips, since they're my voice.
5. If you want to try it
I run this work behind a small batch transcription API, because I wanted it for my own tools. You can try it in the browser with no API key at [https://www.google.com/url?q=http://readaloudai.org/transcribe%5D(https://readaloudai.org/transcribe?utm_source%3Ddevto%26utm_medium%3Dpost%26utm_campaign%3Ds1-stt-20261006)&source=gmail&ust=1791397468821000&sa=E. With an API key it's $0.11 per hour of audio, billed by the second. A new account gets free credits worth about 54 minutes of transcription, and you don't need a card to start.
Straight talk, taken from the docs:
- Batch only: no live streaming, no speaker labels.
- Accuracy is measured on English only.
- The first request after a quiet period can take 10 to 18 seconds while a GPU starts, then a short clip takes about a second.
- Whisper occasionally invents text over silence or noise.
Disclosure: I build it. If it gets your terms wrong on a clip, I'd like to hear the example.
Top comments (0)