What I Benchmarked
I've been fine-tuning Whisper on Nigerian speech for a while, and I built this benchmark to measure my model against the standard ones. The question behind it: when a speech model hears Nigerian English or Pidgin, does it write what was said, or does it rewrite it as standard English?
It evaluates speech recognition rather than text LLMs: two off-the-shelf Whisper models (large-v3 and small) and a version of large-v3 that I fine-tuned myself on Nigerian speech.
I use 'dialect' for any Nigerian English variety or Pidgin form, even though Pidgin is technically its own language.
Standard automatic speech recognition (ASR) accuracy tests mark a model wrong when it correctly transcribes dialect words, slang, or Pidgin. Most evaluation relies on word error rate (WER) computed against references written in standard English spelling. If a model transcribes "pipo" (what the speaker said) but the reference says "people", standard WER counts a substitution error. That penalizes a spelling difference, not a listening mistake.
The benchmark separates real recognition errors from dialect-spelling mismatches and reports 5 metrics:
- Raw WER: standard WER.
- Faithful WER: WER after discounting faithful dialect variants (spellings like "pipo" that match what was spoken).
- Dialect penalty: Raw WER minus Faithful WER, in percentage points.
- Hallucination rate: the share of word errors that aren't faithful dialect variants. In this benchmark's code, "hallucination" means any such error, not only invented words.
- Normalization overcorrection: the share of hallucinations where the reference contains a dialect word and the model "corrects" it to standard English.
For example, if a model makes 100 hallucinations (errors that aren't faithful dialect variants), a 10.6% overcorrection rate means about 11 of them are dialect words the model replaced with standard English, like "dey" → "they". The rest are other errors.
I selected 500 samples across 5 speaker backgrounds (Igbo, Yoruba, Hausa, Pidgin, multilingual), ran 3 Whisper models on them, and built an auto-classifier with 120 known dialect substitution rules to compute dialect-aware metrics without waiting for full human annotation.
Models Tested
Accent Labs is a project I started as a data cleaning pipeline for speech datasets. I then used that pipeline to prepare the ~24K samples I fine-tuned Whisper large-v3 on. The fine-tuned model was trained on Nigerian English and Pidgin speech with accent-weighted sampling.
| Model | Parameters | Role |
|---|---|---|
| OpenAI Whisper large-v3 | 1,540M | Zero-shot general-purpose baseline |
| OpenAI Whisper small | 244M | Lightweight baseline for dialect-gap comparison |
| Accent Labs fine-tuned Whisper | 1,540M | My fine-tune of large-v3 (private Hugging Face repo: nadinev/naija-voice-model) |
I chose Whisper because it's widely used and openly available, and comparing zero-shot and fine-tuned models on the same architecture shows how much dialect-focused training helps.
Setup details:
- Test data: The 500 test samples were drawn from the same combined manifest used for fine-tuning. 340 of them (68%) overlap with the training data by sample ID, which leaves 160 clean samples that I report separately. The zero-shot baselines never saw this data.
-
Inference settings: Language was forced to
en(not auto-detected), task wastranscribe. Punctuation is stripped and text is lowercased before WER alignment. Generation usednum_beams=10,no_repeat_ngram_size=3,max_new_tokens=225.
A note on the first run: My original run didn't strip punctuation before WER alignment. Whisper outputs capitalization and punctuation by default, so trailing periods and commas were counted as errors ("people.," counted as an error against "people"). After fixing this, large-v3 dropped from 41.92% to 35.05%, small from 48.00% to 41.19%, and the fine-tuned model from 11.11% to 8.30%.
The baselines dropped the most, so the gap narrowed (30.8 to 26.8 percentage points), but the relative reduction held (74% before, 76% now). Normalization overcorrection went up for all three models (from 18.5% / 14.7% / 7.7% to 23.5% / 18.5% / 10.6%), because punctuation-only mismatches no longer count as errors.
Auto-Classifier Design
Rather than manually labeling roughly 6,659 word errors across the three models (645 for the fine-tuned model alone), I built a rule-based auto-classifier in compute_benchmark_metrics.py. It labels two error types from word-level alignment:
FAITHFUL_DIALECT_PAIRS (38 pairs):
(ref_word, hyp_word)where the reference is standard English and the hypothesis is a known dialect variant.
These are not real errors and are discounted in Faithful WER.DIALECT_NORMALIZATION_PAIRS (82 pairs):
(ref_word, hyp_word)where the reference contains a dialect word and the hypothesis over-corrects to standard English.
These are errors, caused by the model "fixing" dialect.
| Reference | Model output | Class | Real error? |
|---|---|---|---|
| people | pipo | Faithful variant | No |
| the | di | Faithful variant | No |
| dey | they | Overcorrection | Yes |
| dem | them | Overcorrection | Yes |
| na | is | Overcorrection | Yes |
Any error that isn't a faithful variant is labeled a hallucination, meaning output that doesn't match what was spoken. Each one gets a cause: dialect normalization (an overcorrection pair), noise (an insertion), or phonetic (everything else).
Substitutions that match neither list default to phonetic and are logged to unknown_dialect_pairs.csv for review.
Benchmark Results
| Metric | Whisper large-v3 | Whisper small | Fine-tuned |
|---|---|---|---|
| Raw WER (all 500 samples) | 35.05% | 41.19% | 8.30% |
| Faithful WER | 35.05% | 41.19% | 8.00% |
| Dialect penalty (pts) | 0.00 | 0.00 | 0.30 |
| Hallucination rate (% of errors) | 100% | 100% | 96.4% |
| Normalization overcorrection (% of hallucinations) | 23.5% | 18.5% | 10.6% |
| Raw WER (160 held-out samples) | 39.50% | 44.52% | 7.11% |
Per-dialect raw WER:
| Dialect | Whisper large-v3 | Whisper small | Fine-tuned |
|---|---|---|---|
| Hausa | 35.41% | 41.86% | 9.81% |
| Igbo | 35.38% | 42.92% | 7.75% |
| Multilingual | 27.65% | 33.69% | 3.84% |
| Pidgin | 35.02% | 40.23% | 7.46% |
| Yoruba | 36.51% | 41.49% | 9.34% |
Cause distribution (auto-classified, % of hallucinations):
- Phonetic: the model substituted a wrong word, deleted a word, or produced output not matching the audio signal. Also the default catch-all for any substitution that doesn't match a known dialect pair.
- Dialect normalization: the model replaced a dialect word with its standard-English equivalent.
- Noise: inserted words with no counterpart in the reference transcript.
| Model | Phonetic | Dialect Normalization | Noise |
|---|---|---|---|
| Whisper large-v3 | 71.1% | 23.5% | 5.4% |
| Whisper small | 75.8% | 18.5% | 5.7% |
| Fine-tuned | 81.0% | 10.6% | 8.4% |
What the results show
Fine-tuning on Nigerian English and Pidgin cuts WER by about 80%. On the 160 held-out samples, large-v3 scores 39.50% and the fine-tune 7.11%, an 82% relative reduction. Across all 500 samples the figures are 35.05% and 8.30%. The fine-tune does better on the held-out set than on the full 500, even though that set was harder for the baselines, so overlap doesn't appear to be inflating the result.
The held-out samples still come from the same collection as the training data, so these numbers show what in-domain fine-tuning can do. They can't separate better hearing of Nigerian accents from learning this dataset's spelling conventions.Whisper large-v3 rewrites dialect words as standard English. Overcorrection accounts for 23.5% of its hallucinations, roughly 8 points of its 35.05% WER. For the fine-tune it is 10.6% of hallucinations, under 1 point of WER. The fine-tune is far less likely to turn "dey", "dem" and "na" into "they", "them" and "is". These figures use all 500 samples and are lower bounds, since the classifier checks only 82 known pairs and hasn't been validated against human annotation.
Both baselines have a hallucination rate of 100% because they never write faithful spellings like "pipo", so none of their errors are discounted. The fine-tune's 96.4% reflects the 23 faithful variants among its 645 errors. The rate comes from how errors are classified and doesn't measure invented words. Inserted words are the "noise" category.
Standard WER slightly understates the fine-tuned model. Counting faithful spellings like "pipo" as correct moves its WER from 8.30% to 8.00% (23 of 645 errors). The effect is small and exists only for the fine-tune, because zero-shot models never write these forms. It still shows that standard WER penalizes exactly the models that start transcribing dialect. The per-group penalties (Igbo 0.43, Yoruba 0.42) rest on too few events to rank.
Multilingual speech has the lowest WER for all three models, about 6–9 points lower than the other groups for the baselines. The other four groups are close (large-v3 ranges from 35.02% to 36.51%), too close to rank without confidence intervals. One possible reason for the multilingual result is that code-switched speech has more standard English.
What this changed for me
I built this expecting standard WER to be unfairly penalizing correct dialect spellings. The penalty turned out to be only 0.30 points, but overcorrection accounts for about a quarter of Whisper large-v3's errors. I now look at what a model does with dialect words, not only how many words it gets wrong.
Limitations
- Overcorrection and dialect penalty are computed on all 500 samples, including the 340 that overlap with training.
- The classifier uses 120 hand-curated dialect pair rules and hasn't been validated against human annotation.
- 68% of the test samples (340 of 500) overlap the fine-tuned model's training data. I identified the overlap by sample ID, so same-speaker or same-sentence overlap under a different ID would not be caught. The 160 clean samples give the fine-tuned model 7.11% WER.
- With 500 samples across 5 groups, per-dialect numbers come from small samples and have no confidence intervals.
- Audio files and the fine-tuned model weights are private, so transcription can't be re-run. Hypothesis outputs and alignments are in the repository, so anyone can recompute the metrics.
- Hallucination here means any error not matched as a faithful dialect variant, so it includes ordinary substitutions and deletions as well as invented words.
What's Next
Expand the rule set using the pairs logged in
unknown_dialect_pairs.csv, and validate the classifier with human annotation using the included annotation guide.Test whether text LLMs rewrite dialect words the same way when asked only to fix punctuation.
Reproduce It
The benchmark results and analysis are in this Kaggle notebook: Naija ASR Benchmark Results. It loads the published results file and regenerates the results tables and charts in this post.
Because the benchmark audio and fine-tuned model weights are private, transcription can't be re-run here. Everything after transcription is public in the repository, so you can recompute the metrics from the published hypotheses.
What's published
- Test samples with references, without audio paths
- Hypothesis outputs from Whisper large-v3, Whisper small, and the fine-tuned model
- Word-level alignments and the annotation guide
- Results JSON and per-dialect metrics
- Scoring scripts,
compute_benchmark_wer.pyandcompute_benchmark_metrics.py, which also define the 38 + 82 dialect pairs - Logged unknown dialect pairs (
unknown_dialect_pairs.csv) for expanding the rule set
Repository: https://github.com/nadinev6/naija-asr-benchmark
Note: This post carries the Kaggle Benchmarking Challenge label, but it may not meet the challenge format: Whisper isn't available in Kaggle Benchmarks, and the fine-tuned model behind it predates the challenge. I'm sharing it as a write-up either way.

Top comments (0)