I built Podium, a tool that compares your reading of a speech with a great delivery of the same text (JFK, Reagan) and tells you exactly where you drift and why: "14.3 syllables/s here vs 7.7 in the reference (+85%)", pinned to a time range.
The first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged 41 "flaws" per minute on a perfectly clean recording.
This post covers why that happened, the one-line fix, and the evaluation set-up that caught it.
Step 1: put both recordings on the same word grid
Comparing two deliveries frame by frame is hopeless: they never line up. Comparing them word by word is easy, as long as both are force-aligned to the same transcript. I used torchaudio's MMS_FA (a wav2vec2 CTC aligner):
bundle = torchaudio.pipelines.MMS_FA
model, tokenizer, aligner = bundle.get_model(with_star=False), bundle.get_tokenizer(), bundle.get_aligner()
emission, _ = model(wav)
spans = aligner(emission[0], tokenizer(words)) # one span list per word
Now word 17 of your reading is word 17 of JFK's. For every word I compute speaker-normalised features with Praat (via Parselmouth) and librosa:
- pitch in semitones relative to your own median (a deep and a high voice with the same intonation give the same curve);
- loudness in dB relative to your own level;
- articulation rate, syllables per second of actual speech, pauses excluded, over a 5-word window;
- the pause before the word, and voiced sound inside that pause (more on that below);
- high-frequency energy above 2.5 kHz, which drops when consonants get swallowed.
Step 2: build a dataset where the answers are exact
There's no public dataset that pairs a good delivery with bad deliveries of the same words. So I made one: 207 recordings, starting from three public-domain speeches.
The trick is that I inject the flaws myself, so every label is sample-accurate:
| Flaw | How it's injected | Severity 1 / 2 / 3 |
|---|---|---|
| rushed / dragged | phase-vocoder time-scale of 4–9 words | ×1.3/1.6/2.0 · ×0.8/0.65/0.5 |
| monotone | WORLD vocoder, pitch contour squashed toward its mean | 50/75/95% removed |
| mumbled | gain down + low-pass | −6 dB @ 3 kHz … −16 dB @ 1.1 kHz |
| awkward pause | room-tone silence inserted mid-phrase | 0.7 / 1.3 / 2.2 s |
| missing breath | a natural pause squeezed out | 50 / 20 / 0% kept |
| stutter | word onset repeated | 1 / 2 / 3 repeats |
Edits are joined with 5 ms fades that preserve length, so the label times never drift. My first version used 15 ms crossfades, which quietly shifted every later label by 15 ms per edit.
I also had the same texts read by open Piper TTS voices, with flaws injected into those too. That's the "different speaker" test.
And I split it honestly: thresholds are tuned only on JFK, then frozen and tested on Reagan plus an unseen voice.
Step 3: the bug that wasn't a bug
Version 1 scored each word by how far it departed from the reference, as a robust z-score (median/MAD) after removing your overall offset. On Reagan's own audio with injected flaws it got F1 0.64 with zero false alarms.
On a TTS voice reading the same text:
| F1 | False alarms on clean audio | |
|---|---|---|
| Same speaker | 0.64 | 0 / min |
| Different speaker | 0.03 | 41 / min |
A synthetic voice doesn't breathe where Reagan breathes or lift the words he lifts. Measured against Reagan, everything it does is a deviation. Technically correct, but useless as feedback: nobody wants to hear that every sentence is wrong because they aren't Reagan.
The fix was to separate style from flaws. A flaw is something that's off compared with the reference and off compared with your own delivery. So every word gets two scores:
z_ref = robust_z(participant - reference) # departs from the reference, beyond your overall style
z_self = robust_z(participant_feature) # stands out within your own recording
flaw = np.fmin(z_ref, z_self) # a soft AND
A consistently different voice moves z_ref everywhere but z_self almost nowhere, so it's reported once as style ("33% faster and flatter than the reference") instead of 40 times as flaws.
False alarms on clean different-speaker recordings went from 41/min to 3/min.
Step 4: a stutter that looked like a pause
Injected stutters kept being reported as "awkward pause". The cause: the aligner usually places the word at its last restart, so "w- w- we" becomes a long gap followed by a normal "we".
The fix was one feature: seconds of voiced audio inside the gap. A real pause is silent; a gap full of voiced sound is a restart.
gap_voiced = np.sum(~np.isnan(f0[prev_end:word_start])) * HOP
Long pauses now require gap_voiced < 0.15 s, and stutters can trigger on gap_voiced > 0.12 s. Stutter recall went from 0.11 to 0.44 on the held-out set.
Results on the held-out set (Reagan + an unseen voice)
- detected regions overlap the true flaw with mean IoU 0.81; overall F1 0.51;
- awkward pauses, missing breaths, mumbling and stutters are located within 0.00–0.06 s;
- 0 false alarms per minute on untouched audio of the reference speaker;
- the rubric score falls monotonically from L0 to L4 on every excerpt (Spearman ρ = −1.00).
What doesn't work yet:
- Pacing is weakest (F1 0.29–0.35). Slowing a phrase also stretches its pauses, so "dragged" is often reported as "awkward pause"; the confusion matrix shows exactly that.
- Different speakers are still hard (F1 0.21). Fixed noise floors are the next thing to replace with learned ones.
Three things I'd tell anyone building an evaluator
- Make your labels exact before tuning anything. Injecting the flaws yourself turns an argument about subjective judgement into a number.
- Hold out a speaker, not just files. My calibration F1 (0.62) and held-out F1 (0.51) are close because the split was honest. Same-speaker numbers alone would have hidden the 41-per-minute problem.
- Decide what shouldn't count before deciding what should. The hardest part wasn't detecting deviations, it was deciding which deviations matter.
Code, dataset (207 labelled recordings) and the full evaluation are open source:
JayPokale
/
podium
Contrastive speech delivery analytics: find exactly where a delivery drifts from a great speech, and why. Multimodal AI Hackathon 2026, Track C.
🎙️ Podium: contrastive speech delivery analytics with temporal flaw grounding
Read a great speech. Podium shows exactly where your delivery drifts from it, to the tenth of a second and explains each flaw with the numbers behind it.
Built for the Multimodal AI Hackathon 2026, Track C: Contrastive Speech Analytics & Temporal Flaw Grounding.
What it does
- Pick a reference: an excerpt of a great public-domain speech (JFK, Reagan).
- Give your delivery: record yourself reading the same text in the browser, upload a file, or try one of the 207 dataset samples.
-
Get grounded feedback
- a deterministic 0-10 rubric for pacing, pauses, pitch, volume and fluency;
- flaw regions with exact start/end times, shaded on time-series overlays of your pitch, loudness and pace against the reference (time-warped onto your timeline word by word);
- a causal explanation for every region, built from the measured delta ("14.3 syllables/s here…
Built for the Multimodal AI Hackathon 2026 (Track C), with heavy help from Claude Code, an AI coding agent, for implementation and evaluation runs.


Top comments (0)