DEV Community

Cover image for faster-whisper int8 dropped up to 60 s of speech. float32 didn't
Philipp Primisser
Philipp Primisser

Posted on AI-assisted

faster-whisper int8 dropped up to 60 s of speech. float32 didn't

I run a small podcast transcription tool, and while testing German audio I noticed that some transcripts were missing whole passages. Not wrong words. Twenty, thirty, sometimes close to sixty seconds of speech, gone. The timestamps simply jumped ahead, the text still read fine, and nothing in the logs said anything was wrong.

Here's one of them. A one-minute excerpt of a LibriVox reading of Theodor Storm's Der Schimmelreiter, transcribed with faster-whisper base, int8, 8 threads:

1
00:00:30,690 --> 00:00:32,110
Enkel singelit.
Enter fullscreen mode Exit fullscreen mode

The first subtitle starts at 0:30. The narrator starts reading at 0:01.

What I measured

Setup: faster-whisper 1.2.1, CTranslate2 4.8.2, model base, beam size 1, VAD on, condition_on_previous_text=False, on one 8-core x86-64 Linux server (os.cpu_count() = 8). Five public clips of about 60 seconds (three English, two German: LibriVox recordings and one Hacker Public Radio episode), three runs per configuration and clip, because the problem doesn't happen every time. A run counts as "with a gap" if at least 5 seconds of detected speech have no transcript segment.

Config Runs with a gap ≥ 5 s Longest gap Avg time per clip
int8, 1 thread 1/15 6.0 s 4.7 s
int8, 4 threads 1/15 28.0 s 4.0 s
int8, 8 threads 13/15 60.5 s 2.4 s
float32, 8 threads 0/15 – 3.8 s

(9 October 2026. The machine wasn't idle during the run, so treat the times as rough.)

int8 with 8 threads lost audio in almost every run, and when it did, it lost a lot: the 60.5-second gap was in a 61.5-second clip. Fewer threads made it much rarer, but not zero: one run with 4 threads missed 28 seconds, and one run with a single thread missed 6 seconds. The only configuration without a single gap was float32.

The dropping is random. A day earlier I ran the same clips with the same tool and got gaps in 4 of 15 int8/8-thread runs and none anywhere else. I didn't record the exact flags of that run, so the table above is the one to go by; the raw output of both is in the repository.

I don't know the root cause. My guess would be something in CTranslate2's int8 path under higher thread counts on this CPU, but that's a guess and I haven't dug into it. This is one machine. I haven't tested other CPUs, other model sizes or other CTranslate2 versions.

What I'd do now

If you can't check your transcripts, use float32:

from faster_whisper import WhisperModel

model = WhisperModel("base", device="cpu", compute_type="float32")
Enter fullscreen mode Exit fullscreen mode

In my run it was the only clean configuration, and at 8 threads it took 3.8 s per one-minute clip, against 2.4 s for int8 with 8 threads and 4.0 s for int8 with 4 threads. So it costs some speed compared to int8 with many threads, but not compared to int8 with a capped thread count.

If you stay on int8, capping the threads (cpu_threads=min(4, os.cpu_count() or 1)) cut the gaps from 13 of 15 runs to 1 of 15 here. That's a big improvement, not a guarantee. Check the output.

Check your own transcripts

The nasty part is that you don't notice. A transcript with a missing half minute looks like a normal transcript. So I turned the checking into a tool: faster-whisper-gap-check (MIT).

It runs Silero VAD (which ships with faster-whisper) over the audio, finds where people are speaking, and lists every stretch of speech that no transcript segment covers:

python gapcheck.py check s3_schimmelreiter.mp3 results/schimmelreiter_int8_8threads.srt
Enter fullscreen mode Exit fullscreen mode
Audio 1:01.0 | speech detected 56 s | transcript segments 20
Possibly dropped: 1 stretch(es), 29 s of speech (52 %):
  0:00.9 - 0:29.7  (29 s)
Enter fullscreen mode Exit fullscreen mode

It reads .srt, .vtt and JSON with start/end segments, so it works on output from any Whisper flavour or API, not only faster-whisper. Exit code 1 if something's missing, so you can drop it into a batch job.

The core is just interval arithmetic:

def uncovered(speech, covered, pad=1.0):
    cov = merge([(max(0.0, a - pad), b + pad) for a, b in covered])
    result = []
    for s0, s1 in speech:
        cur = s0
        for c0, c1 in cov:
            if c1 <= cur or c0 >= s1:
                continue
            if c0 > cur:
                result.append((cur, c0))
            cur = max(cur, c1)
            if cur >= s1:
                break
        if cur < s1:
            result.append((cur, s1))
    return result
Enter fullscreen mode Exit fullscreen mode

And there's a repro mode that transcribes the same file several times per configuration and checks every run. The table above comes from bash bench/run_table.sh, which cuts the five clips from archive.org with ffmpeg and runs, per clip:

python gapcheck.py repro CLIP.mp3 --language en --threads 1 4 8 --runs 3 --json results/my-run/gap_CLIP.json
python gapcheck.py repro CLIP.mp3 --language en --threads 8 --compute-types float32 --runs 3 --json results/my-run/gap_CLIP_f32.json
Enter fullscreen mode Exit fullscreen mode

(--language de for the German clips.) If you run faster-whisper on a many-core CPU, I'd be curious what you get. Use a clip with continuous speech (an audiobook chapter is ideal) and do at least three runs per config; one clean run proves nothing.

Caveat on the checker: music with vocals, laughter or crosstalk counts as speech for the VAD, so a "gap" over a jingle can be fine, and a short gap like the 6-second one above is worth listening to before you call it a bug. Treat the output as "go listen to these timestamps", not as a verdict.

What I took away from it

  1. Test with the thread count and compute type you deploy with. A single thread also had one short gap in my test, so a low thread count alone is no guarantee.
  2. Run quality tests more than once. faster-whisper output varies between runs. My two runs of the same experiment gave 4/15 and 13/15 for the same configuration.
  3. Check coverage, not only accuracy. Word error rate against a reference catches this, but most people don't have a reference. Comparing against VAD speech regions needs nothing but the audio.

The free podcast-to-text script caps threads at 4 by default and has --compute-type float32 if you want to be safe.


AI-assisted: a model helped with the wording. The measurements are from my own server (8 and 9 October 2026), and the raw data and the script are in the repository. Test audio: LibriVox recordings (public domain) and Hacker Public Radio (CC BY-SA 4.0). Not affiliated with SYSTRAN or OpenAI.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more