DEV Community

Nikita Iakovlev
Nikita Iakovlev

Posted on

Speaker Labels Burned a Full Hour of Compute and Returned Nothing. One Voice Print Per Phrase Fixed It

I run a visa agency in Bali. The part of my work that has nothing to do with visas is building Actors on Apify, and one of them is a transcriber: you give it an audio or video URL, it runs Whisper large v3 and gives back text with timecodes.

Transcription was the easy half. The feature that nearly did not ship was the one that sounds like a checkbox — "label who is speaking."

The run that returned an empty array

Here is the log of a real run from 11 October 2026. One input: a 62-minute recorded lecture from archive.org, accurate model, diarize: true.

01:40:37  Transcribing 1 file(s) with the accurate model; labelling speakers.
02:12:51  speaker labelling: window 1/3 done
02:40:34  ACTOR: The Actor run has reached the timeout of 3600 seconds, aborting it.
Enter fullscreen mode Exit fullscreen mode

The audio is processed in 30-minute windows so memory stays flat. Window one — the first half hour of audio — took 32 minutes and 14 seconds. The run was killed in the middle of window two, and the dataset at the end of it was exactly this:

[]
Enter fullscreen mode Exit fullscreen mode

Two full compute units spent, nothing written. The status line said TIMED-OUT, which is at least honest, but a buyer whose hour-long interview comes back empty does not care about status lines.

Why the obvious pipeline is unaffordable

The first version used the pyannote-style segmentation pipeline that ships with sherpa-onnx, and the shape of its work is the problem. It computes a speaker embedding for every 10-second window at a one-second step. That means every second of audio is embedded roughly ten times, and then every print is compared against every other print to find the clusters.

It is a fine design on a GPU, or on a 30-second voice memo. On a 2 GB Apify Actor it ran slower than real time — slower than Whisper itself, the part everyone assumes is expensive.

And here is the trap I walked into while tuning it. Apify gives you one CPU core per 4 GB of memory. My code politely asked for two threads:

threads = max(1, min(4, math.ceil(int(os.environ.get('ACTOR_MEMORY_MBYTES') or 2048) / 1024)))
Enter fullscreen mode Exit fullscreen mode

On a 2 GB run that resolves to 2, and the diarizer then clamps it to the cores it actually has: max(1, 2048 // 4096) → 1. I had been reading my own thread count as if it were parallelism. It was a wish.

The rewrite: cut at the pauses first

The fix was to stop embedding time and start embedding speech. Two small ONNX models, both on the CPU, no external API calls:

  1. Silero VAD finds the stretches of speech and cuts the audio at the pauses.
  2. 3D-Speaker CAM++ gives each phrase one voice print.
  3. Phrases whose prints are close are the same person.
config.silero_vad.min_silence_duration = 0.3
config.silero_vad.min_speech_duration = 0.25
config.silero_vad.max_speech_duration = 10.0   # MAX_PHRASE_SECONDS
Enter fullscreen mode Exit fullscreen mode

One print per phrase instead of ten per second is about a ninefold reduction in embedding work, and the clustering input shrinks with it. Same file, same input, same 2 GB:

05:51:33  Transcribing 1 file(s) with the accurate model; labelling speakers.
05:55:49  speaker labelling: window 1/3 done
05:58:11  speaker labelling: window 2/3 done
05:58:19  speaker labelling: window 3/3 done
05:58:44  OK — 62 min billed, 8365 words, language English
Enter fullscreen mode Exit fullscreen mode

437 seconds end to end against a run that had not finished half the file in 3,600. Compute units: 0.24 against 2.0 — and the 2.0 bought an empty dataset.

The price of that shortcut is specific and worth stating: a speaker change with no pause at all inside one phrase does not get split. The 10-second phrase cap bounds the damage to a few seconds of wrong label, but if your audio is two people talking over each other continuously, this pipeline is the wrong tool.

The threshold is the entire algorithm

Clustering is cosine distance with a cut-off, and that single float decides how many people exist:

DEFAULT_THRESHOLD = 0.75
Enter fullscreen mode Exit fullscreen mode

I measured it on a multi-reader LibriVox play, where I knew the answer. At 0.5 it invented 16 speakers for 6 voices. Tight thresholds do not fail by merging people; they fail by splitting one person into a crowd, because the same voice sounds different across a sentence.

If you know the count, pass it — it goes straight to the clusterer as num_clusters and removes the guess:

{ "urls": ["https://example.com/interview.m4a"], "diarize": true, "speakerCount": 2 }
Enter fullscreen mode Exit fullscreen mode

Read speakerCount with suspicion

On that lecture the output said speakerCount: 7. There was one lecturer, an introduction and some audience questions. So I counted the segments per label:

Counter({'Speaker 4': 555, 'Speaker 1': 41, 'Speaker 2': 37,
         'Speaker 7': 18, 'Speaker 6': 3, 'Speaker 3': 1, 'Speaker 8': 1})
Enter fullscreen mode Exit fullscreen mode

That reads correctly once you stop treating the labels as equals. Speaker 4 holds 555 of 656 segments — that is the lecturer. Speakers 1 and 2 are the host and the first questioner. Speakers 3, 6 and 8, with one to three segments between them, are applause, a cough and a microphone handover getting their own voice prints.

Note what is missing from that list: there is no Speaker 5, and the highest number is 8 while the count is 7. Numbering runs over the VAD turns in order of first appearance, but a label only reaches the output if it wins the overlap on an actual transcript segment. The maximum label number is not the speaker count, and the numbers are not contiguous. If you are building UI from this, read the speakers array; do not generate range(1, speakerCount + 1).

The practical rule: sort labels by segment count and treat the long tail as noise.

One field quietly changes meaning

With labels on, the transcript string is rebuilt as one paragraph per turn, prefixed with the speaker:

Speaker 1: CHAPTER II. The Princess and the Goblin. This is a LibriVox recording…
Enter fullscreen mode Exit fullscreen mode

Which means wordCount is no longer only words from the audio. The same lecture, same accurate model, twice:

Run wordCount
diarize: false 8,335
diarize: true 8,365

The difference is 30, and there are exactly 15 turn labels in the diarized transcript: Speaker plus 4: counted as two words apiece. Nothing is wrong with the recognition — but do not compare wordCount across diarization settings.

The setting I expected to matter, and didn't

Turbo is advertised as several times faster. On this file, with both runs started in the same second so the clock is comparable:

Model Wall clock Words
fast (large v3 turbo) 475 s 8,139
accurate (large v3) 486 s 8,335

Eleven seconds apart. For a long single file, inference is not the bottleneck — downloading it is, and Groq returns an hour of audio in well under a minute either way. What the turbo model does cost you is visible in the very first caption:

  • accurate: To be transmitted live to Middle East Technical University, we must ask you to refrain from…
  • fast: We will be transmitted live to Middle East Technical University. We must ask you to refrain from…

Fluent, plausible, and not what was said. Choose fast for price, not for speed.

One more caveat before you trust the timings

Segment durations come straight from Whisper's verbose_json. On sparse audio they get strange — the opening segment of an MIT lecture pulled from YouTube is the single word "Hi." with duration: 18.4. Good enough to find a passage in a two-hour file, useless as word-level timing — which is why the SRT and WebVTT exports redistribute time proportionally by character count instead.

Calling it

curl -X POST "https://api.apify.com/v2/acts/lergassy~whisper-transcriber/run-sync-get-dataset-items?token=<TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://example.com/interview.m4a"],"quality":"accurate","diarize":true,"speakerCount":2,"includeSegments":true}'
Enter fullscreen mode Exit fullscreen mode

The Actor is Whisper Transcriber. Labels are an add-on of $0.006 per started minute on top of $0.009 for the accurate model, so turn diarize on only when who-said-what is the thing you need — on a single-speaker podcast it is pure cost.

And when the source already has captions, do not pay for recognition at all: a YouTube video with subtitles can be read with a transcript scraper for a fraction of the price. Whisper is for audio nobody has transcribed yet.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to