<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Philipp Primisser</title>
    <description>The latest articles on DEV Community by Philipp Primisser (@prime619).</description>
    <link>https://dev.to/prime619</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172250%2F698c9f2c-6440-4600-ad7a-dce942c1fc6e.png</url>
      <title>DEV Community: Philipp Primisser</title>
      <link>https://dev.to/prime619</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/prime619"/>
    <language>en</language>
    <item>
      <title>Free EU tender alerts from the TED API (one Python file, no account)</title>
      <dc:creator>Philipp Primisser</dc:creator>
      <pubDate>Fri, 09 Oct 2026 01:54:23 +0000</pubDate>
      <link>https://dev.to/prime619/free-eu-tender-alerts-from-the-ted-api-one-python-file-no-account-3bph</link>
      <guid>https://dev.to/prime619/free-eu-tender-alerts-from-the-ted-api-one-python-file-no-account-3bph</guid>
      <description>&lt;p&gt;The EU publishes thousands of public procurement notices every week on &lt;a href="https://ted.europa.eu" rel="noopener noreferrer"&gt;TED (Tenders Electronic Daily)&lt;/a&gt;. Commercial tender alert services usually charge a subscription. If you already know which CPV codes you bid on, the official Search API is enough for a personal alert.&lt;/p&gt;

&lt;p&gt;I put that into one Python file with no dependencies beyond the standard library: &lt;a href="https://github.com/philippprimisser-max/ted-tender-alerts" rel="noopener noreferrer"&gt;ted-tender-alerts&lt;/a&gt; (MIT).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python ted_alerts.py &lt;span class="nt"&gt;--cpv&lt;/span&gt; 72 48 &lt;span class="nt"&gt;--country&lt;/span&gt; AUT &lt;span class="nt"&gt;--days&lt;/span&gt; 7 &lt;span class="nt"&gt;--lang&lt;/span&gt; de
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;22 matching notices, 22 new.
- Commvault
  … | AUT | 21,000,000 EUR | deadline 2026-11-09 (32 d)
  https://ted.europa.eu/de/notice/-/detail/696492-2026
- MA 9 – Digitalisierung 2027-2031
  Magistrat der Stadt Wien … | AUT | deadline 2026-11-17 (40 d)
  …
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(Real output from 8 October 2026. Run it again and you get &lt;code&gt;22 matching notices, 0 new.&lt;/code&gt;)&lt;/p&gt;

&lt;h2&gt;
  
  
  What the API gives you
&lt;/h2&gt;

&lt;p&gt;TED Search API v3 is anonymous. One POST to &lt;code&gt;https://api.ted.europa.eu/v3/notices/search&lt;/code&gt; with a query string in TED's expert-search language:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;publication-date&amp;gt;=20261001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; AND buyer-country IN (AUT DEU)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; AND notice-type IN (cn-standard cn-social cn-desg subco qu-sy)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; AND (classification-cpv=72* OR classification-cpv=48*)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; SORT BY publication-date DESC&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fields&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[...],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;page&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;limit&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paginationMode&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PAGE_NUMBER&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scope&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ALL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CPV prefixes work: &lt;code&gt;72&lt;/code&gt; matches every code under 72000000 (IT services). Countries use ISO alpha-3 (&lt;code&gt;AUT&lt;/code&gt;, &lt;code&gt;DEU&lt;/code&gt;). &lt;code&gt;--print-query&lt;/code&gt; prints the string so you can paste it into TED's expert search and compare.&lt;/p&gt;

&lt;p&gt;The script is gentle with the API: it pauses 0.5 s between result pages and stops after 20 pages (100 notices each) by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Only the new ones
&lt;/h2&gt;

&lt;p&gt;A state file (&lt;code&gt;ted_state.json&lt;/code&gt;) remembers publication numbers already reported. Every run compares the current result to that set and prints only the difference. Notices older than 120 days drop out of the state, so the file stays small.&lt;/p&gt;

&lt;p&gt;Same idea for an Atom feed of the latest 100, a CSV you append to, a Markdown table, or an e-mail through your own SMTP server (&lt;code&gt;SMTP_HOST&lt;/code&gt;, &lt;code&gt;MAIL_TO&lt;/code&gt;, … — the password is never printed or written).&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it daily for free on GitHub Actions
&lt;/h2&gt;

&lt;p&gt;The repository includes an example workflow (&lt;code&gt;examples/daily-ted-alerts.yml&lt;/code&gt;). Copy it to &lt;code&gt;.github/workflows/&lt;/code&gt; in your own copy of the repo and it runs every weekday morning, commits the state and the Atom feed back, and optionally e-mails you. Edit the CPV/country filters in the YAML; add SMTP secrets if you want mail. Point a feed reader at the raw URL of &lt;code&gt;data/feed.xml&lt;/code&gt; (public repos) or just open &lt;code&gt;data/latest.md&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding your CPV codes
&lt;/h2&gt;

&lt;p&gt;Look at five tenders you would have bid on and note their CPV codes (they're on every TED notice). Shorter prefixes catch more: &lt;code&gt;72&lt;/code&gt; is all IT services, &lt;code&gt;7222&lt;/code&gt; only IT consultancy. Start broad, then narrow once you see the noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Legal note
&lt;/h2&gt;

&lt;p&gt;Data comes from TED – Tenders Electronic Daily, © European Union. Notices may be reused free of charge with attribution (&lt;a href="https://ted.europa.eu/en/legal-notice" rel="noopener noreferrer"&gt;TED legal notice&lt;/a&gt;). The script adds the attribution line to every output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/philippprimisser-max/ted-tender-alerts
&lt;span class="nb"&gt;cd &lt;/span&gt;ted-tender-alerts
python ted_alerts.py &lt;span class="nt"&gt;--cpv&lt;/span&gt; 72 &lt;span class="nt"&gt;--country&lt;/span&gt; AUT &lt;span class="nt"&gt;--days&lt;/span&gt; 3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;&lt;em&gt;AI-assisted: a model helped with the wording. The sample run and the 22-notice count are from my own machine on 8 October 2026. Not affiliated with the EU Publications Office.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>opendata</category>
      <category>automation</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Word-by-word captions with Python and ffmpeg, no CapCut watermark</title>
      <dc:creator>Philipp Primisser</dc:creator>
      <pubDate>Fri, 09 Oct 2026 01:39:29 +0000</pubDate>
      <link>https://dev.to/prime619/word-by-word-captions-with-python-and-ffmpeg-no-capcut-watermark-10a</link>
      <guid>https://dev.to/prime619/word-by-word-captions-with-python-and-ffmpeg-no-capcut-watermark-10a</guid>
      <description>&lt;p&gt;Short-form video lives or dies on captions. CapCut and a dozen web tools will do the "word lights up as you say it" style for you, usually with a watermark or a free-minute limit. I wanted the same look without an account, so I wrote a small offline script: &lt;a href="https://github.com/philippprimisser-max/whisper-word-captions" rel="noopener noreferrer"&gt;whisper-word-captions&lt;/a&gt; (MIT). This post is about how the trick works, so you can change it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6gv68xqfxn8j2har85nk.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6gv68xqfxn8j2har85nk.gif" alt="Word-by-word caption demo from the repository" width="540" height="300"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python wordcaptions.py clip.mp4
&lt;span class="c"&gt;# -&amp;gt; clip_captioned.mp4  and  clip.ass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Three pieces
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;faster-whisper&lt;/strong&gt; with &lt;code&gt;word_timestamps=True&lt;/code&gt; → every word gets its own start and end.&lt;/li&gt;
&lt;li&gt;Group words into short captions (max 3 by default; also break after &lt;code&gt;.?!&lt;/code&gt; or a pause over 0.6 s).&lt;/li&gt;
&lt;li&gt;Write an &lt;a href="https://github.com/libass/libass/wiki/ASS-File-Format-Guide" rel="noopener noreferrer"&gt;Advanced SubStation Alpha&lt;/a&gt; (&lt;code&gt;.ass&lt;/code&gt;) file with &lt;strong&gt;one Dialogue event per spoken word&lt;/strong&gt;, then burn it into the picture with ffmpeg's &lt;code&gt;subtitles&lt;/code&gt; filter (libass).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The interesting part is step 3.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ASS trick
&lt;/h2&gt;

&lt;p&gt;ASS lets you override style mid-line with curly-brace tags. &lt;code&gt;\c&lt;/code&gt; changes the text colour (ASS colours are &lt;code&gt;&amp;amp;HAABBGGRR&lt;/code&gt;, blue-green-red), &lt;code&gt;\fscx&lt;/code&gt;/&lt;code&gt;\fscy&lt;/code&gt; scale it, &lt;code&gt;\r&lt;/code&gt; resets to the style. For a three-word caption you write three events, one per word:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Dialogue: 0,0:00:03.20,0:00:03.52,Word,,0,0,0,,{\c&amp;amp;H0000D4FF\fscx115\fscy115}THAT{\r} ALL MEN
Dialogue: 0,0:00:03.52,0:00:03.68,Word,,0,0,0,,THAT {\c&amp;amp;H0000D4FF\fscx115\fscy115}ALL{\r} MEN
Dialogue: 0,0:00:03.68,0:00:03.90,Word,,0,0,0,,THAT ALL {\c&amp;amp;H0000D4FF\fscx115\fscy115}MEN{\r}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each event lasts until the next word starts. That way nothing flickers during short pauses, and the viewer always sees the whole caption with only the current word highlighted. In Python that's just a nested loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;group&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;groups&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;texts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;escape_ass&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upper&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;group&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;group&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;group&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;group&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;0.05&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;parts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;texts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;c&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;hl&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;fscx115&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;fscy115}}&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;{{&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;r}}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Dialogue: 0,&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;ass_time&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;,&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;ass_time&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;,Word,,0,0,0,,&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PlayResX/PlayResY in the &lt;code&gt;.ass&lt;/code&gt; header match the video size, so the same file works for 1080×1920 Shorts and for landscape. Font size is a percentage of the shorter side (6.5 % by default), with a bit more margin at the bottom so the caption sits above the like/comment buttons.&lt;/p&gt;

&lt;h2&gt;
  
  
  Burning it in
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; clip.mp4 &lt;span class="nt"&gt;-vf&lt;/span&gt; &lt;span class="s2"&gt;"subtitles=clip.ass"&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt;:v libx264 &lt;span class="nt"&gt;-crf&lt;/span&gt; 20 &lt;span class="nt"&gt;-c&lt;/span&gt;:a copy clip_captioned.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Audio is copied, video is re-encoded once. If you use a non-system font, pass a fonts directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;-vf&lt;/span&gt; &lt;span class="s2"&gt;"subtitles=clip.ass:fontsdir=./fonts"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Paths with backslashes or drive letters need escaping for the filter (&lt;code&gt;C:\…&lt;/code&gt; → &lt;code&gt;C\:/…&lt;/code&gt;). The script does that for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Edit without re-transcribing
&lt;/h2&gt;

&lt;p&gt;Speech recognition gets names and jargon wrong. The &lt;code&gt;.ass&lt;/code&gt; is plain text, so open it, fix the words, and re-burn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python wordcaptions.py clip.mp4 &lt;span class="nt"&gt;--render-only&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That takes seconds because Whisper doesn't run again. &lt;code&gt;--srt&lt;/code&gt; also writes a normal subtitle file if you want something for YouTube's upload form.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install
&lt;/h2&gt;

&lt;p&gt;Python 3.10–3.13 and ffmpeg (with libass, which the common builds include).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Windows: winget install Gyan.FFmpeg&lt;/span&gt;
&lt;span class="c"&gt;# macOS:   brew install ffmpeg&lt;/span&gt;
&lt;span class="c"&gt;# Debian:  sudo apt install ffmpeg&lt;/span&gt;
git clone https://github.com/philippprimisser-max/whisper-word-captions
&lt;span class="nb"&gt;cd &lt;/span&gt;whisper-word-captions
python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first run downloads the Whisper model (&lt;code&gt;small&lt;/code&gt; ≈ 490 MB). &lt;code&gt;requirements.txt&lt;/code&gt; pins &lt;code&gt;av&amp;lt;19&lt;/code&gt;, because PyAV 19 broke faster-whisper 1.2.1 on a fresh install in October 2026.&lt;/p&gt;

&lt;p&gt;Defaults that matter: &lt;code&gt;--model small&lt;/code&gt; (better for non-English than &lt;code&gt;base&lt;/code&gt;) and highlight colour &lt;code&gt;#FFD400&lt;/code&gt;. &lt;code&gt;--threads&lt;/code&gt; is capped at 4. On the bench from the gap checker (&lt;code&gt;base&lt;/code&gt;, int8, 5 clips × 3 runs, &lt;code&gt;bench/run_table.sh&lt;/code&gt;, 9 October 2026) 8 threads had gaps in 13 of 15 runs (longest 60.5 s), 4 threads in 1 of 15 (longest 28 s) and 1 thread in 1 of 15 (longest 6 s). float32 at 8 threads was 0 of 15. Fewer threads reduce the gaps. They don't remove them. Check the &lt;code&gt;.ass&lt;/code&gt;, or use &lt;a href="https://github.com/philippprimisser-max/faster-whisper-gap-check" rel="noopener noreferrer"&gt;faster-whisper-gap-check&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Measured on an 8-core Linux server: a 15-second 1080×1920 clip took 6.1–6.7 s end to end with &lt;code&gt;small&lt;/code&gt; (model load, transcription and burn; two runs, model already downloaded). The clip is made from the repo's test audio, so you can repeat it; the commands are in the README. Your laptop will differ.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;AI-assisted: a model helped with the wording. The script was tested and the timing is from my own server (9 October 2026). Demo audio: LibriVox recording of the Gettysburg Address (public domain). Not affiliated with CapCut, TikTok, YouTube, Instagram, OpenAI or the ffmpeg project.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>ffmpeg</category>
      <category>whisper</category>
      <category>video</category>
    </item>
    <item>
      <title>Podcast to text in Python: transcribe any episode from its RSS feed</title>
      <dc:creator>Philipp Primisser</dc:creator>
      <pubDate>Fri, 09 Oct 2026 01:16:30 +0000</pubDate>
      <link>https://dev.to/prime619/podcast-to-text-in-python-transcribe-any-episode-from-its-rss-feed-3e8f</link>
      <guid>https://dev.to/prime619/podcast-to-text-in-python-transcribe-any-episode-from-its-rss-feed-3e8f</guid>
      <description>&lt;p&gt;I wanted transcripts of a few podcast episodes for notes and search. The options I found were either a web service with a monthly plan, or a Whisper tutorial that starts with "first, download the MP3". I didn't want to do the download part by hand every time, so I wrote a small script that takes whatever link I have, an RSS feed, an Apple Podcasts link or a direct audio URL, and gives me &lt;code&gt;.txt&lt;/code&gt;, &lt;code&gt;.srt&lt;/code&gt;, &lt;code&gt;.vtt&lt;/code&gt; and &lt;code&gt;.json&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It's on GitHub: &lt;a href="https://github.com/philippprimisser-max/podcast-to-text" rel="noopener noreferrer"&gt;podcast-to-text&lt;/a&gt; (MIT). This post walks through the parts that turned out to be interesting.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python podcast_to_text.py &lt;span class="s2"&gt;"https://podcasts.apple.com/us/podcast/hacker-public-radio/id281699640"&lt;/span&gt; &lt;span class="nt"&gt;--search&lt;/span&gt; &lt;span class="s2"&gt;"wl-copy"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Apple Podcasts link -&amp;gt; RSS feed: https://hackerpublicradio.org/hpr_rss.php
Episode: HPR4743: wl-copy (2026-10-07)
Transcribing with faster-whisper 'base' (4 threads) ...
Done: 15.7 min audio in 18 s (0.02 s per audio second), language en.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 0: maybe you don't need to transcribe at all
&lt;/h2&gt;

&lt;p&gt;Podcasting 2.0 added a &lt;code&gt;&amp;lt;podcast:transcript&amp;gt;&lt;/code&gt; tag to RSS. Some hosting platforms generate transcripts and put the link right into the feed. So before burning CPU, check the feed.&lt;/p&gt;

&lt;p&gt;I was curious how common that is, so I looked at the Apple top-chart podcasts for the US, UK, Germany and Austria (109 feeds I could read, checked on 8 October 2026; the script and the raw results are in the repo's &lt;a href="https://github.com/philippprimisser-max/podcast-to-text/tree/main/research" rel="noopener noreferrer"&gt;&lt;code&gt;research/&lt;/code&gt;&lt;/a&gt; folder):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Chart&lt;/th&gt;
&lt;th&gt;Newest episode has a transcript tag&lt;/th&gt;
&lt;th&gt;At least one episode has one&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;US&lt;/td&gt;
&lt;td&gt;5 / 49&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UK&lt;/td&gt;
&lt;td&gt;2 / 25&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Germany&lt;/td&gt;
&lt;td&gt;8 / 25&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Austria&lt;/td&gt;
&lt;td&gt;3 / 10&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;18 / 109&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;28&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So for roughly one in six popular shows the transcript is already there. German shows were ahead in this sample, mostly because many of them are hosted on Podigee: 18 of the 28 hits were Podigee feeds (a few shows chart in both Germany and Austria, so they count twice). The others were on Omny, Buzzsprout, Flightcast, Captivate and two smaller hosts. Most of them were WebVTT, some JSON, a few SRT or plain text. The script prefers JSON, then SRT, then VTT, then plain text, downloads it and stops. &lt;code&gt;--force&lt;/code&gt; transcribes anyway.&lt;/p&gt;

&lt;p&gt;Reading the tag is a few lines with the standard library:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;PODCAST_NS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://podcastindex.org/namespace/1.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;item&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;transcripts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()}&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;PODCAST_NS&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;}}transcript&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 1: Apple Podcasts links are just a pointer to an RSS feed
&lt;/h2&gt;

&lt;p&gt;Most people share Apple Podcasts links, not RSS URLs. Apple's public lookup API returns the show's own feed URL for the numeric ID in the link:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;show_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/id(\d{5,})&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;urlparse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;http_get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://itunes.apple.com/lookup?id=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;show_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;&amp;amp;entity=podcast&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;feed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;results&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;feedUrl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the link points to a single episode (&lt;code&gt;?i=1000…&lt;/code&gt;), a second lookup with &lt;code&gt;entity=podcastEpisode&lt;/code&gt; gives you that episode's GUID, which you match against the feed. The audio then comes from the publisher's server, exactly as in any podcast app. Apple-exclusive subscriber shows have no public feed, and the script says so instead of trying anything clever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: pick the episode
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;--list&lt;/code&gt; prints the latest 30 episodes, with a marker for those that already have a transcript. &lt;code&gt;-e 2&lt;/code&gt; takes the second newest, &lt;code&gt;--search "some words"&lt;/code&gt; takes the newest episode whose title contains all of them. Parsing is plain &lt;code&gt;xml.etree&lt;/code&gt;; I didn't add &lt;code&gt;feedparser&lt;/code&gt; because the script only needs title, date, GUID, enclosure and the transcript tag.&lt;/p&gt;

&lt;p&gt;One thing I kept: an honest user agent (&lt;code&gt;podcast-to-text/0.1 (+github URL)&lt;/code&gt;). A few hosts block unknown clients. Pretending to be a browser would get around that, but if a publisher doesn't want automated downloads, that's their call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: transcribe with faster-whisper
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;WhisperModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cpu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;compute_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;int8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                     &lt;span class="n"&gt;cpu_threads&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cpu_count&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;segments&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;info&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;beam_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                                  &lt;span class="n"&gt;vad_filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;condition_on_previous_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few choices here, and why:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;vad_filter=True&lt;/code&gt;&lt;/strong&gt; skips silence and music beds, which saves time and avoids some of Whisper's "Thank you for watching" hallucinations on silent parts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;condition_on_previous_text=False&lt;/code&gt;&lt;/strong&gt; makes it less likely that one bad segment drags the following ones into a repetition loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;cpu_threads&lt;/code&gt; capped at 4.&lt;/strong&gt; On one 8-core machine, int8 with the &lt;code&gt;base&lt;/code&gt; model dropped speech in 13 of 15 runs at 8 threads (longest 60.5 s), 1 of 15 at 4 threads (longest 28 s) and 1 of 15 with a single thread (longest 6 s). That's 5 clips × 3 runs, from &lt;code&gt;bench/run_table.sh&lt;/code&gt; on 9 October 2026. Fewer threads make the gaps rarer. They don't get rid of them. float32 at 8 threads was clean, 0 of 15, which is why the script has &lt;code&gt;--compute-type float32&lt;/code&gt;. I wrote the check up separately: &lt;a href="https://github.com/philippprimisser-max/faster-whisper-gap-check" rel="noopener noreferrer"&gt;faster-whisper-gap-check&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;How fast is it? On an 8-core Linux server with 4 threads, not counting model loading: the 15.7-minute HPR episode above took 18 s with &lt;code&gt;base&lt;/code&gt;. The one-minute Gettysburg Address from the repo's tests took 1 s with &lt;code&gt;base&lt;/code&gt; and 9–10 s with &lt;code&gt;small&lt;/code&gt; (&lt;code&gt;python podcast_to_text.py tests/fixtures/gettysburg.mp3 -m small --language en&lt;/code&gt;). A laptop will be slower; I'm not going to guess by how much.&lt;/p&gt;

&lt;p&gt;On quality: for clear English, &lt;code&gt;base&lt;/code&gt; was fine in my tests; for German and other languages I'd start with &lt;code&gt;small&lt;/code&gt;. Pass &lt;code&gt;--language&lt;/code&gt; if you know it. Auto-detection can guess wrong: in one run without it, &lt;code&gt;small&lt;/code&gt; took the English test clip for Russian and wrote the whole transcript in Russian.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: four output formats
&lt;/h2&gt;

&lt;p&gt;Segments come back as &lt;code&gt;start&lt;/code&gt;, &lt;code&gt;end&lt;/code&gt;, &lt;code&gt;text&lt;/code&gt;. From that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;.txt&lt;/code&gt; with a paragraph break after pauses over 2 seconds, which makes long transcripts much easier to read&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;.srt&lt;/code&gt; and &lt;code&gt;.vtt&lt;/code&gt; for players and video editors&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;.json&lt;/code&gt; with the segments plus metadata (model, detected language, audio length, processing time)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The SRT timestamp function is the only fiddly bit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sep&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;,&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;divmod&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3_600_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;divmod&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;60_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;divmod&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;02&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;02&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;02&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}{&lt;/span&gt;&lt;span class="n"&gt;sep&lt;/span&gt;&lt;span class="si"&gt;}{&lt;/span&gt;&lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;03&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Install and run
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/philippprimisser-max/podcast-to-text
&lt;span class="nb"&gt;cd &lt;/span&gt;podcast-to-text
python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
python podcast_to_text.py https://hackerpublicradio.org/hpr_mp3_rss.php &lt;span class="nt"&gt;--list&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No ffmpeg needed, faster-whisper decodes audio through PyAV. One trap: on a fresh install in October 2026, PyAV 19 broke faster-whisper 1.2.1 with &lt;code&gt;open() got an unexpected keyword argument 'metadata_errors'&lt;/code&gt;, so &lt;code&gt;requirements.txt&lt;/code&gt; pins &lt;code&gt;av&amp;lt;19&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't do
&lt;/h2&gt;

&lt;p&gt;No speaker labels, no summaries, no private feeds unless you have a personal feed URL you're allowed to use. Machine transcripts still get names and jargon wrong, so read before you quote.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;AI-assisted: a model helped with the wording. The code was tested and all numbers come from runs on my own server (8 and 9 October 2026) and can be repeated with the commands in the repositories. Test audio: LibriVox (public domain) and Hacker Public Radio (CC BY-SA 4.0). Not affiliated with Apple; "Apple Podcasts" is a trademark of Apple Inc.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>podcast</category>
      <category>whisper</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>faster-whisper int8 dropped up to 60 s of speech. float32 didn't</title>
      <dc:creator>Philipp Primisser</dc:creator>
      <pubDate>Fri, 09 Oct 2026 00:27:09 +0000</pubDate>
      <link>https://dev.to/prime619/faster-whisper-int8-dropped-up-to-60-s-of-speech-float32-didnt-4nba</link>
      <guid>https://dev.to/prime619/faster-whisper-int8-dropped-up-to-60-s-of-speech-float32-didnt-4nba</guid>
      <description>&lt;p&gt;I run a small podcast transcription tool, and while testing German audio I noticed that some transcripts were missing whole passages. Not wrong words. Twenty, thirty, sometimes close to sixty seconds of speech, gone. The timestamps simply jumped ahead, the text still read fine, and nothing in the logs said anything was wrong.&lt;/p&gt;

&lt;p&gt;Here's one of them. A one-minute excerpt of a LibriVox reading of Theodor Storm's &lt;em&gt;Der Schimmelreiter&lt;/em&gt;, transcribed with faster-whisper &lt;code&gt;base&lt;/code&gt;, int8, 8 threads:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1
00:00:30,690 --&amp;gt; 00:00:32,110
Enkel singelit.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first subtitle starts at 0:30. The narrator starts reading at 0:01.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I measured
&lt;/h2&gt;

&lt;p&gt;Setup: faster-whisper 1.2.1, CTranslate2 4.8.2, model &lt;code&gt;base&lt;/code&gt;, beam size 1, VAD on, &lt;code&gt;condition_on_previous_text=False&lt;/code&gt;, on one 8-core x86-64 Linux server (&lt;code&gt;os.cpu_count()&lt;/code&gt; = 8). Five public clips of about 60 seconds (three English, two German: LibriVox recordings and one Hacker Public Radio episode), three runs per configuration and clip, because the problem doesn't happen every time. A run counts as "with a gap" if at least 5 seconds of detected speech have no transcript segment.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Runs with a gap ≥ 5 s&lt;/th&gt;
&lt;th&gt;Longest gap&lt;/th&gt;
&lt;th&gt;Avg time per clip&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;int8, 1 thread&lt;/td&gt;
&lt;td&gt;1/15&lt;/td&gt;
&lt;td&gt;6.0 s&lt;/td&gt;
&lt;td&gt;4.7 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;int8, 4 threads&lt;/td&gt;
&lt;td&gt;1/15&lt;/td&gt;
&lt;td&gt;28.0 s&lt;/td&gt;
&lt;td&gt;4.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;int8, 8 threads&lt;/td&gt;
&lt;td&gt;13/15&lt;/td&gt;
&lt;td&gt;60.5 s&lt;/td&gt;
&lt;td&gt;2.4 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;float32, 8 threads&lt;/td&gt;
&lt;td&gt;0/15&lt;/td&gt;
&lt;td&gt;–&lt;/td&gt;
&lt;td&gt;3.8 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;(9 October 2026. The machine wasn't idle during the run, so treat the times as rough.)&lt;/p&gt;

&lt;p&gt;int8 with 8 threads lost audio in almost every run, and when it did, it lost a lot: the 60.5-second gap was in a 61.5-second clip. Fewer threads made it much rarer, but not zero: one run with 4 threads missed 28 seconds, and one run with a single thread missed 6 seconds. The only configuration without a single gap was float32.&lt;/p&gt;

&lt;p&gt;The dropping is random. A day earlier I ran the same clips with the same tool and got gaps in 4 of 15 int8/8-thread runs and none anywhere else. I didn't record the exact flags of that run, so the table above is the one to go by; the raw output of both is in the repository.&lt;/p&gt;

&lt;p&gt;I don't know the root cause. My guess would be something in CTranslate2's int8 path under higher thread counts on this CPU, but that's a guess and I haven't dug into it. &lt;strong&gt;This is one machine.&lt;/strong&gt; I haven't tested other CPUs, other model sizes or other CTranslate2 versions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do now
&lt;/h2&gt;

&lt;p&gt;If you can't check your transcripts, use float32:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;faster_whisper&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;WhisperModel&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;WhisperModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;base&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cpu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;compute_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;float32&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In my run it was the only clean configuration, and at 8 threads it took 3.8 s per one-minute clip, against 2.4 s for int8 with 8 threads and 4.0 s for int8 with 4 threads. So it costs some speed compared to int8 with many threads, but not compared to int8 with a capped thread count.&lt;/p&gt;

&lt;p&gt;If you stay on int8, capping the threads (&lt;code&gt;cpu_threads=min(4, os.cpu_count() or 1)&lt;/code&gt;) cut the gaps from 13 of 15 runs to 1 of 15 here. That's a big improvement, not a guarantee. Check the output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check your own transcripts
&lt;/h2&gt;

&lt;p&gt;The nasty part is that you don't notice. A transcript with a missing half minute looks like a normal transcript. So I turned the checking into a tool: &lt;a href="https://github.com/philippprimisser-max/faster-whisper-gap-check" rel="noopener noreferrer"&gt;faster-whisper-gap-check&lt;/a&gt; (MIT).&lt;/p&gt;

&lt;p&gt;It runs Silero VAD (which ships with faster-whisper) over the audio, finds where people are speaking, and lists every stretch of speech that no transcript segment covers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python gapcheck.py check s3_schimmelreiter.mp3 results/schimmelreiter_int8_8threads.srt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Audio 1:01.0 | speech detected 56 s | transcript segments 20
Possibly dropped: 1 stretch(es), 29 s of speech (52 %):
  0:00.9 - 0:29.7  (29 s)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It reads &lt;code&gt;.srt&lt;/code&gt;, &lt;code&gt;.vtt&lt;/code&gt; and JSON with &lt;code&gt;start&lt;/code&gt;/&lt;code&gt;end&lt;/code&gt; segments, so it works on output from any Whisper flavour or API, not only faster-whisper. Exit code 1 if something's missing, so you can drop it into a batch job.&lt;/p&gt;

&lt;p&gt;The core is just interval arithmetic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;uncovered&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;speech&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;covered&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pad&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;cov&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;merge&lt;/span&gt;&lt;span class="p"&gt;([(&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;pad&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;pad&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;covered&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;s1&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;speech&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s0&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c1&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;cov&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;c1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;c0&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;s1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;c0&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c0&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;s1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;break&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;s1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;s1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And there's a &lt;code&gt;repro&lt;/code&gt; mode that transcribes the same file several times per configuration and checks every run. The table above comes from &lt;code&gt;bash bench/run_table.sh&lt;/code&gt;, which cuts the five clips from archive.org with ffmpeg and runs, per clip:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python gapcheck.py repro CLIP.mp3 &lt;span class="nt"&gt;--language&lt;/span&gt; en &lt;span class="nt"&gt;--threads&lt;/span&gt; 1 4 8 &lt;span class="nt"&gt;--runs&lt;/span&gt; 3 &lt;span class="nt"&gt;--json&lt;/span&gt; results/my-run/gap_CLIP.json
python gapcheck.py repro CLIP.mp3 &lt;span class="nt"&gt;--language&lt;/span&gt; en &lt;span class="nt"&gt;--threads&lt;/span&gt; 8 &lt;span class="nt"&gt;--compute-types&lt;/span&gt; float32 &lt;span class="nt"&gt;--runs&lt;/span&gt; 3 &lt;span class="nt"&gt;--json&lt;/span&gt; results/my-run/gap_CLIP_f32.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(&lt;code&gt;--language de&lt;/code&gt; for the German clips.) If you run faster-whisper on a many-core CPU, I'd be curious what you get. Use a clip with continuous speech (an audiobook chapter is ideal) and do at least three runs per config; one clean run proves nothing.&lt;/p&gt;

&lt;p&gt;Caveat on the checker: music with vocals, laughter or crosstalk counts as speech for the VAD, so a "gap" over a jingle can be fine, and a short gap like the 6-second one above is worth listening to before you call it a bug. Treat the output as "go listen to these timestamps", not as a verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took away from it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Test with the thread count and compute type you deploy with.&lt;/strong&gt; A single thread also had one short gap in my test, so a low thread count alone is no guarantee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run quality tests more than once.&lt;/strong&gt; faster-whisper output varies between runs. My two runs of the same experiment gave 4/15 and 13/15 for the same configuration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check coverage, not only accuracy.&lt;/strong&gt; Word error rate against a reference catches this, but most people don't have a reference. Comparing against VAD speech regions needs nothing but the audio.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The free &lt;a href="https://github.com/philippprimisser-max/podcast-to-text" rel="noopener noreferrer"&gt;podcast-to-text&lt;/a&gt; script caps threads at 4 by default and has &lt;code&gt;--compute-type float32&lt;/code&gt; if you want to be safe.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;AI-assisted: a model helped with the wording. The measurements are from my own server (8 and 9 October 2026), and the raw data and the script are in the repository. Test audio: LibriVox recordings (public domain) and Hacker Public Radio (CC BY-SA 4.0). Not affiliated with SYSTRAN or OpenAI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>whisper</category>
      <category>machinelearning</category>
      <category>debugging</category>
    </item>
  </channel>
</rss>
