I wanted transcripts of a few podcast episodes for notes and search. The options I found were either a web service with a monthly plan, or a Whisper tutorial that starts with "first, download the MP3". I didn't want to do the download part by hand every time, so I wrote a small script that takes whatever link I have, an RSS feed, an Apple Podcasts link or a direct audio URL, and gives me .txt, .srt, .vtt and .json.
It's on GitHub: podcast-to-text (MIT). This post walks through the parts that turned out to be interesting.
python podcast_to_text.py "https://podcasts.apple.com/us/podcast/hacker-public-radio/id281699640" --search "wl-copy"
Apple Podcasts link -> RSS feed: https://hackerpublicradio.org/hpr_rss.php
Episode: HPR4743: wl-copy (2026-10-07)
Transcribing with faster-whisper 'base' (4 threads) ...
Done: 15.7 min audio in 18 s (0.02 s per audio second), language en.
Step 0: maybe you don't need to transcribe at all
Podcasting 2.0 added a <podcast:transcript> tag to RSS. Some hosting platforms generate transcripts and put the link right into the feed. So before burning CPU, check the feed.
I was curious how common that is, so I looked at the Apple top-chart podcasts for the US, UK, Germany and Austria (109 feeds I could read, checked on 8 October 2026; the script and the raw results are in the repo's research/ folder):
| Chart | Newest episode has a transcript tag | At least one episode has one |
|---|---|---|
| US | 5 / 49 | 5 |
| UK | 2 / 25 | 3 |
| Germany | 8 / 25 | 15 |
| Austria | 3 / 10 | 5 |
| Total | 18 / 109 | 28 |
So for roughly one in six popular shows the transcript is already there. German shows were ahead in this sample, mostly because many of them are hosted on Podigee: 18 of the 28 hits were Podigee feeds (a few shows chart in both Germany and Austria, so they count twice). The others were on Omny, Buzzsprout, Flightcast, Captivate and two smaller hosts. Most of them were WebVTT, some JSON, a few SRT or plain text. The script prefers JSON, then SRT, then VTT, then plain text, downloads it and stops. --force transcribes anyway.
Reading the tag is a few lines with the standard library:
PODCAST_NS = "https://podcastindex.org/namespace/1.0"
for item in channel.findall("item"):
transcripts = [
{"url": t.get("url"), "type": (t.get("type") or "").lower()}
for t in item.findall(f"{{{PODCAST_NS}}}transcript")
if t.get("url")
]
Step 1: Apple Podcasts links are just a pointer to an RSS feed
Most people share Apple Podcasts links, not RSS URLs. Apple's public lookup API returns the show's own feed URL for the numeric ID in the link:
show_id = re.search(r"/id(\d{5,})", urlparse(url).path).group(1)
data = json.loads(http_get(f"https://itunes.apple.com/lookup?id={show_id}&entity=podcast"))
feed = data["results"][0]["feedUrl"]
If the link points to a single episode (?i=1000…), a second lookup with entity=podcastEpisode gives you that episode's GUID, which you match against the feed. The audio then comes from the publisher's server, exactly as in any podcast app. Apple-exclusive subscriber shows have no public feed, and the script says so instead of trying anything clever.
Step 2: pick the episode
--list prints the latest 30 episodes, with a marker for those that already have a transcript. -e 2 takes the second newest, --search "some words" takes the newest episode whose title contains all of them. Parsing is plain xml.etree; I didn't add feedparser because the script only needs title, date, GUID, enclosure and the transcript tag.
One thing I kept: an honest user agent (podcast-to-text/0.1 (+github URL)). A few hosts block unknown clients. Pretending to be a browser would get around that, but if a publisher doesn't want automated downloads, that's their call.
Step 3: transcribe with faster-whisper
model = WhisperModel(model_name, device="cpu", compute_type="int8",
cpu_threads=min(4, os.cpu_count() or 1))
segments, info = model.transcribe(str(audio), language=language, beam_size=1,
vad_filter=True, condition_on_previous_text=False)
A few choices here, and why:
-
vad_filter=Trueskips silence and music beds, which saves time and avoids some of Whisper's "Thank you for watching" hallucinations on silent parts. -
condition_on_previous_text=Falsemakes it less likely that one bad segment drags the following ones into a repetition loop. -
cpu_threadscapped at 4. On one 8-core machine, int8 with thebasemodel dropped speech in 13 of 15 runs at 8 threads (longest 60.5 s), 1 of 15 at 4 threads (longest 28 s) and 1 of 15 with a single thread (longest 6 s). That's 5 clips × 3 runs, frombench/run_table.shon 9 October 2026. Fewer threads make the gaps rarer. They don't get rid of them. float32 at 8 threads was clean, 0 of 15, which is why the script has--compute-type float32. I wrote the check up separately: faster-whisper-gap-check.
How fast is it? On an 8-core Linux server with 4 threads, not counting model loading: the 15.7-minute HPR episode above took 18 s with base. The one-minute Gettysburg Address from the repo's tests took 1 s with base and 9–10 s with small (python podcast_to_text.py tests/fixtures/gettysburg.mp3 -m small --language en). A laptop will be slower; I'm not going to guess by how much.
On quality: for clear English, base was fine in my tests; for German and other languages I'd start with small. Pass --language if you know it. Auto-detection can guess wrong: in one run without it, small took the English test clip for Russian and wrote the whole transcript in Russian.
Step 4: four output formats
Segments come back as start, end, text. From that:
-
.txtwith a paragraph break after pauses over 2 seconds, which makes long transcripts much easier to read -
.srtand.vttfor players and video editors -
.jsonwith the segments plus metadata (model, detected language, audio length, processing time)
The SRT timestamp function is the only fiddly bit:
def ts(seconds, sep=","):
ms = int(round(max(seconds, 0) * 1000))
h, ms = divmod(ms, 3_600_000)
m, ms = divmod(ms, 60_000)
s, ms = divmod(ms, 1000)
return f"{h:02d}:{m:02d}:{s:02d}{sep}{ms:03d}"
Install and run
git clone https://github.com/philippprimisser-max/podcast-to-text
cd podcast-to-text
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python podcast_to_text.py https://hackerpublicradio.org/hpr_mp3_rss.php --list
No ffmpeg needed, faster-whisper decodes audio through PyAV. One trap: on a fresh install in October 2026, PyAV 19 broke faster-whisper 1.2.1 with open() got an unexpected keyword argument 'metadata_errors', so requirements.txt pins av<19.
What it doesn't do
No speaker labels, no summaries, no private feeds unless you have a personal feed URL you're allowed to use. Machine transcripts still get names and jargon wrong, so read before you quote.
AI-assisted: a model helped with the wording. The code was tested and all numbers come from runs on my own server (8 and 9 October 2026) and can be repeated with the commands in the repositories. Test audio: LibriVox (public domain) and Hacker Public Radio (CC BY-SA 4.0). Not affiliated with Apple; "Apple Podcasts" is a trademark of Apple Inc.
Top comments (0)