DEV Community

Cover image for Podcast to multilingual subtitles in one step
Tidy Tools
Tidy Tools

Posted on

Podcast to multilingual subtitles in one step

Subtitles in one language are easy to get now: run Whisper, get an SRT. Subtitles in three languages still usually mean three steps: transcribe, split the text into subtitle-sized cues, then translate each cue without breaking the timings. Most translation tools do not know what an SRT file is, so the timestamps get mangled or the cues get merged.

I added a subtitleLanguages option to my audio transcriber on Apify so that the whole thing is one run. Here is a real run on a NASA podcast episode, with the timings and the bill.

The input

Give it the podcast's RSS feed, pick the episode, and list the languages:

{
  "podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"],
  "titleIncludes": ["Launching Soon"],
  "maxEpisodesPerFeed": 1,
  "usePublisherTranscripts": false,
  "subtitleLanguages": ["Spanish", "Traditional Chinese"]
}
Enter fullscreen mode Exit fullscreen mode
  • podcastFeeds also accepts Apple Podcasts show links. Instead of a feed you can pass direct audio or video URLs (MP3, MP4, M4A, WEBM…) or upload files.
  • titleIncludes picks episodes by title. Without it you get the newest maxEpisodesPerFeed episodes.
  • usePublisherTranscripts: false forces a real transcription. By default, if the feed already has the publisher's own transcript (<podcast:transcript>), the Actor uses that and charges no audio minutes.
  • subtitleLanguages takes up to 5 languages, by name or code (es, ja, zh-TW…).

The run

Episode NASA's Curious Universe, "Launching Soon: NASA's Roman Space Telescope" (4 min 57 s, 7.5 MB MP3)
Feed 108 episodes; 107 filtered out by the title
Transcript 754 words, English detected automatically, Whisper large-v3-turbo
Subtitles 87 cues in English, 87 in Spanish, 87 in Traditional Chinese
Processing time 75 seconds for transcription and both translations (81 s for the whole run)
Price 5 audio minutes × $0.006 + 5 minutes × 2 languages × $0.002 = $0.05

Every cue in the translated files has exactly the same timestamps as the English one. I compared all 87 start and end times in the three SRT files: identical. The longest cue was 5.6 seconds.

Two cues from the three files:

00:00:19,400 --> 00:00:22,740
the latest in a legacy of exploration.

00:00:19,400 --> 00:00:22,740
el último de una larga
tradición de exploración.

00:00:19,400 --> 00:00:22,740
這是探索遺產中的最新成果。
Enter fullscreen mode Exit fullscreen mode
00:00:23,840 --> 00:00:25,160
First, there was Hubble.

00:00:23,840 --> 00:00:25,160
Primero, estuvo Hubble.

00:00:23,840 --> 00:00:25,160
首先,有哈勃。
Enter fullscreen mode Exit fullscreen mode

One thing to watch: "Traditional Chinese" gives you Traditional characters, but word choices can follow mainland usage. Hubble became 哈勃, while Taiwan usually writes 哈伯. And one proper name came out half-translated: "Discovery" (the Space Shuttle) became 迪斯covery. Machine-translated subtitles are a strong first draft, not a finished product: if your audience is in one specific region, have a native speaker check names and terms.

What you get back

One dataset row per episode:

  • text, segments (with start and end times) and paragraphs
  • with "wordTimestamps": true, a words array (word, start, end) in every segment, at no extra cost
  • srtUrl and vttUrl for the original language
  • translatedSubtitles: for each language, srtUrl, vttUrl, the number of cues and whether it succeeded
  • episode metadata from the feed: podcastTitle, episodeTitle, pubDate, episodeUrl

Subtitles are cut to subtitle size using word timings: at most 7 seconds and two lines per cue. Raw Whisper segments can run for 20+ seconds, which video platforms reject.

How accurate and how fast?

On longer test files (in the same week, from our own tests):

  • Compared with NASA's official captions on two ScienceCasts videos, the median start-time error was 0.40–0.51 seconds.
  • A 62-minute file took 144 seconds, so roughly 17–26× faster than real time.
  • Whisper sometimes skips a stretch of speech in the middle of a long file. The Actor now re-transcribes any gap of 5 seconds or more without text; on the 62-minute file that recovered 9 passages.

Cost

Price
Transcription $0.006 per audio minute ($0.36 per hour)
Speaker labels (optional) $0.015 per minute instead of $0.006
Translated subtitles $0.002 per audio minute per language, only for languages that succeed
SRT/VTT, paragraphs, segment and word timestamps included

So a 30-minute episode in Spanish and German is $0.18 for the transcript plus $0.12 for the two subtitle languages. No start fee.

Schedule it

For a podcast that releases a new episode every week, turn on newEpisodesOnly and give the schedule a monitorName. The Actor remembers which episodes it already transcribed and only takes episodes released after the newest one it has done, so a scheduled run on a day without a new episode transcribes, and charges, nothing. Old episodes are left alone unless you also set backfillOlderEpisodes: true, which works through the back catalogue maxEpisodesPerFeed episodes per run.

{
  "podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"],
  "maxEpisodesPerFeed": 1,
  "newEpisodesOnly": true,
  "monitorName": "curious-universe",
  "subtitleLanguages": ["Spanish", "Japanese"]
}
Enter fullscreen mode Exit fullscreen mode

Or call it from code (pip install apify-client):

from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("tidytools/audio-transcriber").call(run_input={
    "podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"],
    "maxEpisodesPerFeed": 1,
    "subtitleLanguages": ["Spanish", "Japanese"],
})
for ep in client.dataset(run["defaultDatasetId"]).iterate_items():
    if not ep.get("success"):
        print("failed:", ep.get("input"), ep.get("error"))
        continue
    print(ep["episodeTitle"], ep["srtUrl"])
    for sub in ep.get("translatedSubtitles", []):
        print(" ", sub["language"], sub.get("srtUrl"))
Enter fullscreen mode Exit fullscreen mode

Or with plain HTTP, which waits for the run and returns the rows (fine for short episodes; use the asynchronous run endpoint for long ones):

curl -X POST "https://api.apify.com/v2/acts/tidytools~audio-transcriber/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"], "maxEpisodesPerFeed": 1, "subtitleLanguages": ["Spanish"]}'
Enter fullscreen mode Exit fullscreen mode

It does not download from YouTube, Vimeo or Loom pages: those links come back as a free error row explaining how to get the file. Use files and feeds you have the right to process.

Try it here: Audio & Video to Text Transcription - Whisper, Podcasts, SRT

Disclosure: I built this Actor and earn money when people run it on Apify.

Top comments (0)