Subtitles in one language are easy to get now: run Whisper, get an SRT. Subtitles in three languages still usually mean three steps: transcribe, split the text into subtitle-sized cues, then translate each cue without breaking the timings. Most translation tools do not know what an SRT file is, so the timestamps get mangled or the cues get merged.
I added a subtitleLanguages option to my audio transcriber on Apify so that the whole thing is one run. Here is a real run on a NASA podcast episode, with the timings and the bill.
The input
Give it the podcast's RSS feed, pick the episode, and list the languages:
{
"podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"],
"titleIncludes": ["Launching Soon"],
"maxEpisodesPerFeed": 1,
"usePublisherTranscripts": false,
"subtitleLanguages": ["Spanish", "Traditional Chinese"]
}
-
podcastFeedsalso accepts Apple Podcasts show links. Instead of a feed you can pass direct audio or video URLs (MP3, MP4, M4A, WEBM…) or upload files. -
titleIncludespicks episodes by title. Without it you get the newestmaxEpisodesPerFeedepisodes. -
usePublisherTranscripts: falseforces a real transcription. By default, if the feed already has the publisher's own transcript (<podcast:transcript>), the Actor uses that and charges no audio minutes. -
subtitleLanguagestakes up to 5 languages, by name or code (es,ja,zh-TW…).
The run
| Episode | NASA's Curious Universe, "Launching Soon: NASA's Roman Space Telescope" (4 min 57 s, 7.5 MB MP3) |
| Feed | 108 episodes; 107 filtered out by the title |
| Transcript | 754 words, English detected automatically, Whisper large-v3-turbo |
| Subtitles | 87 cues in English, 87 in Spanish, 87 in Traditional Chinese |
| Processing time | 75 seconds for transcription and both translations (81 s for the whole run) |
| Price | 5 audio minutes × $0.006 + 5 minutes × 2 languages × $0.002 = $0.05 |
Every cue in the translated files has exactly the same timestamps as the English one. I compared all 87 start and end times in the three SRT files: identical. The longest cue was 5.6 seconds.
Two cues from the three files:
00:00:19,400 --> 00:00:22,740
the latest in a legacy of exploration.
00:00:19,400 --> 00:00:22,740
el último de una larga
tradición de exploración.
00:00:19,400 --> 00:00:22,740
這是探索遺產中的最新成果。
00:00:23,840 --> 00:00:25,160
First, there was Hubble.
00:00:23,840 --> 00:00:25,160
Primero, estuvo Hubble.
00:00:23,840 --> 00:00:25,160
首先,有哈勃。
One thing to watch: "Traditional Chinese" gives you Traditional characters, but word choices can follow mainland usage. Hubble became 哈勃, while Taiwan usually writes 哈伯. And one proper name came out half-translated: "Discovery" (the Space Shuttle) became 迪斯covery. Machine-translated subtitles are a strong first draft, not a finished product: if your audience is in one specific region, have a native speaker check names and terms.
What you get back
One dataset row per episode:
-
text,segments(with start and end times) andparagraphs - with
"wordTimestamps": true, awordsarray (word,start,end) in every segment, at no extra cost -
srtUrlandvttUrlfor the original language -
translatedSubtitles: for each language,srtUrl,vttUrl, the number of cues and whether it succeeded - episode metadata from the feed:
podcastTitle,episodeTitle,pubDate,episodeUrl
Subtitles are cut to subtitle size using word timings: at most 7 seconds and two lines per cue. Raw Whisper segments can run for 20+ seconds, which video platforms reject.
How accurate and how fast?
On longer test files (in the same week, from our own tests):
- Compared with NASA's official captions on two ScienceCasts videos, the median start-time error was 0.40–0.51 seconds.
- A 62-minute file took 144 seconds, so roughly 17–26× faster than real time.
- Whisper sometimes skips a stretch of speech in the middle of a long file. The Actor now re-transcribes any gap of 5 seconds or more without text; on the 62-minute file that recovered 9 passages.
Cost
| Price | |
|---|---|
| Transcription | $0.006 per audio minute ($0.36 per hour) |
| Speaker labels (optional) | $0.015 per minute instead of $0.006 |
| Translated subtitles | $0.002 per audio minute per language, only for languages that succeed |
| SRT/VTT, paragraphs, segment and word timestamps | included |
So a 30-minute episode in Spanish and German is $0.18 for the transcript plus $0.12 for the two subtitle languages. No start fee.
Schedule it
For a podcast that releases a new episode every week, turn on newEpisodesOnly and give the schedule a monitorName. The Actor remembers which episodes it already transcribed and only takes episodes released after the newest one it has done, so a scheduled run on a day without a new episode transcribes, and charges, nothing. Old episodes are left alone unless you also set backfillOlderEpisodes: true, which works through the back catalogue maxEpisodesPerFeed episodes per run.
{
"podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"],
"maxEpisodesPerFeed": 1,
"newEpisodesOnly": true,
"monitorName": "curious-universe",
"subtitleLanguages": ["Spanish", "Japanese"]
}
Or call it from code (pip install apify-client):
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("tidytools/audio-transcriber").call(run_input={
"podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"],
"maxEpisodesPerFeed": 1,
"subtitleLanguages": ["Spanish", "Japanese"],
})
for ep in client.dataset(run["defaultDatasetId"]).iterate_items():
if not ep.get("success"):
print("failed:", ep.get("input"), ep.get("error"))
continue
print(ep["episodeTitle"], ep["srtUrl"])
for sub in ep.get("translatedSubtitles", []):
print(" ", sub["language"], sub.get("srtUrl"))
Or with plain HTTP, which waits for the run and returns the rows (fine for short episodes; use the asynchronous run endpoint for long ones):
curl -X POST "https://api.apify.com/v2/acts/tidytools~audio-transcriber/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"podcastFeeds": ["https://www.nasa.gov/feeds/podcasts/curious-universe"], "maxEpisodesPerFeed": 1, "subtitleLanguages": ["Spanish"]}'
It does not download from YouTube, Vimeo or Loom pages: those links come back as a free error row explaining how to get the file. Use files and feeds you have the right to process.
Try it here: Audio & Video to Text Transcription - Whisper, Podcasts, SRT
Disclosure: I built this Actor and earn money when people run it on Apify.
Top comments (0)