Getting text out of a video is a solved problem. Getting subtitles out of it is a different one, and most first attempts look like this:
1
00:00:00,000 --> 00:00:11,400
We choose to go to the moon in this decade and do the other things, not because they are easy, but because they are hard, because that goal will serve to organize and measure the best of our energies and skills
That is a correct transcript and an unusable subtitle: one block, eleven seconds on screen, far too long to read.
What makes a subtitle readable
Broadcasters converged on a few rules, and they are a good default:
- About 42 characters per line for landscape video, at most two lines.
- A few seconds on screen, not ten.
- Break at natural points: the end of a sentence, or a pause in speech.
- For vertical video (Shorts, Reels, TikTok), much shorter lines: around 20 to 25 characters.
To follow these rules you need to know when each word is spoken, not just each sentence. That is the part people miss.
Doing it yourself with Whisper
OpenAI's open-source Whisper can write SRT files directly, and it has options for exactly this:
pip install -U openai-whisper
whisper video.mp4 --model small --output_format srt \
--word_timestamps True --max_line_width 42 --max_line_count 2
--word_timestamps True is what makes --max_line_width and --max_line_count work. Without it, Whisper falls back to one subtitle per segment, which is how you get the wall of text above.
To burn the subtitles into the video afterwards:
ffmpeg -i video.mp4 -vf subtitles=video.srt output.mp4
This route is free and private. Its costs are time and setup: a long video on a laptop without a GPU is slow, and it is one more thing to install and keep working if you want to run it from a server or an automation.
Doing it with one call
I packaged the same idea as a tool: Subtitle Generator on Apify. You give it a video or audio link (or a Google Drive or Dropbox share link, or an upload) and it returns an SRT file and a VTT file.
It builds subtitles from word timings, keeps lines within the length you choose, and breaks at sentence ends and at pauses. This is what it produced for a 20-second NASA clip:
1
00:00:00,120 --> 00:00:04,080
We choose to go to the moon in this decade
and do the other things.
2
00:00:04,580 --> 00:00:07,600
Not because they are easy, but because
they are hard.
3
00:00:07,760 --> 00:00:12,920
3, 2, 1, 0. All engines running.
4
00:00:13,920 --> 00:00:15,980
Liftoff. We have a liftoff.
From Python:
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("spokentext/subtitle-generator").call(run_input={
"urls": ["https://example.com/video.mp4"],
"maxCharsPerLine": 42,
"maxLines": 2,
})
for item in client.dataset(run.default_dataset_id).iterate_items():
print(item["fileName"], item["subtitleCount"], "subtitles")
print("SRT:", item["srtUrl"])
print("VTT:", item["vttUrl"])
For vertical video, set maxCharsPerLine to about 22 and maxLines to 1 or 2.
It costs $0.003 per minute of video, so a one-hour video is $0.18. Files that fail are not charged.
SRT or VTT?
Both hold the same text and timings.
- SRT is what video editors and players import: Premiere Pro, DaVinci Resolve, CapCut, VLC, and YouTube's upload page.
-
VTT is what web players and the HTML
<track>element expect.
What the tool does not do
- No translation. Subtitles are in the language spoken in the video.
- No speaker labels.
- No burning in. You get subtitle files, not a new video. Use the ffmpeg command above for that.
- Files only. It takes a link to a video file, not a YouTube or TikTok page.
- Videos up to about two hours for now.
- Your audio leaves your machine. It is sent to a hosted speech-to-text provider. For sensitive recordings, use the Whisper route.
Which route to pick
| Situation | Route |
|---|---|
| One video, privacy matters, you have time | Whisper on your own machine |
| Many videos, or part of an automation | The API |
| You need translated subtitles | Neither; look for a tool that translates |
- I could not compare with the app. I'm based in France, where the Muse app isn't available to me, while Meta's developer API works fine from here. So everything above was measured through the API only. If you can open the app, try a few of your questions there and compare, and please keep me update in the comment section !
Disclosure: I built the tool described in "Doing it with one call". This article was drafted with AI assistance and checked by me.
Top comments (0)