Short-form video lives or dies on captions. CapCut and a dozen web tools will do the "word lights up as you say it" style for you, usually with a watermark or a free-minute limit. I wanted the same look without an account, so I wrote a small offline script: whisper-word-captions (MIT). This post is about how the trick works, so you can change it.
python wordcaptions.py clip.mp4
# -> clip_captioned.mp4 and clip.ass
Three pieces
-
faster-whisper with
word_timestamps=True→ every word gets its own start and end. - Group words into short captions (max 3 by default; also break after
.?!or a pause over 0.6 s). - Write an Advanced SubStation Alpha (
.ass) file with one Dialogue event per spoken word, then burn it into the picture with ffmpeg'ssubtitlesfilter (libass).
The interesting part is step 3.
The ASS trick
ASS lets you override style mid-line with curly-brace tags. \c changes the text colour (ASS colours are &HAABBGGRR, blue-green-red), \fscx/\fscy scale it, \r resets to the style. For a three-word caption you write three events, one per word:
Dialogue: 0,0:00:03.20,0:00:03.52,Word,,0,0,0,,{\c&H0000D4FF\fscx115\fscy115}THAT{\r} ALL MEN
Dialogue: 0,0:00:03.52,0:00:03.68,Word,,0,0,0,,THAT {\c&H0000D4FF\fscx115\fscy115}ALL{\r} MEN
Dialogue: 0,0:00:03.68,0:00:03.90,Word,,0,0,0,,THAT ALL {\c&H0000D4FF\fscx115\fscy115}MEN{\r}
Each event lasts until the next word starts. That way nothing flickers during short pauses, and the viewer always sees the whole caption with only the current word highlighted. In Python that's just a nested loop:
for group in groups:
texts = [escape_ass(w.text.upper()) for w in group]
for i, w in enumerate(group):
end = group[i + 1].start if i + 1 < len(group) else max(w.end, w.start + 0.05)
parts = []
for j, t in enumerate(texts):
if j == i:
parts.append(f"{{\\c{hl}\\fscx115\\fscy115}}{t}{{\\r}}")
else:
parts.append(t)
events.append(f"Dialogue: 0,{ass_time(w.start)},{ass_time(end)},Word,,0,0,0,,{' '.join(parts)}")
PlayResX/PlayResY in the .ass header match the video size, so the same file works for 1080×1920 Shorts and for landscape. Font size is a percentage of the shorter side (6.5 % by default), with a bit more margin at the bottom so the caption sits above the like/comment buttons.
Burning it in
ffmpeg -i clip.mp4 -vf "subtitles=clip.ass" -c:v libx264 -crf 20 -c:a copy clip_captioned.mp4
Audio is copied, video is re-encoded once. If you use a non-system font, pass a fonts directory:
-vf "subtitles=clip.ass:fontsdir=./fonts"
Paths with backslashes or drive letters need escaping for the filter (C:\… → C\:/…). The script does that for you.
Edit without re-transcribing
Speech recognition gets names and jargon wrong. The .ass is plain text, so open it, fix the words, and re-burn:
python wordcaptions.py clip.mp4 --render-only
That takes seconds because Whisper doesn't run again. --srt also writes a normal subtitle file if you want something for YouTube's upload form.
Install
Python 3.10–3.13 and ffmpeg (with libass, which the common builds include).
# Windows: winget install Gyan.FFmpeg
# macOS: brew install ffmpeg
# Debian: sudo apt install ffmpeg
git clone https://github.com/philippprimisser-max/whisper-word-captions
cd whisper-word-captions
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
The first run downloads the Whisper model (small ≈ 490 MB). requirements.txt pins av<19, because PyAV 19 broke faster-whisper 1.2.1 on a fresh install in October 2026.
Defaults that matter: --model small (better for non-English than base) and highlight colour #FFD400. --threads is capped at 4. On the bench from the gap checker (base, int8, 5 clips × 3 runs, bench/run_table.sh, 9 October 2026) 8 threads had gaps in 13 of 15 runs (longest 60.5 s), 4 threads in 1 of 15 (longest 28 s) and 1 thread in 1 of 15 (longest 6 s). float32 at 8 threads was 0 of 15. Fewer threads reduce the gaps. They don't remove them. Check the .ass, or use faster-whisper-gap-check.
Measured on an 8-core Linux server: a 15-second 1080×1920 clip took 6.1–6.7 s end to end with small (model load, transcription and burn; two runs, model already downloaded). The clip is made from the repo's test audio, so you can repeat it; the commands are in the README. Your laptop will differ.
AI-assisted: a model helped with the wording. The script was tested and the timing is from my own server (9 October 2026). Demo audio: LibriVox recording of the Gettysburg Address (public domain). Not affiliated with CapCut, TikTok, YouTube, Instagram, OpenAI or the ffmpeg project.

Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more