DEV Community

Cover image for Word-by-word captions with Python and ffmpeg, no CapCut watermark
Philipp Primisser
Philipp Primisser

Posted on AI-assisted

Word-by-word captions with Python and ffmpeg, no CapCut watermark

Short-form video lives or dies on captions. CapCut and a dozen web tools will do the "word lights up as you say it" style for you, usually with a watermark or a free-minute limit. I wanted the same look without an account, so I wrote a small offline script: whisper-word-captions (MIT). This post is about how the trick works, so you can change it.

Word-by-word caption demo from the repository

python wordcaptions.py clip.mp4
# -> clip_captioned.mp4  and  clip.ass
Enter fullscreen mode Exit fullscreen mode

Three pieces

  1. faster-whisper with word_timestamps=True → every word gets its own start and end.
  2. Group words into short captions (max 3 by default; also break after .?! or a pause over 0.6 s).
  3. Write an Advanced SubStation Alpha (.ass) file with one Dialogue event per spoken word, then burn it into the picture with ffmpeg's subtitles filter (libass).

The interesting part is step 3.

The ASS trick

ASS lets you override style mid-line with curly-brace tags. \c changes the text colour (ASS colours are &HAABBGGRR, blue-green-red), \fscx/\fscy scale it, \r resets to the style. For a three-word caption you write three events, one per word:

Dialogue: 0,0:00:03.20,0:00:03.52,Word,,0,0,0,,{\c&H0000D4FF\fscx115\fscy115}THAT{\r} ALL MEN
Dialogue: 0,0:00:03.52,0:00:03.68,Word,,0,0,0,,THAT {\c&H0000D4FF\fscx115\fscy115}ALL{\r} MEN
Dialogue: 0,0:00:03.68,0:00:03.90,Word,,0,0,0,,THAT ALL {\c&H0000D4FF\fscx115\fscy115}MEN{\r}
Enter fullscreen mode Exit fullscreen mode

Each event lasts until the next word starts. That way nothing flickers during short pauses, and the viewer always sees the whole caption with only the current word highlighted. In Python that's just a nested loop:

for group in groups:
    texts = [escape_ass(w.text.upper()) for w in group]
    for i, w in enumerate(group):
        end = group[i + 1].start if i + 1 < len(group) else max(w.end, w.start + 0.05)
        parts = []
        for j, t in enumerate(texts):
            if j == i:
                parts.append(f"{{\\c{hl}\\fscx115\\fscy115}}{t}{{\\r}}")
            else:
                parts.append(t)
        events.append(f"Dialogue: 0,{ass_time(w.start)},{ass_time(end)},Word,,0,0,0,,{' '.join(parts)}")
Enter fullscreen mode Exit fullscreen mode

PlayResX/PlayResY in the .ass header match the video size, so the same file works for 1080×1920 Shorts and for landscape. Font size is a percentage of the shorter side (6.5 % by default), with a bit more margin at the bottom so the caption sits above the like/comment buttons.

Burning it in

ffmpeg -i clip.mp4 -vf "subtitles=clip.ass" -c:v libx264 -crf 20 -c:a copy clip_captioned.mp4
Enter fullscreen mode Exit fullscreen mode

Audio is copied, video is re-encoded once. If you use a non-system font, pass a fonts directory:

-vf "subtitles=clip.ass:fontsdir=./fonts"
Enter fullscreen mode Exit fullscreen mode

Paths with backslashes or drive letters need escaping for the filter (C:\… → C\:/…). The script does that for you.

Edit without re-transcribing

Speech recognition gets names and jargon wrong. The .ass is plain text, so open it, fix the words, and re-burn:

python wordcaptions.py clip.mp4 --render-only
Enter fullscreen mode Exit fullscreen mode

That takes seconds because Whisper doesn't run again. --srt also writes a normal subtitle file if you want something for YouTube's upload form.

Install

Python 3.10–3.13 and ffmpeg (with libass, which the common builds include).

# Windows: winget install Gyan.FFmpeg
# macOS:   brew install ffmpeg
# Debian:  sudo apt install ffmpeg
git clone https://github.com/philippprimisser-max/whisper-word-captions
cd whisper-word-captions
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
Enter fullscreen mode Exit fullscreen mode

The first run downloads the Whisper model (small ≈ 490 MB). requirements.txt pins av<19, because PyAV 19 broke faster-whisper 1.2.1 on a fresh install in October 2026.

Defaults that matter: --model small (better for non-English than base) and highlight colour #FFD400. --threads is capped at 4. On the bench from the gap checker (base, int8, 5 clips × 3 runs, bench/run_table.sh, 9 October 2026) 8 threads had gaps in 13 of 15 runs (longest 60.5 s), 4 threads in 1 of 15 (longest 28 s) and 1 thread in 1 of 15 (longest 6 s). float32 at 8 threads was 0 of 15. Fewer threads reduce the gaps. They don't remove them. Check the .ass, or use faster-whisper-gap-check.

Measured on an 8-core Linux server: a 15-second 1080×1920 clip took 6.1–6.7 s end to end with small (model load, transcription and burn; two runs, model already downloaded). The clip is made from the repo's test audio, so you can repeat it; the commands are in the README. Your laptop will differ.


AI-assisted: a model helped with the wording. The script was tested and the timing is from my own server (9 October 2026). Demo audio: LibriVox recording of the Gettysburg Address (public domain). Not affiliated with CapCut, TikTok, YouTube, Instagram, OpenAI or the ffmpeg project.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more