DEV Community

yudong
yudong

Posted on AI-assisted

Build a Python Subtitle Generator with ffmpeg: A Step-by-Step Guide

A step-by-step guide to converting narration scripts into broadcast-ready SRT files — zero external dependencies.


The SRT format

An SRT file is simple text. Each subtitle block contains four parts: a sequential number, a time range in HH:MM:SS,mmm format, the subtitle text, and a blank line separator.

1
00:00:00,000 --> 00:00:05,200
First subtitle line.

2
00:00:05,200 --> 00:00:10,800
Second line.

3
00:00:10,800 --> 00:00:15,000
Third line, and so on.
Enter fullscreen mode Exit fullscreen mode

That is it. No XML, no JSON. Sequential numbers, time ranges, text, blank lines. The blank line between blocks matters — parsers use it as a delimiter.

The algorithm

A typical implementation follows four steps:

  1. Split narration into sentences using punctuation (。!?.!?...)
  2. Estimate duration per sentence (a common default is 5 characters per second, adjustable for speaking pace)
  3. Accumulate timestamps — each sentence starts where the previous one ended
  4. Write SRT blocks in spec format

The accumulation step requires tracking a running time counter. Each sentence's start time equals the previous sentence's end time. Gaps and overlaps should be avoided.

Example implementation:

def script_to_srt(text, total_seconds=None, chars_per_second=5):
    sentences = split_sentences(text)
    if not total_seconds:
        total_seconds = len(text) / chars_per_second

    srt_blocks = []
    current_time = 0.0

    for i, sentence in enumerate(sentences, 1):
        duration = len(sentence) / chars_per_second
        start = format_time(current_time)
        end = format_time(current_time + duration)
        srt_blocks.append(f"{i}\n{start} --> {end}\n{sentence}\n")
        current_time += duration

    return "\n".join(srt_blocks)

def format_time(seconds):
    hours = int(seconds // 3600)
    minutes = int((seconds % 3600) // 60)
    secs = int(seconds % 60)
    millis = int((seconds % 1) * 1000)
    return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"
Enter fullscreen mode Exit fullscreen mode

The split_sentences function typically uses regex on Chinese and English punctuation: [。!?.!?]+\s*. For bilingual content, both character sets need handling without mixing sentence boundaries.

Burning with ffmpeg

Once the SRT file is ready, ffmpeg can burn it directly into the video:

ffmpeg -i input.mp4 -vf "subtitles=subtitle.srt" -c:a copy output.mp4
Enter fullscreen mode Exit fullscreen mode

For styling — font size, color, outline, vertical margin:

ffmpeg -i input.mp4 -vf "subtitles=subtitle.srt:force_style='FontSize=24,PrimaryColour=&H00FFFFFF&,Outline=1,MarginV=40'" -c:a copy output.mp4
Enter fullscreen mode Exit fullscreen mode

The &H00FFFFFF& notation is BGR hex — in this case white text. Outline=1 draws a thin border so subtitles remain readable against light backgrounds.

To verify the result:

ffprobe -v quiet -print_format json -show_streams output.mp4 | grep -E '"codec_type"|"height"|"duration"'
Enter fullscreen mode Exit fullscreen mode

Edge cases

CJK characters. ffmpeg's default font often does not include Chinese glyphs. The fix is adding FontName=Noto Sans CJK SC to force_style. Without it, Chinese text renders as empty boxes or tofu characters.

ffmpeg -i input.mp4 -vf "subtitles=subtitle.srt:force_style='FontSize=24,FontName=Noto Sans CJK SC,PrimaryColour=&H00FFFFFF&'" -c:a copy output.mp4
Enter fullscreen mode Exit fullscreen mode

Long sentences. The 5 chars/sec rate works for typical narration. Some sentences run longer and stay on screen too long — over ~20 characters per sentence, readability suffers. A manual break step can split sentences over a character threshold at commas or conjunctions, giving each piece its own time slice.

Overlapping timestamps. If the rate estimate is too fast (8+ chars/sec for fast-talking narration), shorter sentences end up with the same start time as the previous end time, or worse — they overlap. Monotonicity validation after generation catches this: verify start[i] >= end[i-1] for every block. If violated, adjust the rate down and regenerate.

Bilingual files. When a script mixes Chinese and English paragraphs, character-based timing skews because English text has higher information density per character. One approach is splitting the file and processing language segments separately with different rates.

Limitations

This character-count approach assumes narration timing correlates with character count. It does not work for:

  • Music lyrics — timing comes from audio, not text length
  • Dialogue with pauses — silence gaps break the accumulation model
  • Pre-timed content — when timecodes already exist, use them directly instead of estimating

For those cases, speech detection (WebVTT + Whisper) or manual timecode entry is more appropriate. The script described here targets the simple case: single voice, continuous narration, text-only input.

A complete implementation handling bilingual output, monotonicity validation, and debug mapping typically runs around 150 lines with zero external dependencies.


The version of this script used in production is part of dijily Skill Station, a set of open-source skills that run locally — no API keys, no uploads to external servers.

Top comments (0)