A step-by-step guide to converting narration scripts into broadcast-ready SRT files — zero external dependencies.
The SRT format
An SRT file is simple text. Each subtitle block contains four parts: a sequential number, a time range in HH:MM:SS,mmm format, the subtitle text, and a blank line separator.
1
00:00:00,000 --> 00:00:05,200
First subtitle line.
2
00:00:05,200 --> 00:00:10,800
Second line.
3
00:00:10,800 --> 00:00:15,000
Third line, and so on.
That is it. No XML, no JSON. Sequential numbers, time ranges, text, blank lines. The blank line between blocks matters — parsers use it as a delimiter.
The algorithm
A typical implementation follows four steps:
- Split narration into sentences using punctuation (。!?.!?...)
- Estimate duration per sentence (a common default is 5 characters per second, adjustable for speaking pace)
- Accumulate timestamps — each sentence starts where the previous one ended
- Write SRT blocks in spec format
The accumulation step requires tracking a running time counter. Each sentence's start time equals the previous sentence's end time. Gaps and overlaps should be avoided.
Example implementation:
def script_to_srt(text, total_seconds=None, chars_per_second=5):
sentences = split_sentences(text)
if not total_seconds:
total_seconds = len(text) / chars_per_second
srt_blocks = []
current_time = 0.0
for i, sentence in enumerate(sentences, 1):
duration = len(sentence) / chars_per_second
start = format_time(current_time)
end = format_time(current_time + duration)
srt_blocks.append(f"{i}\n{start} --> {end}\n{sentence}\n")
current_time += duration
return "\n".join(srt_blocks)
def format_time(seconds):
hours = int(seconds // 3600)
minutes = int((seconds % 3600) // 60)
secs = int(seconds % 60)
millis = int((seconds % 1) * 1000)
return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"
The split_sentences function typically uses regex on Chinese and English punctuation: [。!?.!?]+\s*. For bilingual content, both character sets need handling without mixing sentence boundaries.
Burning with ffmpeg
Once the SRT file is ready, ffmpeg can burn it directly into the video:
ffmpeg -i input.mp4 -vf "subtitles=subtitle.srt" -c:a copy output.mp4
For styling — font size, color, outline, vertical margin:
ffmpeg -i input.mp4 -vf "subtitles=subtitle.srt:force_style='FontSize=24,PrimaryColour=&H00FFFFFF&,Outline=1,MarginV=40'" -c:a copy output.mp4
The &H00FFFFFF& notation is BGR hex — in this case white text. Outline=1 draws a thin border so subtitles remain readable against light backgrounds.
To verify the result:
ffprobe -v quiet -print_format json -show_streams output.mp4 | grep -E '"codec_type"|"height"|"duration"'
Edge cases
CJK characters. ffmpeg's default font often does not include Chinese glyphs. The fix is adding FontName=Noto Sans CJK SC to force_style. Without it, Chinese text renders as empty boxes or tofu characters.
ffmpeg -i input.mp4 -vf "subtitles=subtitle.srt:force_style='FontSize=24,FontName=Noto Sans CJK SC,PrimaryColour=&H00FFFFFF&'" -c:a copy output.mp4
Long sentences. The 5 chars/sec rate works for typical narration. Some sentences run longer and stay on screen too long — over ~20 characters per sentence, readability suffers. A manual break step can split sentences over a character threshold at commas or conjunctions, giving each piece its own time slice.
Overlapping timestamps. If the rate estimate is too fast (8+ chars/sec for fast-talking narration), shorter sentences end up with the same start time as the previous end time, or worse — they overlap. Monotonicity validation after generation catches this: verify start[i] >= end[i-1] for every block. If violated, adjust the rate down and regenerate.
Bilingual files. When a script mixes Chinese and English paragraphs, character-based timing skews because English text has higher information density per character. One approach is splitting the file and processing language segments separately with different rates.
Limitations
This character-count approach assumes narration timing correlates with character count. It does not work for:
- Music lyrics — timing comes from audio, not text length
- Dialogue with pauses — silence gaps break the accumulation model
- Pre-timed content — when timecodes already exist, use them directly instead of estimating
For those cases, speech detection (WebVTT + Whisper) or manual timecode entry is more appropriate. The script described here targets the simple case: single voice, continuous narration, text-only input.
A complete implementation handling bilingual output, monotonicity validation, and debug mapping typically runs around 150 lines with zero external dependencies.
The version of this script used in production is part of dijily Skill Station, a set of open-source skills that run locally — no API keys, no uploads to external servers.
Top comments (0)