TL;DR
We'll convert a 4-rung CBR ladder to capped CRF, write a small Python sweep that encodes a set of clips at several CRF and cap values, score each with VMAF, and report average VMAF, lowest-frame VMAF, bitrate, and the share of seconds that hit the cap. Then we'll port the winning settings to NVENC and SVT-AV1.
A fixed-bitrate ladder spends the same on every second of every title. Capped CRF lets the encoder spend less on easy content and holds a ceiling on hard content. The catch is that it has two knobs (the CRF value and the cap) and most people only tune one.
We'll measure instead of guessing. You need ffmpeg 7.x or 8.x built with libx264, libvmaf, and optionally libsvtav1 and NVENC; ffprobe; and python 3.11+. Check your build:
ffmpeg -hide_banner -version | head -1
ffmpeg -hide_banner -filters | grep -E "libvmaf"
ffmpeg -hide_banner -encoders | grep -E "libx264|libsvtav1|h264_nvenc"
ffmpeg version 8.1.2 Copyright (c) 2000-2026 the FFmpeg developers
..C libvmaf VV->V Calculate the VMAF between two video streams.
V..... libx264 libx264 H.264 / AVC / MPEG-4 AVC / MPEG-4 part 10
V..... libsvtav1 SVT-AV1(Scalable Video Technology for AV1) encoder
V....D h264_nvenc NVIDIA NVENC H.264 encoder
If libvmaf is missing, the static builds from BtbN (Linux/Windows) and Homebrew's ffmpeg (macOS) include it.
1. 📐 The baseline: a CBR-ish ladder
Here's the kind of ladder most teams inherit. Four rungs, each a fixed target:
# scripts/ladder_cbr.sh
IN=$1
common="-c:v libx264 -preset slow -profile:v high -g 48 -keyint_min 48 -sc_threshold 0 -c:a aac -b:a 128k"
ffmpeg -y -i "$IN" -vf scale=-2:1080 $common -b:v 6000k -maxrate 6000k -bufsize 12000k out_1080_cbr.mp4
ffmpeg -y -i "$IN" -vf scale=-2:720 $common -b:v 3000k -maxrate 3000k -bufsize 6000k out_720_cbr.mp4
ffmpeg -y -i "$IN" -vf scale=-2:480 $common -b:v 1500k -maxrate 1500k -bufsize 3000k out_480_cbr.mp4
ffmpeg -y -i "$IN" -vf scale=-2:360 $common -b:v 800k -maxrate 800k -bufsize 1600k out_360_cbr.mp4
Keep the GOP flags (-g 48 -keyint_min 48 -sc_threshold 0 for 24 fps, 2-second segments). Capped CRF changes rate control, not segmenting, and your HLS packager still needs aligned keyframes across rungs.
2. 🎚️ Convert it to capped CRF
Capped CRF is three flags: a quality target, a ceiling, and a VBV buffer that enforces the ceiling.
# scripts/ladder_capped.sh
IN=$1; CRF=${2:-23}
common="-c:v libx264 -preset slow -profile:v high -g 48 -keyint_min 48 -sc_threshold 0 -c:a aac -b:a 128k"
ffmpeg -y -i "$IN" -vf scale=-2:1080 $common -crf $CRF -maxrate 6000k -bufsize 6000k out_1080_crf.mp4
ffmpeg -y -i "$IN" -vf scale=-2:720 $common -crf $CRF -maxrate 3000k -bufsize 3000k out_720_crf.mp4
ffmpeg -y -i "$IN" -vf scale=-2:480 $common -crf $CRF -maxrate 1500k -bufsize 1500k out_480_crf.mp4
ffmpeg -y -i "$IN" -vf scale=-2:360 $common -crf $CRF -maxrate 800k -bufsize 800k out_360_crf.mp4
What each flag does:
-
-crf 23: hold this quality; spend whatever bitrate that takes. -
-maxrate 6000k: but never sustain more than this. -
-bufsize 6000k: the VBV buffer the encoder uses to enforcemaxrate.-maxratedoes nothing without it. A buffer equal to the maxrate (1 second of video at the cap) is a common starting point; a larger buffer allows more short-term variability.
💡 Tip:
-crfwithout-maxrate/-bufsizeis "unbounded CRF". It's great for archives and useless for streaming, because one explosion scene can spike past anything your ladder planned for.
Run both scripts on a sample and compare file sizes:
./scripts/ladder_cbr.sh samples/lecture.mp4 && ./scripts/ladder_capped.sh samples/lecture.mp4 23
ls -l out_1080_*.mp4
On easy content (a lecture, a screen recording) the capped-CRF file will be noticeably smaller. On a sports clip it'll be about the same size, because the cap is doing the work. That difference is the whole point, and the next step turns it into numbers.
3. 📊 The sweep: encode, score, count cap hits
We'll sweep CRF values and cap values across a folder of representative clips, score each encode with VMAF against the source, and record four things per encode: mean VMAF, lowest-frame VMAF, average bitrate, and the fraction of one-second windows that sit at or near the cap.
# sweep.py
import csv, json, subprocess, sys
from pathlib import Path
CLIPS = sorted(Path("samples").glob("*.mp4"))
CRFS = [21, 23, 25, 27]
CAPS_K = [6000, 9000, 12000]
HEIGHT = 1080
FPS = 24 # match your source; used for the per-second bitrate windows
def run(cmd):
return subprocess.run(cmd, check=True, capture_output=True, text=True)
def encode(src, crf, cap_k):
out = Path("out") / f"{src.stem}_crf{crf}_cap{cap_k}.mp4"
out.parent.mkdir(exist_ok=True)
run(["ffmpeg", "-y", "-hide_banner", "-loglevel", "error", "-i", str(src),
"-vf", f"scale=-2:{HEIGHT}", "-c:v", "libx264", "-preset", "slow",
"-g", "48", "-keyint_min", "48", "-sc_threshold", "0",
"-crf", str(crf), "-maxrate", f"{cap_k}k", "-bufsize", f"{cap_k}k",
"-an", str(out)])
return out
def vmaf(src, enc):
log = enc.with_suffix(".vmaf.json")
# distorted first, reference second; scale the reference to match
run(["ffmpeg", "-hide_banner", "-loglevel", "error",
"-i", str(enc), "-i", str(src),
"-lavfi", f"[1:v]scale=-2:{HEIGHT}[ref];[0:v][ref]libvmaf=log_fmt=json:log_path={log}:n_threads=8",
"-f", "null", "-"])
d = json.loads(log.read_text())
frames = [f["metrics"]["vmaf"] for f in d["frames"]]
return d["pooled_metrics"]["vmaf"]["mean"], min(frames)
def bitrate_profile(enc, cap_k):
# per-second bitrate from packet sizes; count seconds at >= 95% of cap
out = run(["ffprobe", "-v", "error", "-select_streams", "v:0",
"-show_entries", "packet=pts_time,size", "-of", "csv=p=0", str(enc)]).stdout
buckets = {}
for line in out.strip().splitlines():
pts, size = line.split(",")
buckets[int(float(pts))] = buckets.get(int(float(pts)), 0) + int(size)
kbps = [b * 8 / 1000 for b in buckets.values()]
avg = sum(kbps) / len(kbps)
at_cap = sum(1 for k in kbps if k >= 0.95 * cap_k) / len(kbps)
return avg, at_cap
with open("results.csv", "w", newline="") as f:
w = csv.writer(f)
w.writerow(["clip", "crf", "cap_k", "vmaf_mean", "vmaf_min_frame", "avg_kbps", "pct_seconds_at_cap"])
for src in CLIPS:
for crf in CRFS:
for cap in CAPS_K:
enc = encode(src, crf, cap)
mean, low = vmaf(src, enc)
avg, at_cap = bitrate_profile(enc, cap)
w.writerow([src.name, crf, cap, f"{mean:.2f}", f"{low:.2f}", f"{avg:.0f}", f"{at_cap:.0%}"])
print(f"{src.name:24} crf={crf} cap={cap}k vmaf={mean:.2f} low={low:.2f} {avg:.0f} kbps at_cap={at_cap:.0%}", file=sys.stderr)
python sweep.py
lecture.mp4 crf=21 cap=6000k vmaf=97.1 low=88.4 2410 kbps at_cap=0%
lecture.mp4 crf=23 cap=6000k vmaf=96.3 low=86.9 1680 kbps at_cap=0%
lecture.mp4 crf=25 cap=6000k vmaf=95.0 low=84.7 1190 kbps at_cap=0%
football.mp4 crf=23 cap=6000k vmaf=92.8 low=61.2 5840 kbps at_cap=91%
football.mp4 crf=26 cap=12000k vmaf=92.9 low=74.0 4720 kbps at_cap=12%
...
Those lines show the shape of the output, not numbers you should expect; your clips will differ. But the two patterns are the ones to look for:
-
Easy clip,
at_cap=0%. CRF is in control the whole time. Raising CRF from 21 to 25 drops bitrate a lot and VMAF a little. This is the content you've been overpaying for. -
Hard clip,
at_caphigh. The cap is in control; the CRF value barely matters. Lowest-frame VMAF is poor because a hard ceiling starves hard frames. A higher cap with a higher CRF can lower the average bitrate while raising the lowest-frame score, because the encoder stops smearing bits evenly across frames that need them unevenly. This counterintuitive result is documented in Jan Ozer's February 2026 case study, where a sports clip went from 5542 to 4593 kbps average by moving from CRF 24 at a 6 Mbps cap to CRF 26 at 12 Mbps, with lowest-frame VMAF rising from 63.81 to 75.21.
⚠️ Note:
pooled_metrics.vmaf.meanis the headline, butvmaf_min_frameis where viewer complaints live. A 93 average with a 60 floor means one scene looks broken. Track both.
4. 🎯 Pick the target, then pick per genre
A VMAF of 93 is the widely quoted "good enough" line (most viewers judge the encode indistinguishable from the source or noticeably but not annoyingly different). Scores well above 95 are bandwidth nobody notices.
Open results.csv and, per clip, find the cheapest (CRF, cap) row with vmaf_mean >= 93 and an acceptable floor. You'll likely find they cluster by genre, not by file. Ozer's advice is 10 to 20 test files per genre before you give a genre its own profile. Then the production config is a tiny table:
# app/encode_profiles.py
PROFILES = {
# 1080p 720p 480p 360p
"talking_head": [(25, 4000), (25, 2000), (26, 1000), (27, 600)],
"sports": [(26, 12000), (26, 6000), (26, 3000), (27, 1500)],
"animation": [(24, 5000), (24, 2500), (25, 1200), (26, 700)],
"default": [(23, 6000), (23, 3000), (24, 1500), (25, 800)],
}
# (crf, maxrate_kbps); bufsize = maxrate unless you've measured a reason otherwise
Those numbers are placeholders shaped like real ones. Fill them from your own results.csv.
5. 🔁 Same idea, other encoders
The decision transfers across codecs; the flags don't.
| Encoder | Capped-CRF form | Notes |
|---|---|---|
libx264 / libx265
|
-crf N -maxrate Xk -bufsize Yk |
Single pass. |
h264_nvenc / hevc_nvenc / av1_nvenc
|
-cq N -rc:v vbr -maxrate Xk -bufsize Yk |
CQ scale differs from CRF; re-sweep. |
libsvtav1 |
-crf N -maxrate Xk -bufsize Yk |
CRF + maxrate/bufsize is supported by FFmpeg's SVT-AV1 wrapper. |
libvpx-vp9 |
-crf N -b:v 0 -maxrate Xk -bufsize Yk (or low -b:v + -qmax as a floor) |
Ozer's client used the -qmax approach for VP9 and found it cheaper than capped CRF at equal quality. |
# NVENC H.264, same shape
ffmpeg -y -i in.mp4 -c:v h264_nvenc -preset p5 -cq 32 -rc:v vbr -maxrate 6000k -bufsize 6000k -an out_nvenc.mp4
# SVT-AV1
ffmpeg -y -i in.mp4 -c:v libsvtav1 -preset 6 -crf 35 -maxrate 4000k -bufsize 4000k -g 48 -an out_av1.mp4
Change encode() in sweep.py to take an encoder string and you have the same report for every codec you ship.
6. ✅ Sanity checks before you ship the ladder
# keyframes still aligned across rungs (first 10)
for f in out_1080_crf.mp4 out_720_crf.mp4; do
ffprobe -v error -select_streams v:0 -skip_frame nokey -show_entries frame=pts_time -of csv=p=0 "$f" | head -10 | tr '\n' ' '; echo
done
0.000000 2.000000 4.000000 6.000000 8.000000 10.000000 12.000000 14.000000 16.000000 18.000000
0.000000 2.000000 4.000000 6.000000 8.000000 10.000000 12.000000 14.000000 16.000000 18.000000
And a bandwidth check: the HLS BANDWIDTH attribute per rung should be the measured peak, not the old nominal target. With capped CRF the peak is bounded by maxrate (plus VBV slack), so set BANDWIDTH to your cap and AVERAGE-BANDWIDTH to the measured average from results.csv.
What's next
- Wire
PROFILESinto your upload pipeline with a genre tag at ingest, and keep the default profile for untagged content. - Add a weekly job that re-runs
sweep.pyon a random sample of new uploads, so your CRF values track the catalogue as it changes. - If you're already on a managed encoder, check whether it exposes a quality target plus cap (some call it "constrained quality" or "content-adaptive"); the measuring approach here still applies to what it produces.
Top comments (0)