DEV Community

Mason K
Mason K

Posted on

Test your video pipeline in CI for free: synthetic fixtures, ffprobe goldens, and webhook replay with pytest

TL;DR

We'll give a video pipeline a real test suite that runs on every pull request, costs nothing, and stores no video files in git. Fixtures are generated with FFmpeg's lavfi source, outputs are asserted with ffprobe JSON snapshots (not byte comparisons), and webhook handlers are tested by replaying recorded payloads. pytest, FFmpeg 8.1 pinned in GitHub Actions, Python 3.12.

What we're building

A tests/ directory for a pipeline that looks like most of them: a probe() step, a build_ffmpeg_args() step that picks a ladder, a transcode() step that shells out, and a webhook handler that advances state when the encoder (local or managed) reports back.

The four tiers, in order of how often they run:

Tier Needs FFmpeg? Needs network? Runs on
1. Pure logic (arg builders, ladder policy) no no every commit
2. FFmpeg in the loop on synthetic fixtures yes no every commit
3. Webhook replay from recorded payloads no no every commit
4. One real provider call yes yes nightly, budgeted

This post builds tiers 2 and 3 in full and shows the skeleton for 1 and 4.

1. Project layout

pipeline/
  probe.py          # ffprobe wrapper
  ladder.py         # pure: probe result -> list of renditions
  transcode.py      # shells out to ffmpeg
  webhooks.py       # FastAPI handler
tests/
  conftest.py       # fixture factory
  test_ladder.py    # tier 1
  test_transcode.py # tier 2
  test_webhooks.py  # tier 3
  fixtures/webhooks/*.json
  goldens/*.json
.github/workflows/ci.yml
Enter fullscreen mode Exit fullscreen mode
python -m venv .venv && source .venv/bin/activate
pip install pytest pytest-snapshot fastapi httpx
Enter fullscreen mode Exit fullscreen mode

2. The fixture factory (no binaries in git) 🎬

FFmpeg can synthesise a test pattern and a tone from nothing. A two-second 320x240 clip is a few hundred KB and takes well under a second to make.

# tests/conftest.py
import subprocess
from pathlib import Path
import pytest

def make_fixture(
    out: Path,
    seconds: float = 2,
    size: str = "320x240",
    fps: int = 24,
    audio: bool = True,
    vcodec: str = "libx264",
) -> Path:
    cmd = ["ffmpeg", "-y", "-hide_banner", "-loglevel", "error",
           "-f", "lavfi", "-i", f"testsrc2=duration={seconds}:size={size}:rate={fps},format=yuv420p"]
    if audio:
        cmd += ["-f", "lavfi", "-i", f"sine=frequency=440:duration={seconds}"]
    cmd += ["-c:v", vcodec, "-preset", "ultrafast", "-crf", "28"]
    if audio:
        cmd += ["-c:a", "aac", "-shortest"]
    cmd += [str(out)]
    subprocess.run(cmd, check=True)
    return out

@pytest.fixture
def fixture_factory(tmp_path):
    def _make(name="in.mp4", **kw):
        return make_fixture(tmp_path / name, **kw)
    return _make
Enter fullscreen mode Exit fullscreen mode

Now every test can ask for exactly the input it needs:

fixture_factory()                                   # 2s, 320x240, 24fps, with audio
fixture_factory("vertical.mp4", size="240x426")     # portrait source
fixture_factory("silent.mp4", audio=False)          # no audio stream
fixture_factory("hfr.mp4", fps=60)                  # high frame rate
Enter fullscreen mode Exit fullscreen mode

💡 Tip: testsrc2 produces a pattern with motion and text, which is closer to real content than plain colour bars and makes encoder behaviour (keyframes, bitrate) more realistic.

3. Tier 1: pure logic, no FFmpeg

The refactor that makes everything else possible: keep "decide what to run" separate from "run it".

# pipeline/ladder.py
from dataclasses import dataclass

@dataclass(frozen=True)
class Rendition:
    height: int
    video_kbps: int

LADDER = [Rendition(1080, 5000), Rendition(720, 2800), Rendition(480, 1400), Rendition(360, 800)]

def pick_ladder(width: int, height: int) -> list[Rendition]:
    """Never upscale. Portrait sources are judged by their short edge."""
    short_edge = min(width, height)
    return [r for r in LADDER if r.height <= short_edge]
Enter fullscreen mode Exit fullscreen mode
# tests/test_ladder.py
from pipeline.ladder import pick_ladder

def test_720p_source_gets_three_rungs():
    assert [r.height for r in pick_ladder(1280, 720)] == [720, 480, 360]

def test_portrait_source_uses_short_edge():
    assert [r.height for r in pick_ladder(1080, 1920)] == [1080, 720, 480, 360]

def test_tiny_source_gets_nothing_above_itself():
    assert pick_ladder(320, 240) == []
Enter fullscreen mode Exit fullscreen mode

These run in milliseconds and catch a surprising share of real bugs, because ladder logic is where the edge cases live.

4. Tier 2: FFmpeg in the loop, asserted with ffprobe

Encoders are not bit-exact across versions, so never compare output bytes. Probe the output and assert on the properties your pipeline is responsible for.

# pipeline/probe.py
import json, subprocess
from pathlib import Path

def probe(path: Path) -> dict:
    out = subprocess.run(
        ["ffprobe", "-v", "error", "-print_format", "json", "-show_streams", "-show_format", str(path)],
        capture_output=True, text=True, check=True,
    ).stdout
    return json.loads(out)

STABLE_STREAM_KEYS = ("codec_type", "codec_name", "width", "height", "pix_fmt", "r_frame_rate", "channels", "sample_rate")

def stable_view(p: dict) -> dict:
    """Keep only fields that survive FFmpeg upgrades. Drop bit_rate, size, tags."""
    return {
        "streams": [{k: s.get(k) for k in STABLE_STREAM_KEYS if k in s} for s in p["streams"]],
        "duration": round(float(p["format"]["duration"]), 1),
        "nb_streams": p["format"]["nb_streams"],
    }
Enter fullscreen mode Exit fullscreen mode
# pipeline/transcode.py
import subprocess
from pathlib import Path
from .ladder import Rendition

def transcode(src: Path, dst: Path, r: Rendition) -> Path:
    cmd = ["ffmpeg", "-y", "-hide_banner", "-loglevel", "error", "-i", str(src),
           "-vf", f"scale=-2:{r.height}", "-c:v", "libx264", "-preset", "veryfast",
           "-b:v", f"{r.video_kbps}k", "-g", "48", "-keyint_min", "48", "-sc_threshold", "0",
           "-c:a", "aac", "-b:a", "128k", "-movflags", "+faststart", str(dst)]
    subprocess.run(cmd, check=True)
    return dst
Enter fullscreen mode Exit fullscreen mode

Now the tests. One direct assertion, one golden snapshot.

# tests/test_transcode.py
import json
from pipeline.ladder import Rendition
from pipeline.probe import probe, stable_view
from pipeline.transcode import transcode

def test_360p_rendition_has_expected_shape(fixture_factory, tmp_path):
    src = fixture_factory(size="640x360", seconds=2)
    out = transcode(src, tmp_path / "out.mp4", Rendition(360, 800))
    view = stable_view(probe(out))

    video = next(s for s in view["streams"] if s["codec_type"] == "video")
    audio = next(s for s in view["streams"] if s["codec_type"] == "audio")
    assert video["codec_name"] == "h264"
    assert (video["width"], video["height"]) == (640, 360)
    assert video["pix_fmt"] == "yuv420p"
    assert audio["codec_name"] == "aac"
    assert view["duration"] == 2.0

def test_360p_rendition_matches_golden(fixture_factory, tmp_path, snapshot):
    src = fixture_factory(size="640x360", seconds=2)
    out = transcode(src, tmp_path / "out.mp4", Rendition(360, 800))
    snapshot.snapshot_dir = "tests/goldens"
    snapshot.assert_match(json.dumps(stable_view(probe(out)), indent=2, sort_keys=True), "360p.json")
Enter fullscreen mode Exit fullscreen mode

First run creates the golden; later runs diff against it.

pytest tests/test_transcode.py --snapshot-update   # once, to create tests/goldens/360p.json
pytest tests/test_transcode.py
# tests/test_transcode.py ..                                              [100%]
# 2 passed in 1.84s
Enter fullscreen mode Exit fullscreen mode

⚠️ Note: if a golden ever includes bit_rate or size, it will break on the next FFmpeg upgrade and someone will delete the test. stable_view() exists to stop that.

The keyframe cadence you set with -g 48 is stable across builds and worth asserting too:

def keyframe_indices(path):
    out = subprocess.run(["ffprobe", "-v", "error", "-select_streams", "v:0",
                          "-show_entries", "frame=key_frame", "-of", "csv=p=0", str(path)],
                         capture_output=True, text=True, check=True).stdout.split()
    return [i for i, k in enumerate(out) if k == "1"]

def test_keyframes_every_48_frames(fixture_factory, tmp_path):
    src = fixture_factory(size="640x360", seconds=4, fps=24)
    out = transcode(src, tmp_path / "out.mp4", Rendition(360, 800))
    assert keyframe_indices(out) == [0, 48]
Enter fullscreen mode Exit fullscreen mode

5. Tier 3: webhook replay from recorded payloads

The provider is not part of your test suite. Its payloads are. Capture one real payload per event type once (from a dev environment, with a request logger), strip anything sensitive, and commit the JSON.

// tests/fixtures/webhooks/media_ready.json
{
  "type": "media.ready",
  "id": "evt_01",
  "data": { "id": "media_abc123", "status": "ready", "playback_url": "https://cdn.example.com/media_abc123/index.m3u8" }
}
Enter fullscreen mode Exit fullscreen mode
// tests/fixtures/webhooks/media_failed.json
{
  "type": "media.failed",
  "id": "evt_02",
  "data": { "id": "media_abc123", "status": "failed", "error": { "message": "unsupported codec" } }
}
Enter fullscreen mode Exit fullscreen mode

Field names will differ per provider; the shape of the tests does not.

# pipeline/webhooks.py
from fastapi import FastAPI, Request

app = FastAPI()
STATE: dict[str, str] = {}          # media_id -> status; swap for your DB
SEEN: set[str] = set()              # event ids, for idempotency

@app.post("/webhooks/video")
async def video_webhook(req: Request):
    evt = await req.json()
    if evt["id"] in SEEN:
        return {"ok": True, "duplicate": True}
    SEEN.add(evt["id"])
    media_id = evt["data"]["id"]
    if evt["type"] == "media.ready":
        STATE[media_id] = "ready"
    elif evt["type"] == "media.failed":
        STATE[media_id] = "failed"
    return {"ok": True}
Enter fullscreen mode Exit fullscreen mode
# tests/test_webhooks.py
import json
from pathlib import Path
import pytest
from httpx import ASGITransport, AsyncClient
from pipeline import webhooks

FIX = Path("tests/fixtures/webhooks")
load = lambda name: json.loads((FIX / f"{name}.json").read_text())

@pytest.fixture(autouse=True)
def reset_state():
    webhooks.STATE.clear(); webhooks.SEEN.clear()

@pytest.fixture
async def client():
    async with AsyncClient(transport=ASGITransport(app=webhooks.app), base_url="http://t") as c:
        yield c

async def test_ready_marks_media_ready(client):
    r = await client.post("/webhooks/video", json=load("media_ready"))
    assert r.status_code == 200
    assert webhooks.STATE["media_abc123"] == "ready"

async def test_duplicate_ready_is_idempotent(client):
    await client.post("/webhooks/video", json=load("media_ready"))
    r = await client.post("/webhooks/video", json=load("media_ready"))
    assert r.json()["duplicate"] is True
    assert webhooks.STATE == {"media_abc123": "ready"}

async def test_failed_after_ready_is_a_decision_you_made(client):
    await client.post("/webhooks/video", json=load("media_ready"))
    await client.post("/webhooks/video", json=load("media_failed"))
    # Pick your policy and pin it. Here: last event wins.
    assert webhooks.STATE["media_abc123"] == "failed"

async def test_missing_optional_field_does_not_crash(client):
    evt = load("media_ready"); del evt["data"]["playback_url"]
    r = await client.post("/webhooks/video", json=evt)
    assert r.status_code == 200
Enter fullscreen mode Exit fullscreen mode

You just tested three things a manual staging upload can never produce on demand: a duplicate delivery, an out-of-order failure, and a payload with a field missing.

6. CI: pin FFmpeg, run everything

# .github/workflows/ci.yml
name: ci
on: [push, pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - uses: FedericoCarboni/setup-ffmpeg@v3
        with:
          ffmpeg-version: "8.1"      # pin it; unpinned runners can drift between OSes
      - run: pip install -r requirements.txt
      - run: pytest -q
Enter fullscreen mode Exit fullscreen mode

setup-ffmpeg@v3 installs and caches the requested build and exposes both ffmpeg and ffprobe. Its README is explicit that without a pinned version, different operating systems may get different FFmpeg versions, which is exactly the drift your goldens would catch and you'd rather not debug.

# expected CI output
tests/test_ladder.py ...                                                  [ 33%]
tests/test_transcode.py ...                                               [ 66%]
tests/test_webhooks.py ....                                               [100%]
10 passed in 3.41s
Enter fullscreen mode Exit fullscreen mode

7. Tier 4 skeleton: the one paid call

# .github/workflows/nightly-contract.yml
on:
  schedule: [{ cron: "0 3 * * *" }]
jobs:
  contract:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    steps:
      - uses: actions/checkout@v4
      - uses: FedericoCarboni/setup-ffmpeg@v3
        with: { ffmpeg-version: "8.1" }
      - run: pytest -q -m contract
        env: { PROVIDER_TOKEN: "${{ secrets.PROVIDER_TOKEN }}" }
Enter fullscreen mode Exit fullscreen mode

Mark those tests @pytest.mark.contract, exclude the marker from the default run, upload one two-second generated fixture, poll until ready, fetch the playback manifest, assert a 200. One asset, one run a night, a hard timeout. That's the whole budget.

Things to know

  • Generated fixtures are deterministic in shape, not in bytes. Assert on probe output, never on checksums of media.
  • -preset ultrafast for fixtures, your real preset for the code under test. The fixture just needs to be a valid input; the transcode is what you're testing.
  • Record webhooks from a dev environment, not production. Sanitise IDs anyway.
  • Keep goldens small. stable_view() should produce a few dozen lines, not the full ffprobe dump.
  • Re-baseline on purpose. When you bump the pinned FFmpeg, run --snapshot-update in the same PR and review the diff. That diff is the changelog for your outputs.

What's next

  • Add a tier-2 test per rendition in your ladder and one per "weird input" (VFR, no audio, portrait, HDR) as you meet them in production.
  • If your webhook handler is the idempotent Postgres-backed one, point the replay tests at a test database instead of the dict and keep the same fixtures.
  • Wire stable_view() into your production pipeline as a post-encode QC check; it's the same function.

Post your fixture factory variants under #python or #devops; the "weird input" list is the part every team builds from scratch.

Top comments (0)