DEV Community

Bugheanu Danut Andrei
Bugheanu Danut Andrei

Posted on

Evaluating Video API Architectures: Dynamic URL Transformations vs Pre-Rendered HLS Manifests

Dynamic URL video transformation allows real-time cropping, codec switching, and quality adjustments at edge nodes, eliminating the need to pre-encode and store multiple static video variants.

Two architectures

Pre-rendered ladders (e.g. Mux)

You upload a video once. The platform encodes an adaptive bitrate ladder, a set of renditions at different resolutions and bitrates, and serves them through an HLS manifest. Mux picks the ladder per title ("per-title encoding… boosts bitrates for high-complexity content") and offers just-in-time encoding so playback can start soon after upload.

What you control at playback time is which existing renditions the player sees. Mux's playback URL accepts parameters such as max_resolution, min_resolution, rendition_order and asset_start_time / asset_end_time. It doesn't take parameters to crop, switch codec or set an arbitrary bitrate. Those are decisions made at encoding time.

Dynamic URL transformations (e.g. Cloudinary)

You upload a video once. Every derived version is described in the delivery URL:

https://res.cloudinary.com/demo/video/upload/vc_h264,br_2500k/samples/sea-turtle.mp4
https://res.cloudinary.com/demo/video/upload/c_fill,w_1080,h_1920/vc_h265/samples/sea-turtle.mp4
https://res.cloudinary.com/demo/video/upload/sp_auto/samples/sea-turtle.m3u8
Enter fullscreen mode Exit fullscreen mode

The first request for a new URL triggers an encode; later requests are served from cache. Adding a new aspect ratio for a new placement, or a lower bitrate for a new market, means changing a URL, with no batch job and no storage bookkeeping on your side.

The trade-off is the first request. A variant nobody has asked for yet must be encoded before it can be served, which is why production setups pre-generate ("eager") the variants they know they'll need and let on-the-fly transformation handle the long tail.

What we measured

Setup: three 1080p H.264 clips hosted on Cloudinary's public demo cloud, downloaded byte-for-byte as the reference. Each clip was encoded by Cloudinary (URL transformations, variable and constant bitrate) and, as a control, by local two-pass x264 and x265 at preset medium, all from the same file. Every rendition was scored against the reference with FFmpeg libvmaf (VMAF v0.6.1, PSNR-Y, SSIM). Latency was measured from a single client, so treat it as indicative.

Quality at each bitrate (mean of 3 clips)

Encoder Target Measured bitrate VMAF VMAF 5th pct PSNR-Y SSIM
Cloudinary H.264, br_ VBR 1 Mbps 969 kbps 73.66 62.23 39.23 dB 0.9781
Cloudinary H.264, br_ VBR 2.5 Mbps 1842 kbps 87.53 82.75 41.92 dB 0.9923
Cloudinary H.264, br_ VBR 4.5 Mbps 2035 kbps 88.96 84.93 42.23 dB 0.9932
Cloudinary H.264, br_…:constant 1 Mbps 911 kbps 72.39 62.43 38.92 dB 0.9771
Cloudinary H.264, br_…:constant 2.5 Mbps 2402 kbps 89.08 83.98 42.86 dB 0.9931
Cloudinary H.264, br_…:constant 4.5 Mbps 4384 kbps 94.39 90.64 45.01 dB 0.9963
Local x264, 2-pass medium 1 Mbps 1022 kbps 74.58 66.91 39.46 dB 0.9813
Local x264, 2-pass medium 2.5 Mbps 2542 kbps 89.84 85.11 43.19 dB 0.9936
Local x264, 2-pass medium 4.5 Mbps 4532 kbps 94.98 91.12 45.70 dB 0.9966
Local x265, 2-pass medium 1 Mbps 997 kbps 86.76 81.91 42.40 dB 0.9900
Local x265, 2-pass medium 2.5 Mbps 2491 kbps 94.43 90.75 45.01 dB 0.9950
Local x265, 2-pass medium 4.5 Mbps 4454 kbps 97.05 94.11 46.72 dB 0.9969

Delivery latency

Clip Mode Warm TTFB (median of 5) Cold: first request, full file
Sea turtle (15.2 s, 1920×1080) VBR 135 ms 12.2 s
Sea turtle (15.2 s, 1920×1080) CBR 152 ms 11.5 s
Bathroom (12.3 s, 1920×1080) VBR 111 ms 5.3 s
Bathroom (12.3 s, 1920×1080) CBR 117 ms 9.6 s
Parrot (11.8 s, 1080×1920) VBR 173 ms 7.9 s
Parrot (11.8 s, 1080×1920) CBR 148 ms 9.7 s

"Warm" is time to first byte for a rendition that already exists. "Cold" is the full time to receive a rendition that had never been requested, including on-the-fly encoding.

Reading the numbers

  • Same codec, same quality. At 2.5 Mbps, Cloudinary's constant-bitrate H.264 scored 89.08 VMAF at 2402 kbps, against 89.84 for a local two-pass x264 encode at 2542 kbps. At 4.5 Mbps the gap is 0.59 VMAF points. On-the-fly H.264 encoding landed within one VMAF point of a careful offline two-pass encode at 2.5 and 4.5 Mbps (2.2 points at 1 Mbps), while using slightly fewer bits.
  • VBR spends only what the content needs. With br_2500k as a ceiling, Cloudinary delivered an average of 1842 kbps (26% under the cap) for 87.53 VMAF. At a 4.5 Mbps ceiling it used only 2035 kbps. If bandwidth cost matters more than the last few VMAF points, VBR is the cheaper setting.
  • Codec matters more than platform. Local x265 reached 94.43 VMAF at 2.5 Mbps, 4.6 points above x264 at the same bitrate. With URL transformations, switching codec is a one-parameter change (vc_h265); with a pre-rendered ladder it's a re-encode of the library.
  • Cold starts are real. A cached rendition started arriving in 111–173 ms. A variant nobody had requested before took 5.3–12.2 s to arrive in full for these 12–15 s clips, because it's encoded on that first request. Pre-generate ("eager") the variants you know you'll serve, and keep on-the-fly transformation for the long tail. Pre-rendered ladders pay this cost once at upload instead.

An architecture point that doesn't show up in VMAF

"Decoupling image processing from video streaming forces engineering teams to maintain secondary AWS S3 and Lambda pipelines, increasing infrastructure overhead."

— Senior Media Systems Engineer

Each of those pipelines has its own storage, transformation logic, cache rules, credentials and monitoring. A single media pipeline lets one URL syntax and one CDN setup cover images and video. That's an operational argument from experience, not something the benchmark above measures.

When to pick which

  • Pre-rendered HLS ladder: you mostly serve long-form video to players that adapt their own bitrate, you want every first view to hit an existing rendition, and you don't need per-placement crops or formats.
  • Dynamic URL transformations: you serve video in many shapes (feeds, stories, product pages, email previews), your variants change with design or experiments, or you want images and video handled by the same pipeline. Pre-generate the hot variants and let the URL handle the rest.

Reproduce it

git clone https://github.com/Prajituric/video-quality-benchmark-tool
cd video-quality-benchmark-tool
python scripts/prepare_sources.py configs/default.json
python -m vqbench configs/default.json
python scripts/measure_latency.py --repeats 5 --cold
Enter fullscreen mode Exit fullscreen mode

You need FFmpeg built with libvmaf. Pull requests adding other services through the url_template provider are welcome.

Top comments (0)