DEV Community

Cover image for Meicut Engineering #1: Video Compression at Product Scale
Tutu
Tutu

Posted on

Meicut Engineering #1: Video Compression at Product Scale

Part of the Meicut Engineering series — how we build an in-browser media toolkit.

The encoding command — ffmpeg -i in.mp4 -c:v libx264 -crf 23 out.mp4 — is well documented. What is rarely discussed is the engineering that turns this single command into a service: where encoding runs, how long-running jobs behave, how concurrency is bounded, and how output is matched to the user's stated goal. This post covers those decisions.

1. Where encoding runs

The first architectural decision is the execution environment. Three options exist for a web media tool:

Approach Execution Cost profile Constraints
Server-side ffmpeg Encode on infrastructure Upload → encode → download Requires upload; full ffmpeg capability
ffmpeg.wasm In-browser (WebAssembly) No upload, strong privacy ≈ an order of magnitude slower than native (community benchmarks); input+output held in a WASM heap capped near 2–4GB
WebCodecs In-browser (raw codecs) No upload Most control; most implementation work

Browser-native execution (ffmpeg.wasm / WebCodecs) is attractive for privacy — files never leave the device. It is, however, constrained by encode throughput and memory ceiling: a 500MB input can exceed the WASM heap and crash the tab. Server-side execution reliably handles large files and the full filter set (fonts for subtitles, multi-file merges).

We therefore implemented server-side encoding first, deferring browser-native processing. The ordering principle — support the hardest case before the common case — recurs throughout the product.

2. Task pipeline

With encoding server-side, "compress video" becomes a pipeline:

upload → validate → queue → encode → store → downloadable link
Enter fullscreen mode Exit fullscreen mode

Three constraints follow.

Long-tail latency. A short clip encodes in seconds; a 4K 20-minute input at a slow preset runs for minutes. The pipeline must be asynchronous — each request returns a job ID, the client polls, and progress reflects pipeline state rather than a synthetic spinner. A wall-clock cap per job, with a clear path to re-encode at a faster preset, replaces silent failure.

Concurrency. If every request spawned an ffmpeg process, a burst of 4K uploads would exhaust CPU and starve smaller jobs. We bound concurrency by job size and worker pool: small jobs take a fast lane; heavy encodes run under a concurrency limit, so no single input or burst dominates. A queue with two priority tiers sufficed for our volume.

Batch as the same pipeline, reused. Compression is rarely one file at a time. Many users need to compress multiple videos at once — batch compression is the common request behind any "video compressor" tool. Under the hood it is not a separate system: a user submits many files, each becomes a job in the same queue — sharing the concurrency bound, priority tiers, and per-file validation. Progress and results are tracked per file, and a failed item does not abort the batch; it is reported alongside the successful ones. The single-job pipeline, reused across a list of inputs.

Input validation. Users upload arbitrary inputs — truncated files, unusual containers, audio mislabeled as video. ffmpeg tolerates much until it does not; a mid-encode crash is poor UX. We probe inputs first (ffprobe) and return a readable diagnostic before committing to an encode.

3. Encoding parameters

A single CRF value applied to every input optimizes for no use case. We classify each job before encoding:

  • Intended use determines the parameter set: email attachment → aggressive size target with downscaling permitted; social → balanced 1080p; archive → keep resolution at high quality.

  • Content type matters: screen recordings (static) and concert footage (motion-heavy) require different bitrate budgets under the same preset.

Output is CRF with a size-aware backstop: quality-targeted, but if the result exceeds the size implied by the use case, it is re-encoded once at a stricter CRF. Checking the outcome against the goal, rather than the parameter, catches most poor encodes.

4. Playback compatibility

Encoding is only half the task; guaranteeing playback is the other. Three flags are non-negotiable in a product:

ffmpeg -i in.mp4 -c:v libx264 -preset medium -crf 23 -pix_fmt yuv420p -movflags +faststart -c:a aac -b:a 128k -ar 44100 out.mp4
Enter fullscreen mode Exit fullscreen mode
  • -pix_fmt yuv420p — required for broad player and browser support; 4:4:4 or 10-bit input plays on a subset of devices.

  • -movflags +faststart — moves the moov atom forward so playback starts before download completes, reduces perceived load time on web.

  • Explicit audio parameters (aac 128k 44100) — avoid container defaults; inaudible or silent tracks are the most common audio support issue.

5. Privacy as a design constraint

The stated promise — files are used only for the task and auto-deleted — is implemented as a pipeline property rather than a toggle. Files are written to short-TTL storage and destroyed after the job, with a scheduled cleanup as a backstop in case a process terminates abnormally. Default behavior, not user-remembered settings, is what keeps the promise reliable.

Conclusion

Video compression as a product is the pipeline surrounding the encode: execution environment, long-job behavior, bounded concurrency, and output matched to the user's goal. The command itself is the smallest part; the value is in the decisions that make it reliable at scale.

The pipeline described here is implemented in the online video compressor — no install required; files auto-delete after processing.


Meicut Engineering series:

  • #1: Video Compression at Product Scale

  • #2: Trimming Videos Without Re-encoding (when -c copy is and isn't viable)

  • #3: "Convert Format" Is a Misnomer — containers vs codecs

  • #4: Subtitles: soft vs burned, and the font problem

Top comments (0)