If your browser-based video tool feels slow, the model probably isn't your problem — the pipeline is. We run ClearPix, a set of free in-browser media tools, and our video watermark removal used to take 95 seconds end-to-end on ffmpeg.wasm. After migrating the whole pipeline to WebCodecs for hardware decode/encode and MediaBunny for muxing, the same job takes 9 seconds. This post walks through exactly why ffmpeg.wasm was slow for us, the architecture we replaced it with (with real code), the audio passthrough trick that avoids re-encoding entirely, and how a streaming single-pass design made memory O(1) so we could raise our duration limit from 30s to 120s.
Why ffmpeg.wasm was slow: three taxes we kept paying
We didn't abandon ffmpeg.wasm because it's bad software — it's remarkable that it exists at all. We abandoned it because every run paid three structural taxes that no amount of tuning removes:
- The 33MB core download. Before a single frame is touched, the user downloads and compiles a 33MB WebAssembly build of ffmpeg. On a cold cache that's the first several seconds of the session gone, and it's pure overhead — none of it is doing work the user's OS couldn't do natively.
- The MEMFS round-trip. ffmpeg.wasm operates on an in-memory virtual filesystem. Every input file gets copied in, every extracted frame gets written to it, every output gets copied back out. Our frames were making a JPEG round-trip through a fake filesystem purely so a C program from the 2000s could feel at home.
- Software encoding in wasm. The encode step ran ffmpeg's software H.264 encoder inside WebAssembly — no SIMD hardware encoder, no GPU. Meanwhile the machine sitting under the browser has a dedicated silicon block (VideoToolbox on our Macs) whose entire job is encoding H.264 fast.
The profiling rule this taught us: pipeline overhead and model overhead are the same order of magnitude, and pipeline overhead is usually cheaper to fix. JPEG round-trips, MEMFS copies, wasm software codecs, the 33MB download, materializing the whole video in memory — replacing ffmpeg.wasm removed all of them in one move, and that one move was worth roughly 10x.
The dispatcher: two backends, one escape hatch
We didn't delete the ffmpeg path. It's still the fallback for environments without hardware H.264 support, and it's deliberately feature-frozen — streaming it would cost a week of work on a path almost nobody runs, so it keeps its original behavior and its original hard limits.
A thin dispatcher picks the backend. The WebCodecs path probes for hardware encode and decode support with a representative 1080p config:
export async function isWebCodecsVideoSupported(): Promise<boolean> {
if (typeof VideoEncoder === 'undefined' || typeof VideoDecoder === 'undefined') return false;
const enc = await VideoEncoder.isConfigSupported({
codec: 'avc1.640028', width: 1920, height: 1080,
bitrate: 8_000_000, framerate: 30,
});
if (!enc.supported) return false;
const dec = await VideoDecoder.isConfigSupported({
codec: 'avc1.640028', codedWidth: 1920, codedHeight: 1080,
});
return !!dec.supported;
}
There's also a manual override — ?vbackend=ffmpeg|webcodecs in the URL or a localStorage key — which turned out to be essential for testing. Playwright's bundled Chromium ships without an H.264 codec, so our end-to-end tests force vbackend=ffmpeg; real-Chrome runs exercise the WebCodecs path. vbackend=webcodecs intentionally skips the probe and hard-routes, so real errors surface during hardware testing instead of being silently masked by a fallback.
One pleasant surprise on the WebCodecs side: engine initialization is a zero-download no-op. Where the ffmpeg path starts with a 33MB fetch, the WebCodecs "init" just reports ready immediately.
MediaBunny + hardware decode: browser video frames straight off the silicon
This is the closest thing to a WebCodecs video processing tutorial we can offer: the actual loop we ship. MediaBunny wraps the container demuxing; VideoSampleSink gives us decoded frames as an async iterator:
const input = new Input({ source: new BlobSource(file), formats: ALL_FORMATS });
const track = await input.getPrimaryVideoTrack();
if (!(await track.canDecode())) throw new Error('Frame extraction failed (unsupported codec?).');
const sink = new VideoSampleSink(track);
for await (const sample of sink.samples()) {
sample.draw(ctx, 0, 0, w, h); // decoded on hardware, drawn to canvas
sample.close();
await launchEncode(name); // snapshot to JPEG, capped at 8 in flight
}
Two details worth stealing. First, sample.draw() handles rotation and flip metadata for free — phone footage comes out upright. Second, canvas.toBlob snapshots the canvas synchronously and encodes asynchronously, so we don't await each JPEG; we let up to 8 encodes be in flight at once, which overlaps encoding with the next frame's decode. H.264's yuv420p requirement leaks into everything, so both dimensions get rounded to even numbers up front.
Hardware video encoding in JavaScript: CanvasSource, plus a 3-second trap
On the output side, MediaBunny's CanvasSource feeds a canvas into the platform's hardware H.264 encoder:
const videoSource = new CanvasSource(canvas, {
codec: 'avc',
quality: new Quality('high'), // roughly the crf 18 look of our old ffmpeg path
keyFrameInterval: 2,
});
output.addVideoTrack(videoSource);
await output.start();
// per frame: draw, then
await videoSource.add(i * frameDur, frameDur);
Each add() returns a promise that resolves when the encoder has consumed the frame — this becomes important for backpressure later.
Now the trap. The first VideoEncoder session on our machines has a ~3 second cold start (VideoToolbox initialization). If you let that hit at assembly time, the user stares at a frozen progress bar right at the moment they expect a result. Our fix is a warmup: when frame extraction begins, we fire off — without awaiting — a throwaway 64×64 two-frame encode on a BufferTarget. By the time real assembly starts, the encoder is hot. Failure is silent; worst case you just pay the cold start you were going to pay anyway.
One caveat we learned the hard way elsewhere in the stack: warmups must be awaited before the first real use of the same resource. Our model warmup once raced the first real inference against the same WebGPU session and blew up with Session already started. The encoder warmup is safe precisely because the dummy encode fully finalizes its own Output before the real one starts.
Audio without re-encoding: EncodedAudioPacketSource passthrough
Re-encoding audio that's already fine is wasted work and a quality loss. Our audio plan has three tiers, decided before the muxer starts:
-
Passthrough. If the MP4 output container supports the input's audio codec and we can get a
decoderConfig, we stream the encoded packets straight through viaEncodedAudioPacketSource. Zero decode, zero re-encode, bit-identical audio. -
Decode → AAC re-encode. If the codec isn't MP4-friendly but is decodable, we pull
AudioSampleSinkand re-encode to 128kbps AAC. -
Drop. If neither works, the audio track is silently dropped — same semantics as ffmpeg's
-map 1:a?.
The passthrough loop is almost offensively simple:
const sink = new EncodedPacketSink(audio.track);
for await (const packet of sink.packets()) {
if (packet.timestamp >= videoDur) break; // -shortest semantics
await audioSource.add(packet, first ? { decoderConfig } : undefined);
first = false;
}
The decoderConfig rides along with the first packet only; the rest are raw packets copied verbatim. Audio is truncated to the video duration, matching ffmpeg's -shortest behavior, so A/V sync holds.
Backpressure for free: the streaming single pass
The original architecture was three phases: extract all frames to an in-memory Map, process them all, then assemble. That materializes the entire video in memory twice and puts a hard ceiling on duration — it's why the old limit was 30 seconds.
The streaming rewrite (runVideoPipeline) collapses it into one pass: decode a frame, hand it to the per-frame AI callback, encode the result into the muxer, move on. No frame warehouse. The surprise was that we never had to build a backpressure queue — the primitives already compose:
-
sink.samples()is an async iterator, so decoding is pulled one frame at a time. Nothing buffers ahead unless we ask for it. -
await videoSource.add(...)is MediaBunny's encoder/writer backpressure signal. Awaiting it means the in-flight window naturally converges to 1–2 frames.
Total live state: one or two frames, the encoder, and the inference session. That's it.
We measured it: a 20-second clip peaks at 170MB versus 162MB for a 6-second clip. Memory is flat in input length — O(1) holds in practice, not just on the whiteboard. With the memory ceiling gone, the duration limit stopped being an engineering constraint and became a product decision: the watermark remover now allows 120s (soft limit with an estimated-time confirmation, 200MB file size as the only hard guardrail). Frame capping for the length guard happens inside the stream — skipped frames are just closed without ever touching a canvas. Cancellation is immediate, since the loop checks a flag between frames.
Results
| ffmpeg.wasm | WebCodecs + MediaBunny | |
|---|---|---|
| Engine download | 33MB wasm core | none |
| Decode / encode | software, in wasm | hardware (VideoToolbox etc.) |
| Frame storage | MEMFS + JPEG round-trip | 1–2 frame streaming window |
| Audio | re-encoded | packet passthrough when possible |
| End-to-end watermark removal | 95s | 9s |
| Duration limit | 30s | 120s (soft) |
| Peak memory scaling | grows with input | flat (170MB @ 20s vs 162MB @ 6s) |
The 95s→9s number is the full watermark-removal flow measured on 2026-09-29, including the AI inpainting — the pipeline swap is what made the model the fast part.
What we'd tell our past selves
- Profile the pipeline before the model. Our instinct was to optimize inference. The 10x was in transport.
-
Use the silicon. Hardware video encoding in JavaScript is one
isConfigSupportedprobe away. The browser is not a sandboxed CPU; it's a thin layer over a media engine. - Keep the escape hatch. A forced-backend URL parameter and a frozen fallback path made every regression debuggable and every rollout reversible. Fallbacks are a feature, not a TODO.
- Don't give fallback paths new features. Streaming the ffmpeg path would have cost a week to serve a rounding error of users. Freeze it and spend the week on the path 99% of sessions take.
Try it
Everything above runs in production today. Drop a clip into our free video watermark remover to see the pipeline end-to-end, or try the video upscaler, which rides the same WebCodecs + MediaBunny rails. No upload, no signup — the file never leaves your browser.
Part 2 of the ClearPix engineering series — how we build free, private, in-browser media tools at clearpix.org.
Top comments (0)