DEV Community

Umasou
Umasou

Posted on

How We Remove Watermarks from Video Fast: Motion-Adaptive Frame Skipping + Temporal EMA

(and the bug that froze our masks — a regression story about wiring performance switches to product options)

Running an inpainting model on every frame of a video is the obvious way to do video watermark removal — and unaffordable in a browser. This is the write-up of how we got 58–67% of frames to skip the model entirely: a motion-adaptive, anchor-based frame-skipping gate, a temporal EMA for flicker-free output, plus the failures behind the design — a chain-freeze bug (PSNR 38.7 → 16.7), an anchor-age regression (38.30 → 34.69), and a UX "simplification" that silently disabled the whole optimization and cost us 3x. All numbers are from our own benchmarks; all code is from the shipping pipeline.

Per-frame video inpainting is the cost center

ClearPix removes watermarks entirely client-side: ONNX-runtime inpainting on WebGPU or WASM, fed by WebCodecs decode, encoded back with MediaBunny. No server, no upload.

The profile was blunt. Inside one inpaint call, planning cost about 4.4ms while the model run() segment took 97% or more. Per frame, inpainting took 2100–3500ms on WASM; the WebGPU execution provider brought that to roughly 400ms — still the dominant line item. Pipeline overhead was already cut elsewhere (ffmpeg.wasm → WebCodecs + MediaBunny: 95s → 9s per job). The biggest lever left was not a faster model — it was not running the model at all.

Frame skipping for video processing: reuse, don't recompute

The observation that makes skipping viable: most watermarks sit on content that barely changes — a corner logo over a talking head, a username stamp over a screen recording. Under a static mask, the pixels inside the mask bounding box are nearly identical frame to frame. If nothing in the region moved since the last inference, that output is already the correct answer — composite it onto the current frame instead of inferring again.

The per-frame gate, trimmed from RemoveVideoWatermark.tsx:

const small = makeRegionGray(img, maskBbox);   // bbox region → ~64px-wide grayscale
// Compare against the anchor (last real inference frame); drift measured
// exactly against the anchor too
const mae = st.anchorSmall ? regionMAE(small, st.anchorSmall) : Infinity;
const deltaAnchor = st.anchorRaw
  ? estimateRegionDelta(img, st.anchorRaw, maskData, maskBbox)
  : [0, 0, 0];

if (mae < st.skipThreshold && st.anchorOut) {
  // Skip: composite the anchor's feathered mask region onto the current frame
  out = compositePrevRegion(img, st.anchorOut, maskData, maskBbox, deltaAnchor);
  st.skipped++;
} else {
  out = await runInpaint();
  st.anchorRaw = img;   // a real inference becomes the new anchor
  st.anchorOut = out;
  st.anchorSmall = small;
}
Enter fullscreen mode Exit fullscreen mode

The cheap part matters: the diff runs on the mask bbox downsampled to ~64px grayscale — under 5ms per frame, against hundreds for inference. On a skip, compositePrevRegion pastes the anchor's mask region back with the inpainter's own feathering semantics (100% inside the mask, transition band on clean pixels outside), so a reused frame looks like an inferred one.

The current build skips 58% of frames on our watermark e2e clip (threshold 5.00) and 67% on the subtitle e2e (threshold 4.21); the earlier gate measured 94 of 144 (65%).

The threshold is not a magic constant. The first five frames always run full inference; we collect their MAEs and set skipThreshold = median × 3, clamped to [2, 5]/255 — dark and noisy footage get different baselines for free. (The ceiling of 5 is itself a measured fix — keep reading.)

The chain-freeze bug: PSNR 38.7 → 16.7

The first version (v1) compared each frame to its immediate neighbor and reused the previous output. It passed screen-recording tests, then failed catastrophically on a slow-gradient clip.

On gradient footage each frame drifts ~1/255, so per-frame MAE stays below threshold and the gate skips forever — the mask region freezes at the last inference while the rest of the frame changes around it: a visible, motionless rectangle. E2e masked-region PSNR collapsed from 38.7 to 16.7.

Patch one was a cumulative drift budget: accumulate skipped-frame MAE since the last inference and force re-inference when the total crosses the threshold. That recovered PSNR only to 22.9 — reused pixels still lagged the scene's exposure on gradients. Patch two added delta compensation (estimateRegionDelta): the per-channel mean difference between current and previous raw frames over the non-mask pixels in the bbox, added to historical pixels before blending. Budget + delta landed at 38.44 (vs a 38.74 no-skip reference), inside tolerance, with a clean three-clip visual check in real Chrome. That version shipped.

But the budget was always a proxy: the error we care about is "how far is this frame from the frame whose output we're reusing," and chained per-step MAEs only approximate that.

Anchored comparison: the fix that deleted the budget

So we refactored (v2): compare each frame directly against the anchor — the most recent frame that really ran inference — and reuse the anchor's output on skips. Skip error is now exactly what the threshold bounds; the drift budget was deleted outright. Delta compensation became exact too — measured against the anchor itself.

The first anchored run gave us a proper scare: subtitles gained +3–4dB on all four e2e segments, but the gradient watermark clip dropped from 38.30 to 34.69dB masked-region. Root cause: that clip's drift varies with y (a sinusoidal exposure gradient), so one mean delta can't fully compensate it, and the residual grows with anchor age. Two measured fixes:

  • Tighter threshold ceiling: 8 → 5. Anchors refresh more often, bounding how stale a reuse can get.
  • Split the deltas. Compositing reuses the anchor → to-anchor delta; the EMA blends against the previous frame's output → to-neighbor delta. One mean delta serving both timebases was the bug.

Final numbers: e2e-video masked region 39.84dB (+1.5dB over the chained baseline), 58% skipped (threshold 5.00); subtitles 37.3 / 32.3 / 28.7 / 30.1dB across four segments, 67% skipped (threshold 4.21).

Flicker-free video watermark removal with a temporal EMA

Skipping frames solves cost, but video inpainting temporal consistency needs one more layer. Even frames that do run inference can jitter: the model hallucinates slightly different textures each time, and at 30fps that reads as flicker inside the repaired region.

Our answer: a per-pixel exponential moving average, applied only inside the mask dilated by 4px, with a motion-adaptive blend factor (applyTemporalEMA, trimmed):

// Per-pixel motion from raw frames (0..1)
const m = (|dr| + |dg| + |db|) / (3 * 255);
const t = Math.min(1, m / 0.08);   // motion 0→8%
const s = t * t * (3 - 2 * t);     // smoothstep
const a = 0.4 + 0.6 * s;           // alpha 0.4→1.0
// History term: to-neighbor deltaPrev (composite uses to-anchor deltaAnchor)
out = curr * a + (prev + delta) * (1 - a);
Enter fullscreen mode Exit fullscreen mode

Low motion → alpha 0.4 (lean on history, flicker dies); high motion → alpha 1.0 (pure current frame, no ghosting), with a smoothstep between. Three boundaries are deliberate:

  • It runs only inside the dilated mask. The rest of the frame is already bit-identical to the source; touching it would only blur real motion.
  • It runs on both inferred and skipped frames, so reused composites get the same treatment as fresh inferences.
  • Its history term is compensated against the previous frame, not the anchor — the split described above.

Moving watermark removal: sampled re-detection that still skips frames

Everything so far assumes a fixed mask — wrong for a watermark that moves. Our "Detect sampling" setting has three rates: First (detect once, reuse for the whole video), Med (every 5 frames), High (every frame). Moving watermarks need Med or High.

A rebuilt mask means all temporal state belongs to the old mask position — reusing it would smear content across the frame — so a refresh resets everything: the anchor group (anchorRaw/anchorOut/anchorSmall), the prev group (prevRaw/prevOut), and the dilation cache. No special-casing needed: anchorSmall = null makes the next MAE Infinity, forcing full inference on the re-detect frame (which becomes the new anchor), and prevOut = null makes the EMA skip itself. Between re-detects the watermark usually holds still, so skipping stays fully active on Med/High — the anchor gate handles the motion. (If detection fails or finds nothing, we keep the previous mask and leave temporal state untouched.)

The same machinery powers our subtitle-removal tool, which defaults to Med because burned-in subtitles change every few seconds.

The rollback: how a UX cleanup cost us 3x

Now the embarrassing one. In a September UX pass, the "First" sampling option looked like confusing jargon, so we removed it and forced per-frame re-detection for everyone.

What we didn't trace: skipping and the EMA are keyed on the mask's bounding box, and with per-frame re-detection the mask is rebuilt every frame, so maskBbox was never stable — the entire optimization path silently disabled itself, and video removal got ~3x slower. Users reported "streaming feels slower" — wrong; comparing the complaint timeline with the commit timeline pointed at the sampling default. We rolled back the next day.

The never-again rule: before deleting a "confusing" option, check whether a performance optimization is keyed to it. Performance switches must be explicitly bound to product options — the dependency belongs in a comment, a test, or the option's name, not in someone's memory. Our detectEvery state now carries exactly that comment.

Numbers, in one place

Change Measurement Result
Skip gate cost mask-bbox frame diff <5ms per frame
v1 chain-freeze bug (single-frame MAE) e2e masked-region PSNR 38.7 → 16.7
v1 drift budget only e2e masked-region PSNR 22.9
v1 drift budget + delta (first shipped version) e2e masked-region PSNR / skip rate 38.74 → 38.44; 94/144 skipped (65%), pipeline 20s vs 170–370s WASM
v2 anchored, first run gradient clip masked-region PSNR 38.30 → 34.69 (subtitles +3–4dB on all four segments)
v2 anchored, final (ceiling 5, split deltas) e2e-video masked-region PSNR / skip 39.84dB (+1.5dB vs chained baseline), 58% skipped (thr 5.00)
v2 anchored, final e2e-subtitles, four segments 37.3 / 32.3 / 28.7 / 30.1dB, 67% skipped (thr 4.21)
Removing the "First" option (UX pass) video removal speed ~3x regression, rolled back next day

Try it

Everything here ships in our free video watermark remover — in-browser, no upload, the model runs on your GPU. If your real problem is burned-in text, the same pipeline drives our subtitle removal tool. Feedback on weird footage is how most of these fixes happened.


Part 4 of the ClearPix engineering series — how we build free, private, in-browser media tools at clearpix.org.

Top comments (0)