DEV Community

Umasou
Umasou

Posted on

Automatic Watermark Detection in the Browser: YOLO + OCR, and the Glyph-Level Masks That Made It Work

ClearPix removes watermarks entirely client-side — no upload, no server GPU. This post is the detection half of that story: how we get a browser to find watermarks on its own using a YOLO watermark detector in ONNX plus a PP-OCR text detector, why we show low-confidence boxes instead of silently dropping them, and how a seed-and-grow glyph mask pushed detection precision from 41.4% to 43.2% without losing a single text case. Every number below comes from our 12-case internal bench and twelve rounds of logged tuning, including the rounds that failed.

Detection is the ceiling on removal

The uncomfortable truth we learned early: inpainting quality is bounded by mask quality. A state-of-the-art inpainting model given a sloppy mask produces a confident, high-resolution smear. We have bench cases where a heuristic box over-selected a large clean region and the full-image PSNR collapsed to 8.3 dB — not because the inpainter was bad, but because it faithfully repainted a huge area that never needed repainting. In another acceptance test, a false-positive box over a clean star-trail photo made the inpainter hallucinate a ghost logo onto empty sky. Precision failures are user-visible in the worst possible way: they add artifacts.

Recall failures are just as bad, and quieter: a missed corner of a watermark survives removal and reads as "the tool doesn't work."

So we treat detection as its own product surface, with its own budget of engineering rounds. The stack that survived is a dual-engine design with arbitration, plus a glyph-level masking stage for text.

A YOLO watermark detector in ONNX, running in your tab

The first engine is a purpose-trained YOLO11 watermark detection model, exported to ONNX (~11 MB) and executed with onnxruntime-web. Its job is logos and graphic watermarks — the semi-transparent corner badges and channel marks that have no text in them.

The browser-side plumbing is ordinary YOLO fare, which is exactly why ONNX is nice here: letterbox the longest edge to 640, pad to 640×640 with neutral gray, normalize to 0–1, run one inference, decode the [cx, cy, w, h, score] rows, threshold at 0.35 confidence, NMS at 0.5 IoU, then map the boxes back through the letterbox into original-image coordinates. Everything runs on an OffscreenCanvas so detection never blocks painting.

One runtime detail mattered more than any model detail. On onnxruntime-web 1.20 the YOLO session failed with "No usable ONNX backend" and the inpainter silently fell back to WASM. Upgrading to 1.30 with providers ordered [webgpu, wasm] turned WebGPU on for real: inpaint time dropped from 2100–3500 ms to roughly 400 ms per tile, and single- vs multi-thread timings converged because the GPU doesn't care about your CPU thread count. For low-end devices without WebGPU we keep a load-shedding path that caps the detector input's longest side (our maxSide option downsamples before detection and scales the boxes back up).

YOLO is an opt-in "deep" tier because of the 11 MB download. When it fires, it also arbitrates: heuristic boxes are only kept if they overlap a YOLO detection by more than 15%. That single rule took the bench from recall 0.886 / precision 0.414 to recall 0.927 / precision 0.480 — both directions at once, which almost never happens. The tiled-light case jumped from 0.54 to 0.89 recall.

Detecting text watermarks automatically with PP-OCR

The second engine answers "where is text?" rather than "what does it say?" We use the PP-OCRv5 mobile DB text-detection model (~4.7 MB ONNX). It's script-agnostic — it emits a probability map of text-likeness, which we binarize at 0.3, morphologically close, and connected-component into boxes. Text watermarks, captions, and burned-in subtitles are its territory, and per our per-source instrumentation the OCR detector is genuinely accurate on its own.

Why two engines instead of one? Because each covers the other's blind spot. YOLO has no opinion about a faint diagonal text tile; OCR has no opinion about a solid logo badge with no glyphs. Alongside them we keep two cheap pixel heuristics (a light-overlay detector for semi-transparent white marks, a dark-badge detector for solid dark logo blocks) as the high-recall, high-noise safety net.

The tuning rule that survived twelve rounds is wide in, strict out: keep candidate generation loose, enforce quality at the box level with validators (fill ratio, internal edge density, mean saturation, an 8%-of-image area cap). Round 3 proved the reverse is fatal — tightening generation-side thresholds knocked recall from 0.912 to 0.852, and tiled/faint watermarks die first every time you do it.

We also learned which direction arbitration is allowed to flow. Gating noisy heuristics by the high-precision YOLO source worked. The mirror image — suppressing heuristic boxes that overlapped OCR boxes by more than 50% — was a disaster: one OCR false positive took a real neighboring watermark down with it (text-br-small recall 1.0 → 0.595). High-confidence sources may veto noisy ones. Noisy or merely-different sources may never veto high-recall ones. We reverted that round within a day and logged it as a do-not-retry.

Glyph-level masks: seed and grow, don't repaint the rectangle

Here is the problem with using detection boxes directly as inpainting masks for text: a bounding box around small glyphs is mostly background. Even a perfect text box has precision of only 0.5–0.9 in pixel terms, because rectangle pixels outnumber glyph pixels. Repainting the whole rectangle means asking the inpainter to hallucinate texture over a large area that was fine — and on complex backgrounds (city texture behind a faint centered watermark), that hallucinated texture is exactly the "it left a visible patch" complaint we got from real users. Our measurements: leaving the watermark alone scored MAE 3.11 against the clean original; full-rectangle inpainting scored 24.79. Eight times worse than doing nothing.

The naive fix — threshold the OCR probability map and use that as the mask — also failed, and it's worth understanding why. The DB detection head is trained against a Vatti shrink kernel, so its high-probability region is deliberately smaller than the true stroke. A pure threshold mask (prob > 0.15, fixed dilation) covered only about 50% of the glyphs; recall and PSNR both collapsed.

What shipped instead is a seed-and-grow glyph mask, built per text box from the cached probability map:

  • Seeds: pixels with prob > 0.5 — small but trustworthy.
  • Candidates: pixels inside the box whose color distance from the box's surrounding ring-band median exceeds 30 (the ring band, the box expanded by 3 px, is our local estimate of "what the background here looks like").
  • Growth: flood-fill from the seeds through candidate pixels only. Background noise without a seed never enters the mask.
  • Finish: dilate by 2 px to catch anti-aliased stroke edges.

And one escape hatch: if the grown mask covers less than 8% of its box, the glyph path has failed (this happens with very faint watermarks, whose color distance never clears 30) and that box falls back to a plain rectangle. That fallback alone keeps our tiled-light case above the recall gate.

Result: precision 41.4% → 43.2%, with recall held at 100% across every text case. Two smaller changes rode along: text-line clusters (union aspect ratio over 3, long side over 384 px) now get refined in native-resolution tiles of at most 340 px instead of being squeezed into one 512² crop — squeezing was a visible source of soft blur over text lines — and the mask feather radius adapts to region size, 6–22 px at long-side/40, so large removals blend their seams more gradually.

Low-confidence boxes are shown, not silently dropped

Detection will never be perfect, so the last design decision is about who absorbs the residual error. Our answer: the user, but with good defaults and full visibility.

In the image tool, boxes from the model sources (YOLO, OCR) are painted pre-selected. Heuristic boxes — the noisy ones — are shown as unselected suggestions and only when no model source fired at all, so a stray heuristic guess can't be swept into the mask by one click and repainted onto clean background. In the video tool, scrubbing the timeline shows the sampled re-detection as blue dashed preview boxes, so you can verify what the detector will do at that moment before committing. Nothing is discarded silently; the cost of a false positive is a click, not an artifact baked into your export.

Subtitle detection in video is a moving target

Still images get one detection pass. Burned-in subtitles are a different beast: the text changes every few seconds, so a static mask from frame one is useless by frame two hundred.

Our subtitle mode therefore reconfigures the whole pipeline. It drops YOLO and the heuristics entirely and runs text detection only — subtitles are text by definition, and every extra detector is pure false-positive risk. It caps the detection input at 640 px because per-segment re-detection has to be cheap. It filters boxes to the lower half of the frame (box center below 50% of frame height), which is where burned-in subtitles live, then expands each kept box with margin — 6% of width and 30% of height, floored at 4 px — because subtitle strokes carry outlines and shadows that extend past the tight box. And it re-detects on a sampling cadence: every 5 frames by default. (For static watermarks we instead detect once on the first frame and reuse the mask with frame-skipping — measured about 3× faster than per-frame re-detection — but sampling is the only correct strategy for a target that moves and changes content.)

Try it

Everything described here runs in your browser tab today, free and without sign-up:


Part 6 of the ClearPix engineering series — how we build free, private, in-browser media tools at clearpix.org.

Top comments (0)