We run super-resolution and watermark removal entirely in the browser at ClearPix, using ONNX Runtime Web on WebGPU. This post is the writeup of one optimization round that produced our single biggest inference speedup — 4.8x at page level from fp16 quantization — and, in the same week, two of our worst quality regressions when we applied the exact same trick to the wrong model.
If you're fighting slow browser GPU inference with onnxruntime-web, you'll get three things here: the performance model that explains why fp16 works and parallelism doesn't, the real numbers behind choosing SRVGGNetCompact over the "stronger" RRDBNet, and a concrete failure case study of fp16 quantization quality loss in mask-gated models (with PSNR before/after). At the end: the decision order we now follow before touching any model.
Why browser GPU inference is slow: bandwidth, not FLOPs
The mental model most engineers bring from CUDA land is wrong for the browser. On a desktop GPU with a fat model, you're often compute-bound and the answer to slowness is "more parallelism." In the browser, running through onnxruntime-web's WebGPU execution provider, our profiling kept saying the same thing: we are memory-bandwidth bound.
Two pieces of evidence forced this conclusion on us.
First, worker pools did nothing. We built a frame-level worker pool for parallel video inference — multiple workers, each with its own ONNX session, processing frames concurrently. The result: 396 seconds versus 393 seconds for the serial path. That's within noise. Adding workers only added contention for the same memory bus. If more parallelism can't help, FLOPs were never the bottleneck.
Second, IOBinding did nothing. We profiled the WASM path expecting the JS↔wasm tensor copies to be a meaningful tax, and planned to adopt IOBinding to shave it. The instrumented timings showed the transfer tax was 1–2% of the total; the run call itself was over 97%. Optimizing anything outside the run segment was rearranging deck chairs.
Once you accept "bandwidth-bound," the optimization list writes itself: move fewer bytes. Halve the weight precision and you halve the bytes read per inference. That's the entire theory behind what happened next.
Architecture first: SRVGGNetCompact (x4v3) vs RRDBNet x4plus
Before quantization, though, we made a bigger decision that fp16 later rode on top of: swapping the model architecture itself.
We started with Real-ESRGAN x4plus, the well-known RRDBNet — 1187 nodes, 67MB of weights. On paper it's the stronger model. In the browser it was a disaster in a specific, instructive way: on ORT's WebGPU EP it showed no speedup at all over CPU. Roughly 2.5 seconds per tile on both backends. A big RRDBNet pushes so many bytes through the graph that even the GPU just waits on memory. Runtime-side tuning space was, for practical purposes, exhausted.
So we switched to realesr-general-x4v3, which is SRVGGNetCompact — a compact architecture designed for general and video content. The numbers from our September benchmarks:
| RRDBNet x4plus | SRVGGNetCompact x4v3 | |
|---|---|---|
| Graph size | 1187 nodes | 71 nodes |
| Weight file | 67MB | 4.9MB |
| CPU inference speed | baseline | 14–27x faster |
| Quality (3-case PSNR head-to-head) | baseline | wins every case, +0.1 to +1.5dB |
The quality result deserves emphasis: the compact model didn't just hold even, it won all three benchmark cases, by up to 1.5dB. The heavier RRDBNet's texture hallucinations actually hurt fidelity on real content. Per-frame, the swap took us from about 75 seconds to about 7 seconds — 10x — before we touched precision at all.
Lesson one of model engineering for the browser: the first selection criterion is "designed for the deployment environment," not "SOTA in the paper." If you're shipping a Real-ESRGAN in browser code, picking the compact variant matters more than anything you can do to the runtime.
What fp16 actually bought us on onnxruntime-web WebGPU
With x4v3 in place, we quantized the weights to fp16 and measured carefully before shipping:
- fp16 vs fp32 outputs compared against each other: PSNR 57dB+. For practical purposes the outputs are identical.
- Single-tile inference on WebGPU: 403ms → 113ms. That's 3.5x — better than the naive 2x you'd predict from halving bandwidth, because Apple GPUs also get an fp16 ALU throughput bonus.
- Page level on our video upscaler (24 frames of 720p): 31.4s → 6.6s. 4.8x end-to-end.
- The fp16 file is 2.4MB versus 4.9MB for fp32, which also halves the download.
The implementation detail that kept this painless: we exported with keep_io_types, so the JavaScript side keeps its fp32 tensor contract and zero application code changed. The WASM EP also runs the fp16 model fine, so our Chromium e2e path needed no special-casing.
We still ship fp32 as a fallback spec. If every fp16 URL fails, the loader falls through to the original model; the feature keeps working, just at fp32 speed. This is a habit we'd recommend to anyone shipping models to browsers: every production incident we've had got closed by a fallback chain, not by the primary path.
export const UPSCALE_MODEL: ModelSpec = {
key: 'realesr_general_x4v3_fp16_v1',
urls: [
`${STATIC_BASE}/models/realesr-general-x4v3-fp16.onnx`,
`${BASE}models/realesr-general-x4v3-fp16.onnx`,
],
sha256: '7e7b7f58...',
expectedSize: 2444975,
fallback: {
key: 'realesr_general_x4v3_v1',
urls: [/* R2, local, HF mirrors for the fp32 original */],
expectedSize: 4866417,
},
};
So fp16 was a free 4.8x. Naturally, emboldened, we reached for the same hammer on the other model in our stack. That's where the week went sideways.
When fp16 quantization quality loss is real: the LaMa failure
Our watermark removal pipeline uses LaMa, an inpainting model. Inpainting is a different beast from super-resolution in one structurally important way: LaMa is mask-gated. The mask doesn't just mark where to work — it gates the network's behavior. The model has to learn "trust these pixels, hallucinate those."
We quantized LaMa to fp16. The benchmark came back: masked-region PSNR 21.1 → 11.9dB. That's not a subtle regression — the watermark was left almost fully intact in the output. We tried int8 quantization too, both static and weight-only, because maybe it was an fp16-specific quirk. It wasn't: int8 gave 24.5 → 15.6dB on our dark-text case, with the text watermark surviving completely.
The mechanism, once we thought about it, is almost obvious. A super-resolution network has no gating structure; its job is a smooth, local, well-conditioned transform, and half-precision noise mostly washes out. A mask-gated inpainter has a decision boundary inside the network: which regions are trustworthy context and which must be regenerated from scratch. Numeric precision degradation pushes the model across that boundary — the gating fails softly, and the model starts "copying" the input through instead of repairing it. The watermarked pixels in the masked region get treated as content to preserve rather than a hole to fill. Output looks "fine" at a glance, except the watermark is still there.
Both failures were caught by the benchmark before shipping, and both directions are now on our internal do-not-retry list. The takeaway isn't "never quantize." It's that quantization feasibility is a property of the model's architecture: does it have precision-sensitive structure — mask gating, large-magnitude outputs, long error-accumulation chains? If yes, fp16/int8 will almost certainly kill it. If no, fp16 is free bandwidth.
The decision order we now follow
After this round, we wrote down the order of operations for any model change, and it's deliberately architecture-first:
- Architecture. Pick a model designed for the deployment environment. Lightweight and purpose-built beats paper SOTA. Our x4plus → x4v3 swap (10x, quality up) was worth more than every other inference optimization combined.
- Bit width. Only after the architecture is fixed, ask whether the model has precision-sensitive internals. No gating, no long accumulation chains — fp16 is on the table, and it's the cheapest 3–5x you'll ever get. Gating present — walk away.
- Runtime. Last, and with measured expectations. Worker pools, IOBinding, EP tuning: in our profiling these were the 1–2% layer. Profile first, write down the expected share of the total before touching anything, and be ready to be told "0x."
And one meta-rule that saved us repeatedly: quality admission is always the benchmark, never the paper's numbers and never "it looks fine." The LaMa fp16 output looked plausible. The PSNR said 11.9dB. Benchmarks saw it; eyes didn't.
Try it
Everything above runs in production, client-side, with no uploads — your media never leaves the tab. You can feel the fp16 x4v3 path yourself in our image upscaler or the video upscaler (the 6.6-second page from this post). The LaMa side — kept safely at fp32 — powers our free video watermark remover and the image tools on clearpix.org.
Part 3 of the ClearPix engineering series — how we build free, private, in-browser media tools at clearpix.org.
Top comments (0)