DEV Community

Umasou
Umasou

Posted on

Running LaMa Inpainting and Real-ESRGAN in the Browser: Tiles, Memory Walls, and the Optimizations We Threw Away

We build ClearPix, a set of media cleanup tools that run entirely in the browser — no uploads, no server GPUs. That means when a user removes a watermark, a 198MB LaMa inpainting model runs inside their tab, and when they upscale an image, Real-ESRGAN runs next to it in the same JavaScript heap. This post is about what it actually takes to ship that: why we pin a flagship inpainting model to WASM on purpose, how tiled inference keeps large images from blowing up the tab, which "obvious" optimizations we benchmarked and threw away, and the one rule that saved us repeatedly — profile before you touch anything.

The real constraint is memory, not FLOPS

Server-side model selection is mostly about quality-per-FLOP. In the browser that framing breaks down completely. Our constraints, in order of how often they bite:

  1. Memory. Model weights, WASM linear memory, tensors, source image, and output canvas all live in one renderer process. A 4K image alone is ~33MB as RGBA. Hold a large ONNX session open at the wrong moment and Chrome kills the tab — no exception, the page just vanishes.
  2. Driver stability. A WebGPU session that creates successfully says nothing about whether session.run() will return. We've seen big models deadlock mid-inference on specific GPU/driver combinations with no error at all.
  3. FLOPS. Barely matters. Our profiling keeps showing browser inference is memory-bandwidth-bound, not compute-bound.

This is why our model choices look odd from a papers-first perspective: we run Real-ESRGAN's compact general-x4v3 variant instead of the classic RRDBNet x4plus — the "weaker" model won every acceptance case while running an order of magnitude faster, because "fast in the browser" is a property of the architecture, not the runtime. The full model-selection benchmark (and what happened when we later quantized the winner to fp16) is part 3 of this series; here we focus on what those models cost once they're loaded.

An honest inpainting model comparison: MI-GAN vs LaMa

Our watermark remover ships two inpainting engines, and we picked them on a bench suite rather than paper reputation:

MI-GAN LaMa (big-lama ONNX port)
Size ~29.5MB 198MB (fp32)
Input fixed 512x512, single 4-channel float32 tensor fixed 512x512, image + mask tensors
WASM speed per tile ~0.6s ~4.3s (about 7x slower)
Large-hole structural reconstruction weak strong — our big-hole text case gained +11.6dB over MI-GAN
Small isolated regions slightly better flat (-0.8dB on our small-logo case)

LaMa's fast Fourier convolutions give it a global receptive field, which is exactly what you need when a watermark sits on top of a face or a building edge and the fill has to reconstruct structure, not just smear texture. MI-GAN is the better tool for small, isolated regions — faster and marginally cleaner there. So the product defaults to MI-GAN and offers LaMa as an explicit "HQ" opt-in. 198MB is never prefetched; the user clicks HQ, sees the download size, and chooses.

If you want to run LaMa in a browser yourself, the tensor conventions that cost us a probe script to pin down: image is [1,3,512,512] float32 in 0–1, mask is [1,1,512,512] with 1 = inpaint, and the output is already on a 0–255 scale — the official demo casts straight to uint8, so don't normalize it again.

Why our LaMa runs on WASM on purpose

This is the decision people question most. Our MI-GAN path tries WebGPU first and falls back to WASM — and it's great when it works: after upgrading onnxruntime-web to 1.30, the WebGPU execution provider took inpainting from 2100–3500ms down to ~400ms. But LaMa at 198MB fp32 kept hitting a nastier failure class: session creation or inference would deadlock on certain GPU/driver combinations, with no error and no timeout that could rescue it. A fallback you can't trigger is not a fallback.

Meanwhile, LaMa on WASM does one tile in ~4.3s — slow, but fine for an opt-in quality mode. So we pinned it:

// initLama — reliability over peak speed: 198MB fp32 has repeatedly
// deadlocked under WebGPU (creation *and* inference), so we never try.
const s = await createSession(bytes, onProgress, ['wasm']);
Enter fullscreen mode Exit fullscreen mode

For the smaller MI-GAN we still gamble on WebGPU, but with a runtime safety net, because creation succeeding proves nothing. Every WebGPU run() is wrapped in a watchdog — 120s for MI-GAN, 300s for LaMa-class workloads (WASM gets no timeout; it's the final path and legitimately slow). On timeout or error, we rebuild the session from the cached model bytes (IndexedDB hit, no re-download), retry once on WASM, and then stick to WASM for the rest of the page lifetime instead of re-rolling the dice every pass.

One more production landmine in the same file: warming up models with a dummy forward pass hides 1–5s of shader compilation, but the warmup must be awaited before the first real inference. ORT's WebGPU EP does not tolerate concurrent run() calls on one session — we reproduced Session already started and tensor detach races three times in a row before adding that await.

Where browser ONNX actually runs out of memory

The OOMs we hit were never where we expected. Two real ones:

  • Writing the 198MB model to IndexedDB in one chunk, while a session was being created on top of it, pushed the renderer over the edge and killed the tab. Chunking the IDB write into 32MB pieces fixed it.
  • Keeping inference sessions alive during video assembly. After per-frame processing finishes, we run the encoder — and that assembly step is the memory peak of the whole pipeline. If the ORT session is still holding hundreds of MB of WASM heap at that moment, low-end devices get their renderer OOM-killed: the "page silently crashes" bug report. So we explicitly release sessions before assembly; the model bytes stay in IndexedDB, so re-init later is cheap.

The lesson generalizes: in the browser, your memory budget is a timeline, not a number. What matters is what you're holding at the peak moment.

Tiles and overlap-crop: staying under the memory wall

Both engines take fixed 512x512 inputs (inpainting) or dynamic but memory-hungry inputs (super-resolution), so everything goes through tiles.

For Real-ESRGAN we feed overlapping 256x256 tiles and paste back only the center of each upscaled tile, discarding an 8px border on interior edges — classic overlap-crop. The border pixels of a super-resolved tile are where edge artifacts live, and discarding them makes tile seams invisible. Tiling also bounds peak memory: per-tile tensors are tiny and constant no matter how big the source is, so large inputs can't crash the tab. (We still cap input at 1400px on the long side in the MVP — product guardrail, not a technical one.) The light model even let us grow tiles: going from 128px tiles with 16px margins to 256px tiles with 8px margins cut the redundant overlap pixels from about 1.78x to 1.14x of the image area.

Inpainting tiles are planned from the mask, not a fixed grid. The pipeline: connected components of the mask → group nearby components into clusters (boxes expanded by 80px that intersect are one watermark field) → per-cluster crops with context margin. Fine-detail passes use native-resolution tiles of at most 340px core, which with a 1.5x context margin stays under the 512px model input — zero downscaling, so no blurry patches. LaMa's HQ path does the same with a 340px pixel grid and 1.3x context (crops stay ≤442px).

Two tiling lessons we learned the hard way:

  • Never shatter dense regions into independent tiles. Each tile is an independent hallucination; stitching them produces visible discontinuities. Triggering tiling by cluster size once tanked one bench case from 19.7 to 11.9dB. The fix: trigger on the largest single connected component, plus a special case for long thin clusters — text lines with a union aspect ratio over 3 — because squashing a whole text line into 512x512 is exactly what produced the "you can tell it was edited" smearing. Dense near-square fields still get one global pass: one consistent hallucination beats many stitched ones.
  • Repaint glyphs, not rectangles. For text watermarks we build glyph-level masks — probability map above 0.5 as seeds, candidate pixels by color distance, keep only what's connected to a seed, dilate 2px. If a box ends up under 8% covered we fall back to the plain rectangle (that guard is what keeps faint tiled watermarks from losing recall). This cut over-painting enough to move precision from 41.4% to 43.2% with recall untouched.

Optimizations we benchmarked and killed

The most valuable section of our internal docs is the "do not retry" list. Three model-engineering entries we measured and killed: a frame-level worker pool for parallel inference (no faster than the serial path), IOBinding on the WASM path (the transfer tax it saves is a rounding error next to run() itself), and swapping x4v3 for CUGAN (a lateral move dressed up as a breakthrough by benchmarking against the old baseline). The numbers behind all three are in part 3 of this series — along with the one quantization change that did work, fp16 super-resolution weights, and its catastrophic mirror image when we tried it on LaMa. The short version: browser inference is memory-bandwidth-bound, so you win by moving fewer bytes, and whether quantization is safe is a property of the model's architecture, not of your tooling.

Every entry in that kill list exists because we measured first. The rule we now enforce: no performance work starts without a per-stage timing breakdown and an expected share of the total. If you can't say what fraction of runtime you're attacking, you're guessing.

Try it

Everything described here runs in production, free, with no uploads — the models download to your browser and nothing leaves your machine. If you came here looking for LaMa inpainting online, clearpix.org/remove-watermark is the dual-engine MI-GAN/LaMa pipeline described above. For a Real-ESRGAN web demo with the tiled overlap-crop pipeline, try our image upscaler. Video versions of the same engines power our free video watermark remover.


Part 5 of the ClearPix engineering series — how we build free, private, in-browser media tools at clearpix.org.

Top comments (0)