I run two small tool sites and pay for bandwidth by the gigabyte, so "the AI runs in the browser" always makes me ask the same thing: whose bill do the bytes land on? I've been thinking about adding background removal to one of the sites. Before writing anything I wanted the real first-visit number, not the model size from a README, because those turned out to be two different things.
What does a first visit download?
I used ImgIng (https://imging.ai/) as the reference, with its quick AI cutout tier (ISNet INT8), in a fresh browser profile on an M4 Mac, Chromium 149, and I logged every request. The timing came from calling the page's own TYBG.segment function directly, the same function its start button calls, on a synthetic image I generated. Before the first inference could run, two big files came down. The model was 44,279,201 bytes over the wire. The onnxruntime-web WASM file was 5,955,745 bytes gzipped (24,254,953 unpacked). That's about 50 MB, and it is the same 50 MB whether the page ends up on WebGPU or on WASM, because ORT's WebGPU provider lives inside that same WASM file. I didn't add up the ORT JavaScript on top, so read "about 50 MB" as a floor.
The chart lays out one cold run with WebGPU and one in a Chromium with no usable GPU adapter, which is what you get with hardware acceleration turned off. The blue block is the 44 MB model and the orange block is the 5.95 MB runtime download plus session setup. The orange block is the longer one in both rows, and the small red slice in the second row is where the WebGPU attempt fails and it switches to WASM.
Who serves those bytes?
This is the part that changes the math. The 44 MB model didn't come from the site. It came from ModelScope's CDN (cdn-lfs-cn-1.modelscope.cn). Only the 5.95 MB runtime came from the site's own origin. The code does list fallbacks: ModelScope's CDN first, then a ModelScope backup domain, then a copy under the site's own models/ path. So the site pays for the 44 MB only when both ModelScope sources fail, and in my runs they never did. Whatever arrives gets checked against a byte count and a SHA-256 before it's used, then stored in Cache Storage. So the phrase "no server cost" really means the bytes moved to someone else's CDN. That split is a choice the site made. Running the model in the browser doesn't give it to you automatically. If I self-host the model on my own box, that 44 MB is mine on every first visit.
The visitor pays too, in time and data. That cold run took 14.6 seconds on WebGPU and 16.0 seconds without an adapter. On the second visit it gets cheap: the model is read from Cache Storage and the runtime is served with Cache-Control: max-age=31536000, immutable, so neither is requested again. A new page on a warm cache finished in 0.93 s on WebGPU and 2.64 s without. And the image never left the browser. Across six calls there were zero non-GET requests, no fetch, XHR or beacon uploads.
The bigger tier is another size class. ImgIng's professional tier (BEN2 FP16) is 219,121,675 bytes, and its cold run took 34.8 s with WebGPU. The code only offers the quick tier on mobile. I haven't measured anything on a phone, so I won't guess what 219 MB feels like on one.
What I still don't know is how long a browser keeps a 44 MB Cache Storage entry when disk space gets tight. Every eviction turns a returning visitor back into a first visitor, and I have no numbers for that.
For my own site the plan is simple. Keep the model on a CDN that isn't billed to me, keep the runtime on my origin with an immutable cache header, and load both only when someone actually asks for a cutout. Before you add in-browser AI to your site, open DevTools with the cache disabled, run one inference, and sort the Network panel by size. Then look at the domain column, because that tells you whose bill each row lands on.
Top comments (0)