At ImgIng I own the model loading path: which mirror a model comes from, the hash check, the cache. So when I timed a cold start of our fastest background removal tier, I expected the 44 MB model to be the long bar. It wasn't. The 5.95 MB onnxruntime-web WASM file finished about nine seconds after the model did, and nothing could run until it arrived.
How I timed it
I used ImgIng (https://imging.ai/) with its quick AI tier (ISNet INT8), in a fresh browser context on an M4 Mac, Chromium 149, 16 GB, on the evening of 2026-09-28. I didn't click through the UI. I called the page's own TYBG.segment(img, 'quick', onProgress) directly, which is the same function the start button calls, and logged every network request plus each progress event. The input was a synthetic bottle image I generated, and its content doesn't affect timing because the model resizes everything to 1024×1024. I ran two modes: hardware acceleration on with WebGPU on Metal, and a Chromium with no usable GPU adapter, the state you get with hardware acceleration switched off.
Where the time went on a cold run
With WebGPU, the runtime JS was ready at 32 ms. The model, 44,279,201 bytes from ModelScope's CDN, finished downloading at 4,188 ms, and the SHA-256 check was done by 4,207 ms. Session setup started at 4,250 ms. That is the moment ORT requested ort-wasm-simd-threaded.asyncify.wasm (5,955,745 bytes gzipped) from the site's own origin. Download finished at about +13.2 s, inference started at 13,725 ms, and the result was back at 14,594 ms. Without a GPU adapter the shape was the same: the WebGPU attempt failed, the switch-to-WASM message appeared at 13,551 ms, and the run finished at 16,015 ms. Two things in that sequence are worth separating. The WASM file is requested only when a session is created, so it starts after the model is already downloaded, and the two never overlap. And the WebGPU run needed it too, because ORT's WebGPU execution provider is compiled into that same WASM file.
This is a repro page that logs the start and end of the two downloads on one cold load. Look at the order of the two rows. The runtime request begins only after the model row ends, and it ends later even though the file is about a seventh of the size. It's a single run, so the times on screen differ a little from the ones in the text.
Was it the file or the network?
I checked the per-file speed with curl at the same time of evening. The runtime file took 9.19 s and 10.78 s on two tries, 552 to 648 KB/s. The 44 MB model from ModelScope took 4.20 s, about 10.5 MB/s. That was an ordinary home connection that evening, and I'm not claiming it holds for other regions or other hours. A later cold run saw the runtime finish at +34.2 s with the model at +4.3 s, though that run shared the machine with other jobs, so I only take the ordering from it, not the numbers.
The heavier tier shows the same order. For the professional tier (BEN2 FP16, 219,121,675 bytes) with WebGPU, the model was done at +20.5 s, the runtime at +29.8 s, and the whole call took 34.8 s.
After the first visit none of this matters much. The runtime is served with Cache-Control: max-age=31536000, immutable and isn't requested again, and the model comes out of Cache Storage. On a warm cache, a new page finished the whole call in 930 ms on WebGPU and 2,643 ms without an adapter.
What stays with me is the asymmetry on the first visit. The model has a CDN, a backup domain and a hash check. The runtime is one request to our own origin, and it starts late. Those are the numbers as they stand on the 09-03 build.
If you ship in-browser inference, open the Network waterfall on a cold load. Don't stop at the biggest file. Check when the runtime .wasm request starts relative to the model, and how fast your own origin serves it.
Top comments (0)