I came to this from content work, not a CS degree, so when people say "WebGPU is X times faster than WASM" I used to just remember the X. Then I ran the same model files through both backends myself and got four different X values on one laptop, ranging from 1.4 to 9.4. Nothing was broken. The number is a property of the model and the input size, and after digging through the model files I think I understand most of why.
The setup
I used the models ImgIng (https://imging.ai/) downloads for its background removal and upscaling tools (same byte counts and SHA-256 as the hashes in its code), and ran them through onnxruntime-web 1.27.0 on an M4 Mac, Chromium 149, 16 GB. One benchmark page, one .onnx file, and only the execution provider changes between webgpu and wasm. Every configuration got its own fresh browser process, three repeats, and "steady state" means the median of runs 2 to 6. Four WASM threads unless I say otherwise, because that's what ImgIng picks on this machine. One note: ISNet FP16 below is not something you can choose in ImgIng's interface. Its cutout picker has three tiers, and I only used the FP16 file as a comparison against the fast tier's INT8 version of the same ISNet.
The axis is logarithmic, so compare bar lengths inside each group, not across groups. In both ISNet groups the green WebGPU bar is far shorter than the blue 4-thread bar. In the upscaler group the two are nearly the same length. The grey 1-thread bars show how much of WASM's time is just threading.
Big convolution models get the big numbers
ISNet takes a 1024×1024 input and is mostly convolutions. INT8 steady state was 359 ms on WebGPU against 2,133 ms on WASM, which is 5.9x. The FP16 file was 209 ms against 1,960 ms, or 9.4x. The upscaler, Real-ESRGAN x4v3, is a 4.9 MB model with 175 nodes and 34 convolutions. On a 184×184 tile it ran 331 ms against 485 ms (1.5x), and on a 120×120 tile 150 ms against 211 ms (1.4x). My reading is that a tiny model on a small tile doesn't give the GPU enough parallel work to make up for its fixed overhead, while a million-pixel input does. I haven't profiled per-operator, so treat that as my interpretation of the numbers, not a measurement.
Why INT8 was slower than FP16 on the GPU
This one annoyed me, because I'd assumed INT8 means faster. On WebGPU the INT8 file lost to FP16, 359 ms versus 209 ms, and its first run was 537 ms versus 258 ms. On WASM the two were basically tied (2,133 versus 1,960 ms with 4 threads, 6,313 versus 6,404 ms with one). So I parsed the INT8 file and counted operators. There are 119 DequantizeLinear nodes sitting in front of 119 Conv nodes, and no ConvInteger or QLinearConv at all. The weights are stored as int8 and turned back into float32 before each convolution runs. The quantization saves download size (44 MB instead of 88 MB), not compute. Why that costs extra on the GPU specifically, I can't say without a profile.
The other two things I'd want to know before quoting any number: WebGPU's first run is slower than its steady state (537 against 359 ms for INT8, about 50% more, which I assume is shader compilation), while WASM's first run is within 1 to 10% of steady. And WASM threads flatten out early. Going from 1 to 4 threads made ISNet 3x faster (6,313 to 2,133 ms), and 4 to 9 threads only another 1.3x (1,615 ms), probably because the M4 has four performance cores and the rest are efficiency cores.
All of this is one M4 and one Chromium build. I didn't test Windows, a discrete GPU, a phone, Safari or Firefox, and the ratios could look completely different there.
Before you quote a WebGPU multiplier, run your own model at the input size you really use, and write first run and steady state down as two separate numbers.
Top comments (0)