DEV Community

Convertilo
Convertilo

Posted on AI-assisted

TensorFlow.js in the browser: why one new tensor shape cost 8-17 seconds, and how I cut a 40 s freeze

I run a photo enhancer that does super-resolution fully in the browser: TensorFlow.js on WebGL, nothing is uploaded. The first version used the upscaler library (UpscalerJS). On a 1.2 megapixel photo it took 60-110 seconds and froze the page for about 40 seconds. A user told me "it does not accept any of my photos", which was partly a bad 0.6 MP limit on my side and partly the freeze. Here is what I found while fixing it. Most of it applies to any browser-side ML on WebGL.

1. Every new tensor shape recompiles the shaders

tfjs compiles WebGL programs per operation and per tensor shape. On my machine (an M2 Max) one new shape cost 8-17 seconds of shader compilation. The killer was the edge tiles: I cut the image into tiles, and the last row and column had a different size, so the model was compiled again for them.

What fixed it:

  • Make every tile exactly the same shape. Pad the image once to a whole number of tiles instead of handling a smaller remainder.
  • Set WEBGL_USE_SHAPES_UNIFORMS=true so shapes are passed as uniforms and programs can be reused across shapes.

2. Compilation is synchronous and blocks the page

Even with the same shape, the first compile blocks the main thread. tfjs 4.11 has a way to compile ahead of time without running the model: use ENGINE_COMPILE_ONLY, then backend.checkCompileCompletionAsync() and getUniformLocations(). That moves the wait out of a frozen page and lets you show a progress state.

3. im2col convolutions can eat your video memory

The default convolution path uses im2col. For a 5x5 kernel with 64 channels on a 280x280 tile it needed about 500 MB of GPU memory in my case. Setting WEBGL_CONV_IM2COL=false gave the same speed with a peak of roughly 100-200 MB. That is the difference between working on a phone and crashing the tab.

4. Smaller things that mattered

  • tfjs keeps freed textures around. Set WEBGL_DELETE_TEXTURE_THRESHOLD so the pool does not grow without limit.
  • Read each result tile with await tf.browser.toPixels(...) and draw it on a canvas straight away. Do not stitch tensors together and do not go through base64.
  • upscaler.js never gives control back to the browser without awaitNextFrame, and with it it gets even slower. I dropped the wrapper and load the model directly with tf.loadLayersModel.

Numbers after the fix

Input Time
1 MP photo 4-7 s
12 MP photo (downscaled to 4 MP input, 16 MP output) 16-24 s

The 16 MP output is the ceiling because of the canvas size limit on iOS. I have not tested low-end devices yet, so treat these numbers as "a recent laptop".

If you want to try the result, the photo enhancer is live (the interface is in Russian, but the upload and result are self-explanatory). Has anyone else measured shader compile time on other GPUs?

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

The edge-tile shape change is a useful thing to isolate. For the cross-GPU comparison, could you report cold compilation separately from a second run of the same shape, with the tfjs version and browser/GPU listed? I'd also include repeated images in one tab, then a cancel-and-retry case: a fast first result can still hide retained textures or an incomplete cleanup path. Your caveat about untested low-end devices is helpful.