Our image worker runs in containers with a hard memory limit, and someone upstream wants to switch the product hero images to progressive JPEG. Before signing off I wanted one number I didn't have: how much more memory does a progressive file need while it's being decoded? The answer on my machine was about 14 MiB more for a 2736×1824 photo, and it lines up almost exactly with one buffer you can calculate on paper.
The setup is small. Two images I found online: a landscape photo from Wikimedia Commons scaled to 2736×1824, and a holiday notice template from a Chinese design-template site, 1242×2688. Each was encoded at quality 85 with Pillow 11.3 (libjpeg-turbo underneath), once as baseline and once as progressive. Same pixels in, same quality, only the scan layout differs. The machine is an Apple M4 with 16 GB, macOS 26.5.2.
To measure, I decode each file in a fresh Python process and look at how far the peak resident size moved:
CHILD = '''
import resource, sys
from PIL import Image
before = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss
img = Image.open(sys.argv[1]); img.load()
after = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss
print(before, after)
'''
A new process per file matters. Peak RSS never goes down, so decoding two files in one process would hide the second one. On macOS ru_maxrss is in bytes (on Linux it's kilobytes). I ran each file five times and took the median.
| File | Baseline | Progressive | Extra |
|---|---|---|---|
| landscape 2736×1824 | +20.6 MiB | +35.0 MiB | 14.4 MiB |
| template 1242×2688 | +14.3 MiB | +24.0 MiB | 9.7 MiB |
The baseline numbers are roughly the decoded RGB bitmap plus some overhead. 2736×1824×3 bytes is 14.3 MiB, and the process grew by 20.6.
The extra part is the interesting bit. A baseline JPEG is stored block by block. The decoder reads one block's 64 DCT coefficients, runs the inverse transform, writes pixels, and forgets the coefficients. It only needs a few rows of blocks in flight. A progressive file is split into several scans (Pillow's default script produces 10): first the DC value of every block, then bands of AC coefficients, then refinement passes that add the low bits. No block is complete until the last scan arrives, so the decoder has to keep every coefficient of the whole image around until the end.
You can size that buffer by hand. Both files use 4:2:0 chroma subsampling, and libjpeg stores each coefficient as 2 bytes. For the landscape image that's 2736×1824 luma coefficients plus two chroma planes of 1368×912, times 2 bytes: 14.3 MiB. For the template, with the width padded up to a multiple of 16, it comes to 9.6 MiB. The measured extras were 14.4 and 9.7. I didn't instrument libjpeg itself, so treat this as a match rather than proof, but it's close enough that I stopped looking for another explanation.
So for capacity planning the rule I wrote down is: a progressive JPEG costs roughly one more bitmap-sized buffer at decode time, on top of the output image. In our case the worker occasionally handles several large images in parallel, and that doubles the worst case. That's the part that would have hit the container limit, not the decode time (which also went up, about 2.3 to 2.6 times in Chromium in a separate test, still just milliseconds per image).
One thing that made this easier: files exported in the browser don't need this budget at all. I compress most of my front-end images with ImgIng (https://imging.ai/), and its JPG exports in the Chromium 149 open-source build I tested are baseline (SOF0), because it uses the browser's native encoder. In that build, canvas.toBlob also returns baseline JPEGs. Progressive files only enter our pipeline through server-side scripts.
I haven't measured browser memory for this, only Pillow. If you want to check your own worker, the snippet above plus a loop over five runs is all it takes. Run it on the largest image your service accepts, once as baseline and once as progressive, and compare the difference with width × height × 2 × 1.5 bytes for a 4:2:0 file.
Top comments (0)