Tags: pdf, javascript, performance, webdev, privacy
Title: The PDF compression floor: why "compress to exactly 200KB" usually can't be done
Most PDF tools implement generic compression: apply a preset level, report whatever comes out. That's correct for most use cases and useless when there's a hard upload cap.
I hit this building an exact-size compressor. The interesting part is the constraint, not the code.
The floor is real and it's arithmetic
Four mechanisms reduce PDF size:
| Mechanism | Lossy? | Typical gain |
|---|---|---|
| Image downsampling | Yes | Halving dimensions = 75% fewer pixels |
| JPEG quality reduction | Yes | Large, but degrades below a floor |
| Font/resource deduplication | No | Frequently 20–40% on text documents |
| Object stream optimisation | No | Modest, free, always worth doing first |
The two lossless ones should run first. On text-heavy documents they often get you most of the way with zero quality cost.
The two lossy ones interact. Downsample to hit 400KB, then reduce JPEG quality to reach 200KB, and you've made a signature unreadable.
Why it fails on real documents
Arithmetic for a 10-page scan at 300dpi:
- 300dpi A4 ≈ 2480 × 3508 px ≈ 8.7 megapixels per page
- At ~0.3 bytes/pixel compressed, that's roughly 2.6MB per page
- 10 pages ≈ 26MB
- To reach 200KB total: 20KB per page
- At the same compression ratio: ~65,000 pixels ≈ 255 × 255
At 255×255 an A4 page, body text is ~4 pixels tall. Unreadable.
This is why "compress PDF to 200KB" is often not achievable, and why a tool that claims otherwise is either degrading the document or lying about the result.
The honest behaviour
A correct implementation needs three outcomes, not one:
- Target met — iterate quality/resolution until the budget is satisfied
- Floor reached — report how close, and say the target is unreachable, with the reason
- Not worth degrading — stop when legibility breaks, and explain the tradeoff
Outcome 2 is the one everyone skips and the one that actually matters to a user standing in front of a rejected upload.
Architecture note
All of this is arithmetic, and it can run entirely client-side. pdf-lib plus a canvas pass handles resampling; JPEG re-encoding is canvas.toBlob with a quality parameter.
The commercially interesting part: none of it requires an upload. The same work a server would do happens in the browser, which means a document with financial or identity details never leaves the device.
I wrote up the full method, including per-tool behaviour and the verification step: https://pdff.online/guides/compress-pdf-to-exact-size
Dev.to / Hashnode: same text works. Cross-post to both, canonical to the pdff guide.
Top comments (0)