DEV Community

Rafay Rajput
Rafay Rajput

Posted on

PDF

Tags: pdf, javascript, performance, webdev, privacy
Title: The PDF compression floor: why "compress to exactly 200KB" usually can't be done


Most PDF tools implement generic compression: apply a preset level, report whatever comes out. That's correct for most use cases and useless when there's a hard upload cap.

I hit this building an exact-size compressor. The interesting part is the constraint, not the code.

The floor is real and it's arithmetic

Four mechanisms reduce PDF size:

Mechanism Lossy? Typical gain
Image downsampling Yes Halving dimensions = 75% fewer pixels
JPEG quality reduction Yes Large, but degrades below a floor
Font/resource deduplication No Frequently 20–40% on text documents
Object stream optimisation No Modest, free, always worth doing first

The two lossless ones should run first. On text-heavy documents they often get you most of the way with zero quality cost.

The two lossy ones interact. Downsample to hit 400KB, then reduce JPEG quality to reach 200KB, and you've made a signature unreadable.

Why it fails on real documents

Arithmetic for a 10-page scan at 300dpi:

  • 300dpi A4 ≈ 2480 × 3508 px ≈ 8.7 megapixels per page
  • At ~0.3 bytes/pixel compressed, that's roughly 2.6MB per page
  • 10 pages ≈ 26MB
  • To reach 200KB total: 20KB per page
  • At the same compression ratio: ~65,000 pixels ≈ 255 × 255

At 255×255 an A4 page, body text is ~4 pixels tall. Unreadable.

This is why "compress PDF to 200KB" is often not achievable, and why a tool that claims otherwise is either degrading the document or lying about the result.

The honest behaviour

A correct implementation needs three outcomes, not one:

  1. Target met — iterate quality/resolution until the budget is satisfied
  2. Floor reached — report how close, and say the target is unreachable, with the reason
  3. Not worth degrading — stop when legibility breaks, and explain the tradeoff

Outcome 2 is the one everyone skips and the one that actually matters to a user standing in front of a rejected upload.

Architecture note

All of this is arithmetic, and it can run entirely client-side. pdf-lib plus a canvas pass handles resampling; JPEG re-encoding is canvas.toBlob with a quality parameter.

The commercially interesting part: none of it requires an upload. The same work a server would do happens in the browser, which means a document with financial or identity details never leaves the device.

I wrote up the full method, including per-tool behaviour and the verification step: https://pdff.online/guides/compress-pdf-to-exact-size


Dev.to / Hashnode: same text works. Cross-post to both, canonical to the pdff guide.

Top comments (0)