Most "free online PDF compressor" sites work the same way: you upload the file, a server shrinks it, and you download the result. When I built PDF Perfect, I wanted the opposite. The file should never leave the user's device, and you should be able to prove that by switching your Wi-Fi off halfway through.
Merging, splitting and rotating in the browser are fairly well-trodden: pdf-lib copies pages between documents without re-encoding anything. Compression was the interesting part, so this post covers how it works, the trade-offs I accepted, and what I'd do differently.
Why compression is hard in the browser
Desktop tools like Ghostscript compress a PDF by walking its object tree: they downsample embedded images, subset fonts and strip unused objects. That needs a full PDF rewriter with image codecs, and there's no small, well-maintained JavaScript library that does all of it.
What the browser does have is an excellent PDF renderer (Mozilla's pdf.js, the one inside Firefox) and a fast JPEG encoder (canvas.convertToBlob). So the approach is:
- Render each page to a canvas at a chosen scale.
- Encode the canvas as a JPEG at a chosen quality.
- Build a new PDF where each page is just that image.
It's a rasterising compressor. For scans and photo-heavy PDFs, which are the files people actually need to shrink, it works very well. For text documents it has a real cost, which I'll come back to.
Keeping it off the main thread
Rendering 200 pages blocks the main thread for a long time, so it all runs in a Web Worker. Workers don't have a DOM, but they do have OffscreenCanvas, and pdf.js can render into one.
The core loop, simplified from the real worker:
const pdf = await pdfjsLib.getDocument({ data: arrayBuffer }).promise;
const outDoc = await PDFDocument.create(); // pdf-lib
const canvas = new OffscreenCanvas(1, 1);
const ctx = canvas.getContext("2d");
for (let i = 1; i <= pdf.numPages; i++) {
const page = await pdf.getPage(i);
const viewport = page.getViewport({ scale });
canvas.width = viewport.width;
canvas.height = viewport.height;
await page.render({ canvas, viewport }).promise;
const jpeg = await canvas
.convertToBlob({ type: "image/jpeg", quality })
.then((b) => b.arrayBuffer());
const image = await outDoc.embedJpg(jpeg);
outDoc
.addPage([image.width, image.height])
.drawImage(image, { x: 0, y: 0, width: image.width, height: image.height });
postMessage({ type: "progress", progress: Math.round((i / pdf.numPages) * 100) });
}
const bytes = await outDoc.save();
A few details that matter more than they look:
-
One canvas, reused. The worker creates a single
OffscreenCanvasand resizes it for each page, rather than allocating a new canvas per page, which keeps memory steadier on long documents. -
Transfer, don't copy. The result goes back to the main thread as an
ArrayBufferin the transfer list, so a 300 MB output isn't duplicated. - A timeout. The main thread gives the worker 5 minutes. A tab that silently spins forever on a huge file is worse than a clear error.
-
The pdf.js worker inside my worker. pdf.js wants its own worker script. Importing it with Vite's
?urlsuffix and settingGlobalWorkerOptions.workerSrcworks in both dev and production builds, but the server has to send.mjsfiles astext/javascript. My Apache host didn't, and the result was a very confusing production-only failure.
Quality levels and "compress to 1 MB"
The three presets are just (JPEG quality, render scale) pairs:
| Preset | JPEG quality | Scale |
|---|---|---|
| High Quality | 0.82 | 1.0 |
| Balanced | 0.6 | 0.85 |
| Maximum | 0.4 | 0.75 |
Users often don't care about presets, though. They have a job portal that says "max 2 MB". So there's a Target File Size mode, which walks a ladder of five settings from best to smallest and returns the first result that fits:
for (const settings of ladder) {
const result = await attempt(settings);
if (result.compressedSize <= targetBytes) return result;
smallest = !smallest || result.compressedSize < smallest.compressedSize ? result : smallest;
}
return smallest; // couldn't hit the target; show how close we got
A binary search over quality would need fewer passes. But with only five meaningful steps and output size that doesn't change smoothly with quality, the linear ladder was simpler to reason about and gives the user the best-looking file that fits.
The trade-off I had to be honest about
Rasterising means text in the compressed file is no longer text. You can't select it, search it or have a screen reader read it, and at scale 1.0 very small print looks soft when you zoom in. For a scanned form that's irrelevant. For a 40-page contract exported from Word, it's a real loss, and the file can even get bigger, because vector text is tiny.
So the UI shows before and after sizes up front, the docs say plainly what happens, and the guides recommend re-exporting text documents from the source app instead. I'd rather lose a user to the right tool than hand them a worse file.
Proving the file never leaves
"We don't upload your files" is easy to claim, so I wanted it to be easy to check:
- Offline test: load a tool, turn off Wi-Fi and process a file. It still works. (For compression, pdf.js is lazy-loaded, so run it once while online first.)
- Network tab: process a file with DevTools open. You'll see the page, fonts and analytics, and no request with a body anywhere near your file's size.
The How it works page walks through both tests.
What I'd do next
- A structural compressor for text PDFs: downsample only the embedded images with pdf-lib and leave the text alone. It's harder, but it's the right answer for Word exports.
- A streaming approach for very large files, since the whole document is currently held in memory.
If you want to try it, the compressor is here. I'd love feedback from anyone who's fought pdf.js in a worker too.
I'm the developer of PDF Perfect.
Top comments (0)