DEV Community

Mubashir Ali
Mubashir Ali

Posted on Fully Autonomous

"Compressing a PDF to a Target Size in the Browser", a developer write-up with code from the site's own compressor

an exact target file size (e.g., < 2 MB) using WebAssembly, PDF.js, and Canvas?In this write-up, we’ll explore how browser-based PDF compression works, why reaching a specific target size requires a dynamic binary search loop, and how to build a client-side compressor step by step.Target Keywords & SEO StrategyPrimary Keywords: In-browser PDF compression, compress PDF client-side, PDF.js compress PDF to target size, pdf-lib browser compression.Secondary Keywords: Client-side document processing, JavaScript canvas PDF compression, web worker PDF processing.The Architecture: How In-Browser PDF Compression WorksA PDF is essentially a container. While it can contain vector graphics and text streams, scanned documents and image-heavy PDFs make up over 80% of oversized files. The most effective way to shrink a PDF in the browser without altering text layout is to re-encode its raster images at lower resolutions and higher JPEG compression ratios.Here is the high-level workflow:[ Input PDF File ]
│
▼
[ PDF.js / pdf-lib ] ──► Extract & Render Pages to HTML5 Canvas
│
▼
[ Dynamic Compression Loop (Binary Search) ]
├── Adjust Image Quality (0.1 – 1.0)
├── Adjust Scale / DPI (e.g., 1.0x, 0.75x)
└── Render Page Frame to JPEG Blob
│
▼
[ Assemble New PDF with pdf-lib ]
│
▼
[ Final Size <= Target Size? ] ──YES──► [ Output Compressed PDF ]
│
NO (Refine Quality / Downscale)
Step 1: Loading PDF Pages onto HTML5 CanvasWe use PDF.js to render each PDF page into a hidden standard element. From the canvas, we can export JPEG image buffers with variable quality levels.JavaScriptimport * as pdfjsLib from 'pdfjs-dist';

// Set up worker
pdfjsLib.GlobalWorkerOptions.workerSrc = '//mozilla.github.io/pdf.js/build/pdf.worker.mjs';

async function renderPageToCanvas(pdfDoc, pageNum, scale = 1.0) {
const page = await pdfDoc.getPage(pageNum);
const viewport = page.getViewport({ scale });

const canvas = document.createElement('canvas');
const context = canvas.getContext('2d');
canvas.height = viewport.height;
canvas.width = viewport.width;

await page.render({
canvasContext: context,
viewport: viewport,
}).promise;

return canvas;
}
Step 2: The Target-Size Search AlgorithmStandard compressors ask the user for a static setting (e.g., "Low", "Medium", "High"). But if a user specifically requests "Make this PDF under 2 MB", a static setting might produce 2.4 MB (too big) or 400 KB (unnecessarily degraded quality).To hit a target size accurately, we implement a Binary Search Algorithm over the image quality range [0.1, 0.95]:JavaScriptimport { PDFDocument } from 'pdf-lib';

async function compressPDFToTargetSize(fileBuffer, targetSizeBytes, maxPasses = 5) {
const pdfDoc = await pdfjsLib.getDocument({ data: fileBuffer }).promise;
const pageCount = pdfDoc.numPages;

let minQuality = 0.1;
let maxQuality = 0.95;
let bestBuffer = null;
let scale = 1.0;

for (let pass = 0; pass < maxPasses; pass++) {
const currentQuality = (minQuality + maxQuality) / 2;
const newPdf = await PDFDocument.create();

for (let i = 1; i <= pageCount; i++) {
  const canvas = await renderPageToCanvas(pdfDoc, i, scale);

  // Export canvas frame as JPEG image with target quality factor
  const jpegDataUrl = canvas.toDataURL('image/jpeg', currentQuality);
  const jpegImage = await newPdf.embedJpg(jpegDataUrl);

  const page = newPdf.addPage([canvas.width, canvas.height]);
  page.drawImage(jpegImage, {
    x: 0,
    y: 0,
    width: canvas.width,
    height: canvas.height,
  });
}

const outputPdfBytes = await newPdf.save();
const currentSize = outputPdfBytes.byteLength;

console.log(`Pass ${pass + 1}: Quality=${currentQuality.toFixed(2)}, Size=${(currentSize / 1024 / 1024).toFixed(2)} MB`);

if (currentSize <= targetSizeBytes) {
  bestBuffer = outputPdfBytes;
  // Try to get higher visual quality if room allows
  minQuality = currentQuality; 
} else {
  // Too large, decrease quality threshold
  maxQuality = currentQuality;
}

// Edge case: If quality is already low and file is still too big, downscale dimensions
if (maxQuality < 0.25 && currentSize > targetSizeBytes) {
  scale *= 0.8;
  minQuality = 0.1;
  maxQuality = 0.8;
}
Enter fullscreen mode Exit fullscreen mode

}

return bestBuffer || (await newPdf.save());
}
Key Performance & Memory OptimizationsProcessing high-resolution PDFs entirely in the browser main thread can lead to UI freezes or out-of-memory errors on mobile devices. Here are key performance strategies:1. Offload Heavy Work to Web WorkersRun PDF rendering and page embedding inside a Web Worker or offscreen canvas context. This keeps the UI thread responsive and smooth at 60 FPS.2. Clean Up DOM ObjectsAlways release canvas objects and image memory explicitly to prevent memory leaks during multi-pass compression loops:JavaScriptcanvas.width = 0;
canvas.height = 0;

  1. Handle Hybrid PDFs GracefullyIf a PDF contains crisp vector text mixed with photos, re-rasterizing every page turns sharp text into bitmaps. A hybrid approach checks stream objects:For image streams: Downsample images directly using pdf-lib.For scanned pages: Render full page frames to .Benefits of Client-Side PDF CompressionFeatureServer-Side ProcessingClient-Side (Browser) ProcessingPrivacy & SecurityFiles uploaded to remote servers100% private; zero bytes leave deviceInfrastructure CostsHigh CPU & bandwidth expenses$0 server footprintSpeedNetwork transfer bottleneckInstant local processing speedOffline SupportRequires continuous internetWorks offline via Progressive Web App (PWA)ConclusionBy combining PDF.js, pdf-lib, and an adaptive binary search algorithm, web applications can achieve exact-target PDF compression completely on the client side. This pattern delivers fast, privacy-preserving performance while reducing backend costs to zero.

Top comments (0)