Series: Building PdfWord — a free, no-backend PDF tools site (Part 9)
I shipped a basic PDF compare tool a while back: two PDFs, side by side, scroll and squint. It worked — but "scroll and squint" isn't a feature, it's an apology.
So I rebuilt it with three real view modes. The star is diff highlights: differences glow red, identical content fades, and every page gets a similarity score. Here's how the pixel-diff engine works — all in the browser, no uploads.
Try it: Compare PDFs
Mode 1: Side-by-side (the honest baseline)
Both PDFs render page-by-page with pdf.js at identical scale, paired vertically: page 1 of A next to page 1 of B, then page 2, and so on. Scrolling stays naturally synchronized because the pairs are simply stacked in the DOM — no scroll-sync JavaScript needed. The boring solution wins again.
If the page counts differ, you get an amber warning ("PDF A has 5 pages, PDF B has 7 pages") and unpaired pages show a "missing" placeholder instead of silently misaligning everything downstream.
Mode 2: Overlay (for the "something moved" feeling)
Both pages are superimposed in a single pane with an opacity slider for version B. Nudge the slider and shifted text ghosts between the two positions — your eye catches movement instantly. This is the mode for "did the layout change?" questions. A 2mm margin shift is invisible side-by-side; in overlay mode it shimmers.
Mode 3: Diff highlights (the superpower)
This is the one I'm proud of. The engine:
- Renders both pages to canvas at the same scale with pdf.js
- Walks every pixel, comparing RGB channels
- Flags pixels where any channel differs by more than 32
- Paints a visualization: differences in red, identical content faded
for (let i = 0; i < dataA.length; i += 4) {
const dr = Math.abs(dataA[i] - dataB[i]);
const dg = Math.abs(dataA[i + 1] - dataB[i + 1]);
const db = Math.abs(dataA[i + 2] - dataB[i + 2]);
if (Math.max(dr, dg, db) > 32) {
diffCount++; // mark it — paint red in the diff canvas
}
}
const similarity = 100 * (1 - diffCount / totalPixels);
The threshold of 32 is the entire trick. Render the same PDF twice and the pixels won't be identical — anti-aliasing produces ±10–20 of noise per channel. Without a threshold, every page reports thousands of phantom differences. With 32, real changes (an edited word, a moved logo) pop while rendering noise disappears.
Each page pair gets a similarity badge from that score — green for near-identical, amber for edited, red for heavily changed. And a "Download Diff Report" button assembles all the highlighted pages into a PDF via pdf-lib, so you can send someone proof of what changed.
The honest limitations
- It compares the first 20 pages (a speed cap, stated plainly in the UI). Pixel-walking is O(n) per page and browsers aren't patient.
- Scanned PDFs from different scanners can show ~99% "noise" diffs — different scan, different pixels, even when a human sees the same document. I say this in the guide instead of pretending the score is magic.
- It compares rendering, not text. Reword a paragraph with identical layout? Caught — pixels changed. Change the font but keep the words? Also caught. It's dumb and thorough, which is exactly what you want from a diff.
Why this matters beyond PDFs
The pattern generalizes: render two things to canvas, diff the pixels, visualize. I'm already thinking about an image-compare tool on the same engine. The pdf.js rendering is the PDF-specific part; everything after getImageData is universal.
Try it: Compare PDFs — upload any two PDFs and flip through the three modes.
What would you diff? Contracts, resumes, design revisions? I'm curious what people actually compare.
Top comments (0)