DEV Community

Shahzaib
Shahzaib

Posted on

PDF Compare Got 3 Superpowers: Side-by-Side, Overlay & Red Diff Highlights

Series: Building PdfWord — a free, no-backend PDF tools site (Part 9)

I shipped a basic PDF compare tool a while back: two PDFs, side by side, scroll and squint. It worked — but "scroll and squint" isn't a feature, it's an apology.

So I rebuilt it with three real view modes. The star is diff highlights: differences glow red, identical content fades, and every page gets a similarity score. Here's how the pixel-diff engine works — all in the browser, no uploads.

Try it: Compare PDFs


Mode 1: Side-by-side (the honest baseline)

Both PDFs render page-by-page with pdf.js at identical scale, paired vertically: page 1 of A next to page 1 of B, then page 2, and so on. Scrolling stays naturally synchronized because the pairs are simply stacked in the DOM — no scroll-sync JavaScript needed. The boring solution wins again.

If the page counts differ, you get an amber warning ("PDF A has 5 pages, PDF B has 7 pages") and unpaired pages show a "missing" placeholder instead of silently misaligning everything downstream.

Mode 2: Overlay (for the "something moved" feeling)

Both pages are superimposed in a single pane with an opacity slider for version B. Nudge the slider and shifted text ghosts between the two positions — your eye catches movement instantly. This is the mode for "did the layout change?" questions. A 2mm margin shift is invisible side-by-side; in overlay mode it shimmers.

Mode 3: Diff highlights (the superpower)

This is the one I'm proud of. The engine:

  1. Renders both pages to canvas at the same scale with pdf.js
  2. Walks every pixel, comparing RGB channels
  3. Flags pixels where any channel differs by more than 32
  4. Paints a visualization: differences in red, identical content faded
for (let i = 0; i < dataA.length; i += 4) {
  const dr = Math.abs(dataA[i]     - dataB[i]);
  const dg = Math.abs(dataA[i + 1] - dataB[i + 1]);
  const db = Math.abs(dataA[i + 2] - dataB[i + 2]);
  if (Math.max(dr, dg, db) > 32) {
    diffCount++;            // mark it — paint red in the diff canvas
  }
}
const similarity = 100 * (1 - diffCount / totalPixels);
Enter fullscreen mode Exit fullscreen mode

The threshold of 32 is the entire trick. Render the same PDF twice and the pixels won't be identical — anti-aliasing produces ±10–20 of noise per channel. Without a threshold, every page reports thousands of phantom differences. With 32, real changes (an edited word, a moved logo) pop while rendering noise disappears.

Each page pair gets a similarity badge from that score — green for near-identical, amber for edited, red for heavily changed. And a "Download Diff Report" button assembles all the highlighted pages into a PDF via pdf-lib, so you can send someone proof of what changed.

The honest limitations

  • It compares the first 20 pages (a speed cap, stated plainly in the UI). Pixel-walking is O(n) per page and browsers aren't patient.
  • Scanned PDFs from different scanners can show ~99% "noise" diffs — different scan, different pixels, even when a human sees the same document. I say this in the guide instead of pretending the score is magic.
  • It compares rendering, not text. Reword a paragraph with identical layout? Caught — pixels changed. Change the font but keep the words? Also caught. It's dumb and thorough, which is exactly what you want from a diff.

Why this matters beyond PDFs

The pattern generalizes: render two things to canvas, diff the pixels, visualize. I'm already thinking about an image-compare tool on the same engine. The pdf.js rendering is the PDF-specific part; everything after getImageData is universal.

Try it: Compare PDFs — upload any two PDFs and flip through the three modes.

What would you diff? Contracts, resumes, design revisions? I'm curious what people actually compare.

Top comments (0)