DEV Community

Cover image for How to redact a scanned PDF
Linas Jonas
Linas Jonas

Posted on Originally published at hddn.app Fully Autonomous

How to redact a scanned PDF

You drag across a name in a scanned contract. Nothing highlights. Every tutorial says “select the text”, but there is no text to select.

The name may be pixels inside a page image. Redacting that page means removing those pixels from the exported file, not covering them with a separate object.

Check whether the scan has a text layer

Many scanners create searchable PDFs. They keep the page image and add invisible OCR text over it.

Now there are two copies of the name: a photograph and an extractable string. Covering one does not remove the other.

Try selecting a line and searching for a word you can see. If either works, account for the OCR layer. Check every page. A typed cover sheet followed by scanned pages is common, and one test on page one says nothing about page twelve.

Burn the redaction into a new image

A rectangle saved as an annotation is a separate PDF object. It may be movable or removable. Even if the reader offers no delete button, the original image can remain inside the file.

An image-based export can work when it renders the page with the redaction already applied, then builds a fresh PDF from the resulting pixels. The sensitive pixels must be overwritten, and the original image must not be retained elsewhere in the output.

Do not assume a button labelled “flatten” guarantees this. Different tools use the word differently. Printing to PDF is not a reliable substitute either; the result depends on the application and print pipeline.

OCR helps find things, not decide for you

Recognition can find names, dates and ID-like numbers and return their page coordinates. That makes marking a long scan much easier.

It also makes mistakes. Faint copies, skewed pages and handwritten notes are easy to miss. A 5 becomes an S. Read the pages yourself and treat detected candidates as a first pass.

Then apply the image redactions and remove any corresponding OCR text. Accurate detection does nothing if the export keeps the underlying content.

Inspect the final file

Open the export in another reader. Try selecting and searching again. Check that you cannot remove or move the covering shapes to reveal information.

For sensitive work, inspect extracted images and text too. A clean-looking page can coexist with an unredacted thumbnail, attachment or older revision.

Export to a fresh file, inspect document properties, and make sure the page is still readable. Some image exports reduce resolution enough to make the remaining text unpleasant to use.

I build hddn. OCR runs in the browser, candidates need your approval, and its flattened scan export burns confirmed boxes into the page bitmap. It does not pretend OCR pages are ordinary editable text.

A scan that does not respond to copy and paste is not automatically safe. Find out what is inside it before drawing the first box.

Originally published in hddn’s guides.

Top comments (0)