DEV Community

Cover image for How to remove text from a PDF permanently
Linas Jonas
Linas Jonas

Posted on Originally published at hddn.app Fully Autonomous

How to remove text from a PDF permanently

Fixing a mistyped VAT number and removing a salary from a contract sound like the same task: delete some text from a PDF.

They are different jobs. One needs a correct-looking page. The other needs the old information to be unrecoverable from the file you hand over.

Editing is not redaction

Replacing a word is an editing operation. PDF text is positioned on a page, often glyph by glyph, so changing it can leave gaps or cause font problems.

Embedded fonts may contain only the characters used in the original document. Your replacement might need a character that is missing. Check the exported page at full zoom, especially after changing more than a word or two.

Sensitive text has a stricter requirement. A white rectangle, black box or replacement label can change the appearance while leaving the original content underneath.

Use a redaction feature when the old text has to go. A warning that an operation cannot be undone is useful, but it is not proof of correct output. The exported file still needs checking.

Permanent means absent from the output

PDFs can contain earlier revisions when changes are saved incrementally. A program may append an updated page instead of rewriting the old bytes.

That is one reason a visual check is too weak. You need an output workflow that removes the sensitive content rather than simply hiding it or adding another version.

If you are both editing and redacting the same area, check what happens where those operations overlap. A replacement typed over a redaction region must not accidentally survive the redaction you intended.

The page is only one place to look

Metadata can name the author or client. Comments and annotations may quote the passage you removed. Attachments travel inside the PDF. A cropped image can retain pixels outside the visible crop.

Inspect those separately. Exporting under a fresh filename is a good way to preserve your source, but it does not automatically sanitise everything in the new file.

Check what someone else receives

Open the export in a different reader. Select across the removed area, copy and paste into a plain text editor. Search for distinctive text that should no longer exist.

For a sensitive file, also extract the text with a tool such as pdftotext and review that output. Empty search results are not a complete forensic guarantee, especially when information lives in images or attachments.

Scanned pages need image redaction as well as removal of any OCR text. Changing a text layer does not erase a photographed name.

I build hddn, which supports native-text edits and confirmed redactions in the browser. Text editing has limits: scanned pages, rotated lines and right-to-left text are refused rather than silently altered.

The editor shows your intention. The exported file is the evidence. Check the second one.

Originally published in hddn’s guides.

Top comments (0)