Most "redact PDF" tools draw a black rectangle on top of the text and save. The text is still
there. Select it, copy it, or open the file in any PDF library, and it comes straight back.
Courts, journalists and HR teams have all leaked documents this way.
I build Drible, a free set of file tools where most of the work happens in
the browser. When I added redaction to the PDF editor, I wanted it
to be real redaction, done without uploading the file. Here is how it works and what I learned.
Why the black box fails
A PDF page is a content stream: a list of drawing operators like "move here, show these glyphs,
fill this path". Drawing a black rectangle just appends one more operator. The glyphs underneath
are untouched, so:
- text extraction (and Ctrl+A, Ctrl+C) still returns them
- search still finds them
- the original objects can still be read from the file even if nothing on the page draws them
So removing the visible text is not enough. You have to remove every object that could still
hold it.
The approach: rebuild the page from pixels
Rewriting a content stream to delete exactly the glyphs under a box is hard in general (text can
be split across operators, use custom encodings, sit inside form XObjects, and so on). For a
browser tool I chose a simpler guarantee: a page with redactions is replaced by an image of
itself, with the boxes burned in.
- Render the original page with PDF.js to a canvas at about 200 DPI (longest side capped at 5000 px so phones don't run out of memory).
- Fill each redaction box on the canvas in solid black.
- Encode the canvas as a JPEG and embed it on a fresh page of the same size with pdf-lib.
- Remove the old page.
The new page contains only an image, so there is nothing to select, search or extract under the
boxes. The trade-off is that the rest of that page also becomes an image. Pages without
redactions are left untouched, so their text stays selectable.
The part people miss: the old page is still in the file
Removing a page in pdf-lib takes it out of the page tree, but its objects (content streams,
fonts, images, annotations) can still be sitting in the file. A PDF reader won't show them,
but anyone who opens the file in a text editor or a PDF library can. Two more steps close that:
Scrub the old page. Before it is dropped, its content, resources and annotations are deleted
from the page dictionary. Form fields whose widgets lived on that page are removed too, because a
field's value can hold the very text you redacted.
Garbage-collect unreachable objects. Walk the object graph from the trailer (the document
catalog and the info dictionary), mark everything reachable, and delete everything else:
/** Delete every indirect object not reachable from the trailer. */
function collectGarbage(pdf: PDFDocument): void {
const ctx = pdf.context;
const seen = new Set<string>();
const stack = [ctx.trailerInfo.Root, ctx.trailerInfo.Info].filter(Boolean);
while (stack.length) {
const o = stack.pop();
if (o instanceof PDFRef) {
const key = o.toString();
if (seen.has(key)) continue;
seen.add(key);
const target = ctx.lookup(o);
if (target) stack.push(target);
} else if (o instanceof PDFDict) {
for (const [, v] of o.entries()) stack.push(v);
} else if (o instanceof PDFArray) {
for (const v of o.asArray()) stack.push(v);
} else if (o instanceof PDFStream) {
stack.push(o.dict);
}
}
for (const [ref] of ctx.enumerateIndirectObjects()) {
if (!seen.has(ref.toString())) ctx.delete(ref);
}
}
pdf-lib then writes a clean file without the orphans. A quick way to test any redaction tool:
redact a unique word, save, then run pdftotext out.pdf - | grep yourword or open the file in a
text editor and search for it. If it shows up, the redaction did not work.
Everything else stays real PDF
Redaction is the only feature that turns a page into an image. Everything else the editor adds is
written into the original file as real PDF content:
- text is real text in a standard PDF font, so it stays selectable and searchable
- shapes and freehand ink are vector paths
- images and signatures are embedded once and reused
Because the original file is modified rather than re-created, its existing text, links and form
fields survive.
Why do it in the browser at all
A file you redact is, by definition, a file with something sensitive in it. Uploading it to a
server to remove the sensitive part is backwards. Doing it client-side means the PDF is opened,
edited and saved on the user's device; the server only ever sends the JavaScript.
48 of Drible's 53 tools work like this. The 5 that need desktop software (PDF compression with
Ghostscript and the Office conversions) do upload, say so on the page, and delete the file as
soon as the result is sent back. I wrote up which popular PDF sites upload your files and how to
check if you want the comparison.
If you try the editor and manage to recover redacted text from a
file it saved, I'd really like to hear about it.
Top comments (0)