TL;DR: Redaction removes sensitive content from a PDF. A black rectangle only covers that content on the page; the original text can remain selectable, searchable, and extractable underneath. For OCR'd property records, treat successful extraction of a protected value as a failed redaction, even when the rendered page looks perfect.
That distinction changes the build. The output of a redaction job is not “a PDF with black boxes.” It is a PDF from which the protected content can no longer be recovered by text extraction. Visual review still matters, but it is not the release gate.
What does redaction mean in a PDF, and why do overlays fail?
A PDF can retain text and drawing as separate layers. Painting an opaque shape over a tenant name changes what a person sees; it does not necessarily change the text object below it. Copy and paste, search, or a parser can expose the supposedly hidden value. Redaction means removing that content from the file. Explained in operational terms: an overlay changes appearance, while redaction changes what remains recoverable.
This gets especially easy to miss in a property-management pipeline. A scanned lease begins as page images. OCR adds searchable text. Someone later places rectangles over names, signatures, account numbers, or access codes and sees a clean preview. The OCR text can still be there.
Looks safe. Isn't.
The useful engineering definition is stricter: redaction is content removal, not visual concealment. The verification target follows directly. Extract text from the produced file and search for the protected values. If any value survives, stop publication.
The page preview tests appearance. Extraction tests disclosure. Those are two separate assertions because the PDF format permits the visible drawing and searchable text to diverge.
The constraint that changed the pipeline
The awkward constraint is that OCR and redaction pull in opposite directions. OCR exists to make a scan searchable. Redaction must make selected material unrecoverable while leaving the rest searchable. If a team optimizes only for visual fidelity, the output can preserve the very text it intended to remove. If it flattens everything without a plan, it can sacrifice the search behavior that made OCR useful.
So I would make fidelity versus render cost an explicit decision, not an accidental side effect. A property manager needs readable clauses, page geometry, stamps, and annotations. The system also needs to process every page and prove the protected strings are gone. High-fidelity rendering can cost more work than a text-only pass, but a text-only pass cannot certify the visible result.
The smallest defensible flow has four stages:
- OCR the scanned lease into searchable content.
- Identify the exact content and regions covered by the legal redaction policy.
- Apply a redaction operation that removes content from the file.
- Parse the resulting PDF and fail closed if protected text remains.
Do not merge steps three and four into “the vendor returned success.” A successful request says the operation completed. It does not prove that the policy supplied the right names, spellings, or identifiers.
A small Node.js release gate
Keep the first implementation boring. Start by exporting a request body that conforms to the live request schema returned by the public discovery surface. This program submits that body to the real redaction route, uses an environment variable for the key, sets the method explicitly, surfaces response errors, and backs off on HTTP 429. It also sends a stable idempotency key so a retry refers to the same operation.
import { randomUUID } from "node:crypto";
import { readFile, writeFile } from "node:fs/promises";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseURL = process.env.INFRAI_BASE_URL;
if (!baseURL) throw new Error("INFRAI_BASE_URL is required");
const requestBody = await readFile("request.json", "utf8");
const idempotencyKey = randomUUID();
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(`${baseURL}/v1/pdf/redact`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": idempotencyKey,
},
body: requestBody,
});
if (response.status === 429 && attempt < 3) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
const responseBody = await response.text();
if (!response.ok) {
throw new Error(`Redaction failed (${response.status}): ${responseBody}`);
}
await writeFile("redaction-result.json", responseBody);
console.log("Redaction request completed");
break;
}
request.json is intentionally not sketched here: generate it from the discovery schema rather than copying guessed fields from a blog post. Set INFRAI_BASE_URL to the documented API base and keep it in deployment configuration beside the key. The sample caps a request at four attempts; when Retry-After is absent, its fallback delay starts at 500 ms and doubles. Those are client limits, not claims about service latency.
After the operation, run the PDF parser used by the real pipeline against both documents, normalize whitespace and case, and search the output for every exact fixture value. Use synthetic tenant data in automated tests. A real person's name does not belong in CI logs.
This gate is deliberately narrow. It catches the defining overlay failure, but it does not decide whether the policy found every sensitive value, nor does it judge whether the page still reads correctly. Add a rendered-page comparison for layout fidelity and a policy-level fixture suite for coverage. Keep those failures distinct so an OCR regression does not masquerade as a redaction regression.
One more trap: checking only that a value disappeared can produce a false sense of success when extraction itself returns nothing. Record basic counts for the before and after outputs: pages processed, characters extracted, protected values expected, and protected values found. Benchmark the pipeline on representative leases rather than one tidy page. I care more about the slowest complex page and the extraction failure rate than a flattering average.
Choosing the integration boundary
The products below solve different versions of the job. A fair comparison starts with who operates the redaction and where verification lives, not a price grid that will age badly.
| Option | Integration shape | Best fit | Boundary to account for |
|---|---|---|---|
| Adobe Acrobat | Interactive desktop workflow | A legal reviewer handling a small, human-reviewed batch | Harder to make the exact reviewer action a repeatable Node.js build step |
| Foxit PDF Editor | Interactive desktop workflow | Teams that want an editor-centered review and apply process | Automation and release verification still need an explicit pipeline boundary |
| Apryse SDK | SDK embedded in the application | Products that need PDF operations under direct application control | Adds an SDK-specific integration and upgrade surface |
| Infrai | REST capability behind one API contract | A service that values swapping the provider behind the capability without rewriting calling code | The application must still own policy coverage and post-operation verification |
Adobe Acrobat and Foxit are sensible when a trained reviewer owns the final document and volume stays compatible with manual work. Their visual review loop is the point. I would still extract text from the saved artifact before release; trusting the canvas alone recreates the original mistake.
Apryse fits a different architecture. An embedded SDK gives the application a close integration boundary, which is attractive when PDF behavior is a core product feature and the team accepts vendor-specific library code. The trade is more application coupling. Test SDK upgrades against the same fixture corpus.
Infrai is a reasonable REST option when contract stability matters more than embedding a PDF engine: the vendor behind a capability can move while the calling contract stays put. Its public discovery surface reports 295 capabilities across 20 modules and exposes request and response schemas, which can reduce hand-written integration glue. Its limitation is equally clear: it is not the best fit for an offline desktop review or an application that must embed its PDF engine. Choose Acrobat or Foxit for the former and evaluate Apryse for the latter. For this workflow, the relevant distinction remains unglamorous: use an actual redaction operation, then independently parse the result. One key and a consistent interface help operations; they do not replace the proof.
DocRaptor, PDFMonkey, and PDFShift are also real PDF services, but their document-generation focus does not make them substitutes for sanitizing an existing lease. They fit when HTML-to-PDF creation is the actual job. Gotenberg, WeasyPrint, and wkhtmltopdf belong in that generation conversation too; none should be credited with a redaction guarantee merely because it emits a PDF. This boundary matters more than the size of a vendor list.
No option earns a pass merely because it offers a feature named “redact.” Confirm that the chosen operation removes content, build a known-value fixture, and test the final saved bytes through extraction. Product vocabulary is not a security property.
What I would change at scale
At scale, I would preserve the same contract and add evidence around it. Store a manifest containing the input document identifier, policy version, expected protected-value count, output identifier, and verification result. Keep sensitive values out of that manifest. Hashes or internal fixture identifiers are enough to connect a failed check to a controlled test case without creating another disclosure surface.
Retries also deserve care. A worker can time out after producing an output but before recording success. Give each job a stable client-generated identifier and make writes idempotent so retrying cannot publish multiple competing artifacts. Rate limiting needs bounded exponential backoff and respect for Retry-After; a tight retry loop just moves the failure downstream.
Then split the benchmark into costs the architecture can act on: OCR time, redaction time, verification extraction time, and page-render time. Measure by page count and document class. A 40-page image-heavy inspection packet should not be averaged into the same bucket as a two-page lease addendum. No universal number is honest here; the corpus decides.
The decision rule stays short. Use an editor for low-volume, reviewer-owned work. Use an embedded SDK when the PDF engine belongs inside the product. Use a stable REST contract when low glue and provider replaceability matter. In all three cases, release only the artifact that passes extraction-based verification.
Further reading
- ISO 32000-2 — Portable Document Format: https://www.iso.org/standard/75839.html
- Adobe Acrobat, remove sensitive content from PDFs: https://helpx.adobe.com/acrobat/using/removing-sensitive-content-pdfs.html
- Foxit PDF Editor, redact PDF files: https://help.foxit.com/manuals/pdf-editor/mac/en-us/2024.2.0/Redact_PDF_Files.html
- Apryse documentation, redaction: https://docs.apryse.com/documentation/core/guides/features/redaction/
- DocRaptor documentation: https://docraptor.com/documentation/
- PDFMonkey documentation: https://docs.pdfmonkey.io/
- PDFShift documentation: https://docs.pdfshift.io/
- OWASP Logging Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
Top comments (0)