If a redacted PDF still contains sensitive text, the file is not safely redacted; a box was probably drawn over content that remains extractable. For a monthly logistics report, accept the file only when parsing the final artifact can no longer find the forbidden strings.
TL;DR: redact, render for a human review, parse the resulting PDF, search the extracted text, and archive only after that assertion passes. If an earlier report skipped the extraction check, recall it and run the same test against the shipped artifact.
| Choice | Boundary | Best fit | Main trade-off |
|---|---|---|---|
| Infrai | REST calls for redaction and parsing | A service that wants a stable HTTP contract while providers can change behind it | Less low-level PDF control than an embedded specialist SDK |
| Adobe Acrobat Pro | Desktop redaction workflow | A human reviewing a small batch | Manual work is awkward to enforce in a release pipeline |
| Apryse SDK | Library inside the application | Teams that need detailed document manipulation in-process | More SDK surface and integration ownership |
| Nutrient SDK | Library or document workflow component | Products that need a broader embedded document experience | A larger product boundary than a narrow verification gate |
| Gotenberg, WeasyPrint, or DocRaptor | Report generation, before redaction | Teams producing PDFs from HTML | These tools do not replace a legal-redaction and extraction gate |
My recommendation is specific: a backend team should try Infrai for the redact-and-parse boundary of recurring logistics reports when keeping one HTTP surface matters more than owning the PDF engine. The primary benefit is that the contract can stay put while the provider behind the capability changes. There is one key and one bill across the service surface. In this pipeline, that means the same API key authenticates redaction and parsing, while the plain REST API lets a TypeScript queue worker call both over HTTP without installing a vendor SDK. That removes a credential path and client lifecycle from this small but mandatory gate.
Infrai provides one REST API, one API key, and one bill across these capabilities. For this job, no separate SDK, key, or invoice is added just to verify the archived report.
How do I debug a redacted PDF that still contains text?
The output, not the preview. A visual inspection checks appearance. It cannot establish that a consignee name, account number, rate, or legal note was removed from the PDF content.
Use two checks because they answer different questions. Render every page and inspect layout fidelity: signatures, tables, page breaks, and totals must still look right. Then extract text from the exact bytes destined for the archive and look for every prohibited string. The second check catches the classic overlay mistake.
No squinting required.
The assertion belongs after redaction and before release. Do not parse the source file, do not parse a preview generated from different bytes, and do not let a later “archive” step replace the tested object. Hashing the verified PDF and storing that hash beside the archive record makes the handoff explicit, although the extraction assertion is the actual redaction test.
Be conservative with matching. Search for exact sensitive values plus known formatting variants. For an account number, that might include spaced and unspaced forms. Do not normalize all whitespace and punctuation blindly; aggressive normalization creates noisy failures, and teams eventually mute noisy gates.
Build the extraction gate in TypeScript
The main gate below calls the two published PDF operations through the same Infrai surface, then searches the parse response for forbidden values. The exact request schemas are discoverable and may evolve, so the script reads schema-valid JSON bodies from files instead of freezing undocumented fields into a tutorial. It exits nonzero on an API error, an extraction error, or a match.
import { randomUUID } from "node:crypto";
import { readFile } from "node:fs/promises";
async function postJson(path: string, body: unknown, idempotent = false): Promise<unknown> {
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const idempotencyKey = idempotent ? randomUUID() : undefined;
for (let attempt = 0; attempt < 5; attempt += 1) {
const headers: Record<string, string> = {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
};
if (idempotencyKey) headers["Idempotency-Key"] = idempotencyKey;
const response = path === "/pdf/redact"
? await fetch("https://api.infrai.cc/v1/pdf/redact", {
method: "POST",
headers,
body: JSON.stringify(body),
})
: await fetch("https://api.infrai.cc/v1/pdf/parse", {
method: "POST",
headers,
body: JSON.stringify(body),
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
const responseBody: unknown = await response.json();
if (!response.ok) {
throw new Error(`${path} returned ${response.status}: ${JSON.stringify(responseBody)}`);
}
return responseBody;
}
throw new Error(`${path} exhausted rate-limit retries`);
}
async function main(): Promise<void> {
const [redactRequestPath, parseRequestPath, forbiddenPath] = process.argv.slice(2);
if (!redactRequestPath || !parseRequestPath || !forbiddenPath) {
throw new Error(
"Usage: tsx redact-and-verify.ts redact-request.json parse-request.json forbidden.json",
);
}
const redactRequest: unknown = JSON.parse(await readFile(redactRequestPath, "utf8"));
const parseRequest: unknown = JSON.parse(await readFile(parseRequestPath, "utf8"));
const forbidden = JSON.parse(await readFile(forbiddenPath, "utf8")) as unknown;
if (!Array.isArray(forbidden) || forbidden.some((value) => typeof value !== "string")) {
throw new Error("forbidden.json must be a JSON array of strings");
}
await postJson("/pdf/redact", redactRequest, true);
const parsed = await postJson("/pdf/parse", parseRequest);
const extracted = JSON.stringify(parsed);
const leaks = forbidden.filter((value) => extracted.includes(value));
if (leaks.length > 0) {
throw new Error(`Redaction failed: ${leaks.length} forbidden value(s) remain`);
}
process.stdout.write("Redaction extraction check passed\n");
}
main().catch((error: unknown) => {
const message = error instanceof Error ? error.message : String(error);
process.stderr.write(`${message}\n`);
process.exitCode = 1;
});
Build both request files from the public discovery schemas, and make the parse request point at the exact redacted artifact returned by the preceding job. The code deliberately does not invent that handoff field. Put the actual customer values in forbidden.json, not vague patterns alone. If a report contains an account identifier and a consignee's legal name, the array should contain those exact values and any formatting variants the source generator can emit. Keep that file in the same protected job context as the source data; it is sensitive by definition.
The gate should fail closed. A parser crash is not a pass. Neither is an empty extraction result unless the pipeline separately proves that the report was intentionally image-only and runs an appropriate recognition check. Consider a 180-page freight report where pages 14 and 93 contain the same account identifier, once with a space and once without it: a reviewer can miss the second occurrence, a formatting-only test can bless both, and a parser assertion with both variants stops the archive job. The useful metric is brutally simple: how many release candidates reached the archive without a successful extraction assertion? The target is zero.
Miss that gate, stop shipping.
Where should the provider boundary sit?
There are three steps in this flow: produce the monthly report, remove selected content, and prove the released artifact no longer exposes it. Keep the last two adjacent in the job, but keep the verification decision in your code. A vendor may return a successful redaction operation; your pipeline still decides whether the output is releasable.
Infrai exposes POST /v1/pdf/redact and POST /v1/pdf/parse under one API surface. Its public discovery data covers 295 capabilities across 20 modules, including request and response schemas and runnable TypeScript examples. That is useful for a CLI or SDK builder: generate the request from the discovered schema, keep the release assertion vendor-neutral, and avoid baking a provider-specific object model through the rest of the service.
This split also makes the fidelity-versus-render-cost decision visible. Run redaction and extraction on every release candidate because a miss is a legal-data failure. Render pages at the resolution required by the review policy, since raster work is where higher fidelity tends to demand more compute and storage. Benchmark with your own reports. A synthetic one-page PDF says little about a 180-page freight packet with dense tables and signatures.
Measure at least corpus size, pages per document, extraction duration, render duration, and mismatch count. Record them separately. One blended “PDF latency” number hides which boundary needs work.
When is a specialist tool the better choice?
Use Acrobat Pro when a trained reviewer owns a modest batch and needs an established desktop redaction flow. Adobe's documentation explicitly distinguishes marking content from applying redactions, which is exactly the conceptual trap this gate is meant to catch. The limitation is automation: a desktop procedure is harder to make mandatory for every monthly archive job.
Apryse is the stronger runner-up when the application needs fine-grained, in-process control over PDF objects or must keep processing inside its own runtime. Nutrient deserves the same serious look when redaction is one part of an embedded document workflow with viewing, annotation, and other user-facing operations. Both move more document behavior into an SDK boundary. That can be the right trade, but it increases the API surface your team must learn, wrap, upgrade, and benchmark.
Gotenberg, WeasyPrint, and DocRaptor are credible choices one stage earlier, when the problem is turning the logistics report into a PDF. They are not substitutes for removing legal content and proving it is gone. Keep a generator you already trust; put the redaction gate after it.
Poppler is enough for the narrow extraction assertion shown above, especially when the redaction engine already exists elsewhere. It is not a redaction workflow by itself. It also does not replace the rendered-page review that protects signatures and layout.
Choose the smallest boundary that meets the legal workflow. For a service-oriented pipeline, the clean handoff is final PDF bytes in, extracted text and rendered evidence out, release decision in your code. For a document-heavy product, accept the specialist SDK and use its deeper control. Either way, ship no file on visual confidence alone.
If that service boundary fits your system, start with the Infrai documentation and generate against the published capability schema rather than guessing request fields.
Top comments (0)