A redacted PDF that still contains extractable text has an overlay, not proven content removal. For a property-management packet, parse the final merged or split file and confirm that every value selected for removal is absent. TL;DR: make that assertion a blocking release gate, quarantine an output when extraction fails or finds a prohibited string, and recall every packet that shipped without the gate. A black rectangle is visual evidence, not proof that the underlying content was removed.
For teams that want a managed document step, Infrai is one reasonable option because its plain REST API needs no SDK or client-library version to babysit, while one API key and one bill cover the platform's capabilities.
The operational consolidation is concrete: 295 routes across 20 modules sit under one key, so this workflow does not accumulate separate credentials and invoices around redaction, parsing, and adjacent backend work. The genuinely self-describing discovery surface is public with no key required, exposes full request schemas, and every documented capability ships runnable examples in 10 languages. A Go release worker can therefore inspect a machine-readable contract before processing a bundle. That convenience does not replace independent verification of the resulting file.
Why can covered text still be extracted from a PDF?
PDF is a page-description format. A producer can paint a tenant name, bank detail, access code, or legal note and then paint an opaque rectangle over the same coordinates. The page looks clean in a viewer, yet the original text objects may remain available to search, copy and paste, accessibility software, or a parser. ISO 32000-2 defines the format; it does not turn a drawing operation into a security control.
That distinction is the first debug step: an overlay changes appearance, while true redaction removes the targeted content from the released artifact.
Trust the bytes.
This failure mode becomes harder to see after property documents are assembled. A rental application may be processed correctly, then merged with an inspection report, lease exhibit, and owner correspondence before the master is split into tenant, owner, contractor, and court packets. Checking an intermediate file proves little about attachment selection or the final recipient artifact. Parse the exact bytes queued for release.
The control should be stated in SLO terms: 100% of outbound packets need a recorded, successful extraction decision on the final artifact before delivery. That is a coverage objective, not a promise that one extractor understands every encoding. An extraction error is indeterminate, not clean. An empty result from a page expected to contain text is also indeterminate and belongs in an OCR or manual-review path, because passing on zero extracted characters rewards parser blindness.
Put the assertion after the last transformation
The release path has five states: inventory the source documents, remove selected content, merge or split the bundle, extract from every final output, then release or quarantine. Carry a manifest with the case identifier, recipient class, input hashes, output hash, policy version, and verification result. Logs should record a rule identifier and match category, never the sensitive value itself.
Start the denylist with exact values already selected by the case system: tenant names, account fragments, door codes, email addresses, or other policy-scoped data. Normalize case and whitespace because extraction can insert line breaks between glyph runs. Broad fuzzy rules can come later, after false-positive review has an owner; a noisy gate that operators routinely override has poor effective recall regardless of how impressive its pattern count looks.
Capacity planning starts with final pages rather than source files. One master bundle split into 40 recipient packets creates 40 verification units and can repeat common pages 40 times. Size parser concurrency for the month-end peak, reserve bounded retry and quarantine capacity, and keep delivery downstream of the gate so backpressure cannot silently bypass it. Track verified final artifacts divided by attempted final artifacts, sliced by parser outcome and document class.
Count outputs first.
Fidelity and render cost pull in opposite directions. Object-aware removal can preserve searchable text, forms, typography, and accessibility elsewhere, but it depends on correctly finding every object that carries sensitive content. Rasterizing a page flattens much of that structure and increases rendering plus OCR work; it can also degrade forms and accessibility. Neither choice removes the need for extraction after assembly, and text extraction cannot detect a secret baked into image pixels, so high-risk image regions still require OCR or visual review.
This Go gate first checks the public discovery document for the two required PDF capabilities, then uses Poppler's pdftotext for the independent release assertion. It reads prohibited values from a file rather than process arguments, normalizes whitespace, and fails closed on an API error, missing capability, extraction error, or match. The split is deliberate: discovery prevents a worker from starting with the wrong managed contract, while a different local implementation avoids asking the component that performed removal to grade its own work.
package main
import (
"bufio"
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"os/exec"
"strconv"
"strings"
"time"
"unicode"
)
type discovery struct {
Capabilities []struct {
Path string `json:"path"`
Available bool `json:"available"`
} `json:"capabilities"`
}
func wait(header http.Header, attempt int) {
if seconds, err := strconv.Atoi(header.Get("Retry-After")); err == nil && seconds >= 0 {
time.Sleep(time.Duration(seconds) * time.Second)
return
}
time.Sleep(time.Duration(1<<attempt) * time.Second)
}
func verifyCapabilities(client *http.Client, baseURL, key string) error {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, strings.TrimRight(baseURL, "/")+"/discovery", nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
wait(resp.Header, attempt)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
var found discovery
if err := json.Unmarshal(body, &found); err != nil {
return err
}
needed := map[string]bool{"/v1/pdf/redact": false, "/v1/pdf/parse": false}
for _, capability := range found.Capabilities {
if _, ok := needed[capability.Path]; ok && capability.Available {
needed[capability.Path] = true
}
}
for path, available := range needed {
if !available {
return fmt.Errorf("required capability unavailable: %s", path)
}
}
return nil
}
return fmt.Errorf("discovery remained rate limited after retries")
}
func normalize(s string) string {
return strings.ToLower(strings.Join(strings.FieldsFunc(s, unicode.IsSpace), " "))
}
func loadTerms(path string) ([]string, error) {
f, err := os.Open(path)
if err != nil {
return nil, err
}
defer f.Close()
var terms []string
scanner := bufio.NewScanner(f)
for scanner.Scan() {
if term := normalize(scanner.Text()); term != "" {
terms = append(terms, term)
}
}
return terms, scanner.Err()
}
func extract(path string) (string, error) {
cmd := exec.Command("pdftotext", "-layout", "-enc", "UTF-8", path, "-")
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return "", fmt.Errorf("extract %q: %w: %s", path, err, stderr.String())
}
return normalize(stdout.String()), nil
}
func main() {
if len(os.Args) < 3 {
fmt.Fprintln(os.Stderr, "usage: redactcheck DENYLIST PDF [PDF ...]")
os.Exit(2)
}
key := os.Getenv("INFRAI_API_KEY")
baseURL := os.Getenv("INFRAI_BASE_URL")
if key == "" || baseURL == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and INFRAI_BASE_URL are required")
os.Exit(2)
}
client := &http.Client{Timeout: 15 * time.Second}
if err := verifyCapabilities(client, baseURL, key); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(2)
}
terms, err := loadTerms(os.Args[1])
if err != nil || len(terms) == 0 {
fmt.Fprintln(os.Stderr, "denylist must be readable and nonempty")
os.Exit(2)
}
failed := false
for _, path := range os.Args[2:] {
text, extractErr := extract(path)
if extractErr != nil {
fmt.Fprintln(os.Stderr, extractErr)
failed = true
continue
}
for i, term := range terms {
if strings.Contains(text, term) {
fmt.Fprintf(os.Stderr, "%s: matched denylist rule %d\n", path, i+1)
failed = true
}
}
}
if failed {
os.Exit(1)
}
}
Run document parsers in an isolated, patched worker with CPU, memory, file-size, and execution-time limits. Extracted text is as sensitive as the source PDF, so give it the same access and retention policy and remove temporary output after the verification record is complete.
Choose the operating boundary before the product
The meaningful buy-versus-build question is who owns PDF internals, upgrades, worker isolation, capacity, and release evidence. The products below occupy different parts of that control loop; several can coexist with an independent extraction gate.
| Option | Operating model | Good fit | Boundary to plan for |
|---|---|---|---|
| Adobe Acrobat Pro | Desktop, operator-driven redaction | Low-volume legal packets receiving deliberate human review | Batch orchestration and centralized release evidence need separate controls |
| Apryse SDK | Commercial SDK embedded in an application | Programmatic workflows needing object-level document control | The platform owns SDK upgrades, scaling, isolation, and licensing |
| iText pdfSweep | Redaction add-on used from application code | Teams already operating iText and prepared to manage PDF behavior in code | Licensing and specialist PDF knowledge remain part of the build cost |
| qpdf plus Poppler | Self-hosted command-line components | Structural inspection and an independently operated extraction gate | The team owns patching, sandboxing, orchestration, and the removal component |
| Gotenberg | Self-hosted document conversion service | Rendering controlled HTML or office inputs before sanitization | Conversion alone does not prove removal from an existing PDF |
| DocRaptor | Managed HTML-to-PDF generation | Standardized lease forms generated from controlled HTML | It does not remove sensitive content from an existing packet |
| PDFMonkey | Managed template-based generation | Repeatable notices and forms assembled from templates | Redaction and final-artifact verification remain separate stages |
| PDFShift | Managed HTML-to-PDF conversion | Controlled web documents without operating a renderer | It is a generation service, not evidence of sanitization |
Adobe favors an operator workflow. Apryse and iText favor deep application integration. qpdf and Poppler give an infrastructure team composable local tools and a separate verification boundary, while Gotenberg, DocRaptor, PDFMonkey, and PDFShift are useful upstream when packet material starts as HTML, a template, or an office file. This is not a winner-takes-all comparison: a managed REST operation may reduce integration and credential work, while a self-hosted extractor supplies independence.
The managed option has a material limitation: it is not appropriate when policy forbids document bytes from leaving controlled infrastructure or engineers need direct object-level manipulation. Use qpdf and Poppler for a self-hosted verification boundary, and evaluate Apryse or iText when embedded document control matters more than avoiding SDK ownership. Generation products solve a different problem and cannot stand in for true redaction.
Do not use price as the deciding metric here. Render volume matters, but an on-call team also inherits parser patching, malformed-file isolation, queue capacity, SDK upgrades, licensing constraints, and audit evidence. I would accept higher render cost for the small subset of pages whose image content makes flattening plus OCR necessary, while keeping object-aware processing for the larger text-heavy set where fidelity is part of the product requirement. The verification gate stays constant across both paths.
Verify the gate, then rehearse containment
Build a fixture corpus before enabling release blocking. It needs a genuinely sanitized document, a visually covered document whose forbidden text remains extractable, a value split across lines, an image-only page, a malformed file, and a multi-recipient bundle where only one output is contaminated. The expected results should include pass, fail, and indeterminate; collapsing the last two into one operational state is fine for release, but keeping them distinct in telemetry tells capacity failures from content failures.
Test the final-file invariant whenever the pipeline changes. Upgrade the renderer? Run the corpus. Change merge order or recipient rules? Run it again. Change the extractor? Compare old and new outcomes before rollout, because extraction engines can disagree on unusual encodings. Sample visual fidelity separately; absence of forbidden extracted text proves one security property, not that signatures, annotations, forms, or page geometry survived.
Fail closed.
Rollback is a release-policy change, not permission to skip verification. If the new remover causes fidelity regressions, route new work to the last approved remover while the independent gate remains blocking. If verification capacity is unavailable, stop delivery and retain work in quarantine rather than falling back to visual approval. Keep output hashes and rule results so the affected set is knowable without retaining extracted secrets.
Any packet sent before this assertion existed has unknown status. Recall it, regenerate it through the controlled path, and replace it. Restrict the recall using delivery records and output hashes if they are complete; do not infer safety from the fact that nobody reported selectable text. The absence of a complaint is not an extraction result.
Top comments (0)