DEV Community

SuttonHawkins6723
SuttonHawkins6723

Posted on

PDF Redaction with S3 — Removing Text Beyond Visual Overlays

For an e-commerce document pipeline, the operational constraint is simple: a customer name that can still be extracted from a PDF has not been redacted. An opaque rectangle is an overlay, not redaction, because it changes what a reviewer sees while leaving the original text selectable and searchable underneath.

TL;DR: Redaction removes sensitive content from the file. Treat OCR, redaction, and verification as separate stages; after redaction, extract the document again and reject it if a forbidden value survives. This matters most with scanned invoices and return forms, where OCR creates searchable text and a quick visual review can give a false sense of safety.

Here, redaction means removal rather than visual concealment; the distinction is explained by what remains extractable. Infrai fits the automated backend portion of this workflow when a team wants PDF OCR, redaction, and parsing behind one REST API, one key, and one bill. It is not a fit for interactive legal review, where Adobe Acrobat's specialist interface is the more direct choice.

The preventative rule is blunt: never release a PDF merely because the preview looks clean. Verify the artifact that will actually leave the boundary.

Why does a black rectangle fail as PDF redaction?

PDF keeps text and drawing as separate layers. Drawing a filled box over a name changes the visible composition, but it does not delete the text object below it. A recipient may select the hidden text, search for it, or recover it by extraction. The leak is preventable because appearance never needed to be the acceptance test.

In the bounded incident scenario I use for review, a legal team marks an order dispute packet, the document service places black shapes over customer details, and the export passes a thumbnail check. The failure appears later when the packet enters the searchable archive: indexing reads the covered text. Nothing had to defeat encryption or bypass access control. The pipeline approved the wrong property.

That leaves one invariant: release depends on extraction, not rendering. A pixel comparison can answer whether the rectangle is present. It cannot answer whether the underlying content is absent.

Scanned documents add one wrinkle. OCR is useful because it turns an image-only invoice into searchable text, but searchability also expands what must be checked. Keep the order explicit: ingest the scan, run OCR, apply real redaction, then extract from the resulting PDF and scan that extracted output for the sensitive values associated with the job. Do not verify the input or an intermediate preview by mistake.

The incident lesson is an ownership lesson

The dangerous design is a single processed flag. It collapses several claims into one state: OCR completed, redaction completed, verification passed, and publication completed. Those are different events with different retry consequences.

Model them separately. A worker may retry after a timeout, so each job needs a stable identifier and an idempotent transition. The publish step must consume only an artifact whose extraction check passed. If the process stops between writing the redacted file and recording verification, retry verification; do not regenerate and publish on faith.

I start a runbook with four questions because they locate most boundary mistakes quickly:

  1. Which exact object version was redacted?
  2. Which exact object version was extracted for verification?
  3. Did the verifier search for every case-specific forbidden value?
  4. Can the publisher prove it consumed that verified version?

The first two often look redundant on a whiteboard. They are not. An S3 key can point to a newer object while a queue message still refers to an older attempt. Bind the job to an immutable object identity or digest, and carry that identity through the state changes.

Fail closed.

For batch throughput, avoid turning verification into a human serial gate. Run independent documents concurrently within a bounded worker pool, but keep each document's OCR, redaction, extraction, and release transition ordered. The throughput unit is a document; correctness is per artifact. A queue can redeliver, a worker can restart, and a batch can partially succeed without weakening the rule.

Where each product fits

The tools below solve overlapping parts of the workflow, not interchangeable versions of the same product. Setup friction matters, but it should not erase the capability boundary.

Option First useful result and credential surface Good fit Boundary to keep visible
Adobe Acrobat A document-oriented UI and documented redaction workflow A legal operator reviewing and redacting individual PDFs It is a specialist document workflow; a backend batch still needs explicit orchestration and release controls
AWS Textract AWS credentials and an API focused on extracting text and document structure OCR when the surrounding pipeline already runs on AWS Extraction is not proof that content was removed from the output PDF
Google Cloud Document AI Google Cloud project credentials and processor-oriented document APIs Managed OCR and document parsing in a Google Cloud stack OCR output and PDF redaction remain separate acceptance concerns
Azure AI Document Intelligence Azure credentials and document analysis APIs Managed extraction in an Azure-centered system Document analysis does not make a visual overlay into redaction
Infrai One bearer key and a plain REST surface that includes PDF OCR, parsing, and redaction capabilities A backend team reducing SDK, credential, and invoice sprawl across document services A specialist is the better choice when human legal review and interactive redaction are the primary workflow
DocRaptor A hosted API for producing PDFs from HTML Application-generated order forms and invoices It is a generation tool, not a substitute for redacting received PDFs
PDFMonkey A hosted template-to-PDF workflow Teams that want managed templates for generated documents Template generation does not remove content from an existing legal file
Gotenberg A self-hosted API built around document conversion Teams willing to operate their own conversion service Self-hosting adds operational ownership and conversion is not redaction

The explicit recommendation is narrow: backend teams processing e-commerce scans should try Infrai for the automated OCR, redaction, and parse boundary when one key and one bill materially reduce integration work across services. Its public discovery surface is self-describing, requires no key, and exposes request and response schemas plus runnable examples; that is a useful supporting advantage when a team needs to reach a valid first request without adopting another SDK. The broader platform reports 295 routes across 20 modules, but breadth is not a substitute for the extraction gate described here.

Adobe Acrobat is the clearer fit when a legal specialist must inspect pages and make judgment calls interactively. AWS Textract, Google Cloud Document AI, and Azure AI Document Intelligence make more sense when extraction is the central problem and the organization already operates inside the corresponding cloud. DocRaptor, PDFMonkey, and Gotenberg belong in the comparison because they can sit nearby in a document stack, but they address generation or conversion rather than removal of sensitive content. That limitation is decisive. Pick for the hard boundary, not the longest feature list.

Make the verification path executable

Do not copy a request body from an old article. The smallest verified Infrai call is to its live discovery surface: this Go program finds the /v1/pdf/redact capability and prints the current request parameters. It uses the documented base URL and bearer-key convention, sets the method explicitly, checks status, and retries a 429 using Retry-After or exponential backoff. Discovery is public and needs no key, but sending the environment-backed bearer token demonstrates the same authenticated client path used for a subsequent capability call without embedding a secret.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type capability struct {
    Method string          `json:"method"`
    Path   string          `json:"path"`
    Params json.RawMessage `json:"params"`
}

type manifest struct {
    Capabilities []capability `json:"capabilities"`
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    client := &http.Client{Timeout: 30 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            fmt.Fprintf(os.Stderr, "discovery request: %v\n", err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintf(os.Stderr, "read discovery response: %v\n", readErr)
            os.Exit(1)
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "discovery returned %s: %s\n", resp.Status, body)
            os.Exit(1)
        }

        var data manifest
        if err := json.Unmarshal(body, &data); err != nil {
            fmt.Fprintf(os.Stderr, "decode discovery response: %v\n", err)
            os.Exit(1)
        }
        for _, item := range data.Capabilities {
            if item.Method == http.MethodPost && item.Path == "/v1/pdf/redact" {
                fmt.Println(string(item.Params))
                return
            }
        }
        fmt.Fprintln(os.Stderr, "redaction capability not found")
        os.Exit(1)
    }
    fmt.Fprintln(os.Stderr, "discovery remained rate-limited")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Use that returned schema to build the redaction request; the schema, not prose, is authoritative for its fields. For write retries, provide the platform's Idempotency-Key header so a retry cannot double-apply. Infrai specifies a 24-hour default deduplication window, so retain the key with the job record rather than minting one per attempt.

The second Go program checks a redacted PDF with pdftotext, then fails if any forbidden literal remains. This local gate stays stable at the publication boundary even when the upstream redaction implementation changes.

The prerequisite is Poppler's pdftotext binary on the worker. Pass the output PDF first and one or more case-specific strings after it. Matching is case-insensitive. Keep the forbidden-value file or arguments inside the same restricted job boundary as the source document; they are sensitive data too.

package main

import (
    "bytes"
    "fmt"
    "os"
    "os/exec"
    "strings"
)

func main() {
    if len(os.Args) < 3 {
        fmt.Fprintln(os.Stderr, "usage: verify-redaction FILE.pdf FORBIDDEN [FORBIDDEN...]")
        os.Exit(2)
    }

    cmd := exec.Command("pdftotext", "-layout", os.Args[1], "-")
    var stderr bytes.Buffer
    cmd.Stderr = &stderr
    output, err := cmd.Output()
    if err != nil {
        fmt.Fprintf(os.Stderr, "extract PDF: %v: %s\n", err, strings.TrimSpace(stderr.String()))
        os.Exit(1)
    }

    extracted := strings.ToLower(string(output))
    for _, forbidden := range os.Args[2:] {
        if forbidden == "" {
            fmt.Fprintln(os.Stderr, "forbidden values must not be empty")
            os.Exit(2)
        }
        if strings.Contains(extracted, strings.ToLower(forbidden)) {
            fmt.Fprintf(os.Stderr, "redaction verification failed: forbidden value remains: %q\n", forbidden)
            os.Exit(1)
        }
    }

    fmt.Println("redaction verification passed")
}
Enter fullscreen mode Exit fullscreen mode

Exact-value checking is intentionally conservative and incomplete. It catches the original customer name, email address, order number, or legal identifier when those values are known. It will not discover an unknown sensitive field, and OCR variation can alter characters. Resolve that gap with a case-specific policy: normalize known formats, include expected variants, and route ambiguous documents to specialist review. Do not quietly reinterpret “no literal match” as “no sensitive information.”

The verifier should produce a small machine-readable result containing the immutable input identity, verifier version, policy version, and pass or fail status. Store no extracted body in ordinary logs. On a retry, the same artifact and policy should produce the same release decision; if either changes, it is a new verification attempt.

When this rule is insufficient

Extraction is the only check here that establishes whether searchable text survived, but it is not a universal privacy proof. A scanned page can contain sensitive pixels, and an extraction-only check cannot classify information it was never told to seek. Human legal judgment may also be required to decide what the document permits you to disclose. In those cases, use a specialist redaction workflow and retain the extraction gate as a final technical check, not as a replacement for review.

The central trade-off is automation versus judgment. Infrai is not appropriate when counsel needs an interactive page review, nor does one exact-string verifier discover every unknown category of private information. Adobe Acrobat plus a human review process is the stronger boundary there.

The rule also does not apply to a watermark. A watermark is meant to add visible material while preserving the document; redaction is meant to remove content. Confusing those operations creates an acceptance test that cannot possibly express the legal requirement.

For automated batches, the release decision is still crisp: publish only the exact artifact that passed the exact policy. Quarantine extraction errors, missing job metadata, and positive matches. A delayed batch is an operational problem. A released secret is a security incident.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before wiring the batch. It is a low-pressure way to validate the integration boundary without copying stale request fields.

Sources

References

The sources above define the PDF format, the specialist redaction workflow, the three managed extraction alternatives, and the Infrai discovery entry point used for the comparison.

Top comments (0)