DEV Community

DimitriReed2158
DimitriReed2158

Posted on

Redacted PDF Still Contains Text: Debugging Customer Support Bundle Assembly

Covering sensitive text with a dark rectangle is painting, not redaction. For customer-support evidence bundles, the reliable boundary is the document part that owns the sensitive content: remove that content before pages are merged or split, rebuild the output, and then test the exported bytes as an untrusted recipient would. If copied or extracted text still contains the value, the job has failed even when every page looks correct.

Short answer: a PDF that still yields the hidden text was overlaid, incompletely rewritten, or assembled from an unsafe intermediate. Treat visual review as one check, never as proof of removal.

The production scenario is bounded but common in shape. A support workflow collects chat transcripts, uploaded forms, and agent notes into one legal evidence packet, then splits that packet by case recipient. The template team owns the cover sheet and page furniture; the case-data service owns names, account identifiers, and message content. A scheduled worker owns assembly and delivery. I've been paged by missed jobs and duplicate deliveries, so I care about that last boundary: a retry can deliver the same artifact twice, and a polished black box doesn't make the bytes beneath it less sensitive.

Why does a redacted PDF still contain text after export?

PDF is a page-description format, not a screenshot container. A page can contain text-showing operations and a later operation that paints a rectangle over the same coordinates. The viewer composes both and displays the rectangle on top. Text extraction can still interpret the earlier content, search can still find it, and removing or changing the covering object can expose it again.

That is the first invariant: appearance is not content deletion. A black fill, annotation, drawing, or flattened-looking preview answers only what a particular renderer displayed. It does not establish what objects, streams, metadata, attachments, form values, or earlier revisions remain in the file. ISO 32000-2 defines the PDF format and is the baseline for reasoning about document objects and page content rather than treating the file as pixels.

The merge step widens the failure surface. A clean cover page does not sanitize the uploaded pages appended behind it. A split step can also copy source objects into several recipient files. If sanitization happens after routing, one missed branch becomes a disclosure. If it happens only inside a template, uploaded content sits outside the template's authority.

Template ownership therefore needs a hard limit. The template team may own layout and labels, but it cannot promise removal of values embedded in arbitrary source documents. The service that understands case data should classify the fields; the document transformation should remove marked content; and the delivery worker should reject any artifact that fails independent verification. Three owners, three different claims.

The incident lesson is to verify the artifact

The operational trap is a green pipeline. The render call returned success, the merge completed, the object store accepted the upload, and the delivery event was acknowledged. None of those signals says that sensitive text is absent. They prove transport and execution, not the security property the job was meant to enforce.

My default runbook treats each output as a new artifact with a new decision. The worker records a source digest, policy version, template version, output digest, recipient class, and verification result. It does not log the sensitive match itself. On retry, the idempotency key binds the case, recipient class, source digest, and policy version, so a replay cannot silently substitute a different input under the same delivery record.

One short rule prevents a long postmortem: fail closed.

No exceptions.

A useful state transition is received -> classified -> transformed -> rebuilt -> verified -> released. Only verified may advance to released. A timeout during verification is not a soft success, and a duplicate queue delivery reuses an already verified immutable artifact only when every key component matches. Otherwise it starts a separate attempt.

package redaction

import (
    "context"
    "crypto/sha256"
    "errors"
    "fmt"
)

type Request struct {
    CaseID        string
    Recipient     string
    SourcePDF     []byte
    PolicyVersion string
    Needles       [][]byte
}

type Transformer interface {
    RemoveAndRebuild(context.Context, []byte, [][]byte) ([]byte, error)
}

type Verifier interface {
    Inspect(context.Context, []byte, [][]byte) error
}

type ArtifactStore interface {
    PutVerified(context.Context, string, []byte) error
}

func Build(ctx context.Context, req Request, tx Transformer, verify Verifier, store ArtifactStore) (string, error) {
    sourceSum := sha256.Sum256(req.SourcePDF)
    key := fmt.Sprintf("%s:%s:%x:%s", req.CaseID, req.Recipient, sourceSum, req.PolicyVersion)

    out, err := tx.RemoveAndRebuild(ctx, req.SourcePDF, req.Needles)
    if err != nil {
        return "", fmt.Errorf("transform artifact: %w", err)
    }
    if len(out) == 0 {
        return "", errors.New("transform produced an empty artifact")
    }
    if err := verify.Inspect(ctx, out, req.Needles); err != nil {
        return "", fmt.Errorf("verification rejected artifact: %w", err)
    }
    if err := store.PutVerified(ctx, key, out); err != nil {
        return "", fmt.Errorf("store verified artifact: %w", err)
    }
    return key, nil
}
Enter fullscreen mode Exit fullscreen mode

The interfaces matter more than the implementation choice. Transformation and verification should not be the same assertion from the same code path. A function that reports its own success can repeat its own blind spot. The verifier needs a recipient-view input: the final bytes after merge, split, metadata updates, compression, and save.

Removal, rasterization, and overlays serve different jobs

There are three broad mechanisms, and confusing them causes most bad decisions.

Mechanism What it changes Useful when Main operational risk
Content removal and rebuild Deletes marked content and writes a new document structure Searchable text and vector quality must remain where allowed Incomplete coverage of related objects or alternate representations
Rasterization and reconstruction Renders allowed page appearance into pixels, then creates new pages Source internals are untrusted and loss of document semantics is acceptable OCR, accessibility, links, signatures, and fine detail may be lost or changed
Visual overlay Adds paint above existing content Review markup or non-security presentation Original information may remain extractable

For customer-support records, content removal is usually the precise path because the remaining conversation often needs to stay searchable. Rasterization can create a stronger isolation boundary from complex page internals, but it changes the artifact substantially. It may be wrong when the legal workflow requires preserved digital signatures, accessible text, exact vectors, or original evidence semantics. Those requirements must be resolved before choosing the transformation, not discovered during delivery.

An overlay is still valid for annotation. It is just the wrong control for confidentiality. Label it that way in APIs and schemas; a field named redactions that merely draws rectangles invites an unsafe assumption downstream.

This approach has limits.

Build the test from known secrets and recipient boundaries

Start with fixtures whose sensitive values are unique and known. Put a synthetic account identifier in ordinary text, a form field, metadata, an annotation, and an embedded attachment. Add a page where the same value appears twice. Merge the fixture with a template cover sheet, split it into at least two recipient classes, then inspect every final file. This is a test corpus, not production data.

Verification should cover several independent observations: extracted text, full-file byte search for suitable test tokens, object and attachment enumeration, metadata and form inspection, and rendered-page comparison. No single check is complete. Compression or character encoding can defeat a naive byte search; text extraction can miss image text; rendering can hide surviving objects. Their intersection is useful.

The negative tests deserve equal weight. Confirm that allowed phrases remain searchable, page counts match the routing policy, and the case identifier intended for a recipient has not disappeared. A pipeline that removes every text object passes a simplistic secret-absence test while destroying the evidence packet. That is not success.

I would gate deployment with a small matrix: at least one fixture per input class, recipient class, and transformation policy. In production, sample structural properties and state transitions without putting customer text into logs. Count verification failures by policy and template version; alert on any released artifact lacking a verification record. Track retries and duplicate deliveries separately because they diagnose orchestration, not redaction quality.

The key decision is where this gate runs. Put it after the last byte-changing operation. Verification before merge says nothing about a cover template that injects metadata. Verification before split says nothing about recipient routing. Verification before a final incremental save may inspect a different revision from the one delivered. The release gate must hash and approve the exact byte sequence stored for delivery.

When this method does not apply

Do not rewrite a signed original and then describe it as the same signed evidence. If evidentiary rules require retaining the original, preserve it in a restricted evidence store and create a separately identified disclosure copy under an approved policy. Access control around the original and redaction of the disclosure copy solve different problems. The trade-off is explicit: rebuilding reduces the chance that hidden source objects survive, but it can invalidate signatures and alter document semantics; preserving the signed original retains those properties but cannot produce a disclosure-safe copy by itself.

The same caution applies when a recipient is entitled to the unmodified record, when a legal hold forbids transformation, or when accessibility requirements rule out a raster-only copy. Those are policy decisions with technical consequences. The job should stop at classification until the policy selects an allowed output, rather than guessing from a template name.

Scanned pages also need a distinct threat model. Removing an OCR text layer does not remove characters visible in the image. Covering image pixels in a viewer does not prove the stored image was rewritten. Test the pixel output and any associated text layer, and rebuild the delivered file from approved content.

Choose per policy.

The release rule

A redacted bundle is a derived, verified artifact, not a successfully rendered picture. Remove sensitive content at the data-aware boundary, rebuild before cross-recipient assembly where possible, and inspect the exact final bytes after every merge and split. Keep template ownership narrow: templates control presentation; policy controls what may remain; the release worker controls whether an artifact leaves the system.

This rule costs additional processing and more fixtures. It also gives the on-call engineer a defensible answer to the only question that matters during an incident: which exact artifact was checked, under which policy, before it was delivered?

Sources

Top comments (0)