DEV Community

ValtorMist7692
ValtorMist7692

Posted on

Supplier PDF Handling Explained: 6 Controls to Decrypt and Delete Plaintext

Treat a decrypted supplier PDF as a short-lived capability, not as a file your application owns. The operational rule is: admit a batch only when the service has bounded temporary capacity, decrypt each document into a private job directory, produce the required signed evidence, close every handle, remove the directory, and record the outcome without recording document contents.

TL;DR: make cleanup part of the job state machine and its SLO, rather than a best-effort line at the end of a happy path. Keep encrypted input as the recoverable source, cap concurrency by measured bytes in flight, and do not acknowledge success until the signed result and audit event are durable and the intermediate path is gone. This is less convenient than launching one goroutine per contract. It is also the design that remains understandable during a 4,000-document supplier batch.

This example uses Go because the control flow around resource ownership is unusually visible, but the boundary maps directly to a Node.js worker: a per-job directory, an injected PDF decryptor, bounded admission, finally-style cleanup, and an audit sink. The PDF-specific operation must come from an implementation that supports the encryption used by the supplier document and preserves the document behavior required by ISO 32000-2. The surrounding lifecycle should not depend on that implementation.

How should Node.js decrypt, process, and delete a supplier PDF?

The dangerous failure is not merely "a temp file was left behind." The real failure is ambiguity: the queue says a contract succeeded, the audit system has no corresponding event, and a decrypted copy remains on a worker whose next action is unknown. Retries make that ambiguity worse because the same logical document can acquire several plaintext copies under different job attempts.

Cleanup can fail.

A useful state model is deliberately small:

  1. admitted: capacity has been reserved for the encrypted input and expected working copy.
  2. decrypted: plaintext exists only inside the job directory.
  3. signed: the service has produced the required signature evidence.
  4. committed: the result and audit event are durable.
  5. purged: the job directory no longer exists.

Success is committed plus purged. A job that reached signed but not committed can be retried from encrypted input. A job that reached committed but not purged is a cleanup incident, not a reason to sign the contract again. That distinction prevents an infrastructure retry from becoming a duplicate business action.

Deletion is narrow here. os.RemoveAll removes directory entries visible to the process; it does not prove that every historical byte has been physically overwritten on every storage layer. The stronger control is to minimize plaintext residence in the first place, use storage whose lifecycle matches the threat model, and destroy the temporary directory promptly on every exit path. Do not market a filesystem call as cryptographic erasure.

The audit event should carry a stable document identifier, job attempt, input digest, result digest, signer key identifier, timestamps, terminal state, and a normalized error class. It should not carry the supplier name extracted from the document, contract clauses, passwords, raw paths, or the PDF itself. Logs are another retention system.

This design has real limitations. Disk-backed plaintext is the wrong choice when policy forbids any decrypted bytes on a filesystem, even inside an isolated worker; choose a decrypt-and-process implementation that can keep the entire operation in bounded memory, provided the PDF engine actually supports streaming and the largest accepted document fits the memory budget. The detached signature in the example is also unsuitable when a counterparty requires an embedded PDF signature, visible signature appearance, certificate-chain validation, or another document-level conformance profile. In that case, choose a standards-aware signing component and keep the same outer admission, audit, and cleanup state machine. The trade-off is deliberate: the sample makes lifecycle ownership reviewable, but leaves PDF semantics to a component qualified for the specific contract workflow.

Build one owner for the plaintext

The safest implementation has one function that creates the private directory and owns its removal. Everything that can fail sits beneath that owner. The decryptor is an interface because PDF encryption is not a primitive to improvise; the adapter must be selected and tested against the actual supplier files. The signer below creates detached evidence over a SHA-256 digest. It does not claim to create an embedded, standards-conforming PDF signature field.

package contractjob

import (
    "context"
    "crypto/ed25519"
    "crypto/sha256"
    "encoding/hex"
    "errors"
    "fmt"
    "io"
    "os"
    "path/filepath"
    "time"
)

type Decrypter interface {
    Decrypt(ctx context.Context, encryptedPath, plaintextPath string, password []byte) error
}

type AuditEvent struct {
    DocumentID string
    Attempt    int
    InputHash  string
    ResultHash string
    KeyID      string
    State      string
    OccurredAt time.Time
}

type AuditSink interface {
    Commit(ctx context.Context, event AuditEvent, signature []byte) error
}

type Job struct {
    DocumentID   string
    Attempt      int
    EncryptedPDF string
    Password     []byte
}

type Processor struct {
    TempRoot  string
    Decryptor Decrypter
    Audit     AuditSink
    PrivateKey ed25519.PrivateKey
    KeyID     string
}

func (p *Processor) Run(ctx context.Context, job Job) (err error) {
    if p.Decryptor == nil || p.Audit == nil {
        return errors.New("processor dependencies are required")
    }
    if len(p.PrivateKey) != ed25519.PrivateKeySize {
        return errors.New("invalid Ed25519 private key")
    }

    dir, err := os.MkdirTemp(p.TempRoot, "contract-")
    if err != nil {
        return fmt.Errorf("create job directory: %w", err)
    }
    defer func() {
        removeErr := os.RemoveAll(dir)
        if removeErr != nil {
            err = errors.Join(err, fmt.Errorf("remove job directory: %w", removeErr))
        }
    }()

    plainPath := filepath.Join(dir, "document.pdf")
    if err := p.Decryptor.Decrypt(ctx, job.EncryptedPDF, plainPath, job.Password); err != nil {
        return fmt.Errorf("decrypt document: %w", err)
    }
    if err := os.Chmod(plainPath, 0o600); err != nil {
        return fmt.Errorf("restrict plaintext permissions: %w", err)
    }

    inputDigest, err := fileSHA256(job.EncryptedPDF)
    if err != nil {
        return fmt.Errorf("hash encrypted input: %w", err)
    }
    resultDigest, err := fileSHA256(plainPath)
    if err != nil {
        return fmt.Errorf("hash decrypted result: %w", err)
    }
    signature := ed25519.Sign(p.PrivateKey, resultDigest[:])

    event := AuditEvent{
        DocumentID: job.DocumentID,
        Attempt: job.Attempt,
        InputHash: hex.EncodeToString(inputDigest[:]),
        ResultHash: hex.EncodeToString(resultDigest[:]),
        KeyID: p.KeyID,
        State: "committed",
        OccurredAt: time.Now().UTC(),
    }
    if err := p.Audit.Commit(ctx, event, signature); err != nil {
        return fmt.Errorf("commit audit evidence: %w", err)
    }
    return nil
}

func fileSHA256(path string) ([sha256.Size]byte, error) {
    f, err := os.Open(path)
    if err != nil {
        return [sha256.Size]byte{}, err
    }
    defer f.Close()

    h := sha256.New()
    if _, err := io.Copy(h, f); err != nil {
        return [sha256.Size]byte{}, err
    }
    var sum [sha256.Size]byte
    copy(sum[:], h.Sum(nil))
    return sum, nil
}
Enter fullscreen mode Exit fullscreen mode

There is a sharp edge in this otherwise small function: returning an audit commit error and then failing cleanup produces two errors. errors.Join preserves both. Replacing the first failure with the cleanup failure would make diagnosis look cleaner while discarding the event that explains whether a business action may have occurred.

Passwords need the same ownership discipline. Fetch them as late as possible, do not put them in a job payload or environment dump, and release the backing secret according to the secret provider's contract. Overwriting a Go byte slice can reduce accidental reuse inside the process, but it is not proof that a runtime, kernel, or storage subsystem retained no copy. The honest security boundary is architectural, not ceremonial.

Size the batch from bytes in flight

Batch throughput is constrained by the scarce resource, and for encrypted supplier documents that resource is often temporary bytes or decryptor CPU rather than queue depth. A worker count chosen from average PDF size will fail on the batch with a long tail. Use an admission budget based on observed upper-percentile working size, then validate it under representative documents before setting the production limit.

Suppose a worker has 80 GiB reserved for temporary data, the operational policy keeps 25% free, and the tested working-set allowance is 320 MiB per active document. The initial disk bound is floor(80 GiB * 0.75 / 320 MiB) = 192 concurrent jobs. That is a capacity-planning example, not a universal setting. CPU, memory, key-service quotas, and audit-sink latency can force a lower limit, so start with the minimum of those independently measured bounds. Then put an SLO around the lifecycle, not merely around decryption latency. Useful indicators include end-to-end batch completion, oldest admitted job age, plaintext directory age, cleanup failure count, temporary bytes in use, audit commit latency, and retry count by normalized error class. The most important alert is simple: a plaintext directory is older than the maximum legitimate processing window. One old directory can matter more than a thousand fast jobs. This is where an apparently healthy throughput graph can mislead an operator: 191 jobs may finish within target while one canceled attempt leaves plaintext behind, so the batch latency objective is green while the security lifecycle objective is red. Track both, and give the age alert its own response procedure rather than burying it under a generic worker-health page.

Tail risk wins.

Backpressure must happen before decryption. A full disk after plaintext creation leaves the service with the sensitive artifact and no room to finish the work that justifies it. Reserve capacity conservatively, reject or defer admission when the budget is exhausted, and release the reservation only after purge. Fast queues can wait.

Bytes decide.

Choose the boundary with a buy-versus-build table

The decision is not "library or service" in the abstract. Split it into control planes: PDF interpretation, signing keys, job orchestration, temporary storage, and audit retention. Each boundary has a different failure domain and a different on-call owner.

Boundary Build or self-host favors Managed capability favors Batch-throughput question
PDF decryption Local data custody and direct format testing Less parser maintenance Can concurrency and file-size limits be tested with the real corpus?
Signing key operations Maximum control over deployment topology Reduced key custody burden Are quotas, latency, and retry semantics compatible with the peak batch?
Job orchestration Custom state and admission logic Lower scheduler operations load Can it prevent duplicate business actions after partial completion?
Temporary storage Explicit placement and lifecycle Reduced host management Is purge observable, and what exactly does deletion guarantee?
Audit retention Schema and query control Reduced storage operations load Can evidence commit atomically enough for the chosen retry model?

Lock-in shows up in state, not client syntax. A thin adapter around a decrypt call is easy to replace; years of audit records with provider-specific identity fields are not. Keep the durable event schema independent, store algorithm and key identifiers explicitly, and make verification possible without the worker that created the evidence.

Cost belongs in the table, but it is rarely the first constraint for sensitive contracts. Compare total capacity at the required batch window, on-call ownership, evidence export, incident response, and exit cost. A cheaper unit operation that adds an unstaffed parser or storage system to the platform roadmap is not automatically cheaper.

Verify cleanup before increasing concurrency

Verification needs failure injection at every transition. Use synthetic PDFs whose handling rights and expected rendering are known; include wrong passwords, malformed files, canceled contexts, a full temporary volume, an unavailable audit sink, and a cleanup permission failure. For every case, assert the terminal job state, number of business signatures committed, audit evidence present or absent, and directory presence after the worker exits.

The acceptance checks are intentionally asymmetric. A decrypt failure should produce no signature and no plaintext directory. An audit failure should preserve encrypted input for retry but remove plaintext. A cleanup failure after audit commit should page or quarantine the worker, while the job remains committed and must not be signed again. Process restart tests matter because deferred cleanup runs only during orderly function return; startup reconciliation must inspect the dedicated temporary root and handle abandoned job directories under the same retention policy.

Do not point reconciliation at a shared operating-system temp directory. Give the service a dedicated root, validate that the resolved path is beneath that root, and limit cleanup to directory names and metadata the service created. Broad cleanup code is a deletion incident waiting for a typo.

Before raising a concurrency limit, run a batch large enough to expose queueing and tail behavior, then compare peak temporary bytes, CPU saturation, audit latency, and cleanup age against the reserved margins. Increase one limit at a time. Stop when an SLO margin narrows, not when the worker finally falls over.

Roll back without duplicating a contract action

Rollback means returning to a known worker version and concurrency limit while retaining encrypted inputs and committed audit evidence. It does not mean replaying every job that lacks a final queue acknowledgment. Reconciliation must query the durable document identifier and attempt history first: if the signed evidence is committed, repair cleanup or delivery separately; if it is not committed, retry from encrypted input with a new attempt number.

Keep the previous decryptor adapter deployable until the new version has processed a representative canary batch and verification can reproduce every result required by the business workflow. A format regression may affect only one supplier's producer, so a global success percentage can conceal it. Segment canary results by document producer characteristics that are safe to retain, encryption mode, and failure class, without copying sensitive content into telemetry.

The final operational decision rule is strict: no batch expansion until abandoned-plaintext detection, idempotent audit lookup, capacity admission, and rollback reconciliation have all been exercised. Throughput is useful only after the lifecycle is bounded.

References

Top comments (0)