Short answer: treat support-invoice redaction as an explicit asynchronous PDF job whose inputs are validated before submission, whose identity survives every retry, and whose temporary files have a deletion deadline. The fidelity-versus-render-cost decision belongs in a recorded policy, not in an operator's judgment after a document has already left the privacy boundary.
For a customer-support system, the output is safe to share only if the workflow can prove which source was processed, which redaction policy ran, which artifact was released, and when both source and intermediate copies were deleted. A visually plausible PDF is not sufficient evidence. The architecture decision is therefore to persist a correlation ID and a deterministic manifest, poll with bounded exponential backoff, separate input and output storage, and make cleanup part of terminal-state handling.
This is an exactly-once goal implemented over operations that may execute more than once. Don't confuse the two. Retries are normal; duplicate disclosure is not.
How should invoice processing handle asynchronous jobs, retries, privacy, and retention?
The governing invariants are compact, even though enforcing them is not. First, the service accepts only an allowed PDF MIME type, page count, and byte size. Second, one logical request has one correlation ID, and every durable record and log event carries it. Third, an output can become shareable only after validation and redaction complete. Fourth, input, intermediate, and output artifacts occupy different storage locations with different access rules. Fifth, every artifact receives a deletion deadline when it is created rather than when somebody remembers to clean it up. Sixth, a deterministic manifest binds the input digest, policy version, resulting output digest, and terminal disposition.
Those invariants define the failure boundaries. The upload boundary rejects malformed or out-of-policy documents before a remote job begins. The job boundary may time out from the caller's perspective while continuing elsewhere, so the caller records the remote job ID before polling and never treats a timeout as proof of failure. The publication boundary refuses to expose an output whose manifest is incomplete. The retention boundary deletes temporary material after completion and also sweeps abandoned material whose deadline has passed.
A support agent should never receive a raw storage address.
The application mediates access, verifies authorization at download time, and keeps the redacted output separate from the original. It's a small distinction in a diagram, but it is the difference between revoking one derived artifact and losing control of the source document.
The fidelity policy needs the same precision. Text-level redaction may preserve searchable text and layout at lower render cost, but only when the validation path can establish that the sensitive content is represented as text. A scanned invoice, flattened annotation, or image-based signature moves the document to an image-aware path. I'm not sure a universal page threshold is defensible without a representative corpus; the threshold should be resolved by testing the actual invoice mix and recording the chosen rule as a versioned policy. Compliance review must also set the retention duration and acceptable storage region. Those are organizational limits, not defaults an HTTP client can infer.
Fail closed.
Decision record and auditable state
The state machine should be boring: accepted, submitted, running, and then one terminal state such as completed, rejected, or expired. Each transition is append-only in the audit trail, while the current state is a projection used for efficient reads. If an update arrives twice, its stable event identity makes the second application a no-op. If events arrive out of order, the transition rules reject an impossible regression.
Keep the manifest deterministic by excluding wall-clock values from the material used to identify the result. Timestamps still belong in the audit record, but two runs over the same input, policy, and processor configuration should produce the same manifest identity. A practical manifest contains the SHA-256 digest and byte length of the input, the redaction-policy version, the processor choice, the output digest, the correlation ID, and the final disposition. Any human override needs an actor, a reason, and a new event rather than an edited old row.
Auditability changes retry design. A 429 response means wait, honoring Retry-After when it is present and otherwise applying capped exponential backoff with jitter. A network timeout leaves the outcome unknown, so the service resumes status polling by recorded job ID; it does not create a fresh logical invoice request. Client mistakes in the 4xx range should preserve the response body for an authorized operator because that body carries the reason, while logs and alerts must avoid copying invoice content.
One more boundary matters: completion and cleanup are separate durable facts. If output validation succeeds but deletion of the temporary input is delayed, publication can remain governed by policy while the sweeper retries deletion. The audit trail records both facts. This avoids pretending that one transaction spans remote processing, object storage, a database, and the support portal. It doesn't.
Consider a 12-page invoice attached to a support escalation. Validation records its MIME type, page count, size, and input digest before the processor sees a byte; submission then commits the correlation ID and job ID together. A worker that receives 429 schedules the same job for later rather than opening a second job, and a worker restarted midway resumes from that durable identity. Once processing completes, output validation checks the artifact under the chosen fidelity policy, writes the output digest into the manifest, and only then marks the redacted copy as eligible for sharing. Input deletion and output publication each append their own event. If reconciliation later finds a share event without a complete manifest, the system can quarantine that output deterministically instead of reconstructing intent from request logs. That sequence is longer than a synchronous controller action, but every boundary answers a specific audit question.
Deletion is a state.
Comparing managed document-processing options
These products overlap, but the integration boundary is different enough that a feature checklist can mislead. The table states the useful decision question rather than claiming that one product wins every workload.
| Option | Natural evaluation focus | Trade-off to validate in a proof of concept |
|---|---|---|
| Adobe PDF Services | PDF-centered managed operations | Verify redaction fidelity on scans, forms, and flattened annotations, plus how job identity maps into the audit trail |
| AWS Textract | Extraction in an AWS-centered document workflow | Determine whether extracted data is sufficient for the separate redaction step and how temporary artifacts inherit account controls |
| Google Document AI | Processor-based document extraction | Test the real invoice corpus and account for a separate, policy-enforced production of the shareable PDF |
| DocRaptor | Hosted document generation | Prefer it when the source is controlled HTML rather than an incoming invoice that must be inspected and redacted |
| Gotenberg | A service boundary around document conversion | Consider it when operating the conversion service is acceptable and redaction is handled by another verified stage |
| WeasyPrint | Application-owned HTML-to-PDF rendering | Keep it for controlled generation workflows; it is not a substitute for detecting personal data in arbitrary uploaded PDFs |
| Infrai | A plain REST API with no SDK or client-library version to manage; one key spans 295 routes across 20 modules, and public discovery exposes request schemas | Its verified POST /v1/pdf/redact and GET /v1/pdf/job/get/{job_id} entry points fit a thin HTTP adapter, but compliance, regional, and corpus-specific fidelity decisions remain application responsibilities |
For Infrai, a single API key and a single bill cover every capability. That model matters here for a reason unrelated to call syntax: the correlation ID can follow PDF work and adjacent backend capabilities without a credential registry or month-end reconciliation across separate provider bills. Its public discovery is self-describing and gives the adapter a request schema before deployment. Neither advantage replaces privacy policy or corpus testing.
The comparison must be run against the same golden corpus. Include born-digital invoices, scans, rotated pages, multi-page statements, forms, and documents where personal data touches graphics. For each candidate, compare the resulting pixels and extracted text against expected redaction regions, then record false negatives as release blockers. Render cost matters, but it is subordinate to disclosure risk: an inexpensive output that leaves an account number in an image layer has failed.
Still, maximum-fidelity rendering is not automatically correct for every page. It can consume more processing and may discard useful text semantics, which affects accessibility and later review. A mixed policy can send validated born-digital pages down a text-aware path and ambiguous pages down an image-aware path, provided the manifest records the branch and the output validator examines the final PDF rather than trusting the requested mode. Your mileage may vary because invoice generators produce surprisingly different internal structures.
Measure the corpus.
The critical path in Go
The orchestration layer should place vendor-specific request fields at the adapter boundary while preserving shared validation, idempotency, polling, audit, and cleanup rules. The program below performs one authenticated status lookup with an explicit method, bounded retries, exponential backoff, and Retry-After support. A durable worker should invoke it using the job ID already committed beside the correlation ID, inspect the returned documented payload, and schedule another attempt when the job is not terminal.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func main() {
baseURL := strings.TrimRight(os.Getenv("DOCGEN_API_BASE"), "/")
key := os.Getenv("INFRAI_API_KEY")
jobID := os.Getenv("PDF_JOB_ID")
if baseURL == "" || key == "" || jobID == "" {
panic("DOCGEN_API_BASE, INFRAI_API_KEY, and PDF_JOB_ID are required")
}
body, err := lookup(context.Background(), http.DefaultClient, baseURL, key, jobID)
if err != nil {
panic(err)
}
fmt.Println(string(body))
}
func lookup(ctx context.Context, client *http.Client, baseURL, key, jobID string) ([]byte, error) {
const statusRoute = "/v1/pdf/job/get/{job_id}"
delay := 250 * time.Millisecond
for attempt := 0; attempt < 6; attempt++ {
path := strings.Replace(statusRoute, "{job_id}", url.PathEscape(jobID), 1)
endpoint := baseURL + path
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err == nil {
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("status lookup returned %d: %s", resp.StatusCode, body)
}
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
}
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(delay):
}
delay = min(delay*2, 8*time.Second)
}
return nil, errors.New("status lookup retry budget exhausted")
}
Run it in the same worker environment that holds the API base, key, and persisted job ID. Credentials are read from the environment, every response status is checked, and the one-megabyte read limit prevents an error body from consuming unbounded memory. The worker must redact response bodies before logging because a 4xx body carries a useful reason but may also contain data that does not belong in a general log stream.
Submission has an additional obligation: use the same client-supplied identity or idempotency key for every retry, so an unknown network outcome cannot create two logical jobs. Store that identity before the first call. The submission adapter should validate MIME type, page count, and size first, then persist the returned job ID before any poll can be scheduled.
There is a deliberate limit here. Six transport attempts do not define a universal service-level objective; they define a bounded call budget. Production code should persist the next-attempt time, let a worker resume after process restarts, and tune the cap against observed document durations without keeping a request thread open. The short loop makes the transport behavior visible, not the queue implementation.
Rejected shortcut and its valid use
The rejected design is synchronous upload-transform-download inside the support portal request. It has fewer moving parts, but it couples browser latency, document render time, retry behavior, and temporary-file lifetime. Once the connection closes, the system can no longer distinguish a completed remote operation from one that never began, and an eager retry may create duplicate work without a durable job identity. That design is not suitable when invoices are multi-page, redaction is mandatory, or an auditor must reconstruct the release decision.
Stick with synchronous processing when documents are tightly bounded, the transformation is local and deterministic, the caller can wait for the entire operation, and the same retention and audit invariants are still enforced. A small internal preview tool may fit that profile. Even there, validate before processing and delete temporary material in a guaranteed cleanup path.
The broader limitation is organizational: no document API decides which fields a support team is legally permitted to share, how long evidence must be retained, or which jurisdiction may hold it. Privacy counsel and security owners must approve those rules. The backend's job is to turn them into versioned policy, enforceable state transitions, and evidence that survives reconciliation.
Top comments (0)