DEV Community

AshtonBlake6879
AshtonBlake6879

Posted on

PDF Preview API: Convert or Embed a Browser Viewer in 2026 (Audit-First)

Preserve the submitted PDF as the evidence object, and treat every browser preview as a disposable representation of it. The deciding constraint is the signature and audit trail: a customer-support agent may need a quick visual answer, but a screenshot or reconstructed page cannot stand in for the bytes that were actually submitted.

TL;DR: embed the original PDF when target browsers render the required form and signature appearance correctly. Add a derived image preview for constrained clients or safer isolation, but never let that derivative replace the original. If uniform rendering or annotation is mandatory, use a controlled viewer while keeping its state and telemetry outside the evidence object. This is an evidence-lifecycle decision, not a contest over which preview looks nicest.

What must remain true?

A support workflow has at least three distinct objects: the uploaded form, a browser representation, and an audit record. Combining them creates the dangerous failure mode. An agent sees a plausible page, the ticket says "reviewed," and nobody can later prove which byte sequence produced that page.

The architecture decision record needs five invariants:

  1. The original PDF is immutable after intake. Its content digest, storage identifier, byte length, intake time, and authenticated actor or session identifier belong in the audit event.
  2. A preview is derived data. It points to the original digest and carries its own renderer identity, generation time, page count, and status.
  3. The UI never describes a visual mark as cryptographically valid merely because it is visible. Signature validation is a separate operation with separately recorded results.
  4. Flattening happens only on a copy. It creates a new byte sequence, so it needs a new digest and an explicit relationship to its source.
  5. Authorization is checked at retrieval. Audit events record the decision without copying customer-entered field values into logs.

PDF is standardized by ISO 32000-2. That establishes the document format; it does not make every browser renderer, form implementation, or signature-validation path equivalent.

Should a PDF preview API convert or embed a browser viewer?

"Original" below means the exact accepted byte sequence, not a file re-saved by a PDF library.

The original wins.

Approach Evidence relationship Signature and form behavior Main failure boundary Operational consequence
Embed the original Browser reads the evidence through authorized delivery Rendering depends on the browser; visible appearance is not validation Client capability, content isolation, interactive-form behavior Lowest derivative storage, but browser and authorization outcomes still need measurement
Convert pages to images Images are derivatives linked to the original digest Visual marks appear; interaction and cryptographic validation are absent Renderer, fonts, page limits, conversion timeout Predictable passive display, with CPU, queued work, and derivative storage
Use a controlled viewer Viewer reads the original or a normalized derivative Consistent interaction is possible; validation stays separate Parser, worker isolation, versioning, state persistence More code and telemetry, with tighter release control

For a customer-support form, I would default to original-byte storage plus authorized embedding, then add image derivatives for clients that fail a capability check or for a deliberately passive review mode. A controlled viewer earns its complexity when agents need consistent navigation, redaction overlays, annotations, or controlled form interaction. None of these choices authorizes flattening the original.

One subtle trap is caching. A key based only on ticket ID cannot distinguish a replaced attachment, a renderer upgrade, or a changed output policy. Derive it from the original digest, renderer version, and policy version. Keep tenant and customer identifiers in access-control data, not in high-cardinality metric labels.

How does the critical path preserve the audit chain?

The following curl flow describes an illustrative internal contract, not a public product API. Intake returns an immutable source digest; preview creation names a representation policy; review records exactly what the agent saw. A deployment must authenticate each call and define idempotency, size, page, and timeout limits.

curl --fail-with-body \
  --request POST \
  --header "Authorization: Bearer ${SUPPORT_TOKEN}" \
  --header "Idempotency-Key: case-4827-upload-1" \
  --form "document=@customer-form.pdf;type=application/pdf" \
  https://documents.example.test/v1/pdf/form/fill

curl --fail-with-body \
  --request POST \
  --header "Authorization: Bearer ${SUPPORT_TOKEN}" \
  --header "Content-Type: application/json" \
  --data '{"source_digest":"sha256:SOURCE_DIGEST","mode":"passive-images","policy_version":"2026-01"}' \
  https://documents.example.test/v1/pdf/convert

curl --fail-with-body \
  --request GET \
  --header "Authorization: Bearer ${SUPPORT_TOKEN}" \
  https://documents.example.test/v1/pdf/job/get/JOB_ID
Enter fullscreen mode Exit fullscreen mode

The response contract should distinguish accepted, processing, ready, rejected, and expired states. It should not collapse a conversion failure into "document missing," or a failed signature check into "preview failed." Those are different facts and different support actions. Do not log the bearer token, raw request body, customer name, form values, or signed URL.

Short-lived retrieval authorization deserves care. The audit record identifies the evidence digest and access decision; recording the complete retrieval URL may leak a credential into the logging system. Small detail, large boundary.

This boundary matters.

Observability without document-shaped metric labels

Start with questions. Can intake preserve the original? Can an authorized agent obtain a representation? How long does each stage take? Are failures concentrated by renderer version or representation mode? Bounded labels such as stage, mode, outcome class, and renderer version answer those questions. Ticket ID, digest, user ID, filename, and tenant ID belong in restricted events or deliberately retained traces, not metric dimensions.

Count the series before shipping labels. Suppose a histogram uses 12 buckets and dimensions have 4 stages, 3 modes, 5 outcome classes, 2 regions, and 6 renderer versions. Ignoring sum and count series, the budget is:

printf '%s\n' $((12 * 4 * 3 * 5 * 2 * 6))
Enter fullscreen mode Exit fullscreen mode

That yields 8,640 bucket series for one metric family before replicas or extra dimensions. Add 10 tenants as a label and it becomes 86,400. Add ticket ID and the bound disappears. Correlation belongs in sampled traces and restricted audit events.

Retention math is equally plain. For a planning example, if 40,000 preview attempts per day each produce four structured events averaging 900 stored bytes after envelope fields, daily ingest is 144,000,000 bytes before indexing, replication, and compression. At 30 days, the raw arithmetic is 4.32 billion bytes. These are scenario inputs, not measured production figures; replace each one with an observed value before setting policy.

printf 'daily_bytes=%s\n' $((40000 * 4 * 900))
printf 'thirty_day_bytes=%s\n' $((40000 * 4 * 900 * 30))
Enter fullscreen mode Exit fullscreen mode

Keep audit and diagnostic retention separate. Audit policy follows legal and business requirements, while high-volume renderer diagnostics may expire sooner. Never probabilistically discard the authoritative intake, authorization, signature-validation, or review decision event. Sample repetitive successful spans if volume requires it, retaining errors and slow paths under a documented policy. The trade-off is explicit: lower volume reduces cost but weakens reconstruction of rare latency patterns.

Flattening is a derived publication step. It may help when a recipient needs a stable, non-interactive appearance, but its output gets a new digest, transformation record, tool and policy version, and pointer to the original. Never label it "the signed original."

Before release, build a corpus around the support workload: blank and completed forms, multiple and rotated pages, unusual page boxes, embedded fonts, large images, malformed inputs, password protection, visible signature appearances, and documents with and without cryptographic signatures. Screenshots are not a sufficient oracle. Check page count, expected appearance, source immutability, digest linkage, authorization behavior, and the independently computed validation result.

Canary a renderer or viewer by policy version rather than overwriting old derivatives. Compare bounded outcome and latency metrics, inspect only samples allowed by data policy, and retain metadata identifying the generator version. Rollback changes the active policy while evidence stays untouched.

Why reject universal image conversion?

I reject unconditional server-side image conversion as the primary representation here. It removes browser interaction, multiplies stored objects with page count, introduces a renderer into the critical path, and cannot establish cryptographic signature validity. Making it universal pays those costs even when authorized embedding is sufficient.

The option still has a valid use case. Passive images suit constrained clients, thumbnail grids, deliberately non-interactive review, or isolation policies that prohibit delivering original bytes to the browser. The decision changes when those requirements dominate. The invariant does not: retain the source, link every derivative by digest, and record validation separately from appearance.

A preview is useful evidence about what an agent could see. It is not the evidence object itself. Keeping that boundary sharp lets the system evolve its viewer, renderer, telemetry, and retention policy without silently rewriting a customer case's history.

References

Top comments (0)