Short answer: for a US/EU SaaS merging and splitting large case files, make the signed manifest the primary object and treat PDF rendering as a queued, observable stage. That choice preserves signature evidence when latency spikes, because a slow worker changes completion time rather than the identity of the bytes that were approved. A synchronous endpoint still has a place for a tiny preview; it is a poor boundary for a six-hundred-page filing.
The important design question is not which endpoint sounds fastest. It is what a reviewer can prove six months later: which source pages entered the bundle, which version was split out, and which bytes a signer actually saw.
The constraint is an evidence chain
Start with immutable input objects. For every uploaded document, record its digest, media type, page count, and source timestamp in a manifest. A merge creates a new manifest that points at those inputs; it does not rewrite them. A split stores the parent manifest digest and a range expressed against that manifest version. This prevents a cover-sheet insertion from silently changing “pages 10–12” into different exhibits.
Signature fidelity has two separate checks. Visual fidelity covers fonts, rotation, annotations, transparency, and page boxes. Evidence fidelity binds the exact bytes to the manifest, signer identity, approval event, and time. A visually accurate PDF without that binding is hard to defend. A perfectly logged file with a shifted stamp is still wrong.
Keep the manifest in canonical JSON: deterministic key order, UTF-8, explicit arrays, and timestamps in one agreed representation. Hash the canonical bytes, then store the hash beside the output object. JSON whitespace is not evidence; the canonical representation is. The PDF can contain metadata for humans, but the external record should remain authoritative for merge and split operations.
Here is a deliberately small record. It carries a routing hint for geography without pretending that a failover policy is automatic:
from dataclasses import dataclass
@dataclass(frozen=True)
class BundleOperation:
operation_id: str
tenant_id: str
parent_manifest_digest: str | None
input_digests: tuple[str, ...]
page_ranges: tuple[tuple[int, int], ...]
requested_region: str
idempotency_key: str
The region field is a policy input. US/EU teams still have to document retention, encryption-key ownership, and the behavior of a regional capacity loss. A route label in an API request is not a residency guarantee.
How should PDF endpoints handle case files when latency rises under load?
Separate admission latency from completion latency. The upload or submit call should acknowledge a durable job record quickly; the client can then observe state transitions such as accepted, running, succeeded, failed, or expired. Report p50, p95, and p99 by page-count and byte-size buckets. A 20 MB scan and a 600-page text bundle do not exercise the same path, so one blended percentile hides the queue that users feel.
Queue age is the leading signal. Track the oldest job, per-tenant backlog, worker utilization, memory high-water mark, retry count, and output-publication delay. Set a concurrency ceiling for each tenant and a global ceiling for the renderer. Backpressure should be explicit: return the job identifier and a retry-after hint for status polling, rather than holding a socket open until a renderer happens to finish.
I once assumed a low median meant a healthy service. It did not. A small Monday burst left the median unchanged while the p99 queue wait crossed our user-facing budget, because one tenant consumed every worker with image-heavy bundles. The fix was fair scheduling and a queue-age alarm, not a new PDF switch.
Measure twice.
This test harness records the two clocks separately:
import asyncio
import time
async def submit_and_measure(client, payload, key):
started = time.monotonic()
response = await client.submit(payload, idempotency_key=key)
accepted = time.monotonic()
job_id = response["job_id"]
while True:
state = await client.status(job_id)
if state["state"] in {"succeeded", "failed", "expired"}:
finished = time.monotonic()
return {
"admission_ms": (accepted - started) * 1000,
"completion_ms": (finished - started) * 1000,
"state": state["state"],
}
await asyncio.sleep(0.25)
Do not make a latency target that ignores fidelity. Gate a release on zero structural mismatches in the approved corpus and on a p99 completion budget for each workload bucket. I am not sure a universal pixel threshold can represent every legal stamp, so keep a human-review rule for signatures and record why a difference was accepted.
Where the endpoint contract meets storage semantics
An asynchronous API is useful only when its state is durable. Write the job record before acknowledging submission, and make the output write conditional on the operation identifier. A worker that outlives its lease must not replace a newer result. Notifications are hints; clients should be able to reconcile by reading job state and object metadata.
Idempotency belongs at the operation boundary. The same tenant, payload digest, and idempotency key should return the existing result, while a reused key with a different payload should be rejected. Without that rule, a client timeout can lead to two bundles and two competing audit histories.
Use browser and service APIs that expose bytes as bytes. The web Blob interface, for example, represents immutable raw data and supports reading a response as a Blob; that lets a client calculate a digest or hand the exact bytes to a signature verifier without converting through a lossy text layer. The interface is a building block, not a PDF renderer.
def verify_downloaded_bytes(blob_bytes: bytes, expected_digest: str, digest_fn):
actual = digest_fn(blob_bytes)
if actual != expected_digest:
raise ValueError("download digest does not match the signed manifest")
return blob_bytes
Keep object publication and manifest publication ordered. A consumer should never observe a manifest that points to an unavailable object, and it should never download an object whose digest is absent from the signed record. A short-lived “publishing” state makes that ordering visible to operators.
Fidelity testing is a corpus, not a screenshot
Build fixtures that reflect legal documents: rotated pages, embedded fonts, transparency, right-to-left text, annotations, malformed-but-readable files, and attachments with their own page labels. Compare output against approved references with both pixel tolerances and structural assertions for page count, MediaBox, CropBox, annotations, and extracted text. Save input bytes, renderer version, region, and comparison artifacts whenever a mismatch appears.
Large bundles also need memory tests. Stream uploads and downloads where the platform permits it; avoid constructing several full byte copies while merging. Run a soak test long enough to expose temporary-file leaks, then drain workers during deployment and verify that in-flight jobs resume from a durable state rather than restarting from an unknown page.
Your mileage may vary. A team that only needs a two-page browser preview may accept a synchronous call and a simpler operator surface. A legal-export workflow with bursty tenants, strict signatures, and six-hundred-page bundles should pay the operational cost of a queue, fair scheduling, and reconciliation. The unsuitable choice is the one whose failure mode cannot be explained to an auditor.
A rollout rule for US/EU SaaS teams
Begin with a shadow corpus and a single region, but keep region selection in every job record from day one. Introduce asynchronous finalization behind a feature flag while previews remain synchronous. During the migration, compare admission, queue, render, validation, and publication timings independently; otherwise an endpoint dashboard will blame the wrong stage.
The decision table is intentionally about constraints rather than brands:
| Constraint | Endpoint and storage shape | Main trade-off |
|---|---|---|
| Tiny interactive preview | Synchronous render with a strict timeout | Simple UX, but timeout coupling and duplicate retries |
| Final bundle with hundreds of pages | Durable asynchronous job plus immutable output | More queue operations, far better burst isolation |
| Signature-heavy evidence | Canonical manifest digest beside the PDF | Extra validation work, defensible provenance |
| Spiky multi-tenant traffic | Per-tenant quotas and fair scheduling | Lower single-tenant peak, predictable p99 |
| Regional policy | Region-aware workers and storage with explicit failover | Failover planning is operational work, not a checkbox |
Promote the design only after the audit record, downloaded-byte digest, and latency histograms agree for the same operation IDs. That correlation is the practical definition of fidelity under load.
Top comments (0)