A customer-support archive has one constraint that changes the design: invoice generation must keep up with the order stream without weakening the record needed years later. Short answer: preserve each accepted invoice as an immutable, self-contained archival PDF, compress its assets before rendering, encrypt the stored object rather than rewriting the PDF, and treat batch size as a measured control. Keep the order snapshot and a cryptographic digest beside it.
The tempting design is one request, one render, one upload. Easy. It also leaves throughput to chance. A better mental model is a bounded conveyor: claim a small batch, render with fixed resources, validate, store, then acknowledge. Every stage emits a duration and a count.
How should a small SaaS team compress and encrypt a long-term PDF archive?
Format sounds like a retention concern. It reaches backward into the hot path. Fonts, images, color profiles, metadata, and encryption consume CPU, memory, or bytes in motion. A worker that embeds a 4 MB logo in every invoice does not have a queue problem; it has an asset policy problem repeated thousands of times.
ISO 32000-2 defines PDF itself. PDF/A, standardized in the ISO 19005 family, adds constraints intended for long-term preservation. The operational distinction matters: a PDF that opens today is not automatically an archival PDF. Choose the precise conformance level from the records policy and validator capabilities. Never stamp a conformance claim into metadata merely because a renderer accepted an option. Validate the resulting bytes.
Claims need proof.
Here is the before/after diagram in words. Before: order database to render call to object storage, with success inferred from an HTTP status. After: immutable order snapshot to bounded queue to render to conformance validation to digest to encrypted object storage, with a manifest joining every transition. The second path has more boxes, and each box creates a retry boundary.
That costs work. It removes ambiguous failure.
Compression belongs before or during rendering. Downsample an invoice image only when the retained visual evidence permits it, subset embedded fonts where the format and license allow it, and avoid inserting the same oversized asset per page. Generic ZIP compression around an already compressed PDF adds work without addressing large inputs. Measure bytes per invoice and render time separately; a smaller file that takes much longer to produce can reduce the batch rate.
Encryption has a different boundary. PDF password encryption changes the document and can complicate preservation workflows. Storage-layer authenticated encryption protects the archived object while leaving the validated PDF bytes stable inside that envelope. NIST SP 800-38D specifies GCM, an authenticated-encryption mode. Key identifiers, rotation, authorization, and restore tests still require design. An algorithm name is not a key-management plan.
Keep those boundaries separate.
Make the worker contract boring
The API should accept stable facts, not a request to query mutable business state later. Pass an order snapshot identifier, template revision, locale, currency, and archive policy revision. Return a durable job identifier immediately. Rendering stays asynchronous, while support tooling reads job state instead of holding a connection open for a batch.
Idempotency is the hinge. Derive an operation key from stable business inputs, but keep the archive digest separate because it describes actual output bytes. If the same order and template revision are claimed twice, the worker may reuse the completed manifest. If inputs differ, it creates a new record instead of silently replacing history.
This TypeScript example keeps the renderer, validator, object store, and metrics sink generic on purpose:
import { createHash } from "node:crypto";
type InvoiceTask = {
orderSnapshotId: string;
templateRevision: string;
locale: string;
currency: string;
archivePolicyRevision: string;
};
type Dependencies = {
renderPdf: (task: InvoiceTask) => Promise<Uint8Array>;
validateArchivePdf: (pdf: Uint8Array) => Promise<void>;
putEncryptedObject: (key: string, pdf: Uint8Array) => Promise<void>;
observe: (name: string, value: number, labels: Record<string, string>) => void;
};
export async function archiveInvoice(task: InvoiceTask, dependencies: Dependencies) {
const startedAt = performance.now();
const operationKey = createHash("sha256")
.update(JSON.stringify(Object.values(task)))
.digest("hex");
const pdf = await dependencies.renderPdf(task);
await dependencies.validateArchivePdf(pdf);
const sha256 = createHash("sha256").update(pdf).digest("hex");
await dependencies.putEncryptedObject(`invoices/${operationKey}.pdf`, pdf);
dependencies.observe("invoice_archive_duration_ms", performance.now() - startedAt, {
outcome: "stored",
policy_revision: task.archivePolicyRevision
});
return { operationKey, sha256, bytes: pdf.byteLength };
}
Do not log customer names, addresses, invoice numbers, or object keys. Labels with unbounded values damage metric systems and can leak record data. Put job identifiers in access-controlled structured logs or traces, and keep metric labels bounded to states and revisions. W3C Trace Context defines interoperable propagation; it does not make sensitive attributes safe to collect.
There is deliberately no fixed batch size. A copied number would be theater because image mix, font loading, renderer memory, storage latency, and worker limits determine useful concurrency. Begin with a conservative bound. Raise it while watching queue age, completed documents per minute, p95 stage duration, validation failures, process memory, and retry volume. Stop when a resource saturates or tail latency rises sharply.
Observe stages, not one success rate
A single failure counter cannot tell an operator whether to retry. Split outcomes by claim, render, validate, store, and manifest commit. Alert on customer impact: oldest queued item breaching its service objective, sustained completion rate below arrival rate, or repeated terminal failures. CPU usage alone is context.
Rate tells half the story.
The manifest is the audit spine. Record the operation key, input snapshot reference, template revision, policy revision, output digest, byte count, conformance result, storage version reference, timestamps, and terminal state. Never store a secret key or plaintext sensitive fields there.
| Decision | Default direction | Throughput consequence | Verification |
|---|---|---|---|
| Archival profile | Match a documented PDF/A profile to record needs | Validation adds work | Validate every output independently |
| Asset compression | Normalize reusable assets before peak batches | Fewer bytes; preprocessing costs CPU | Track bytes and render duration |
| Encryption | Encrypt stored objects under an access policy | Adds storage-path latency | Restore and authenticate samples |
| Batch size | Bound concurrency per worker | Prevents memory spikes | Load test with the real invoice mix |
Test the ugly path. Kill a worker after upload but before manifest commit. Replay the same task. Feed the validator a corrupted fixture. Deny storage authorization. Restore an encrypted object with its documented key version and compare its SHA-256 digest with the manifest. These tests prove retry semantics without inventing a production incident.
What about signatures and searchable text?
A digest detects byte changes when compared with a trusted manifest, but it is not a digital signature and does not identify a signer. If policy requires signed invoices, define signature handling separately and preserve its validation material. Signing changes ordering and may limit later transformations. That trade-off deserves its own tested policy.
Treat search indexes as rebuildable derivatives. Keep the authoritative PDF and snapshot independent from the index. Restrict indexed fields, attach the same deletion and access rules, and test that permitted metadata finds an invoice without exposing its contents through metrics. If source data already contains invoice text, avoid adding OCR by habit; it spends CPU and introduces another output to verify.
Can a small team operate all of this?
Yes, if the team limits variation. One archival profile, one manifest schema, one queue contract, and a short list of terminal states are easier to own than per-customer render modes. Template revisions should be immutable. Deploy a renderer revision against representative fixtures, compare validation and visual output, then canary it on bounded traffic.
The approach is not free of limitations. Validation reduces raw throughput. Immutable records use more storage than overwrite-in-place workflows. Storage-layer encryption cannot replace document-level controls when a recipient must carry protection outside the archive. A tiny workload with no retention duty may not justify the full pipeline; a synchronous render plus documented backup policy can be proportionate.
That is a real trade-off.
The team still needs an escape hatch: pause claims without deleting queued work, drain a worker revision, re-run validation without re-rendering, and rebuild a manifest projection from durable records. Those controls beat a button that blindly retries everything.
The durable choice is a process, not a PDF setting. Preserve stable inputs, validate final bytes, store them behind an encryption boundary, record a digest and manifest, and tune bounded batches against queue age plus stage telemetry. That keeps support invoices recoverable while showing operators that the archive is falling behind.
Top comments (0)