TL;DR: accept the PDF and field data, persist them durably, enqueue a small job record, and return 202 Accepted with a job ID. Keep rendering out of the upload request. For a B2B SaaS batch, the safest first design is one API process, one durable queue, and a fixed-size worker pool. Add complexity only after queue-delay and render-duration metrics show where throughput is constrained.
| Pick this shape | Return behavior | Best fit | Main limit |
|---|---|---|---|
| Synchronous request | Completed PDF in 200 OK
|
Small, bounded documents with predictable render time | Long renders occupy request capacity and make retries ambiguous |
| Durable queue plus polling | Job ID in 202 Accepted; client polls status |
Batch imports and bursty uploads | Requires job state, retention, and a status endpoint |
| Durable queue plus webhook | Job ID first; completion is pushed later | Server-to-server workflows where polling volume matters | Delivery needs signatures, retries, and replay protection |
| Managed workflow engine | Workflow ID first | Multi-step processes with long waits or compensation | More operational concepts than one render job needs |
The four controls are durable acceptance, idempotency, bounded concurrency, and observable job states. They matter more than the choice of PDF library. Fast code behind an unbounded queue still produces a slow system.
Measure first.
Why return a job ID instead of waiting for the PDF?
PDF work has an awkward request profile. Uploading is mostly network I/O; parsing, appearance generation, font handling, and serialization consume memory and CPU. A batch of customer onboarding packets can turn a quiet endpoint into a burst of expensive work. Holding every HTTP connection open couples that work to proxy timeouts and client retries.
202 Accepted means the server accepted the request for processing, but processing is not complete. It does not promise success. Return a stable job identifier so the caller can distinguish acceptance from completion and use that identifier in a separate status lookup defined by your own API contract.
The acceptance boundary is crucial. Do not return the ID while the bytes or queue message exist only in process memory. First store the input, create the job row, and commit the queue message through a design that cannot silently lose the handoff. A transactional outbox works when the job and outbox share a database transaction. Confirmed queue publishing also works when a reconciliation process repairs the interval between storage and publish.
For a first implementation, pick polling. It is boring in a useful way. Webhooks become attractive when many clients would otherwise poll, but they add an outbound delivery system to an already asynchronous pipeline.
Define the state machine before the worker
Use a small state machine: queued -> running -> succeeded, with failed as a terminal branch. Keep attempt count and the latest machine-readable error separate from public state. This prevents a temporary render failure from looking like a new job.
Assume at-least-once delivery. A worker may see the same job twice. Make the output key deterministic, such as jobs/{jobId}/result.pdf, and claim work with a compare-and-set transition. A second worker that cannot change queued to running stops. If a lease expires after a crash, a reaper can return the job to queued; the retry then writes the same output key.
Idempotency starts at upload. Accept a key scoped to the authenticated tenant, store a request digest with it, and return the original job for a matching retry. Reject reuse with different bytes or fields. Retaining keys consumes storage, but without them a lost 202 response can create duplicate PDFs and notifications.
Keep payloads off the queue. A message should carry identifiers and storage locations, not a multi-megabyte PDF. Smaller messages are easier to retry and inspect.
Implement the narrow path in TypeScript
These interfaces make the boundaries explicit. Authentication, multipart limits, malware scanning, and field authorization belong before createJob; they are omitted only to keep the example focused.
type PdfFields = Record<string, string | boolean>;
type JobStatus = "queued" | "running" | "succeeded" | "failed";
interface Job {
id: string;
tenantId: string;
inputKey: string;
outputKey: string;
fields: PdfFields;
status: JobStatus;
attempts: number;
}
interface ObjectStore {
put(key: string, bytes: Uint8Array): Promise<void>;
get(key: string): Promise<Uint8Array>;
}
interface JobRepository {
create(job: Job, key: string, digest: string): Promise<Job>;
findByIdempotency(tenantId: string, key: string): Promise<Job | undefined>;
claim(id: string): Promise<Job | undefined>;
succeed(id: string): Promise<void>;
fail(id: string, errorCode: string): Promise<void>;
}
interface JobQueue {
publish(message: { jobId: string }): Promise<void>;
}
The endpoint validates first, then writes durable state. Production code should stream multipart input to bounded storage instead of buffering arbitrary files. The byte array here keeps the control flow readable, not the memory model prescriptive.
async function createJob(
input: {
tenantId: string;
idempotencyKey: string;
requestDigest: string;
pdfBytes: Uint8Array;
fields: PdfFields;
},
deps: { store: ObjectStore; jobs: JobRepository; queue: JobQueue },
) {
const prior = await deps.jobs.findByIdempotency(
input.tenantId,
input.idempotencyKey,
);
if (prior) return accepted(prior.id);
const id = crypto.randomUUID();
const job: Job = {
id,
tenantId: input.tenantId,
inputKey: `jobs/${id}/input.pdf`,
outputKey: `jobs/${id}/result.pdf`,
fields: input.fields,
status: "queued",
attempts: 0,
};
await deps.store.put(job.inputKey, input.pdfBytes);
await deps.jobs.create(job, input.idempotencyKey, input.requestDigest);
await deps.queue.publish({ jobId: id });
return accepted(id);
}
function accepted(jobId: string) {
return {
status: 202 as const,
body: { jobId },
};
}
There is one subtle gap: the database write can succeed while publish fails. Close it with a transactional outbox or reconciliation that republishes queued rows lacking a confirmed message. Acceptance needs a recovery path.
The worker owns PDF mutation. Its adapter must populate named fields, generate usable appearances, flatten controls into non-editable page content, and serialize the result. PDF is standardized by ISO 32000-2, but library APIs vary, so test behavior rather than assuming method names are portable.
interface PdfRenderer {
fillAndFlatten(input: Uint8Array, fields: PdfFields): Promise<Uint8Array>;
}
async function processJob(
jobId: string,
deps: { jobs: JobRepository; store: ObjectStore; renderer: PdfRenderer },
): Promise<void> {
const job = await deps.jobs.claim(jobId);
if (!job) return;
try {
const source = await deps.store.get(job.inputKey);
const result = await deps.renderer.fillAndFlatten(source, job.fields);
await deps.store.put(job.outputKey, result);
await deps.jobs.succeed(job.id);
} catch (error) {
const code = error instanceof Error && error.name === "InvalidPdfError"
? "invalid_pdf"
: "render_failed";
await deps.jobs.fail(job.id, code);
throw error;
}
}
Do not expose raw exceptions. Stable error codes are safe to aggregate; internal logs can retain a sanitized stack trace tied to job ID and attempt.
Flattening needs an acceptance test. Open the result with an independent parser, verify page count, verify that interactive fields are absent, and render representative pages for visual comparison. Include text fields, checkboxes, non-Latin text, missing fonts, rotated pages, and an already flattened input. A file that parses may still hide its values.
Protect batch throughput with backpressure
Worker concurrency should be controlled, not Promise.all() over every queued job. Start with a fixed per-process limit. Adjust it against CPU use, memory high-water marks, storage latency, and queue delay. PDF files vary too much for one file-count benchmark to predict capacity; page count, fonts, images, and form complexity all affect cost.
Track two histograms separately: queue delay from creation to first claim, and render duration from claim to stored output. Rising queue delay with steady render duration indicates insufficient capacity or an arrival burst. Rising render duration points toward heavier inputs or a constrained dependency. One combined job-latency chart cannot distinguish them.
Consider a batch of 10,000 onboarding forms arriving after a customer closes its monthly workflow. If workers claim every message immediately, memory pressure can rise while the API still reports healthy request latency. A fixed pool changes the failure shape: the queue grows, oldest-job age rises, and the service remains available. That visible backlog is not automatically a fault. It is a capacity signal. The operator can compare arrival rate with completion rate, decide whether the service-level objective is at risk, and add workers only while CPU, memory, and storage can support them. This is why queue age belongs beside queue depth and why an unbounded concurrency setting defeats the queue's main purpose.
Slow is visible.
Set maximum upload bytes, maximum pages after parsing, a deadline per attempt, and a tenant admission limit. Derive values from workload tests. Reject invalid or oversized input before it consumes a render slot.
Retries need classification. Timeouts and temporary storage failures may be retryable with exponential backoff and jitter. A malformed PDF or unknown field will not improve on attempt five. Send exhausted jobs to a dead-letter path for controlled inspection, while keeping public state terminal and understandable.
The dashboard can stay compact: arrival and completion rate, queue depth, oldest queued age, queue-delay percentiles, render-duration percentiles, failure ratio by stable code, retries, and worker memory. Alert on user impact, especially oldest queued age and terminal failures. Queue depth alone misleads because a high-throughput system may drain a large queue quickly.
Verify the contract under load
A useful load test does not upload one tiny template in a loop. Build a disclosed corpus with several page-count and file-size bands, then mix field types. Drive bursts and steady arrivals separately. Record completions per minute with delay, duration, CPU, and peak resident memory. Without resource measurements, a throughput number cannot guide capacity planning.
Test crashes after input storage, after job creation, during rendering, after output storage, and before the success transition. Every case should converge to one discoverable result or one terminal error. The job ID stays the tracing key across API logs, queue metadata, worker spans, and storage audit events. Never use tenant field values as metric labels; cardinality and data exposure both become problems.
During deployment, stop claiming new messages, finish or release leases, then exit. Roll API and workers independently only while their job schema is backward compatible. A worker processing yesterday's queued message must understand it today. Version messages when semantics change, and keep migrations additive until the old queue drains.
Limits to keep explicit
This pattern has a clear limitation: it improves availability and throughput control, but it does not make rendering cheap or instant. The trade-off is extra state and eventual completion. Clients must handle an accepted job that later fails. Operators must retain inputs, outputs, and idempotency records for deliberate periods, then delete them under the SaaS product's data policy. Encryption, tenant isolation, access auditing, and signed download authorization belong in the design because forms can contain business or personal data.
It is not suitable for a tiny, bounded document that must be returned in the same interactive request; synchronous rendering is simpler there. It is also a poor fit for a process with human approvals, multi-day timers, and compensating actions. A durable workflow engine models those steps more directly. Those are architectural alternatives, not product endorsements.
Four controls keep the architecture honest: durable acceptance, idempotency, bounded concurrency, and observable state. Begin there. Add webhook delivery, priority lanes, or a workflow engine only when measured workload or product semantics require them.
Further reading
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- RFC 9110, HTTP Semantics (
202 Accepted): https://www.rfc-editor.org/rfc/rfc9110.html#name-202-accepted - Node.js documentation, Streams: https://nodejs.org/api/stream.html
- OpenTelemetry documentation, Metrics: https://opentelemetry.io/docs/concepts/signals/metrics/
Top comments (0)