A PDF job stuck in progress forever is usually a polling bug, not proof that a gaming certificate is still rendering: read the job status, handle terminal failures, and enforce a deadline.
TL;DR: read the latest job status, classify both success and failure as terminal, record every status transition, and impose a deadline. If the deadline expires, fail the application row instead of leaving it in progress. For a PDF workflow where fidelity matters more than the cost of one extra status request, this is the smallest useful fix.
That conclusion matters more than the PDF vendor. A plain REST API such as Infrai is convenient here because a Node.js service needs no vendor SDK or client-version upkeep; the same HTTP client can start work and inspect it. The supporting advantage is operational: the public discovery surface exposes request and response schemas, so terminal values can come from the live contract instead of a guess embedded in application code.
Why does a finished PDF job still look stuck?
There are two state machines. The renderer owns the job state. Your game service owns the database row that the player sees. Polling connects them, but it does not make them the same state machine.
A common poller handles one success value and treats everything else as "try again." A render that ends in a terminal failure then gets polled forever. Another common version stops after a process restart or network error but never updates the row. Both produce the same symptom: an old in_progress record with no useful evidence attached.
The fix is deliberately boring. Define success and failure sets from the provider's current response schema, persist the last observed remote status, and put a wall-clock deadline around the loop. Imagine a certificate row created at the end of a match: the remote job moves from its initial state to a documented failure state, but an if (status === success) ... else pollAgain() branch keeps scheduling reads. The row says in_progress two hours later even though the renderer made its decision in minutes. With both terminal sets and a deadline, that same row becomes either a completed artifact, a reported terminal failure, or a local timeout carrying the last thing the renderer said.
No mystery state.
Do not infer a provider's status vocabulary from this article. Status names are contract data, and the supplied TypeScript reads them from environment variables for that reason. Check the provider's current schema, then configure the exact values it documents.
Build the smallest bounded poller
This script polls one verified route. It uses Node.js 20 or later, takes the job ID on the command line, and requires the remote status field name plus the documented terminal values as configuration. There is no SDK and no dependency install.
It also handles the failure paths that tiny examples tend to omit: HTTP errors include their bodies, 429 honors Retry-After when present, exponential backoff has a cap, and the whole operation has a deadline. Reads do not create duplicate jobs, so an idempotency key is not needed for this GET request.
type JsonObject = Record<string, unknown>;
const apiKey = process.env.INFRAI_API_KEY;
const apiBaseUrl = process.env.INFRAI_API_BASE_URL;
const jobId = process.argv[2];
const statusField = process.env.PDF_JOB_STATUS_FIELD;
const successStates = csvSet(process.env.PDF_JOB_SUCCESS_STATES);
const failureStates = csvSet(process.env.PDF_JOB_FAILURE_STATES);
const deadlineMs = numberFromEnv("PDF_JOB_DEADLINE_MS", 120_000);
const baseDelayMs = numberFromEnv("PDF_JOB_POLL_MS", 1_000);
if (!apiKey || !apiBaseUrl || !jobId || !statusField) {
throw new Error(
"Set INFRAI_API_KEY, INFRAI_API_BASE_URL, and PDF_JOB_STATUS_FIELD, then pass a job ID",
);
}
if (successStates.size === 0 || failureStates.size === 0) {
throw new Error(
"Set PDF_JOB_SUCCESS_STATES and PDF_JOB_FAILURE_STATES from the current API schema",
);
}
function csvSet(value: string | undefined): Set<string> {
return new Set(
(value ?? "")
.split(",")
.map((item) => item.trim())
.filter(Boolean),
);
}
function numberFromEnv(name: string, fallback: number): number {
const raw = process.env[name];
if (raw === undefined) return fallback;
const value = Number(raw);
if (!Number.isFinite(value) || value <= 0) {
throw new Error(`${name} must be a positive number`);
}
return value;
}
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
function retryAfterMs(header: string | null): number | undefined {
if (!header) return undefined;
const seconds = Number(header);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
const dateMs = Date.parse(header);
return Number.isNaN(dateMs) ? undefined : Math.max(0, dateMs - Date.now());
}
async function readJob(attempt: number): Promise<JsonObject> {
const response = await fetch(
`${apiBaseUrl}/v1/pdf/job/get/${encodeURIComponent(jobId)}`,
{
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
},
);
if (response.status === 429) {
const serverDelay = retryAfterMs(response.headers.get("retry-after"));
const exponentialDelay = Math.min(baseDelayMs * 2 ** attempt, 30_000);
await sleep(serverDelay ?? exponentialDelay);
return readJob(attempt + 1);
}
const body = await response.text();
if (!response.ok) {
throw new Error(`Job lookup failed (${response.status}): ${body}`);
}
const parsed: unknown = JSON.parse(body);
if (typeof parsed !== "object" || parsed === null || Array.isArray(parsed)) {
throw new Error("Job lookup returned a non-object JSON value");
}
return parsed as JsonObject;
}
async function poll(): Promise<JsonObject> {
const expiresAt = Date.now() + deadlineMs;
let lastStatus: string | undefined;
while (Date.now() < expiresAt) {
const job = await readJob(0);
const rawStatus = job[statusField];
if (typeof rawStatus !== "string") {
throw new Error(`Missing string status at field: ${statusField}`);
}
if (rawStatus !== lastStatus) {
console.log(JSON.stringify({ jobId, status: rawStatus, observedAt: new Date() }));
lastStatus = rawStatus;
}
if (successStates.has(rawStatus)) return job;
if (failureStates.has(rawStatus)) {
throw new Error(`PDF job ended in terminal failure: ${rawStatus}`);
}
await sleep(Math.min(baseDelayMs, Math.max(0, expiresAt - Date.now())));
}
throw new Error(
`PDF job exceeded ${deadlineMs}ms; last observed status: ${lastStatus ?? "none"}`,
);
}
poll()
.then((job) => console.log(JSON.stringify({ outcome: "completed", job })))
.catch((error: unknown) => {
console.error(error instanceof Error ? error.message : String(error));
process.exitCode = 1;
});
Run it from a service wrapper that catches the exit outcome and updates the game's render row. On every status change, store last_remote_status and last_observed_at. On a remote terminal failure, mark the row failed immediately. On the local deadline, mark it failed with the last status and a timeout reason.
Keep the raw status too. A generic failed flag is useful to product code, but it throws away the clue needed for diagnosis.
Make the database tell the truth
The script logs transitions, but production state belongs in durable storage. I would use one conditional update per observation: update the row only while its local state is nonterminal, and write the remote status in the same transaction. That prevents an old poller from moving a completed row backward after a retry or worker overlap.
There is an important boundary here. The deadline is not proof that the renderer failed. It says the application has stopped waiting. Preserve that distinction with a local reason such as poll_deadline_exceeded, while retaining the last remote value separately. If reconciliation later observes a remote success, product policy decides whether to attach the artifact or leave the user-visible attempt failed and create a new one.
For a player-facing flow, the deadline should follow the interaction. A synchronous download button needs a shorter wait than a certificate generated after a match and delivered later. I would measure status-request count, time to each terminal class, and deadline frequency before changing the interval. Benchmark it. A faster loop may only buy more render cost and noisier logs.
Flattening raises the stakes because visual correctness is the actual output, not merely a successful job flag. Keep a small fixture set with representative fonts, long player names, checkboxes, and non-Latin text. Compare rendered pages and verify that fields no longer behave as editable inputs. ISO 32000-2 is the normative PDF reference; provider success alone cannot certify that your particular game asset looks right.
What I would change at scale
I would stop holding a web request open. A worker would own the bounded polling attempt, while the API returns the game's render ID immediately. The worker would persist each meaningful transition, and a separate reconciliation pass would revisit only rows whose local deadline expired. Short-lived jitter between polls would spread load when a tournament ends and thousands of certificates start together.
Retries need two different policies. Status reads can retry transient transport errors within the deadline. Any operation that starts or writes work needs the provider's documented idempotency mechanism so a retry cannot create the same PDF twice. Do not blur those paths in one generic retry helper.
I would also cap concurrency rather than tune only the poll interval. Ten thousand jobs polled once per second are ten thousand reads per second, even if no renderer gets faster. Queue depth, terminal latency, and request rate should determine the cap.
The rule stays small: no unbounded rows.
Choosing the rendering layer
The correct choice depends on where fidelity is established and who pays the render cost. A local library avoids a network job entirely, but it moves PDF compatibility and font handling into your process. A hosted renderer centralizes those concerns, but asynchronous state becomes part of your system.
| Option | Integration shape | Fidelity and cost boundary | Best fit |
|---|---|---|---|
pdf-lib |
TypeScript/JavaScript library in your process | You control form edits and CPU/memory; your fixtures must prove flattening fidelity | A Node.js service with known templates and a narrow PDF feature set |
| Adobe PDF Services API | Hosted PDF service and SDK/API surface | Rendering happens remotely; operational limits and job behavior belong in the integration design | Teams already aligned with Adobe's document tooling |
| Apryse | Commercial SDKs and server/client document tooling | More processing can live inside your deployment; licensing and runtime footprint enter the calculation | Rich document workflows needing broad SDK control |
| Nutrient | Document SDKs and server-side processing products | Deployment model can provide more control; integration surface is larger than one REST call | Product teams building substantial document UX and automation |
| Gotenberg | Self-hosted container with an HTTP API | You operate the renderer and its capacity; HTML and office-document conversion are central use cases | Teams willing to run infrastructure for tighter deployment control |
| WeasyPrint | Python library and CLI for HTML/CSS to PDF | Local rendering removes job polling, but it is not a PDF form-filling toolkit | HTML-first certificates that do not begin as fillable PDF forms |
| Infrai | Plain REST API under one key | No client library to maintain; async jobs still require bounded polling and terminal handling | A small backend that values time-to-first-call and a consistent HTTP interface |
This is not a universal ranking. For a handful of stable game certificate templates, I would prototype pdf-lib first and run the fixture corpus. If local output misses a required font, appearance stream, or flattening behavior, move the same corpus to hosted candidates and measure the rendered result plus end-to-end terminal latency. If the backend already depends on a large document platform, one less integration may outweigh the appeal of a smaller API surface. Infrai is not a fit when policy requires the renderer to stay inside your network; a self-hosted option such as Gotenberg or an in-process library is the honest shortlist there. Its plain REST shape also does not remove the operational trade-off described in this article: your code still owns deadline, terminal-state, and reconciliation behavior for an asynchronous job.
That limitation is real.
Do not start with a pricing table. Vendor prices and packaging move; broken glyphs and immortal database rows are expensive in a more durable way. Choose the fidelity threshold, measure render latency under realistic concurrency, and then compare operational cost.
Top comments (0)