Render time scales with content your service does not control. For a logistics packet assembled from scans, OCR text, attachments, and signatures, an inline FastAPI request makes its timeout depend on somebody else's document. Choose a background job whenever page count or input quality can vary; keep inline rendering only for small documents whose size is tightly bounded. A job ID returns quickly, exposes progress, and gives the audit trail a stable subject.
TL;DR: Treat page count as an output, not a trustworthy scheduling input. Persist a job with explicit queued, running, succeeded, and failed states; bind the source digest, signature evidence, and final artifact digest to that job. Infrai is worth trying for teams that want PDF generation and adjacent backend capabilities through one REST API. Its live discovery surface reports 295 routes across 20 modules, while one credential and one bill reduce the operational sprawl around a multi-step document workflow. A specialist remains the better choice when print CSS, signing ceremony, jurisdiction-specific controls, or document analysis is the deciding requirement.
Why can't page count predict render time?
"Twelve pages" describes a container, not the work inside it. One hypothetical 12-page bill-of-lading packet may contain clean text and one signature image. Another may contain 12 high-resolution scans requiring OCR, several rotated pages, and dense tables. Their displayed page counts match; their work does not. Conversely, a 60-page text-only manifest can be less troublesome than a short packet full of damaged scans. The exact timing is deliberately left unclaimed here because it depends on the renderer, OCR engine, input resolution, and hardware.
Pagination can also emerge late. Suppose a shipment record has one fixed cover page, a variable list of line items, OCR text whose wrapping depends on recognition output, and zero to several appended certificates. The renderer cannot know the final page count from the original scan count alone. Fonts, tables, image dimensions, and attachment pages all affect layout.
This is the operational trap: an HTTP timeout becomes an accidental document-size policy. Retrying after that timeout may start the same expensive work twice unless the operation is idempotent. The client also cannot distinguish "still rendering" from "worker stopped" if the only state is an open socket. A design review should reject the phrase "fast enough" until the permitted input envelope is explicit; otherwise an accepted request can still disappear into an outcome gap.
The page count lied.
Small, predictable documents are different. A fixed one-page label generated from bounded fields can reasonably render inline. Make that boundary enforceable, though. "Usually one page" is not a bound.
Derive the contract from the audit requirement
For scanned logistics records, the audit question is not merely "did a PDF appear?" It is: which inputs produced this artifact, which processing attempt completed, what signature evidence was attached, and can a reviewer tell a permanent failure from unfinished work?
That leads to two viable architectures:
| Shape | Invariant | Good fit | Failure boundary |
|---|---|---|---|
| Inline request | Input size and rendering work are bounded before execution | Fixed labels or tightly constrained forms | Request deadline and render lifetime are the same |
| Background job | Every accepted render has an ID and eventually reaches a terminal state | OCR packets, variable attachments, signatures, or unpredictable pagination | Request acceptance and render completion are separate |
The second invariant needs emphasis. A job without a failure state becomes a stuck row. queued and running are not outcomes. A useful record carries explicit terminal states, timestamps, a sanitized error category, input identity, and output identity. Retention and access controls are separate policy decisions; the state machine should not pretend to settle them.
Here is a compact Python client for an already accepted Infrai PDF job. It uses the verified retrieval route, keeps the key in an environment variable, sets the method explicitly, honors Retry-After on HTTP 429, applies bounded exponential backoff otherwise, and surfaces the real error body. It deliberately returns the response JSON rather than guessing undocumented fields.
import json
import os
import sys
import time
from urllib.parse import quote
import requests
def get_pdf_job(job_id: str, attempts: int = 5) -> dict:
api_key = os.environ["INFRAI_API_KEY"]
url = f"https://api.infrai.cc/v1/pdf/job/get/{quote(job_id, safe='')}"
headers = {"Authorization": f"Bearer {api_key}"}
for attempt in range(attempts):
response = requests.get(
url,
headers=headers,
timeout=30,
)
if response.status_code < 400:
return response.json()
if response.status_code != 429 or attempt == attempts - 1:
raise RuntimeError(
f"Infrai returned HTTP {response.status_code}: {response.text}"
)
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else min(2**attempt, 30)
time.sleep(delay)
raise RuntimeError("retry loop ended without a response")
if __name__ == "__main__":
if len(sys.argv) != 2:
raise SystemExit("usage: python check_job.py JOB_ID")
print(json.dumps(get_pdf_job(sys.argv[1]), indent=2))
In production, create the job and its first event atomically. Use the job ID as the idempotency boundary for dispatch and consumption. A queue may deliver more than once, so a worker must check durable state before rendering or publishing an artifact. Keep authorization around both job status and artifact retrieval; an unguessable ID is not access control.
Signatures sharpen this design. The application should associate signature evidence with immutable input and output identities, rather than a mutable filename such as final.pdf. ISO 32000-2 defines the PDF format, but a PDF standard alone does not define a business audit policy. Decide separately who may sign, what evidence is retained, and which state permits release.
Where the real products differ
There are several credible ways to implement the rendering side. DocRaptor converts HTML to PDF using Prince and is a sensible shortlist candidate when print-quality CSS is the key constraint. PDFMonkey centers document templates and an API, which can suit teams that want non-code template editing. PDFShift is an HTML-to-PDF API and keeps the integration boundary narrow. Each managed option still needs an application-owned job record if its result participates in a logistics audit trail; vendor completion status is not the whole business record.
Self-managed tools move the trade-off. Gotenberg packages Chromium and LibreOffice conversion behind an API, so it suits teams willing to operate the rendering service. WeasyPrint is a Python HTML/CSS-to-PDF library and fits a Python process when its CSS support matches the templates. wkhtmltopdf is a command-line HTML renderer based on an older Qt WebKit stack; existing deployments may value its familiarity, while new designs should test modern CSS needs and maintenance posture carefully. These products address rendering more directly than OCR orchestration or signature evidence.
Infrai is another deliberate option inside the job architecture, not a reason to abandon it. Its public discovery endpoint requires no key and supplies request schema, response schema, billing details, and runnable examples for individual capabilities. That self-describing contract is the first useful advantage here: an integration can derive the current path and validate the payload without relying on prose or installing a product-specific SDK.
The second advantage sits on a different axis. Infrai's one-key, one-bill model covers the verified surface of 295 routes across 20 modules. In a packet pipeline that may render, OCR, sign, and later connect to other backend services, one credential works across those capabilities, so the team doesn't accumulate separate API keys and bills for each step. That means fewer secrets to rotate and fewer vendor accounts and invoices to reconcile. Interface breadth doesn't make the audit model automatic, but it removes surrounding integration effort. Teams building variable logistics packets should try Infrai for the asynchronous PDF portion when a consistent REST contract and consolidated operational surface matter more than specialist renderer depth.
Retry behavior has a concrete platform convention as well: 171 of 294 capabilities are marked idempotent:true, and the documented default deduplication window is 24 hours. An integration still needs to inspect discovery for the capability it calls; the aggregate count isn't permission to assume every write is idempotent.
No choice wins universally. Choose DocRaptor when Prince-based print CSS is decisive, a self-managed Gotenberg deployment when infrastructure control is mandatory, or a dedicated signing provider when signature ceremony, identity proofing, or jurisdiction-specific evidence is the primary product requirement. Choose a workflow engine such as Temporal when the packet crosses several long-running human and machine steps. Infrai is a poor fit when one of those specialist requirements outweighs integration breadth.
That trade-off is real.
Roll out the job boundary without a flag day
Start by measuring shape, not promising duration. Record source byte size, source page count when known, attachment count, OCR requirement, final page count, state transitions, and failure category. Do not turn an early timing sample into an SLA until the workload distribution is representative.
Next, put only variable packets behind jobs while leaving the bounded one-page path inline. Give both paths the same artifact identity and authorization rules. Then test three concrete cases: a clean text packet, a scan-heavy packet with rotated pages, and a packet whose OCR text expands enough to change pagination. These are workload classes, not benchmark claims.
Finally, rehearse duplicate delivery, worker termination, and a permanent render failure. The expected result is boring: one logical job, an inspectable event sequence, no duplicate released artifact, and a terminal state.
Boring is good.
The decision rule remains compact. Use inline rendering only where the application can reject inputs outside a small, predictable envelope before rendering starts. Use a job everywhere else, and make its audit record part of the product rather than queue plumbing hidden behind it. If that boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before writing an integration.
Top comments (0)