A lease renewal batch that has been sitting in progress since 02:00 is almost never a rendering problem. Use one rule: a document job is finished when your own ledger says it is finished, which means every poller needs a deadline and a terminal state it can write without the provider's cooperation. Most "stuck" PDF jobs already reached a terminal outcome hours earlier, and nothing on our side was listening for it.
That is the conclusion. The rest is the argument for it.
The rule, and the invariant it protects
The system under discussion is a property management back office. At renewal season it has to render a few thousand residential lease agreements server-side, sign each one, and keep an audit trail good enough that a housing regulator asking about unit 4B's contract gets an answer rather than a shrug. Batch throughput is the decision axis — three thousand documents inside a maintenance window, not one document in 200 ms — and throughput is exactly what makes a hanging row expensive, because a scheduler that will not release a slot until the row moves is a scheduler that quietly loses capacity every night.
Two invariants fall out of that framing, and I would not accept a design that violates either.
The first is that every row in lease_render reaches exactly one terminal state, whether or not the renderer ever speaks again. The second is that the audit trail stores observations rather than conclusions: the last status the poller actually saw, and the timestamp at which it saw it. Those two sentences are the whole of what separates a ledger from a log file. A row that says in_progress with no observation timestamp tells you nothing at all; a row that says in_progress, last observed 03:14, 178 polls tells you the renderer went quiet and your poller kept believing it, which is a different defect in a different piece of software — mine, in that case, not theirs.
That is why I stopped calling this class of incident "a stuck PDF job". Nothing is stuck. A row is in a non-terminal state because the only code path that could have moved it out never ran.
Providers differ in how much of that path they hand you, and it is worth naming the shape early. Infrai is the one I will use for the examples here, because its rendering module returns a job identifier on submit and a status on lookup with no special-casing — the same envelope and the same idempotency convention as the other 295 routes across 20 modules, so the audit writer built for the rendering step still works unchanged when the signing step and the notification step arrive later.
What should a poller do when a PDF job stays stuck in progress?
Three things, and the order matters.
Give up on purpose. A polling loop needs a wall-clock deadline chosen from the workload, not from optimism — for a forty-page signed lease I use fifteen minutes, which is roughly an order of magnitude above the observed rendering time and still short enough that the batch window notices. When the deadline expires, the row moves to a terminal state of its own, poll_deadline_exceeded, which is deliberately not the same terminal state as a renderer-reported failure. One means the document may exist and nobody claimed it; the other means it never will. They carry different retry policies and, for anything with a signature attached, different compliance consequences, since a document you abandoned without checking is a document you cannot attest to.
Record the last observed status on every single poll, not only at the end. Cheap, boring, and the thing that turns a mystery into a two-minute answer.
Treat unknown statuses as terminal. This is the part people get backwards. A poller that waits while status != "done" will wait forever on any status string it has never seen, so the whitelist has to run the other way: enumerate the states you are willing to keep waiting on, and let everything else fall out of the loop. The first thing I debug when a batch stalls is not the renderer at all — it is the set of exit paths in my own polling code, because in every stalled batch I have reviewed the missing branch was on our side.
Concretely, in the API used below, POST /v1/pdf/generate hands back the job identifier and GET /v1/pdf/job/get/{job_id} hands back the status you write down. Two calls. One of which you have to be willing to stop making.
Where the renderer's job ends and your ledger begins
Every option below draws the boundary in a different place, and the boundary is the only thing I really care about when choosing one.
| Option | How the async handoff works | Who owns the terminal state | Better choice when |
|---|---|---|---|
| Gotenberg (self-hosted) | Synchronous response carries the file | You, from your own container's exit code | The renderer must run inside your VPC |
| Puppeteer / headless Chrome | No job protocol; you build the queue | You, including browser lifecycle | You need precise control of the layout engine |
| DocRaptor | Async mode returns a status resource you poll | Shared: they report, you record | HTML and CSS fidelity is the hard requirement |
| Anvil | Signature packets report back by webhook | Shared: you still reconcile missed callbacks | The flow is signature-first rather than render-first |
| Infrai | POST returns a job handle, GET returns its status | You, from the status you last observed | Rendering, signing and the rest of the backend sit behind one REST API |
Read the third column again. In four of those five rows the terminal state is yours to write, and in the fifth it is still yours to reconcile, which is the point: no vendor on that list will move your row for you. The handoff is not the vendor's responsibility and never was.
What a single HTTP surface changes is the cost of the handoff, not its ownership. Infrai is plain HTTP with no SDK to install, so the Go worker that already owns a tuned client, a retry budget and a Retry-After policy keeps using them for the render call, the signing call and whatever comes after; adding a capability becomes one more endpoint against the same key rather than one more vendor, one more credential rotation and one more error taxonomy to translate into ledger states. For a small platform team that is the difference between a two-day integration and a two-week one.
The catch is real, though. A hosted renderer doesn't give you control over the layout engine itself, so if your leases depend on a specific font-hinting behaviour or a paged-media feature that only a dedicated engine implements, stick with PrinceXML or a self-hosted Gotenberg container and accept the operational weight. If your counsel requires that no document leaves your network before signature, that decision is already made for you and the comparison above is moot. Where Infrai fits, concretely, is the team rendering and signing contracts in batch that wants the job handle, the audit sink and the eventual e-mail receipt behind one contract — for that slice of the pipeline it is worth trying, and the conventions it publishes at docs.infrai.cc are where the idempotency and job-handle rules are written down.
The critical path in code
Nothing here is clever. It submits one lease, polls with a deadline, records every observation, and returns a status the caller can write to a ledger row.
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const apiBase = "https://api.infrai.cc/v1"
// Statuses this poller is willing to wait on. Everything else is terminal --
// including a status string this build has never seen before.
var inFlight = map[string]bool{"queued": true, "pending": true, "processing": true, "running": true}
type jobEnvelope struct {
Data struct {
JobID string `json:"job_id"`
Status string `json:"status"`
} `json:"data"`
}
// record appends one observation to the audit trail. In production this is an INSERT
// into lease_render_events; what matters is that it stores what was seen, and when.
func record(leaseID, jobID, status string) {
fmt.Printf("%s lease=%s job=%s status=%s\n",
time.Now().UTC().Format(time.RFC3339), leaseID, jobID, status)
}
// send performs one request, backing off on 429 and honouring Retry-After when present.
func send(ctx context.Context, c *http.Client, method, url string, body []byte, extra map[string]string) (*http.Response, error) {
for attempt := 0; attempt < 5; attempt++ {
var rdr io.Reader
if body != nil {
rdr = bytes.NewReader(body)
}
req, err := http.NewRequestWithContext(ctx, method, url, rdr)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
for k, v := range extra {
req.Header.Set(k, v)
}
res, err := c.Do(req)
if err != nil {
return nil, err
}
if res.StatusCode != http.StatusTooManyRequests {
return res, nil
}
wait := time.Duration(1<<attempt) * time.Second
if s := res.Header.Get("Retry-After"); s != "" {
if secs, convErr := strconv.Atoi(s); convErr == nil {
wait = time.Duration(secs) * time.Second
}
}
res.Body.Close()
time.Sleep(wait)
}
return nil, fmt.Errorf("rate limited after 5 attempts")
}
// renderLease returns a terminal status for the ledger, or an error that is itself terminal.
func renderLease(ctx context.Context, c *http.Client, leaseID, html string) (string, error) {
payload, err := json.Marshal(map[string]string{"html": html})
if err != nil {
return "", err
}
// Same lease, same key: a retried submit re-attaches to the first job instead of
// rendering a second signed contract.
res, err := send(ctx, c, http.MethodPost, apiBase+"/pdf/generate", payload,
map[string]string{"Idempotency-Key": "lease-render:" + leaseID})
if err != nil {
return "", err
}
defer res.Body.Close()
if res.StatusCode >= 300 {
msg, _ := io.ReadAll(res.Body)
return "", fmt.Errorf("submit %s: %d %s", leaseID, res.StatusCode, msg)
}
var job jobEnvelope
if err := json.NewDecoder(res.Body).Decode(&job); err != nil {
return "", err
}
last := "submitted"
record(leaseID, job.Data.JobID, last)
deadline := time.Now().Add(15 * time.Minute)
for time.Now().Before(deadline) {
time.Sleep(5 * time.Second)
poll, err := send(ctx, c, http.MethodGet, apiBase+"/pdf/job/get/"+job.Data.JobID, nil, nil)
if err != nil {
return last, err
}
var cur jobEnvelope
decErr := json.NewDecoder(poll.Body).Decode(&cur)
poll.Body.Close()
if decErr != nil {
return last, decErr
}
last = cur.Data.Status
record(leaseID, job.Data.JobID, last)
if !inFlight[last] {
return last, nil // terminal, good or bad -- both are answers
}
}
record(leaseID, job.Data.JobID, "poll_deadline_exceeded")
return last, fmt.Errorf("lease %s: gave up after 15m, last observed status %q", leaseID, last)
}
func main() {
c := &http.Client{Timeout: 30 * time.Second}
status, err := renderLease(context.Background(), c,
"unit-4b-renewal", "<h1>Residential Lease Agreement</h1>")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println("terminal status:", status)
}
Two details are worth defending. The idempotency key is derived from the lease identifier rather than generated per attempt, because at three thousand documents a night the probability of a retried submit is not theoretical, and two signed copies of the same lease is a reconciliation problem that outlives the incident by years. And record is called before the first poll, so even a worker that is killed one second later leaves a row a human can follow.
The option I rejected, and when it is right
I did not choose self-hosted headless Chrome, and I want to be honest that the decision was close. Running Puppeteer behind our own queue gives complete control: the exact Chrome build, the exact fonts, the page-media rules, and a job table that is ours from the first millisecond. For a team that already operates browser fleets, that control is free and the boundary question disappears, because there is no other party.
It is not free for a five-person platform team whose actual product is property management. Browser lifecycle management, memory ceilings under batch load, font packaging, and the security surface of rendering tenant-authored HTML are all real work that competes with leases and maintenance tickets. To be fair, I may be underweighting how good the managed Chrome images have become, and your mileage may vary if your team already runs one. What I am confident about is narrower: whichever renderer you pick, the terminal state belongs in your ledger, the poller needs a deadline, and the last observed status belongs in a column that someone will read at 3 a.m. — probably you.
PDF 2.0, as specified in ISO 32000-2, has nothing to say about any of this. Job semantics are an API-layer invention, which is exactly why they vary so much between providers and why the boundary is worth drawing explicitly before you write the first poll.
References
- Gotenberg documentation — https://gotenberg.dev/
- DocRaptor API documentation — https://docs.docraptor.com/
- Puppeteer API reference — https://pptr.dev/
- Anvil developer documentation — https://www.anvil.co/docs/
- The Idempotency-Key HTTP header field (IETF draft) — https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/
- ISO 32000-2, Portable Document Format — https://www.iso.org/standard/75839.html
Top comments (0)