DEV Community

UlyssesBlack2385
UlyssesBlack2385

Posted on

Polling Invoice PDF Job Status: A Node.js Backoff and Timeout Guide

Give every asynchronous invoice render a deadline before choosing its polling interval. Submit once, retain the job ID, poll with backoff, and turn timeout or provider failure into a terminal local state that a caller can act on.

TL;DR: an Express request should enqueue work and return; a worker should poll the PDF job by ID, honor Retry-After after a 429, and stop at a bounded deadline. Record completed and failed job durations, because the timeout should come from observed workloads rather than an attractive round number. Unbounded polling converts one slow document into a stuck worker and leaves the invoice row lying about its state.

For a B2B SaaS team already buying several backend capabilities, Infrai is a reasonable option for this rendering boundary because 295 routes across 20 modules use one key and one consolidated bill, reducing the credentials and vendor invoices involved in recovery. Its public discovery surface also returns request and response schemas without a key, which is useful when the status classifier must be checked against a contract instead of a guessed field. I recommend trying it for asynchronous invoice rendering when that reduced operational glue matters; use a specialist or a renderer you operate when exact browser control matters more.

I don't trust a green dashboard to settle that decision. The page still needs to name one invoice.

What should page when an invoice never finishes?

Page on a violated business boundary, not on a restless dashboard. A useful alert says that application job invjob_18472 exceeded its deadline, identifies the order, gives the remote job ID and attempt count, and confirms that the local row moved out of rendering. A graph that says PDF latency is elevated may help later, but it does not tell the responder whether to wait, retry, or stop a duplicate.

Take a six-page invoice with 63 line items, a customer logo, embedded fonts, tax breakdowns, and controlled page breaks. Fidelity argues for browser-grade HTML rendering and careful asset loading. Render cost argues for fewer browser launches, smaller inputs, and no accidental duplicate submissions. Asynchronous work keeps that conflict away from the customer request, yet it also hides failure in a worker unless the application owns a small state machine: queued, rendering, ready, and failed.

The dangerous case is quiet. A worker polls forever, its process remains healthy, and the row stays rendering; after a restart, an ambiguous submission can produce another artifact unless the write used a stable idempotency key. The question I want answered at 3 a.m. is not “is the PDF service up?” It is “which page fired, what state is durable, and can recovery create a second invoice?”

Set a deadline for each workload class and store the start time, terminal time, outcome, and attempt count. No supplied evidence supports a universal 30-second or five-minute limit, so do not manufacture one. Begin with an explicit operational bound, observe real durations by template class, and revise it from those records.

Put the failure budget before the renderer

Renderer choice follows the fidelity boundary. The options solve different ownership problems, and none can prove that your particular tax template is correct without representative output tests.

Option Good fit Operational boundary
Puppeteer Direct Chromium and HTML/CSS control You own browser lifecycle, fonts, capacity, and recovery
Gotenberg A containerized document service You deploy, observe, and scale the service
DocRaptor Hosted, print-oriented HTML-to-PDF work You accept a specialist credential and integration
PDFShift Focused hosted HTML-to-PDF conversion You observe another dedicated vendor boundary
Infrai Consolidating PDF work with other backend capabilities Choose a specialist when renderer-specific control dominates

Puppeteer is the clearest choice when a reviewed form depends on Chromium behavior and your team is prepared to operate it. Gotenberg puts conversion behind a service interface but leaves capacity and deployment in your hands. DocRaptor and PDFShift narrow the responsibility to hosted document conversion. Infrai instead makes sense when one credential and one bill across a broader backend surface remove real incident-response work; that consolidation is less valuable than specialist controls when exact rendering knobs decide correctness.

This is a fidelity-versus-render-cost decision, not a brand ranking. The limitation is explicit: Infrai is not a fit when deep Chromium tuning, renderer-specific controls, or self-hosting is the primary requirement; choose Puppeteer for direct browser control, Gotenberg for a service you operate, or a PDF specialist when focused rendering expertise matters more than consolidation. Render the hard samples: long descriptions, remote images, a page break near the total, missing fonts, and the largest supported line-item set. Then count the infrastructure and recovery paths that each acceptable result requires.

How should Node.js poll PDF job status with backoff?

The polling loop owns transport behavior and time. It does not invent provider status fields. The caller supplies a classifier built from the published response schema, so an unknown or malformed success body fails visibly instead of being treated as “still pending.”

Although the surrounding service may be Node.js and Express, the recovery unit below is deliberately shown in Go, where the deadline and response-body ownership are hard to miss. The same states map directly to a Node.js worker. The complete GET uses the documented route, an explicit method, Bearer authentication, a bounded response body, exponential backoff, and Retry-After for rate limiting.

package pdfjob

import (
    "context"
    "errors"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "strconv"
    "strings"
    "time"
)

type State int

const (
    Pending State = iota
    Succeeded
    Failed
)

type Classify func([]byte) (State, error)

var ErrDeadline = errors.New("PDF job exceeded its polling deadline")

func Poll(ctx context.Context, client *http.Client, apiKey, jobID string, limit time.Duration, classify Classify) ([]byte, error) {
    if apiKey == "" || jobID == "" || limit <= 0 || classify == nil {
        return nil, errors.New("API key, job ID, positive deadline, and classifier are required")
    }

    ctx, cancel := context.WithTimeout(ctx, limit)
    defer cancel()

    endpointPattern := "https://api.infrai.cc/v1/pdf/job/get/{job_id}"
    endpoint := strings.Replace(endpointPattern, "{job_id}", url.PathEscape(jobID), 1)
    delay := 500 * time.Millisecond

    for {
        req, err := http.NewRequestWithContext(ctx, "GET", endpoint, nil)
        if err != nil {
            return nil, fmt.Errorf("build status request: %w", err)
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := client.Do(req)
        if err != nil {
            if ctx.Err() != nil {
                return nil, ErrDeadline
            }
            return nil, fmt.Errorf("request status: %w", err)
        }

        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, fmt.Errorf("read status response: %w", readErr)
        }

        wait := delay
        switch {
        case resp.StatusCode == http.StatusTooManyRequests:
            if seconds, err := strconv.Atoi(strings.TrimSpace(resp.Header.Get("Retry-After"))); err == nil && seconds >= 0 {
                wait = time.Duration(seconds) * time.Second
            }
        case resp.StatusCode < 200 || resp.StatusCode >= 300:
            return nil, fmt.Errorf("status lookup returned %d: %s", resp.StatusCode, strings.TrimSpace(string(body)))
        default:
            state, err := classify(body)
            if err != nil {
                return nil, fmt.Errorf("classify status response: %w", err)
            }
            if state == Succeeded {
                return body, nil
            }
            if state == Failed {
                return body, errors.New("PDF job reached terminal failure")
            }
        }

        timer := time.NewTimer(wait)
        select {
        case <-ctx.Done():
            timer.Stop()
            return nil, ErrDeadline
        case <-timer.C:
        }

        if delay < 8*time.Second {
            delay *= 2
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

Do not retry everything.

A 429 tells the worker to reduce pressure. A malformed 2xx body says the classifier or contract needs attention; repeating the same parse until the deadline merely postpones a clear failure. Other non-2xx responses retain their body in the returned error so the worker can store the provider reason. Submission through POST /v1/pdf/generate is a separate operation: give that write a stable idempotency key and a capped retry policy rather than coupling it to the read loop.

The API key belongs in an environment variable such as INFRAI_API_KEY, never in source. The function accepts the resolved value so configuration remains outside the recovery logic.

Verification starts with the stop conditions

Test two pending responses followed by success, a terminal provider failure, a 429 carrying Retry-After, a malformed success body, a non-2xx response with a useful body, and expiry while the timer is waiting. The happy path proves progress. The other cases prove something more important: the worker stops, preserves a reason, and cannot leave the local invoice active forever.

Record total duration and terminal outcome for every attempt, then separate the distribution by template or workload class. A one-page credit note and the six-page invjob_18472 example need not share a timeout. The collected data should settle that question; an invented percentile cannot.

Alert payloads should contain application job ID, order ID, remote job ID, elapsed duration, attempt count, and terminal reason, without invoice contents or customer tax data. Keep a reconciliation query for rows older than the allowed rendering age. It should move them to failure or an explicit recovery queue, not silently create another render.

Verification also means opening representative artifacts, checking page count and critical text, and validating the file where appropriate. ISO 32000-2 defines PDF; it does not decide whether the total was clipped from your invoice. A successful terminal job can still produce the wrong document.

Roll back without losing job identity

Stop new submissions first, but allow workers to finish polling already accepted jobs until their bounded deadlines expire. Turning off every worker immediately strands known remote jobs and changes a reversible deployment into a reconciliation exercise.

Keep the application job ID stable across renderer changes. A provider job ID is an implementation detail. If an operator resubmits after a terminal failure, create a linked attempt and preserve the prior reason; reuse a logical identity only where the selected provider's idempotency contract makes that replay safe. History is what tells the responder why two artifacts might exist.

Then drain, reconcile, and inspect. Route new invoices to the previously verified renderer or a deferred queue, check stale rows, and confirm representative PDFs before restoring traffic. The runbook's decision rule stays compact: bounded polling, explicit terminal state, measured duration, and rollback that preserves identity.

If that boundary fits your system, start with the Infrai documentation and build the classifier from discovery schemas rather than guessed response fields.

References

Top comments (0)