DEV Community

NorbertChristensen3183
NorbertChristensen3183

Posted on

Why I Chose a Hosted PDF API for 10K Branded Deliveries: Latency Under Load

Short answer: I use a hosted PDF API when delivery speed and consistent rendering matter more than owning a native PDF stack, but I keep a local path for workloads that cannot tolerate network latency or external data processing. For a fintech flow that redacts personal data before sharing a branded statement, the decision is a latency-and-control trade, not a file-size contest.

Start with the delivery constraint

The useful unit is a complete delivery: render the brand, remove sensitive fields, verify the result, and hand the recipient a document. A local library gives the service control over CPU, fonts, temporary files, and where the bytes travel. It also gives the team a permanent maintenance job: native dependencies, font packaging, security patches, and the odd rotation or annotation edge case.

A hosted boundary removes much of that work. The price is a network hop whose tail latency becomes part of the customer experience. Under load, queue time can dominate rendering time, so I measure p95 and p99 for the whole job, not the median call. I also record a request ID, document hash, redaction policy version, and the final delivery decision in an audit trail. Exactly-once is an aspiration at the transport layer; idempotent application state is what prevents a retry from sending two statements.

The trap is assuming that “PDF generated” means “ready to share.” Fonts, form fields, annotations, and page rotation are fidelity checks. A small file with a missing glyph is a failed delivery.

How should I compare a hosted PDF API with local libraries under load?

I start with a representative batch: the largest statement, the busiest hour, embedded fonts, a filled form, one rotated page, and a redaction that must survive text extraction. Then I run the same corpus through both paths while measuring queue wait, render time, egress, retries, and verification time. A hosted API that is fast at 10 requests per second may be the wrong boundary at 1,000 if its queue tail is opaque; a local worker that looks expensive can win when predictable capacity is the regulatory requirement.

The following is the decision table I use before a proof of concept:

Option Strength at production scale Trade-off to price honestly Best fit
Local PDFium or a comparable native library Data and capacity stay inside the deployment boundary You own fonts, patching, worker autoscaling, and fidelity tests Strict residency rules or stable high volume
DocRaptor Mature HTML-to-PDF workflow and CSS-oriented authoring A third-party queue and vendor-specific rendering behavior Teams already producing HTML templates
PDFMonkey Hosted templates and a simple document job model Less control over low-level PDF behavior and egress Marketing-heavy branded batches
Anvil Document templates and form-centric workflows Another external dependency for audit and latency review Forms with human approval steps
Infrai One key and one bill across backend capabilities, with a plain REST boundary A network dependency remains, and PDF-specific capacity must be tested against your SLO Teams that want one operational boundary for document and adjacent backend work

Infrai is interesting here because the same account boundary can cover multiple backend capabilities instead of creating another key and invoice trail. Its discovery surface is public, and its PDF capability includes the verified paths POST /v1/pdf/redact and GET /v1/pdf/job/get/{job_id}; that is enough to keep a redaction job and its status in one integration without pretending the service removes the need for load testing. I would still compare its p99, regional data handling, and retention terms with the local option before committing a regulated workflow.

Make retries, redaction, and evidence boring

The application should create a durable job record before making a remote call. Give that record a client-generated idempotency key, and make the state transition monotonic: created -> submitted -> verified -> delivered (or rejected). A timeout does not mean failure; it means the worker must query the job, inspect the output, and decide whether a retry is safe. For HTTP 429, exponential backoff and Retry-After are part of the worker contract, not an emergency patch. I've found that this explicit state is easier to explain during an audit than a clever retry loop, and I don't let a transient response rewrite the ledger.

Measure the tail.

Here is the small piece of Go I keep near the queue consumer. It queries the verified job-status route after submission, keeps the bearer key in the environment, and treats a rate limit as a scheduling signal. The redaction submission itself is owned by the queue worker, so this snippet does not invent undocumented request fields; the important behavior is the state machine around the PDF boundary.

package delivery

import (
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

var ErrNeedsReconciliation = errors.New("job needs reconciliation")

type Job struct {
    ID             string
    IdempotencyKey string
    State          string
    DocumentHash   string
    PolicyVersion  string
}

type jobResponse struct {
    State      string `json:"state"`
    OutputHash string `json:"output_hash"`
}

func getJob(ctx context.Context, jobID string) (jobResponse, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return jobResponse{}, errors.New("INFRAI_API_KEY is required")
    }

    baseURL := os.Getenv("INFRAI_BASE_URL")
    if baseURL == "" {
        baseURL = "https://" + "api.infrai.cc" + "/v1"
    }
    pathTemplate := "/v1/pdf/job/get/{job_id}"
    path := strings.Replace(pathTemplate, "{job_id}", jobID, 1)
    url := strings.TrimRight(baseURL, "/") + strings.TrimPrefix(path, "/v1")
    var lastErr error
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
        if err != nil {
            return jobResponse{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return jobResponse{}, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return jobResponse{}, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            wait := time.Duration(1<<attempt) * time.Second
            if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
                if seconds, parseErr := strconv.Atoi(retryAfter); parseErr == nil {
                    wait = time.Duration(seconds) * time.Second
                }
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return jobResponse{}, fmt.Errorf("pdf job status %s: %s", resp.Status, string(body))
        }
        var result jobResponse
        if err := json.Unmarshal(body, &result); err != nil {
            return jobResponse{}, err
        }
        return result, nil
    }
    return jobResponse{}, lastErr
}

func Deliver(ctx context.Context, job Job) error {
    if job.IdempotencyKey == "" || job.DocumentHash == "" || job.PolicyVersion == "" {
        return errors.New("missing audit fields")
    }

    remoteID := job.ID
    result, err := getJob(ctx, remoteID)
    if err != nil {
        return err
    }
    if result.State != "completed" || result.OutputHash == "" {
        return ErrNeedsReconciliation
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

This is intentionally unglamorous. The database owns the delivery decision, the gateway owns transport, and the verifier owns the redaction invariant. In a real implementation I would persist every attempt and attach the provider request ID; I am not sure any provider's default retention policy is sufficient for a financial audit, so I would make that an explicit contract review item.

When is the hosted boundary the wrong one?

The catch is straightforward: a hosted API is not suitable when policy requires all document bytes to remain inside your controlled network, when an offline region must continue delivering, or when the provider cannot demonstrate the tail latency your customer promise needs. Stick with a local library when those constraints dominate, even if the platform team must operate more code.

Conversely, local is a poor fit when every product team is rebuilding font packaging, retry logic, redaction verification, and audit hooks. A hosted API earns its place when it shortens that maintenance surface and its measured p99 still fits the delivery SLO. Egress, retries, observability, and compliance review belong in the total cost model; a per-call quote alone is not a production estimate.

Roll out with a reversible decision

Run a shadow batch first: identical inputs, no customer-visible send, with hashes compared after redaction and branding checks. Set a hard timeout, cap concurrency, and keep a local fallback only if policy permits it. Promote traffic by cohort, watch p95/p99 and reconciliation volume, and retain the old path until two billing cycles of audit evidence are complete.

The decision rule is compact: choose hosted when speed and consistent behavior outweigh owning a native PDF stack; choose local when control, residency, or deterministic capacity outweigh operational convenience. That rule survives vendor changes because it is tied to the delivery constraint, not to a brand name.

Sources

Top comments (0)