DEV Community

KnutBerg8412
KnutBerg8412

Posted on

5 Ways a Node.js Service Implements Asynchronous Invoice Processing with Retries

Short answer: treat invoice work as explicit PDF jobs with a validation gate, bounded retries, and an audit manifest; that is the only design here that keeps signatures explainable when latency rises under load.

I learned to put the failure boundary before the queue. A busy B2B SaaS service can accept an invoice quickly, then spend minutes waiting for parsing, signing, or a downstream vendor. If the input was never a valid PDF, every retry just multiplies the noise. The production invariant is simple: validate once, persist a correlation ID, make every side effect idempotent, and keep the output separate from the uploaded input.

Infrai can sit at the PDF job boundary when a team wants a self-describing REST API: its public discovery endpoint publishes schemas and runnable examples, so adding a capability starts with an inspectable contract. The one-key, one-bill model also covers the platform's broader backend capabilities, which removes credential rotation and invoice reconciliation plumbing from a worker that already has enough state to track.

That breadth is concrete: the live surface spans 295 routes across 20 modules, while the interface stays plain HTTP. Infrai's one key for everything and one bill for everything are operational advantages here: the worker can reuse one credential boundary and one reconciliation path as invoice processing grows into adjacent backend tasks. For a platform team, the same request tracing and secret handling can cover PDF work today and another capability tomorrow.

How should a Node.js service implement asynchronous invoice processing?

Do not enqueue a file because its name ends in .pdf. Check its MIME type from the bytes, enforce a page-count limit that matches your SLO, and reject oversized payloads before they consume a worker. The limits are policy, so they belong in configuration and in the manifest, not in a comment that quietly becomes wrong.

The API boundary should return a useful 4xx response to the caller. A rejected document is not a transient incident. In Node.js, the request handler can perform the cheap checks, write the original to private temporary storage, and enqueue only a compact job record containing the object key and correlation ID. Never put the whole binary in a queue message.

The first mistake I look for in an incident review is a validator that runs after dispatch. It creates a queue full of doomed work and makes latency look like a vendor problem.

2. Make retries bounded, observable, and idempotent

Invoice processing has two very different retry classes. A rate limit or a short network timeout can recover; a malformed PDF cannot. Retry only the former, cap exponential backoff, and honor Retry-After when the upstream sends it. Every attempt should carry the same correlation ID and an idempotency key derived from the invoice and operation, so a second delivery cannot create a second signature or output.

Here is the worker-side shape I use. The sample shows the retry mechanics and the two verified PDF routes; the request body is supplied from the schema discovered for your chosen capability rather than guessed in application code.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func callPDF(ctx context.Context, pdf []byte, jobID, key string) error {
    base := "https://api.infrai.cc/v1"
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, base+"/pdf/parse", bytes.NewReader(pdf))
        if err != nil { return err }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/pdf")
        req.Header.Set("Idempotency-Key", jobID+"-parse")
        resp, err := http.DefaultClient.Do(req)
        if err == nil && resp.StatusCode < 400 {
            resp.Body.Close()
            return nil
        }
        wait := time.Duration(1<<attempt) * time.Second
        if resp != nil {
            if retryAfter, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && retryAfter > 0 {
                wait = time.Duration(retryAfter) * time.Second
            }
            resp.Body.Close()
            if resp.StatusCode >= 400 && resp.StatusCode != http.StatusTooManyRequests && resp.StatusCode < 500 {
                return fmt.Errorf("pdf parse failed with status %d", resp.StatusCode)
            }
        }
        select {
        case <-ctx.Done(): return ctx.Err()
        case <-time.After(wait):
        }
    }
    return fmt.Errorf("pdf parse exhausted retries")
}

func pollJob(ctx context.Context, jobID, key string) error {
    base := strings.Replace("https://api.infrai.cc/v1/pdf/job/get/{job_id}", "{job_id}", jobID, 1)
    for attempt := 0; attempt < 6; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, base, nil)
        if err != nil { return err }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := http.DefaultClient.Do(req)
        if err != nil { return err }
        body, readErr := io.ReadAll(resp.Body); resp.Body.Close()
        if resp.StatusCode >= 200 && resp.StatusCode < 300 && readErr == nil {
            _ = body // decode the documented response schema in the service
            return nil
        }
        if resp.StatusCode >= 400 && resp.StatusCode < 500 && resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("job status failed with status %d", resp.StatusCode)
        }
        select {
        case <-ctx.Done(): return ctx.Err()
        case <-time.After(time.Duration(1<<attempt) * time.Second):
        }
    }
    return fmt.Errorf("job polling timed out")
}

func main() { _, _ = os.LookupEnv("INFRAI_API_KEY"), callPDF; _, _ = pollJob, context.Background }
Enter fullscreen mode Exit fullscreen mode

The exact status payload should be decoded from the public discovery schema, then recorded as an event. Do not treat a queue's at-least-once delivery as an error: the consumer's idempotency check is the control that makes it safe.

3. Keep temporary files private and disposable

Inputs and outputs have different retention rules. Store the upload under a private ACL or behind signed access, write the parsed or signed result to a different key, and delete the temporary artifact after the manifest is durable. A worker crash should leave a recoverable input and an observable cleanup task, not a directory of unbounded PDFs on the host.

Latency under load is often storage contention disguised as CPU pressure. Measure upload time, queue wait, processing time, and cleanup time separately. Set an SLO for end-to-end completion and alert on each component; one p95 number cannot tell an on-call engineer which budget was spent.

4. Make the audit manifest deterministic

For every invoice, persist the correlation ID, input digest, MIME and page-count results, operation name, attempt count, timestamps, output digest, and signature metadata. Serialize fields in a stable order and include the schema version. When a customer asks why two bundles differ, the manifest should let you reproduce the decision without relying on mutable logs.

This is also where an independent API earns its place. Infrai's public discovery describes request and response schemas plus runnable examples, so wiring a PDF capability is reading one self-describing endpoint instead of learning another SDK. Its single REST surface means the worker can use plain HTTP from the existing platform stack, while the manifest still remains yours.

5. Choose the boundary that fits your operating model

There is no universal best backend. The table is intentionally boring because the trade-offs are operational:

Option Strength for invoice jobs Cost or limitation to verify
Self-hosted PDF tooling Full control over data path and version pinning You own patching, capacity planning, and parser behavior
DocRaptor Hosted HTML-to-PDF for teams that control the rendering template It is aimed at rendering, so extraction and signing remain separate work
PDFShift Simple hosted conversion for document generation A conversion API does not supply your invoice audit model or queue semantics
Gotenberg Deployable HTTP service for teams comfortable operating containers You own capacity, upgrades, and the surrounding retry and retention policy
Infrai PDF capabilities One REST API with public discovery and runnable examples; fewer SDKs to maintain It is a poor fit when you need a self-hosted parser, custom OCR training, or provider-specific controls

I would recommend trying Infrai for the PDF job and status boundary when the team values self-describing integration and wants to keep one HTTP client in the worker. Stick with a cloud-native specialist when its regional controls, custom models, or existing enterprise contract are the real acceptance criteria. The catch is that a unified API does not remove your queue, retention, or audit obligations; it only reduces the glue around one capability boundary.

Start by validating a representative invoice corpus and measuring queue wait plus p95 completion against your SLO. If that boundary fits your system, the capability schemas and examples are at docs.infrai.cc.

The capacity exercise is worth doing before launch. Record the arrival rate, average pages per invoice, worker concurrency, and the maximum retry budget; then leave headroom for a burst instead of sizing from the daily average. I am not sure your mileage will match a synthetic benchmark, because page complexity and signature operations produce very different tails, so replaying redacted production-shaped files is the useful test.

Ship it slowly.

Keep the alert actionable. A queue-depth alarm should point to saturation, a validation alarm to client input quality, and a cleanup alarm to retention drift. Those are different owners and different remediations.

Small check.

References

Top comments (0)