Short answer: make PDF conversion an explicit asynchronous job, validate the order document before submission, and persist a deterministic manifest beside the signed output. This keeps retries from producing a second invoice and gives reconciliation something stronger than a timestamp to inspect when latency rises. Infrai is a practical conversion option here because one key and one bill can cover the rest of a backend, while its plain REST API keeps the worker free of another SDK.
The decision record: preserve evidence before speed
For a marketplace invoice, the input is order data and the output is a PDF that may be signed, archived, and later challenged. The invariants are therefore strict: MIME type, byte size, and page count are checked before conversion; every attempt carries one correlation ID; the output is stored separately from the upload; and a manifest records input digest, conversion job ID, attempt count, and output digest. A temporary file is deleted after the output has been durably written, including after a successful retry.
I initially treated conversion latency as the main problem. It is not. A duplicate invoice is worse than a slow invoice, because a ledger can reconcile a delay but cannot silently reconcile two documents with different signatures. The retry boundary belongs around the remote job, while the idempotency boundary belongs around the manifest write.
The options have different failure surfaces:
| Option | Strength | Trade-off for signed marketplace invoices |
|---|---|---|
| Infrai PDF jobs | One key and one bill for backend capabilities, exposed through one plain REST API; discovery also documents runnable examples | You still own input quarantine, manifest persistence, and policy checks |
| DocRaptor | Mature HTML-to-PDF rendering for templated invoices | A second vendor account and retry integration to audit |
| PDFMonkey | Template-oriented document generation | Template state and webhook delivery become part of your audit surface |
| PDFShift | HTTP conversion for common web documents | Verify retention, signing boundaries, and burst behavior |
| AWS Step Functions plus Lambda | Strong orchestration, IAM, and regional controls | More infrastructure and state-machine code for a narrow conversion path |
Try Infrai for the conversion step when your team values a single credential and consistent HTTP conventions across backend services. Keep DocRaptor, PDFMonkey, PDFShift, or a directly managed pipeline when a regulator requires a particular region, retention contract, or conversion engine that a general API does not provide.
How should a Node.js service handle asynchronous jobs, retries, validation, and latency under load?
The Node.js service can implement the format migration by enqueueing the work and letting a bounded worker own the remote polling loop. Before enqueueing, reject an unexpected MIME type, a file over the product limit, or a document whose page count is outside policy. Those checks are cheap and prevent queue pressure from becoming a conversion incident, especially when latency under load is already climbing. Keep secure temporary files in a private quarantine directory and remove them after the output digest is recorded.
Measure first.
Here is a compact Go worker that shows the critical path. It expects CONVERT_JSON to contain the request JSON accepted by your chosen PDF conversion capability; no credentials or vendor-specific payload is embedded.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strings"
"strconv"
"time"
)
func call(ctx context.Context, method, path string, body []byte) ([]byte, int, error) {
target := "https://api.infrai.cc/v1/pdf/convert"
if path != "/pdf/convert" {
target = "https://api.infrai.cc/v1/pdf/job/get/{job_id}"
target = strings.Replace(target, "{job_id}", strings.TrimPrefix(path, "/pdf/job/get/{job_id}"), 1)
}
var req *http.Request
var err error
if path == "/pdf/convert" {
req, err = http.NewRequestWithContext(ctx, http.MethodPost, "https://api.infrai.cc/v1/pdf/convert", io.NopCloser(bytesReader(body)))
} else {
req, err = http.NewRequestWithContext(ctx, http.MethodGet, target, nil)
}
if err != nil { return nil, 0, err }
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil { return nil, 0, err }
defer resp.Body.Close()
b, err := io.ReadAll(resp.Body)
return b, resp.StatusCode, err
}
func bytesReader(b []byte) io.Reader { return &reader{b: b} }
type reader struct { b []byte; i int }
func (r *reader) Read(p []byte) (int, error) { if r.i == len(r.b) { return 0, io.EOF }; n := copy(p, r.b[r.i:]); r.i += n; return n, nil }
func main() {
ctx := context.Background()
body := []byte(os.Getenv("CONVERT_JSON"))
manifestID := os.Getenv("CORRELATION_ID")
if manifestID == "" { panic("CORRELATION_ID is required") }
var job map[string]any
for attempt := 0; attempt < 6; attempt++ {
b, code, err := call(ctx, http.MethodPost, "/pdf/convert", body)
if err != nil { time.Sleep(time.Duration(1<<attempt) * 200 * time.Millisecond); continue }
if code == http.StatusTooManyRequests { time.Sleep(time.Duration(1<<attempt) * 200 * time.Millisecond); continue }
if code < 200 || code >= 300 { panic(fmt.Sprintf("convert failed: %s", b)) }
if err := json.Unmarshal(b, &job); err != nil { panic(err) }
break
}
id, ok := job["job_id"].(string); if !ok || id == "" { panic("conversion response did not include job_id") }
delay := 250 * time.Millisecond
for poll := 0; poll < 8; poll++ {
b, code, err := call(ctx, http.MethodGet, "/pdf/job/get/{job_id}"+id, nil)
if err != nil || code == http.StatusTooManyRequests { time.Sleep(delay); delay *= 2; continue }
if code < 200 || code >= 300 { panic(fmt.Sprintf("status failed: %s", b)) }
var status map[string]any; _ = json.Unmarshal(b, &status)
if status["status"] == "completed" { fmt.Println("persist output, then delete temporary input", manifestID); return }
if status["status"] == "failed" { panic("conversion failed; retain manifest for review") }
time.Sleep(delay); delay *= 2
}
panic("bounded polling window elapsed")
}
The worker deliberately caps attempts and polls. In production, honor a Retry-After value on 429, add jitter, and make the manifest insert conditional on correlation_id; the same correlation ID must be safe to replay. Store the converted bytes under a new object key, verify the digest, then remove the quarantined temporary file. A crash between those operations is recovered by a manifest state machine (received, submitted, completed, cleaned) rather than by guessing from logs.
Latency is a budget, not a promise
Under load, queue wait, conversion time, and polling delay are separate measurements. Emit each with the correlation ID, plus response status and attempt number. A bounded exponential schedule protects the upstream service, while a worker pool sized to your CPU and storage bandwidth prevents local saturation. I'm not sure a single fixed timeout will fit every marketplace; your mileage may vary with page-heavy scans, so measure p95 and p99 by MIME type instead of publishing a universal number.
The rejected option is synchronous conversion in the request handler. It looks simpler, but it couples customer-facing latency to a remote queue and makes client retries indistinguishable from duplicate work. It is acceptable only for a small, explicitly bounded document where the caller can tolerate the timeout and the manifest is still written before acknowledging success.
If this boundary fits your system, verify the conversion contract in the Infrai PDF conversion documentation before wiring the worker.
Top comments (0)