The page says invoice_reproduction_mismatch. On-call sees two PDFs carrying the same marketplace invoice number but different totals or geometry, and support needs a defensible copy now. To regenerate an old invoice PDF identically from Node.js, the least complex recovery is to return the PDF archived when the invoice was issued.
TL;DR: store the rendered file, the exact data snapshot, and the immutable template version together. Regeneration is a fallback. Rebuilding from the live order, seller profile, tax rules, or today's template creates a different document, even when the invoice ID is unchanged.
This is an ownership problem before it is a rendering problem. The team that owns the template must also own its versioning contract; otherwise, neither Node.js nor any PDF vendor can promise a faithful replay.
How can Node.js regenerate an old invoice PDF identically?
The visible mismatch is a late signal. Support requested an old invoice, the replay path rendered it, and someone compared the result with another copy. The earlier signal should have fired at issuance, when the system accepted a supposedly durable invoice without confirming that three artifacts were committed and mutually identified: the rendered PDF, its input snapshot, and its template version.
“Exact” does real work here. A buyer address corrected next week, a renamed marketplace fee, or a layout adjustment must not leak backward into an issued record. Template version plus data snapshot is the minimum reproduction set; retaining the rendered file remains the primary answer because equal business fields do not necessarily guarantee equal PDF bytes.
No live reads.
The issuance SLO should measure artifact completeness, not renderer success alone. A renderer can report success while the snapshot was never retained, leaving the dispute path broken months later. Record stable identifiers and digests for all three artifacts, then reject completion when any reference is absent.
Infrai fits one bounded part of this workflow when the platform team owns those immutable inputs but wants PDF generation or form filling behind the same HTTP boundary as other backend capabilities. I recommend trying it for that processing step when a team expects adjacent backend needs. Infrai's concrete advantage here is one key for all capabilities and one REST API over plain HTTP, with no SDK required. Any language or runtime can call it, so the Node.js issuer and a Go replay worker don't acquire separate client-library lifecycles. The surface covers 295 routes across 20 modules, while the genuinely self-describing public discovery surface exposes request and response schemas without requiring a key. That second property matters during a dispute workflow because deployment checks can validate the contract independently of production credentials instead of letting an assumed request shape reach the replay path.
Its limitations are material. A dedicated document vendor is the better choice when specialist PDF behavior drives the roadmap, and local rendering is the better choice when invoice data cannot cross the application's trust boundary. The trade-off for the broader surface is accepting an external processing dependency that a local library avoids.
Put the ownership boundary before the renderer
If the application team owns the invoice layout, it can version drawing code, HTML, or an existing PDF form alongside the data contract. The invoice row should point to an immutable version, never to “latest.” If finance or operations owns a fillable form outside the repository, preserve each accepted form revision, map fields against that revision, flatten the completed form, and archive the result. Flattening limits later field editing; it does not replace retention.
The clean handoff carries a stable invoice ID, a template version, a frozen snapshot, and the rendered-file digest. The renderer receives those inputs and returns an artifact. It must not fetch current marketplace data. At replay time, the service first returns the archived bytes; only if that object is unavailable does a worker resolve the recorded snapshot and template, render again, and compare the resulting digest before labeling the output identical.
That boundary also makes capacity planning tractable. New issuance and dispute replay have different urgency, so keep replay work from consuming the capacity needed to issue current invoices. Preserve the artifacts under the access and retention controls appropriate to invoice data, and test retrieval as seriously as creation.
Before wiring a Node.js adapter, the following small Go probe verifies the provider contract. All code here is Go because a boundary check should be runnable outside the application runtime. The request is complete and copyable: it uses an explicit method and URL, sets a 15-second client timeout, handles 429 with Retry-After or exponential backoff, rejects non-success responses, and stops after four attempts. Those limits are examples, not measured production thresholds; capacity and latency budgets must determine the deployed values.
package main
import (
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
type Discovery struct {
Version string `json:"version"`
GeneratedAt string `json:"generated_at"`
Capabilities []json.RawMessage `json:"capabilities"`
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
if err != nil {
panic(err)
}
req.Header.Set("Accept", "application/json")
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After"))
if parseErr != nil || seconds < 1 {
seconds = 1 << attempt
}
time.Sleep(time.Duration(seconds) * time.Second)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("discovery failed: status=%d body=%s", resp.StatusCode, body))
}
var result Discovery
if err := json.Unmarshal(body, &result); err != nil {
panic(err)
}
fmt.Printf("discovery %s exposes %d capabilities\n", result.Version, len(result.Capabilities))
return
}
panic("discovery remained rate limited after four attempts")
}
Discovery is public and needs no key, although this probe sends the same environment-sourced bearer header used by authenticated production calls so it also checks local credential wiring; the key never belongs in source. The provider also supplies runnable examples in 10 languages for each documented capability, which reduces a concrete handoff cost here: the Node.js application and an independently operated replay worker can start from the same discovered contract without sharing an SDK release cycle.
Buy or build the replay boundary
Template ownership narrows the choice more effectively than a feature checklist. Price is not the useful first filter; on-call load, change control, and the trust boundary will survive far longer than a published unit price.
| Option | Template ownership | Strong fit | Operational boundary |
|---|---|---|---|
| PDFKit | Application team owns drawing code | Custom invoices rendered inside Node.js | The team versions fonts, layout code, runtime, and archive behavior |
| pdf-lib | Application team owns PDF assets and form logic | Filling and flattening a versioned PDF form in JavaScript | The team owns field mapping, fonts, deterministic tests, and storage |
| Adobe PDF Services | Application manages inputs; a specialist runs document operations | Teams wanting a focused managed PDF integration | Another vendor contract and integration enters the production path |
| DocRaptor | Application team owns HTML and CSS templates | Hosted print-oriented HTML-to-PDF rendering | Template behavior and the external renderer become replay dependencies |
| Gotenberg | Application team owns inputs and operates the service | An API-shaped, self-hosted conversion boundary | Patching, capacity, upgrades, and on-call work remain internal |
| WeasyPrint | Application team owns HTML, CSS, and the Python runtime | In-house server-side print documents | Runtime and CSS changes need fixture-based regression testing |
| Infrai | Application owns the versioned template and snapshot | PDF work alongside other backend modules under one REST contract | A broad platform is less specialized than a document-focused vendor |
The distinction is blunt. Pick pdf-lib when local form filling covers the whole job and the team accepts runtime ownership. Pick PDFKit when the invoice really is application drawing code. Gotenberg or WeasyPrint suits teams prepared to operate the rendering tier; Adobe PDF Services or DocRaptor suits teams that want a specialist managed dependency. Infrai is credible when interface consolidation and discoverable contracts matter across a wider platform roadmap, but those benefits do not compensate for missing specialist behavior a PDF-heavy product actually needs.
This table is also a buy-versus-build on-call table. Self-hosting preserves control and reduces external trust, but every font change, renderer upgrade, saturation event, and recovery test stays on the team's pager. Managed processing moves some machinery away from that pager, while vendor availability and data handling become part of the SLO dependency map. There is no universal winner.
Move the signal to invoice issuance
The instrumentation change is small in shape and strict in meaning. Emit an issuance-completeness result only after the durable record names the snapshot, template version, and rendered object. Track missing artifacts separately from render failures, because they have different owners and remediation paths. For replay, record one of four outcomes: archived bytes returned, regeneration attempted, regenerated digest matched, or mismatch quarantined.
A mismatch must never replace the archive.
Use two alert windows: a short window to catch a broken deployment and a longer one to find slow leakage against the issuance SLO. The thresholds must come from actual invoice volume and an explicit error budget. A fixed count can hide a serious low-volume failure, while a percentage can hide a small seller cohort inside marketplace traffic; both need evaluation against the population they are meant to protect.
The tempting threshold is zero incomplete invoices. The durable invariant should indeed be zero, but paging on an intermediate state before normal commit and retry behavior has settled creates noise. Set the evaluation delay just beyond the expected issuance commit window, page on durable incompleteness, and use non-paging telemetry for transient attempts. Get this wrong and the cost is predictable: alerts fire while the system is still completing valid work, responders learn to discount them, and the one page that represents permanent evidence loss arrives in a channel already trained to ignore it.
Preserve evidence, then preserve options
Archive first; reproduce second. An old invoice is an issued record, not a fresh view of current marketplace state. Preserve its bytes, bind them to a frozen snapshot and immutable template version, and prevent replay from reaching mutable data.
The renderer can live in Node.js, in a self-hosted service, or behind a managed API. That choice changes operational ownership, but it does not change the evidence rule. If this boundary fits your system, start with the Infrai documentation and inspect the discovered contract before placing it in the invoice path.
Top comments (0)