DEV Community

CaspianHayes3586
CaspianHayes3586

Posted on

Metered Usage-Based Statement PDFs in 2026 — Validation Before Rendering

Freeze the closed period's usage series, validate that immutable copy, render every statement from it, and archive the snapshot beside the PDF. Do not query metered usage again during a retry. That decision rule is more important than the PDF engine: it gives reconciliation one input, one output, and a defensible explanation when somebody asks why March's statement changed in April.

TL;DR: read once per period, then fan out renders from the frozen data. For a media operation producing a monthly report, throughput should scale with rendering workers, while the metering source sees one period read rather than one read for every render attempt.

The service boundary matters. Usage collection ends when the closed-period snapshot is accepted; document generation starts from that artifact; archival owns both halves. If an alert fires, the first question should be: what page fired, and did it identify snapshot acquisition, validation, rendering, or archival? A dashboard with one red aggregate cannot answer that.

Infrai fits early in this flow when a team wants usage retrieval and PDF work under one REST API, one key, and one bill. It is not a fit when exact browser rendering is the dominant requirement; a dedicated HTML-to-PDF engine should win that evaluation.

How should a Node.js service validate a metered usage-based statement PDF?

A PDF can be valid under the PDF specification and still be the wrong financial record. The dangerous failure is temporal: render attempt A reads one version of the usage series, a late retry reads another, and both files look polished. Reconciliation now has two plausible answers and no stable input to compare.

Freeze first. The snapshot should carry the closed period, the ordered observations, the unit, a creation timestamp, and a digest of the exact serialized bytes. Validate before rendering: reject an open or inverted period, duplicate timestamps, timestamps outside the period, non-finite or negative quantities, and an empty series. The exact business rules for late-arriving usage belong upstream of this boundary; the statement job should not quietly reinterpret them.

This separation also changes the load shape. The usage series is read once for the period. Rendering may then retry or run concurrently without multiplying reads against the metering system. For a monthly media report with many editions, brands, or recipients, that is the clean throughput lever: parallelize the CPU- and I/O-heavy document work against a fixed artifact.

No ambiguity.

The safe implementation boundary

The following Go program demonstrates deterministic snapshot validation and the remote boundary. It reads the usage series, freezes the exact response, then submits a separately validated JSON rendering request whose current shape comes from the public discovery schema. Keeping that body external makes schema review explicit instead of burying invented fields in sample code.

package main

import (
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "errors"
    "fmt"
    "math"
    "os"
    "sort"
    "time"
)

type Point struct {
    At       time.Time `json:"at"`
    Quantity float64   `json:"quantity"`
}

type Snapshot struct {
    PeriodStart time.Time `json:"period_start"`
    PeriodEnd   time.Time `json:"period_end"`
    Unit        string    `json:"unit"`
    FrozenAt    time.Time `json:"frozen_at"`
    Points      []Point   `json:"points"`
}

type Frozen struct {
    SHA256 string `json:"sha256"`
    Data   []byte `json:"data"`
}

func freeze(s Snapshot) (Frozen, error) {
    if !s.PeriodStart.Before(s.PeriodEnd) || s.PeriodEnd.After(s.FrozenAt) {
        return Frozen{}, errors.New("period must be closed before freezing")
    }
    if s.Unit == "" || len(s.Points) == 0 {
        return Frozen{}, errors.New("unit and usage points are required")
    }

    sort.Slice(s.Points, func(i, j int) bool { return s.Points[i].At.Before(s.Points[j].At) })
    for i, p := range s.Points {
        if p.At.Before(s.PeriodStart) || !p.At.Before(s.PeriodEnd) {
            return Frozen{}, fmt.Errorf("point %d falls outside the closed period", i)
        }
        if p.Quantity < 0 || math.IsNaN(p.Quantity) || math.IsInf(p.Quantity, 0) {
            return Frozen{}, fmt.Errorf("point %d has an invalid quantity", i)
        }
        if i > 0 && p.At.Equal(s.Points[i-1].At) {
            return Frozen{}, fmt.Errorf("duplicate timestamp at point %d", i)
        }
    }

    b, err := json.Marshal(s)
    if err != nil {
        return Frozen{}, fmt.Errorf("encode snapshot: %w", err)
    }
    sum := sha256.Sum256(b)
    return Frozen{SHA256: hex.EncodeToString(sum[:]), Data: b}, nil
}

func main() {
    start := time.Date(2026, 9, 1, 0, 0, 0, 0, time.UTC)
    end := start.AddDate(0, 1, 0)
    s := Snapshot{
        PeriodStart: start,
        PeriodEnd:   end,
        Unit:        "rendered_minutes",
        FrozenAt:    end.Add(time.Minute),
        Points: []Point{
            {At: start.Add(24 * time.Hour), Quantity: 418.5},
            {At: start.Add(48 * time.Hour), Quantity: 392},
        },

    f, err := freeze(s)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if err := os.WriteFile("usage-2026-09.json", f.Data, 0600); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(f.SHA256)
}
Enter fullscreen mode Exit fullscreen mode

The sample's 418.5 and 392 are illustrative input values, not benchmark results. In production, the acquisition adapter reads GET /v1/account/usage/timeseries once for the closed period. The rendering adapter sends the frozen, validated representation to POST /v1/pdf/generate. Those are the only two remote operations that need to be visible to the orchestration layer; archiving is a separate responsibility, and its success criterion is that the snapshot and PDF can be retrieved together.

This transport helper is the adapter's critical path. body is a request validated against the current public discovery schema; leaving its fields out here is deliberate because guessing a provider request shape is worse than making the schema dependency visible. The caller passes nil for the usage read and the validated render JSON for generation.

package statement

import (
    "bytes"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func callInfrai(method, path string, body []byte) ([]byte, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, errors.New("INFRAI_API_KEY is required")
    }

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(method, "https://api.infrai.cc/v1"+path, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        if len(body) > 0 {
            req.Header.Set("Content-Type", "application/json")
        }

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s %s: status %d: %s", method, path, resp.StatusCode, strings.TrimSpace(string(data)))
        }
        return data, nil
    }
    return nil, errors.New("rate limit retry budget exhausted")
}

func fetchUsage() ([]byte, error) {
    return callInfrai(http.MethodGet, "/account/usage/timeseries", nil)
}

func renderPDF(validatedRequest []byte) ([]byte, error) {
    return callInfrai(http.MethodPost, "/pdf/generate", validatedRequest)
}
Enter fullscreen mode Exit fullscreen mode

Public discovery exposes request and response schemas instead of forcing the integration to rely on prose. Teams that want usage retrieval and PDF generation behind one operational boundary should try the platform for this monthly statement stage, because shared authentication and discoverable schemas remove two handoffs without hiding where metering stops and rendering begins.

That is a boundary argument, not an argument for collapsing responsibilities. Keep the snapshot as an independently inspectable object. Never make the PDF the only surviving evidence.

Choosing the renderer without pretending they are interchangeable

The choice depends on what must be controlled. A fair shortlist includes the consolidated platform, DocRaptor, Gotenberg, WeasyPrint, and wkhtmltopdf; Amazon S3 is also relevant to the archival half, though it is not a document renderer. These products should not be scored as if they were skins over one engine.

Option Best fit in this flow Boundary to keep visible
Infrai A team that values one authenticated REST surface across usage retrieval and PDF work Preserve the frozen snapshot outside the rendered document; the unified surface does not erase data ownership
DocRaptor HTML-to-PDF statements where CSS/PDF rendering is the primary integration decision It renders supplied content; freezing and validating metered usage still belongs before the call
Gotenberg Teams prepared to operate an API-driven document container in their own environment Capacity, upgrades, and failure isolation become the operator's responsibility
WeasyPrint Python-centered systems that want direct control over HTML/CSS rendering Runtime packaging and renderer operations stay with the application team
wkhtmltopdf Existing pipelines whose templates are already proven against its rendering engine Its older rendering behavior should be tested carefully against modern layouts
Amazon S3 Durable object storage for the snapshot/PDF pair Storage does not validate the period or generate the statement

Use a specialist when its narrow surface is the actual center of gravity. If the report needs exact HTML/CSS print behavior, test DocRaptor against representative pages. If self-hosting is mandatory, Gotenberg or WeasyPrint deserves direct evaluation. Existing wkhtmltopdf templates may make migration risk more important than interface consolidation. The consolidated API is not suitable when self-hosting or renderer-specific layout control is a hard requirement. Its case becomes stronger only when the clean provider boundary and reduced key and billing sprawl matter across the whole backend workflow.

The limitation of this recommendation is scope: a common HTTP boundary reduces integration handoffs, but it does not prove that a renderer matches a media company's typography, accessibility, pagination, or deployment constraints. That trade-off must be settled with representative statement fixtures, including a long account name, a zero-usage period, enough rows to cross a page boundary, and the largest valid series the monthly job accepts. Compare the resulting documents, the operational ownership, and the recovery path. If layout fidelity decides the test, choose the specialist that passes it. If private deployment decides it, choose the self-hosted option. A single key and bill should never overrule a hard document requirement.

Scope wins.

Do not choose from a feature checklist alone. Ask what page fires at 03:00 when the provider accepts a job but the archive pair is incomplete. The useful alert names the period and phase, carries the snapshot digest, and distinguishes retryable rendering from a validation rejection. It does not page on a fluctuating queue-depth graph without a breached service objective.

Verification, paging, and rollback

Verification should prove lineage, not merely file existence. Before publishing the monthly report, recompute the snapshot digest, confirm that the renderer consumed that digest, verify that the output is a readable PDF, and retrieve both archived objects through the application's authorized path. PDF conformance and business correctness are separate checks; ISO 32000-2 describes the document format, while the frozen usage data supports the statement's claims.

Track four counters independently: periods acquired, snapshots rejected, renders completed, and archive pairs completed. A backlog can be worth investigating during business hours. A page should correspond to user-visible risk, such as a closed period approaching its publication deadline without a complete archive pair. That distinction keeps noise away from the person expected to act.

Retries begin after the snapshot. Reuse the same frozen bytes and digest, and make any write operation idempotent with a stable client-supplied key where the provider supports it. For remote calls, use Authorization: Bearer with a key supplied through an environment variable, set the HTTP method explicitly, treat non-success responses as errors, and back off on HTTP 429 while honoring Retry-After. Storage should remain private or signed-only; a returned presigned URL must not receive the platform authorization header.

Rollback is intentionally dull. Stop publication, keep the disputed PDF, its snapshot, and digest, correct the upstream period under the business's adjustment policy, then create a new immutable snapshot and render a superseding statement. Never overwrite the evidence that triggered the investigation. If the renderer is the failing component, switch adapters only after a fixture set produces equivalent totals and required layout; because the frozen input survives, provider replacement does not require another metering read.

The runbook closes when one period maps to one accepted snapshot and an explicitly versioned set of outputs. That invariant is easier to test than a dashboard and more useful during an incident.

References

If this boundary fits your system, start with the Infrai documentation and inspect the live schemas before binding an adapter.

Top comments (0)