A monthly healthtech statement is reproducible only when the dashboard, live query, and PDF renderer use the same data cutoff and the same calculation and template versions. If those inputs differ, matching totals would be luck. Freeze the reporting inputs first; investigate rendering only after the frozen records and computed totals agree.
TL;DR: treat a statement as an immutable release, not a fresh view of mutable data. Record four keys with every run: reporting cutoff, dataset identity, calculation version, and template version. Re-run from those keys, compare intermediate totals, and archive the PDF with its manifest.
I have been paged for missed jobs and duplicate deliveries, so my first question is operational: which exact execution produced the artifact? Consider a monthly patient account report scheduled at 00:10 UTC on the first day of a month. The dashboard debug snapshot was captured before a late adjustment became visible, while an engineer's live query ran afterward. Both numbers can be internally correct. They answer different questions.
That is the invariant: one statement run gets one immutable input identity.
Why don't statement numbers match the dashboard debug snapshot?
A wall-clock label such as August report does not define a database view. A cutoff does. The job must specify an instant, an inclusion rule, and the data scope: records with an effective timestamp before the exclusive cutoff, belonging to the intended account and statement period. Store timestamps in an unambiguous format such as RFC 3339 and make the inclusive or exclusive rule explicit [1].
Transaction visibility is another boundary. Under multiversion concurrency control, a query observes a snapshot determined by its transaction semantics; two queries that start at different times can legitimately see different committed rows [2]. A dashboard cache or materialized aggregate adds another observation time. The PDF adds a fourth if its renderer fetches data independently.
Start with identifiers, not screenshots. Locate the statement run ID, snapshot ID or export digest, cutoff, calculation version, template version, and archive checksum. Then recompute each stage in order: selected record count, grouped subtotals, adjustments, final total. The first divergent stage narrows the fault domain. A pixel comparison cannot do that.
Numbers need lineage.
Keep effective time distinct from ingestion time. A correction received after the cutoff may describe an earlier event, but that does not mean a closed statement should silently change. Business policy must say whether the correction creates a revised statement or belongs to a later cycle. This is not a renderer setting.
Make the run manifest replayable
A manifest is small enough to inspect during an incident and precise enough to drive a replay. It should point to immutable data rather than contain sensitive health records. Access controls, retention, and audit policy still apply because identifiers and totals can themselves be sensitive.
package statement
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"time"
)
type Manifest struct {
RunID string `json:"run_id"`
AccountID string `json:"account_id"`
Period string `json:"period"`
Cutoff time.Time `json:"cutoff"`
DatasetDigest string `json:"dataset_digest"`
CalculationVersion string `json:"calculation_version"`
TemplateVersion string `json:"template_version"`
}
func (m Manifest) IdempotencyKey() (string, error) {
if m.RunID == "" || m.DatasetDigest == "" ||
m.CalculationVersion == "" || m.TemplateVersion == "" {
return "", fmt.Errorf("incomplete statement manifest")
}
payload, err := json.Marshal(m)
if err != nil {
return "", fmt.Errorf("marshal manifest: %w", err)
}
sum := sha256.Sum256(payload)
return hex.EncodeToString(sum[:]), nil
}
Define the dataset digest over a canonical export or immutable object with documented serialization. Do not hash an unordered ad hoc dump and expect stable results. Persist the key behind a uniqueness constraint before rendering. A retry with identical inputs should return the existing result; a changed cutoff, calculation, dataset, or template creates a distinct revision.
| Key | What it fixes | Mismatch it exposes |
|---|---|---|
| Cutoff | Observation boundary | Late commit or cache refresh |
| Dataset identity | Exact selected records | Filter, scope, or export drift |
| Calculation version | Aggregation rules | Rounding or adjustment-rule change |
| Template version | Presentation logic | A helper recomputes a displayed total |
If a key is absent, label the old artifact unreproducible rather than inventing certainty from current database state.
Stop there.
Template ownership sets the arithmetic boundary
Template ownership determines who can change the last transformation before a number reaches the page. A design team may own layout, but the service that owns statement semantics should own monetary or account calculations. The template receives typed, already-computed display fields. It may format a timestamp or choose a page break; it should not refetch rows, apply adjustment policy, or independently sum line items.
There are three workable models. Service-owned templates put code review and release controls near calculations, at the cost of making visual changes depend on that team. A document team can own templates if the interface is versioned and contract-tested. User-editable templates offer flexibility, but executable helpers must be constrained so presentation cannot redefine authoritative totals. None is universally best. The non-negotiable part is one owner for each calculation and a recorded version at render time.
The trade-off is real: tighter ownership slows some layout changes, while looser ownership widens the set of places where arithmetic can drift.
I initially look for duplicate delivery because queue incidents train that reflex. It is still the wrong first diagnosis when two artifacts have different input identities. Duplicate execution with one idempotency key should collapse to one archived artifact. Two legitimate revisions need different keys and clear labels; suppressing the second would hide a business event.
Stop drift before rendering
The safer pipeline computes once, renders once, and publishes only after verification. The renderer takes a sealed render model. It has no database credentials and no callback that can turn a replay into a live query.
package statement
import (
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
)
type RenderModel struct {
Manifest Manifest
Currency string
TotalMinorUnits int64
LineCount int
}
type Renderer interface {
PDF(context.Context, RenderModel) ([]byte, error)
}
type Archive interface {
PutIfAbsent(context.Context, string, []byte, map[string]string) error
}
func RenderAndArchive(ctx context.Context, r Renderer, a Archive, m RenderModel) error {
key, err := m.Manifest.IdempotencyKey()
if err != nil {
return err
}
pdf, err := r.PDF(ctx, m)
if err != nil {
return fmt.Errorf("render statement: %w", err)
}
sum := sha256.Sum256(pdf)
metadata := map[string]string{
"run_id": m.Manifest.RunID,
"manifest_key": key,
"pdf_sha256": hex.EncodeToString(sum[:]),
}
if err := a.PutIfAbsent(ctx, key+".pdf", pdf, metadata); err != nil {
return fmt.Errorf("archive statement: %w", err)
}
return nil
}
PutIfAbsent expresses the required atomic behavior; its implementation must reject overwriting an existing key. Mark success only after the archive acknowledges the write. Queue acknowledgement before that point can turn an archive failure into a missed report. Unbounded retry after that point can create duplicate delivery unless delivery has its own idempotency record.
Ack last.
Observe each boundary with low-cardinality dimensions: outcome, pipeline stage, calculation version, and template version. Keep account IDs and run IDs in access-controlled structured logs or traces, not metric labels. Alert on stuck age and missing expected completions, then use run IDs for drill-down. Render-error counts cannot detect a job that was never enqueued.
Tests should pin fixture exports and expected intermediate totals. Contract-test each template against the render-model schema. Exercise retry after render, retry after archive, concurrent duplicate workers, a correction on either side of the cutoff, and a template rollout during the schedule window. PDF conformance matters for interchange, but the PDF standard defines the document format, not the business meaning of a statement total [3]. Test semantic totals before page bytes.
When should a live query stay authoritative?
For an exploratory dashboard explicitly defined as current state, a live query is appropriate. Do not burden it with statement-style immutability. Preview documents can also use live data when they are clearly marked as previews and cannot be mistaken for archived records.
A finalized monthly report is different. Once delivered or archived, a later correction should follow an explicit revision policy with lineage to the prior artifact. Never overwrite history under the same identity. Short-lived operational views and durable statements serve different readers, so forcing both onto one freshness rule creates confusion rather than consistency.
Fresh is not frozen.
The decision rule is compact: if a reader may later ask why a number appeared, freeze the inputs and versions needed to answer. Give templates data, not authority. Then the dashboard can stay fresh while the PDF stays explainable, and an incident begins with a manifest instead of a guess.
References
- RFC 3339, Date and Time on the Internet: Timestamps: https://www.rfc-editor.org/rfc/rfc3339
- PostgreSQL documentation, Concurrency Control: https://www.postgresql.org/docs/current/mvcc.html
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
Top comments (0)