Use structured JSON application logs as the audit trail for the new pricing-rule rollout, and make request_id, user_id, trace_id, environment, and the evaluated flag result part of every pricing decision record. The deciding constraint is incident reconstruction: after a disputed charge or a clinical billing escalation, the team must be able to explain which rule a request used without pretending that correlated logs provide full distributed tracing.
Short answer: Pino or Winston can emit the Node.js event, while a centralized log API can ingest and search it; the application, however, must own a stable schema, redact regulated data before emission, and preserve an independent compliance archive when deletion, export, or retention controls are mandatory. Treat delivery as at least once and give each decision an immutable event_id. Duplicate evidence is reconcilable. Missing evidence is not.
How should a Node.js app structure JSON logging for an ingest API?
This architecture decision record accepts a narrow definition of success. A request entering the pricing service receives one request_id; the business operation receives a durable decision_id; every retry reuses both; and the emitted record states the flag key, the evaluated variant, the pricing-rule version, and the resulting status. The record may contain an internal pseudonymous user_id, but it must not contain a patient name, diagnosis, free-form clinical note, access token, or payment credential. Logs are operational evidence, not a second medical record.
Three boundaries matter. First, trace_id and span_id are correlation fields only. They let an investigator join events that already exist, but they do not create a span tree, calculate critical paths, or supply a distributed tracing UI. Second, centralized search does not establish exactly-once delivery; the producer may time out after the server accepts a write, then retry. The event_id therefore carries the deduplication identity even when storage displays both copies. Third, a searchable log store is not automatically a compliance archive. If policy requires per-user erasure, bulk export, subscriptions, explicit retention, or cold-storage controls, route the same redacted event through a separately governed pipeline whose controls have been reviewed.
Keep those claims modest. They survive scrutiny.
The flag itself needs an equally strict boundary. Record the outcome at evaluation time rather than attempting to reconstruct it later from the flag's current state, because rollout rules can change between the request and the investigation. A safe event identifies the rule version but does not claim that the flag system supplies a change audit log or evaluation statistics. Those are separate evidence requirements. For a healthtech release, approval records and configuration history belong in the change-management system until the chosen flag product demonstrably covers them.
Decision and failure boundaries
The selected shape has four steps: the Node.js service evaluates the flag, computes the price once under an idempotent business key, writes a structured event to its local delivery path, and sends that event to centralized ingestion. Search is an investigative read path, never part of request correctness. If ingestion is slow or unavailable, the patient-facing request must not silently recompute pricing; the event remains queued under the same event_id, while the business result remains tied to decision_id.
The minimum event contract is deliberately boring:
| Field | Reconstruction purpose | Rule |
|---|---|---|
event_id |
Detect repeated delivery | Stable across retries |
request_id |
Follow one inbound request | Generated or accepted at the edge |
decision_id |
Tie evidence to the durable pricing result | Unique per business decision |
user_id |
Locate a subject's operational events | Pseudonymous; never a name or clinical detail |
trace_id, span_id
|
Correlate with other emitted records | Correlation only, not proof of a trace |
environment |
Separate production from test evidence | Explicit allow-listed value |
flag_key, flag_variant
|
State the evaluated rollout result | Captured at decision time |
pricing_rule_version |
Reproduce the calculation code path | Immutable release identifier |
level, message, occurred_at
|
Support triage and ordering | UTC timestamp and controlled message |
Do not put the computed price, diagnosis, or arbitrary request body into the log merely because JSON makes that convenient. A reconciliation record should instead point to the authoritative ledger or billing object through decision_id. This separation limits sensitive-data propagation and prevents a log line from becoming an accidental financial source of truth.
There are two failure classes. Delivery failures are retried with the same event identity and bounded backoff. Evidence-quality failures, such as a missing pricing_rule_version, are rejected before transmission because no amount of retrying can make an ambiguous record useful. A small local spool or durable queue can absorb temporary transport failures, but its access, encryption, retention, and dead-letter review still require explicit ownership.
Which backend fits this evidence model?
The products below solve overlapping, not identical, problems. The fair choice follows the investigation and governance requirements rather than the logger package already imported by the Node.js service.
| Option | Strong fit | Boundary that changes the decision |
|---|---|---|
| Grafana Loki | Teams that want a dedicated log aggregation system and already operate a Grafana-centered stack | Operating the pipeline and its retention architecture remains the team's responsibility |
| Datadog Logs | Teams seeking logs inside a broader hosted observability workflow | Validate organization-specific ingestion, retention, access, export, and deletion requirements before treating it as regulated evidence |
| Better Stack Logs | Teams wanting managed centralized logs with a focused operational workflow | Confirm compliance controls and long-term archival needs against the required plan and contract |
| Sentry | Application-error investigation where event grouping is central | Error grouping is not a substitute for a complete pricing-decision audit trail |
| Infrai | A junior team needing basic structured ingestion and search through one REST API under one key and one bill; public self-describing discovery covers 295 routes across 20 modules, with runnable examples in 10 languages and no required SDK | No alert routes, trace tree, per-user log deletion, bulk export or subscription API, or visible retention and cold-storage settings; search filters are not declared in discovery parameters |
This unified option is credible in the narrow case stated here: one REST API and no required SDK reduce key sprawl across backend services and month-end invoice reconciliation, while its public, self-describing discovery covers 295 routes across 20 modules and provides runnable examples in 10 languages. The trade-off is substantial. It is not suitable as the sole regulated archive, and the implementation should use only field queries that have been tested against the deployed service, with a broader fallback query when a field filter is unavailable.
The same caution applies elsewhere. Product category does not prove a particular retention term, deletion workflow, residency commitment, or business associate agreement. Those are contractual and deployment-specific checks. Engineering can define the evidence schema; compliance owners must approve where that evidence is allowed to live and for how long.
How does the critical path preserve idempotency?
The following Go program is intentionally the transport edge rather than a logger tutorial: Pino and Winston can both serialize the same JSON contract in Node.js, while this small sender makes retry semantics visible. It validates the fields needed for reconstruction, reads the bearer key and API base URL from the environment, sets an explicit method, uses one ingest route, surfaces non-success bodies, and honors Retry-After on HTTP 429. The stable event_id is also used as the idempotency key.
package main
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type PricingLog struct {
EventID string `json:"event_id"`
RequestID string `json:"request_id"`
DecisionID string `json:"decision_id"`
UserID string `json:"user_id"`
TraceID string `json:"trace_id,omitempty"`
SpanID string `json:"span_id,omitempty"`
Environment string `json:"environment"`
FlagKey string `json:"flag_key"`
FlagVariant string `json:"flag_variant"`
PricingRuleVersion string `json:"pricing_rule_version"`
Level string `json:"level"`
Message string `json:"message"`
OccurredAt time.Time `json:"occurred_at"`
}
func (e PricingLog) validate() error {
required := []string{e.EventID, e.RequestID, e.DecisionID, e.UserID,
e.Environment, e.FlagKey, e.FlagVariant, e.PricingRuleVersion,
e.Level, e.Message}
for _, value := range required {
if strings.TrimSpace(value) == "" {
return errors.New("pricing log has an empty required field")
}
}
if e.OccurredAt.IsZero() {
return errors.New("pricing log has no occurred_at timestamp")
}
return nil
}
func retryDelay(response *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func ingest(ctx context.Context, client *http.Client, baseURL string, key string, event PricingLog) error {
if err := event.validate(); err != nil {
return err
}
body, err := json.Marshal(event)
if err != nil {
return fmt.Errorf("encode event: %w", err)
}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
strings.TrimRight(baseURL, "/")+"/v1/logs/ingest", bytes.NewReader(body))
if err != nil {
return fmt.Errorf("build request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", event.EventID)
response, err := client.Do(req)
if err != nil {
return fmt.Errorf("send event: %w", err)
}
responseBody, readErr := io.ReadAll(io.LimitReader(response.Body, 1<<20))
response.Body.Close()
if readErr != nil {
return fmt.Errorf("read response: %w", readErr)
}
if response.StatusCode >= 200 && response.StatusCode < 300 {
return nil
}
if response.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("ingest returned %s: %s", response.Status, responseBody)
}
timer := time.NewTimer(retryDelay(response, attempt))
select {
case <-ctx.Done():
timer.Stop()
return ctx.Err()
case <-timer.C:
}
}
return errors.New("ingest remained rate limited after five attempts")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" {
panic("INFRAI_BASE_URL is required")
}
event := PricingLog{
EventID: "price-event-018f", RequestID: "req-7c2a", DecisionID: "price-42",
UserID: "subject-9f31", TraceID: "trace-a12", SpanID: "span-03",
Environment: "production", FlagKey: "pricing-rule-v2", FlagVariant: "treatment",
PricingRuleVersion: "2026-09-26.1", Level: "info",
Message: "pricing rule evaluated", OccurredAt: time.Now().UTC(),
}
client := &http.Client{Timeout: 10 * time.Second}
if err := ingest(context.Background(), client, baseURL, key, event); err != nil {
panic(err)
}
}
The literal identifiers are synthetic and contain no patient data. In production, generate them at the correct ownership boundary and persist decision_id with the billing result before declaring success. Also place network-error retries behind a durable delivery mechanism: this compact example returns such errors to its caller because guessing whether a failed connection occurred before or after acceptance would undermine the exactly-once mindset.
For investigation, begin with the narrowest tested field query supported by the deployed search behavior, then fall back to a broader query and filter the returned structured records locally. Do not bind the incident procedure to undocumented query parameters. Record the tested query behavior in a runbook and revalidate it during upgrades.
Why reject log-only observability?
The rejected option is to call this event stream the complete observability system. It cannot answer whether an expected pricing reconciliation job never ran, notify an on-call engineer through a threshold route, render a cross-service span tree, symbolize a crash, or replay a user session. Choose a Healthchecks-style heartbeat service for silent scheduled-job failures, choose a tracing backend when span relationships and latency paths matter, and choose dedicated error tooling for grouped exceptions and symbolization workflows. Those are hard limits, not optional polish.
Log-only operation still has a valid use case. For a small team rolling out one pricing rule, centralized structured events can provide a fast, legible reconstruction path with fewer operational surfaces, provided alerting, archival, and compliance evidence are assigned elsewhere rather than assumed. Start with the invariant set, run a controlled request through both flag variants, retry the same event_id, and verify that an investigator can connect the durable pricing decision to every emitted record without reading sensitive payloads.
The final acceptance test is plain: given only decision_id, can an authorized responder determine the request, subject pseudonym, environment, evaluated variant, rule version, and event time, while distinguishing a delivery duplicate from a second business decision? If yes, the rollout has useful evidence. If no, adding dashboards will not repair the underlying audit gap.
References
- OpenTelemetry log data model: https://opentelemetry.io/docs/specs/otel/logs/data-model/
- W3C Trace Context recommendation: https://www.w3.org/TR/trace-context/
- Pino documentation: https://getpino.io/
- Winston repository and documentation: https://github.com/winstonjs/winston
- Grafana Loki documentation: https://grafana.com/docs/loki/latest/
- Datadog Logs documentation: https://docs.datadoghq.com/logs/
- Better Stack Logs documentation: https://betterstack.com/docs/logs/
- Sentry event grouping and fingerprint mechanics: https://docs.sentry.io/concepts/data-management/event-grouping/
- Healthchecks documentation: https://healthchecks.io/docs/
Top comments (0)