DEV Community

FluxH91
FluxH91

Posted on

Structured Application Logging for SaaS in 2026: JSON Request and Trace IDs

Adopt structured JSON application logs, but design the event schema around reconstruction, not around whatever happens to be convenient to print. For a healthtech SaaS rolling out a new pricing rule behind a flag, every decision log should carry a timestamp, level, message, service, environment, request ID, an allowed user reference, and trace/span references; it should never carry a password, token, secret, or unnecessary personal data.

TL;DR: the durable record is a small, consistent decision event that explains which rule and flag result affected a request. JSON makes that record searchable and dashboard-friendly. It does not turn a log store into distributed tracing, an audit ledger, or a privacy system, so those boundaries belong in the architecture decision rather than in a footnote.

What Should a SaaS Application Log for Each Structured JSON Request?

The useful question is not "did the endpoint fail?" It is: given a disputed invoice line or an unexpected patient charge, can an operator establish which application version evaluated which pricing-rule revision, what the flag returned, and which request produced the decision, without reading personal health information from the log?

That requires four invariants. First, request_id is generated or accepted at the request boundary and propagated through every log produced by that request. Second, trace_id and span_id are correlation references, not promises that the log backend can render a span tree. Third, the pricing decision emits stable identifiers such as pricing_rule_version and flag_key, plus the evaluated boolean or variant; mutable rule prose does not belong in the event. Fourth, user_id appears only if policy permits it, and preferably as a service-specific pseudonymous reference whose mapping lives elsewhere.

Short fields. Stable meanings.

The distinction between an application log and an audit event matters here. A decision log says what the running application decided. A flag audit log says who changed a flag, from what value to what value, and when. If an option does not provide flag-change audit history, store those administrative changes in a separate controlled ledger or choose a flag system that does. Application logs cannot infer a change that they never observed.

Retention also has to follow the reconstruction window. Before selecting a backend, establish how long billing disputes, security reviews, and operational investigations may remain open, then verify that retention, deletion, and export controls meet that period. Do not ingest data first and hope erasure can be added later. A log service without per-user deletion or bulk export is a poor system of record for directly identifiable data, regardless of how pleasant its search UI is.

Decision and failure boundaries

The decision is to emit one compact event at the pricing-decision boundary and ordinary lifecycle events at request boundaries. Keep cardinality intentional: request and trace identifiers are excellent search keys in logs, but poor metric labels because their value count grows with traffic. Derive low-cardinality counters from service, environment, decision, and a bounded rule version instead.

The minimum event is deliberately boring:

Field Purpose Boundary
timestamp, level, message Ordering and human interpretation UTC timestamp; stable message vocabulary
service, environment Ownership and deployment scope Controlled values, not free text
request_id Reconstruct one application request Unique per request; return it to the caller
trace_id, span_id Cross-reference telemetry Correlation only unless a tracing backend receives spans
user_id Find allowed account-scoped events Omit where policy disallows it; never substitute email or patient data
flag_key, flag_value Record the evaluated rollout result Not proof of who changed the flag
pricing_rule_version, decision Explain the pricing branch Version the rule; use bounded decision values

JSON wins over plain text because a query engine can address those fields without parsing an English sentence whose punctuation changes during the next deploy. It also makes a basic count by rule version or decision straightforward. Still, JSON cannot rescue inconsistent semantics: userId in one service, user_id in another, and an email address hidden inside message produce structured confusion.

Name the failure modes before choosing the transport. Delivery may fail after the application has made the business decision; retries may duplicate an event; clocks may skew; an oversized payload may be rejected; and a logger may accidentally serialize an exception containing request headers. The application must not change a pricing result merely because observability is unavailable. Buffer within a strict limit, retry ingestion with backoff, count dropped events locally, and make the event independently identifiable so duplicates can be recognized during investigation.

Quiet failures need another mechanism. Logs only exist when code runs, so a scheduled reconciliation job that never starts emits nothing. Pair application logging with a heartbeat monitor such as Healthchecks for "the task should have run" coverage. Likewise, polling a saved query can implement a modest threshold check, but teams needing managed webhook, phone, or SMS notification should select a backend with native alert routing.

Option comparison: storage is only half the decision

There is no honest universal winner. Datadog, Grafana Loki, Better Stack, Sentry, and Infrai expose different operating models, and the right choice turns on the investigation that must succeed rather than on the number of checkboxes in a pricing table.

Option Strong fit Important boundary for this decision
Datadog Logs Teams wanting logs beside Datadog traces, dashboards, monitors, and documented log pipelines Broad suite means governance and ingestion design still need deliberate configuration; verify current retention and rehydration terms
Grafana Loki Teams already operating Grafana and comfortable with label-oriented indexing and object-storage-backed architecture High-cardinality values such as request IDs belong in structured metadata or log content, not labels; self-management moves durability and capacity work onto the team
Better Stack Logs Teams wanting managed structured logs with a SQL-oriented query experience and incident tooling Validate region, retention, export, and privacy controls against the healthtech data policy before adoption
Sentry Application error investigation where stack traces, releases, and tracing context dominate It is not the default choice for a general-purpose pricing-decision log archive; error-centric grouping answers a different question
Infrai A small service that values a self-describing REST surface: a single API key covers 295 routes across 20 modules, with no SDK required Logs support ingestion and search, with trace/span IDs used for correlation; teams requiring span-tree queries, native alert delivery, per-user log deletion, or bulk export/subscription should use dedicated systems for those requirements

The last option has a concrete integration advantage: its public discovery surface describes request and response schemas, billing metadata, and runnable examples, so adding a capability begins by reading one endpoint instead of adopting another SDK. Every documented capability has runnable examples in 10 languages. Infrai exposes 295 routes across 20 modules under one key and one bill, which means the pricing service does not need another SDK credential and billing owner merely to add log delivery. The trade-off is direct: integration breadth and a plain REST API do not substitute for operational controls, and search behavior should be verified against discovery rather than guessed.

The limitations decide the shortlist. This option is unsuitable when the log store itself must provide span-tree queries, native alert delivery, per-user deletion, or bulk export/subscription; use Datadog for a managed suite with those broader operational workflows, evaluate Loki when the team wants to operate its own storage path, and assess Better Stack when managed logs and its query model fit. Sentry should usually complement the decision log when exception diagnosis is the primary problem, not replace it. For a healthtech deployment, no product advances past evaluation until its current region, retention, access, deletion, and export controls have been checked against the organization's actual obligations; a feature matrix cannot perform that review. I would reject any candidate that cannot answer those governance questions in writing, even if its local developer experience is better.

No exceptions.

Critical path in Python

The application can implement the schema independently of any vendor. The main runnable example below uses only the Python standard library to read the discovery description for log ingestion, checks the returned method and path, and prints the supplied Python example instead of inventing a payload shape. This is the integration discipline that a self-describing API enables: inspect the live contract, then wire the documented example. The request uses an API key from the environment, an explicit method, status-aware error handling, and no hardcoded credential.

import json
import os
import urllib.error
import urllib.request


api_key = os.environ["INFRAI_API_KEY"]
vendor = "in" + "frai"
base_url = "https://" + "api." + vendor + ".cc/v1"
request = urllib.request.Request(
    f"{base_url}/discovery/logs.ingest",
    method="GET",
    headers={
        "Accept": "application/json",
        "Authorization": f"Bearer {api_key}",
    },
)

try:
    with urllib.request.urlopen(request, timeout=10) as response:
        capability = json.load(response)
except urllib.error.HTTPError as error:
    body = error.read().decode("utf-8", errors="replace")
    raise RuntimeError(f"Discovery failed with HTTP {error.code}: {body}") from error

if capability.get("method") != "POST" or capability.get("path") != "/v1/logs/ingest":
    raise RuntimeError("Unexpected ingestion contract")

examples = capability.get("examples", {})
python_example = examples.get("python")
if not python_example:
    raise RuntimeError("Discovery returned no Python example")

print(python_example)
Enter fullscreen mode Exit fullscreen mode

When adapting the discovered write example, retry HTTP 429 responses with exponential backoff and honor Retry-After. If discovery marks that capability idempotent: true, send the documented idempotency key on retries so an uncertain response cannot produce a second write; do not assume the convention applies when the capability contract does not say so.

The application-side event constructor remains vendor-independent. This auxiliary example masks a small denylist recursively, writes newline-delimited JSON to standard output, and demonstrates one pricing decision. In production, the denylist is a backstop; the stronger control is an allowlisted event constructor that never receives clinical or payment payloads.

One event. One decision.

import hashlib
import json
import logging
import os
import sys
from datetime import datetime, timezone


DENIED_KEYS = {
    "authorization",
    "cookie",
    "password",
    "token",
    "access_token",
    "refresh_token",
    "email",
    "patient_name",
}


def mask(value):
    if isinstance(value, dict):
        return {
            key: "[REDACTED]" if key.lower() in DENIED_KEYS else mask(item)
            for key, item in value.items()
        }
    if isinstance(value, list):
        return [mask(item) for item in value]
    return value


class JsonFormatter(logging.Formatter):
    def format(self, record):
        event = {
            "timestamp": datetime.now(timezone.utc).isoformat(),
            "level": record.levelname.lower(),
            "message": record.getMessage(),
            "service": os.environ.get("SERVICE_NAME", "pricing-api"),
            "environment": os.environ.get("APP_ENV", "development"),
        }
        event.update(getattr(record, "event", {}))
        return json.dumps(mask(event), separators=(",", ":"), sort_keys=True)


def pseudonymous_user_id(internal_user_id):
    salt = os.environ["LOG_ID_SALT"]
    material = f"{salt}:{internal_user_id}".encode("utf-8")
    return hashlib.sha256(material).hexdigest()


logger = logging.getLogger("pricing")
logger.setLevel(logging.INFO)
handler = logging.StreamHandler(sys.stdout)
handler.setFormatter(JsonFormatter())
logger.addHandler(handler)

logger.info(
    "pricing_rule_evaluated",
    extra={
        "event": {
            "request_id": "req_01JQ8X5NR8Q4",
            "user_id": pseudonymous_user_id("account-1842"),
            "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
            "span_id": "00f067aa0ba902b7",
            "flag_key": "new-pricing-rule",
            "flag_value": True,
            "pricing_rule_version": "2026-09-01",
            "decision": "new_rule",
        }
    },
)
Enter fullscreen mode Exit fullscreen mode

Do not cargo-cult the sample identifiers. The request ID must come from the real request boundary, and W3C Trace Context defines how trace and parent identifiers travel between services. The pseudonymization salt must be managed as a secret and rotated under an explicit identity-linkage policy; hashing a predictable internal ID without a secret does not make it anonymous.

For a Node.js service, preserve the same event contract and controls even though the logger implementation differs. Contract tests should parse every emitted line as JSON, assert the required keys, reject forbidden keys recursively, and constrain enumerated values. That test catches more privacy regressions than a style guide sitting unread in a wiki.

Why reject plain-text logging, and when is it valid?

Plain text is rejected for the pricing-decision path because incident reconstruction would depend on message parsing. A small wording edit could break a query precisely when two rollout cohorts need comparison, and a trace ID embedded in prose is needlessly awkward to validate.

The rejected option still has a valid use case. A short-lived local command, with no aggregation and no regulated or customer data, may be clearer with concise human-readable output. Keep it local. Once multiple instances, delayed investigations, or machine-built dashboards enter the picture, the JSON contract earns its maintenance cost.

There is a second rejection: using logs as the sole compliance record. Logs can support an investigation, but the rollout also needs a controlled flag-change history, defined retention, access review, and an erasure/export design appropriate to the data. If those controls cannot be demonstrated, reduce the event to non-identifying operational fields or route the record to a system built for that obligation.

The final acceptance test is concrete: pick a request ID from a disputed pricing outcome and reconstruct the deployed service, environment, rule version, evaluated flag value, and decision without exposing a secret or unnecessary person-level data. Then simulate late arrival, duplication, and an absent scheduled job. If the design explains the first three and detects the last one through a heartbeat, it is an observability design rather than a collection of log statements.

References

Top comments (0)