TL;DR: The expensive part of structured application logging is usually the multiplication of event volume, event size, and retention time. Measure those three terms before choosing fields or storage. For a SaaS product whose nightly data pipeline must be searchable, I use four retention layers: compact operational events, short-lived diagnostic detail, quarantined security records, and aggregate metrics. Every event carries a request ID and trace ID when available; actor IDs are pseudonymous, and secrets or direct identifiers are removed before serialization. The schema is versioned so a logging change can be rolled back without making last night's records unreadable.
This is a rollback decision, not a formatting contest. JSON helps machines parse an event, but JSON alone does not control cardinality, prove that masking happened, or preserve meaning after a deploy is reversed. The useful design starts with the bill, then works backward to the minimum evidence needed to answer: which pipeline run failed, which tenant was affected, which stage changed state, and can the previous application version still read the record?
What should a structured SaaS application log for each request?
Start with a measurement from production-like traffic. For each event class, record events per pipeline run, mean encoded bytes per event, runs per day, and retained days. The rough storage input is:
events per run x runs per day x bytes per event x retained days
Indexing, replication, compression, and query work add system-specific costs, so they belong in separate measured columns rather than a universal multiplier. This distinction matters. Trimming 80 bytes from a rare deployment event may accomplish less than stopping a verbose per-row success event emitted millions of times.
For a nightly pipeline, the dominant term is often discovered by grouping volume by event_name and schema_version, then comparing encoded bytes rather than row counts. Do that before debating retention. I would keep one completion event per stage and counters for routine row outcomes; I would reserve detailed per-record events for failures or an explicitly sampled diagnostic window. That choice gives up effortless reconstruction of every successful row. It preserves the evidence needed for rollback and failure triage without treating the log store as a shadow database.
A practical four-layer policy looks like this:
| Layer | Contents | Retention rule | Rollback purpose |
|---|---|---|---|
| Operational | Run, stage, outcome, duration, deploy and schema versions | Long enough to span the rollback and investigation window | Compare behavior before and after a release |
| Diagnostic | Bounded error details and sampled debug context | Short, explicit expiry | Explain a failed stage without retaining routine noise |
| Security | Authentication and authorization outcomes with restricted access | Set by the applicable security and legal policy | Investigate abuse while limiting exposure |
| Metrics | Counts, rates, and duration distributions | Independent aggregate policy | Detect change without indexing event-level identifiers |
The durations are intentionally absent. A seven-day rollback window may be sensible for one release process and reckless for another. Derive each duration from deployment cadence, incident response time, legal obligations, and the time needed to rerun or restore a batch.
Context fields that survive a rollback
An event should describe one state transition in stable language. For this pipeline, a compact envelope needs a timestamp, severity, event name, schema version, service, environment, pipeline run ID, stage, outcome, and deployment version. Add request_id for the triggering HTTP boundary and W3C trace_id/span_id values when the work participates in a distributed trace. In Trace Context version 00, a traceparent carries a 16-byte trace ID encoded as 32 lowercase hexadecimal characters and an 8-byte parent ID encoded as 16. A request ID is useful for local correlation; it is not a substitute for trace context.
Keep the meanings separate. Reusing a user ID as a request ID makes support searches tempting, but it couples identity to transport and repeats a high-value identifier across unrelated events. The same warning applies to one-time passcodes, session tokens, email addresses, phone numbers, message bodies, and authorization headers: they should never become convenient context fields. OWASP recommends excluding or masking data such as access tokens, authentication passwords, sensitive personal data, and connection strings from logs.
For actor correlation, emit a scoped, pseudonymous subject key produced with a keyed digest. Scope it by environment or tenant boundary, manage the key outside the log, and include a key version so rotation is possible. Plain hashing is weak when inputs come from a small, guessable space such as email addresses or phone numbers. Also remember that pseudonymization reduces exposure; it does not automatically remove every privacy obligation.
Do not put request IDs, user IDs, trace IDs, or pipeline run IDs into metric label values. Prometheus naming guidance warns that every unique label combination creates a new time series. Logs can carry those identifiers for targeted lookup; metrics should aggregate bounded dimensions such as stage and outcome.
Mask before the JSON encoder
Masking at the storage destination is too late because the sensitive value has already crossed a process boundary. The safer boundary is a typed logging adapter that accepts an allowlisted event schema, recursively redacts forbidden keys, caps free text, and serializes once. Application code should not concatenate JSON or pass arbitrary request objects.
The following Python example shows that boundary. The values and retention policy remain application decisions; the important property is that every call travels through the same normalizer.
import hashlib
import hmac
import json
import logging
from datetime import datetime, timezone
from typing import Any
FORBIDDEN_KEYS = {
"authorization",
"cookie",
"email",
"otp",
"password",
"phone",
"set-cookie",
"token",
}
def redact(value: Any) -> Any:
if isinstance(value, dict):
return {
str(key): "[REDACTED]" if str(key).lower() in FORBIDDEN_KEYS else redact(item)
for key, item in value.items()
}
if isinstance(value, list):
return [redact(item) for item in value]
return value
def subject_key(raw_user_id: str, secret: bytes, key_version: int) -> str:
digest = hmac.new(secret, raw_user_id.encode("utf-8"), hashlib.sha256)
return f"hmac-sha256:v{key_version}:{digest.hexdigest()}"
def write_event(logger: logging.Logger, *, event_name: str, context: dict[str, Any]) -> None:
event = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"severity": "INFO",
"event_name": event_name,
"schema_version": 4,
**redact(context),
}
logger.info(json.dumps(event, separators=(",", ":"), sort_keys=True))
write_event(
logging.getLogger("pipeline"),
event_name="pipeline.stage.completed",
context={
"pipeline_run_id": "run_01JEXAMPLE",
"stage": "normalize_records",
"outcome": "success",
"records_processed": 1842,
"request_id": "req_01JEXAMPLE",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"deploy_version": "2026.10.6.1",
},
)
This example deliberately does not log a raw user object, exception object, or HTTP headers. A recursive denylist is defense in depth, not the primary contract; an allowlisted schema is easier to review. Free-form exception messages need equal care because upstream libraries can echo payloads. Store a bounded error code and exception type by default, then permit sanitized detail only in the short-lived diagnostic layer.
The transport should also define failure behavior. Logging must not turn a successful pipeline stage into a failed one merely because an asynchronous sink is unavailable, yet silent loss is unacceptable for required audit records. Separate operational telemetry from mandatory audit trails, expose dropped-event counts as bounded metrics, and test queue saturation. Delivery semantics need names: best effort, at least once, and durable audit capture imply different code paths.
Schema evolution is part of deployment
A rollback-safe event contract is append-only within a schema version. Readers ignore unknown fields, writers do not silently change the type or meaning of an existing field, and incompatible changes receive a new schema_version. During a deployment transition, queries should tolerate both the old and new versions. That makes rollback ordinary: the previous binary resumes writing its known schema while searches still understand records produced by the reverted release.
Test this contract with fixtures. Serialize representative success, retry, partial failure, and permission-denied events; assert required fields and types; scan the encoded output for seeded secrets and direct identifiers; then run saved searches against both schema versions. Include newline, Unicode, nested arrays, oversized exception text, and missing trace context. Edge cases are where masking rules tend to become wishful thinking.
There is one more trap: retries. A retried stage may legitimately emit the same logical event more than once, so include attempt and a stable operation ID where deduplication matters. Do not promise exactly-once logging unless the entire delivery path can uphold it. Search results should distinguish duplicate delivery from duplicate business execution.
The release gate I would use
Before shipping a logging change, compare bytes per event class and projected retained bytes against the previous version. Then verify that no unbounded identifier has slipped into metric labels, every high-cardinality log field has a concrete investigation use, and access controls match the sensitivity of each retention layer. A privacy review should cover deletion and access requirements, not merely the redaction function.
The deploy gate also needs two rehearsals: read new events with the old query set, and roll the writer back while mixed schema versions remain searchable. If either rehearsal fails, the logging change is not rollback-safe.
This approach has a real limitation. It is a poor fit when the log is expected to be the authoritative replay record, because deliberate payload removal makes exact reconstruction impossible. Use a governed event journal or source-of-record snapshot for that job and keep observability events separate. I favor that split because the trade-off is visible: operators lose ad hoc payload inspection, while access control, deletion, and replay semantics stop depending on a general-purpose log index.
Finally, state what you are choosing not to keep. In this design, routine per-record successes, raw payloads, direct contact details, credentials, OTP values, and unrestricted exception text are discarded. During an incident, that can mean you cannot replay the exact input from logs or inspect every successful row. The proper recovery source is the governed system of record or a controlled replay mechanism, while logs retain correlation and state-transition evidence. That loss of convenience is the price of lower exposure, bounded volume, and a rollback path that stays understandable.
Further reading and References
- W3C Trace Context: https://www.w3.org/TR/trace-context/
- OWASP Logging Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
- OpenTelemetry logs data model: https://opentelemetry.io/docs/specs/otel/logs/data-model/
- Prometheus metric and label naming practices: https://prometheus.io/docs/practices/naming/
- NIST guidance on protecting personally identifiable information: https://csrc.nist.gov/pubs/sp/800/122/final
Top comments (0)