DEV Community

CrimsonWave9361502
CrimsonWave9361502

Posted on

Small SaaS App Logging Service: Structured JSON for Postgres Rollbacks

Choose an app logging service by testing one mixed-release query through the real export path before committing to it. The deciding constraint for a small US-and-EU logistics SaaS is rollback safety: after a notification release is reversed, can an engineer still identify failed deliveries emitted by both versions without reading prose or guessing which field survived?

TL;DR: define one narrow, versioned failure event; deploy a compatible change; roll it back; and make every candidate ingest and query the resulting records. This experiment is more revealing than a feature checklist because it exercises the boundary between application code, runtime output, collection, storage, and the query an on-call engineer actually needs.

The tempting shortcut is to compare setup time, dashboards, and an attractive entry price. I wouldn't start there. A logger that takes minutes to install but loses field types during collection creates a nasty failure mode: everything looks healthy until the exact moment two releases have to be searched together.

What should a small SaaS app logging service prove before rollback?

A delivery notification is not one action. The service accepts work, renders a message, attempts a channel, receives or misses an acknowledgement, records state in Postgres, and may retry. During a rollback, queued work from the newer release can finish after older application instances return. Buffered logs can arrive later still.

Three clocks now matter: when the application says the event occurred, when a collector observed it, and when the backend ingested it. OpenTelemetry's logs data model separates Timestamp from ObservedTimestamp; preserving that distinction prevents a delayed export from masquerading as a fresh delivery failure. Its model also provides severity, body, resource, attributes, and optional trace and span identifiers. Those are useful boundaries for evaluating a service even when the application emits plain JSON to stdout. In the drill, I would deliberately pause collection between the second and third records: if the second record arrives last, a query sorted only by ingestion time tells the wrong deployment story. The event timestamp should place the attempt in the release window where it ran, while the observed timestamp exposes collection delay rather than concealing it.

Short messages don't solve this.

Suppose release A writes notification_id, while release B replaces it with delivery_id. Rolling B back does not remove its queued work or its already buffered events. A saved query using either name sees only part of the incident. The safer evolution is additive: keep the established field, add an optional field if needed, and change meaning only behind a new event_version.

Postgres remains the source of truth for delivery state. Logs provide evidence about execution, including duplicate attempts and ambiguous outcomes; they are not an exactly-once business ledger. That distinction matters when an operator decides whether replay is safe.

Keep that boundary sharp.

Start with the query, then design the event

Write the incident question before choosing storage: "For release R in region EU, which terminal notification attempts failed, why did they fail, and were they marked retryable?" If the event cannot answer that question without parsing its message string, fix the event first.

For this experiment, the stable contract needs event_name, event_version, notification_id, shipment_id, attempt, outcome, retryable, failure_class, release, and region. Trace identifiers belong alongside those fields when trace context exists. Recipient addresses, phone numbers, message bodies, access tokens, prompts, model responses, and arbitrary headers do not. Regional storage controls cannot undo sensitive data that was exported in the first place.

Here is a focused Python emitter. JSON is serialized once, domain identifiers remain distinct from telemetry identifiers, and field types are predictable across releases.

import json
import os
from datetime import datetime, timezone
from typing import Literal

Outcome = Literal["delivered", "failed", "deferred"]


def emit_notification_result(
    *,
    notification_id: str,
    shipment_id: str,
    attempt: int,
    outcome: Outcome,
    retryable: bool,
    failure_class: str | None = None,
    trace_id: str | None = None,
    span_id: str | None = None,
) -> None:
    event = {
        "timestamp": datetime.now(timezone.utc).isoformat(),
        "severity_text": "ERROR" if outcome == "failed" else "INFO",
        "event_name": "logistics.notification.result",
        "event_version": 1,
        "notification_id": notification_id,
        "shipment_id": shipment_id,
        "attempt": attempt,
        "outcome": outcome,
        "retryable": retryable,
        "failure_class": failure_class,
        "release": os.environ["APP_RELEASE"],
        "region": os.environ["APP_REGION"],
        "trace_id": trace_id,
        "span_id": span_id,
    }
    print(json.dumps(event, separators=(",", ":"), sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

The code is intentionally boring. That is good. An eval harness can call this boundary with synthetic identifiers, decode the output, and assert required keys and types before a release reaches production. The same harness should reject a terminal result with no notification_id, an attempt serialized as text, or a changed meaning hidden behind the same version number.

One fixture is enough to begin.

This also keeps AI-related telemetry under control. For a notification feature that uses a model to classify a delivery failure, log the model identifier, latency, token counts, and evaluation outcome only when they are operationally necessary. Raw prompts and responses are both high-volume and potentially sensitive. A compact evaluation signal is easier to retain and safer to search.

Run the experiment through the path production uses

Create one synthetic failed-notification fixture. Emit it from the current release, deploy a compatible candidate release and emit it again, then restore the current release and emit a third record. Query all three by stable fields. Do this through stdout, the normal collector or agent, its buffering layer, regional routing, and the destination under evaluation. Posting the fixture directly to a backend API skips most of the system being tested.

Then interrupt export briefly and restore it. The purpose is not to demand exactly-once log delivery. It is to observe and document whether records are lost, delayed, or duplicated, and to prove that notification_id plus attempt makes duplicates recognizable. A candidate fails the rollback test if nested values become an opaque string, numeric attempts change type, event time is replaced by ingestion time, or one release's optional fields make the shared query fail.

I would put the checks in this order:

  1. Can one query find all three terminal records by event_name, release, region, and outcome?
  2. Are timestamps, integers, booleans, nulls, and nested attributes preserved rather than coerced?
  3. Does the US fixture reach only the intended US destination, and the EU fixture only the intended EU destination?
  4. After interrupted export, can the team distinguish a gap from a duplicate?
  5. Can an engineer correlate each attempt with Postgres using non-sensitive identifiers?

The third check deserves real attention. "Supports regions" is not evidence that the application's routing, collector configuration, storage location, and access policy form the intended boundary. Use synthetic events with unmistakable region markers and verify both positive and negative cases.

Group errors by the action they require

Once ingestion works, grouping can still obscure a rollback. Stack-based grouping may combine failures that require different responses, or split one operational problem after a minor code change. Sentry's public grouping documentation explains how stack traces, exception information, messages, and custom fingerprints can influence issue identity. The transferable lesson is about response boundaries, not a particular product.

Use a controlled, low-cardinality failure_class such as a timeout category only when the default grouping fails the team's operational test. Never build a fingerprint from shipment_id, notification_id, or raw exception text containing variable values. That produces a stream of one-off groups instead of a usable signal.

There is a trade-off here. Coarse groups reduce alert noise but can hide a channel-specific regression; fine groups preserve detail but increase triage work. Test grouping with at least two failure classes and two releases, then ask whether each group maps to one owner and one response. If it doesn't, the grouping rule is wrong for this service.

Decide with evidence, not a feature matrix

Only after the rollback drill passes should search ergonomics, retention, access control, export portability, and operational effort decide among candidates. Give an engineer who did not build the test three tasks: isolate the failing release, identify whether failures cluster by region, and list retryable notification attempts. Record the time to a correct answer and any undocumented query knowledge they needed.

Volume still matters, but price is not the thesis. Measure the 50th percentile, 95th percentile, and maximum serialized event sizes in a representative fixture set. Count delivered, deferred, retried, and terminal events. Look for unbounded attributes before estimating retention or ingestion needs; shipment identifiers are necessary for reconciliation, while copied payloads and duplicate framework messages usually are not.

Be honest about what this method does not prove. One clean drill covers one application path. It does not establish behavior for every network partition, collector failure, or retention boundary. Repeat the drill after changes to the runtime, collector, event schema, or regional topology, and keep contract assertions in the release eval suite so drift fails early.

The selection rule is straightforward: accept the service that preserves the versioned event and answers the rollback query through the real production path with an operational burden the team can sustain. The durable asset is the contract. With it, a small logistics SaaS can reverse a release, identify delivery failures across mixed versions, reconcile ambiguous attempts with Postgres, and restore traffic without guessing from free-form lines.

References

Top comments (0)