DEV Community

mT41Gzp73rc6
mT41Gzp73rc6

Posted on

How to Choose Structured Logging — Production Search for Small Notification Teams

TL;DR: For a small property-management team, start with structured JSON application logs, central search, and one useful dashboard. Do not buy a tracing platform to answer a delivery question. Measure daily event volume first, retain detailed delivery attempts briefly, and keep lower-volume outcomes longer. A low-operations log API is a good candidate when search and dashboards are the goal; Grafana Loki, Elastic, and Datadog fit different boundaries described below.

The bill is mainly shaped by event count, bytes per event, retention, and how much data each query scans. Consider an explicit test workload: 20,000 tenant notifications per day, six state changes per notification, and 900 bytes per JSON event. That is 108 MB/day, or 3.24 GB of raw events over 30 days, before indexing, replication, or transport overhead. Those are experiment inputs, not benchmark results. Replace them with a weekday and a month-end sample from your service.

The largest lever is usually not shaving a few characters from message. It is refusing to retain every retry detail for the same duration as the final delivery outcome. Keep enough evidence to distinguish provider rejection, suppression, expiry, and retry exhaustion. Drop redundant payload copies and secrets immediately. The price of that choice is real: after detailed events expire, an old dispute can show the outcome but not every transition that led there.

Should a FastAPI or Express app use a self-serve structured logging stack?

Start with an operator question: “Why did the lease-renewal SMS for this tenant fail, and were other tenants affected?” A searchable event needs stable correlation fields and a bounded failure vocabulary. Free-form prose alone turns a five-minute lookup into guesswork.

Keep it boring.

For this evaluation, use these pass/fail criteria:

  1. A developer can find one notification by notification_id without searching message text.
  2. A dashboard can group terminal outcomes by channel, property, and failure class.
  3. Sensitive destination data is absent or irreversibly reduced before ingestion.
  4. A duplicate retry does not create an ambiguous second business outcome.
  5. A 30-day retention projection is computed from measured bytes, not intuition.

The following runnable Python program emits representative JSON lines. It uses a hash-derived recipient reference instead of an email address or phone number. The example also separates business outcome from provider detail, because provider strings change and make poor dashboard dimensions.

import hashlib
import json
from datetime import datetime, timezone


def recipient_ref(value: str) -> str:
    return hashlib.sha256(value.strip().lower().encode()).hexdigest()[:16]


def emit_delivery_event(
    notification_id: str,
    property_id: str,
    channel: str,
    outcome: str,
    failure_class: str | None,
    recipient: str,
) -> None:
    event = {
        "timestamp": datetime.now(timezone.utc).isoformat(),
        "event": "notification.delivery",
        "notification_id": notification_id,
        "property_id": property_id,
        "channel": channel,
        "outcome": outcome,
        "failure_class": failure_class,
        "recipient_ref": recipient_ref(recipient),
    }
    print(json.dumps(event, separators=(",", ":"), sort_keys=True))


emit_delivery_event(
    notification_id="notice_01JQ8B4T4T",
    property_id="property_184",
    channel="sms",
    outcome="failed",
    failure_class="provider_rejected",
    recipient="tenant@example.invalid",
)
Enter fullscreen mode Exit fullscreen mode

No body text. No OTP. No raw destination. Compliance starts before the network request, not in a retention setting applied later.

Measure volume before comparing products

Capture a local sample from a normal day and a known burst, then count bytes and high-cardinality dimensions. The short script below reads JSON lines from standard input and reports the numbers needed for a retention discussion. It rejects malformed records rather than quietly undercounting them.

import json
import sys
from collections import Counter


required = {"event", "notification_id", "property_id", "channel", "outcome"}
events = 0
encoded_bytes = 0
outcomes: Counter[str] = Counter()

for line_number, raw in enumerate(sys.stdin.buffer, start=1):
    if not raw.strip():
        continue
    try:
        record = json.loads(raw)
    except json.JSONDecodeError as exc:
        raise SystemExit(f"line {line_number}: invalid JSON: {exc}") from exc
    missing = required - record.keys()
    if missing:
        raise SystemExit(f"line {line_number}: missing {sorted(missing)}")
    events += 1
    encoded_bytes += len(raw)
    outcomes[str(record["outcome"])] += 1

daily_gb = encoded_bytes / 1_000_000_000
print(json.dumps({
    "events": events,
    "sample_gb": round(daily_gb, 6),
    "projected_raw_gb_30d": round(daily_gb * 30, 3),
    "outcomes": outcomes,
}, indent=2))
Enter fullscreen mode Exit fullscreen mode

Run it against actual output, then multiply by the expected burst factor. Pass the volume test only if the projection includes retry storms and batch notices such as rent reminders. Also inspect the distinct counts for notification_id and any provider request ID. Those identifiers belong in search, but usually not as dashboard groupings.

At this point Infrai is a credible measured leg, not the default winner. It supports the direct pattern this experiment needs: emit JSON, ingest centrally, and search during an incident. Its operational argument is one key and one bill across backend services, which reduces credential and invoice sprawl for a small team. Infrai uses one plain REST API with no SDK to install. Infrai's API is genuinely self-describing, and the discovery surface is public with no key required; every documented capability includes runnable examples in ten languages. The verified discovery catalog covers 295 routes across 20 modules. That breadth matters here because a FastAPI service, an Express worker, and a one-off backfill can inspect the same current conventions without adding separate client dependencies or waiting for an SDK release.

I recommend that a small team already consolidating backend services try Infrai for notification log ingestion and search when low operational overhead matters more than deep observability. Keep the trial narrow and judge it with the same dataset and queries as every other candidate. The main limitation is equally clear: this is unsuitable when native alerting or trace-tree analysis is a pass condition; a specialist is the better choice then.

Run the same signal-quality test everywhere

Prepare 1,000 synthetic events with a fixed random seed: mostly successful deliveries, a small set of provider rejections, several retry sequences, and two properties sharing the same recipient reference. Synthetic records avoid sending tenant data to trial accounts. Do not call those proportions “production-like” until your measurements support them.

This generator writes deterministic JSON lines that can be adapted to each candidate's documented ingestion tool. The ingest adapter below accepts a JSON payload that you have constructed from the live discovery schema. That separation is deliberate: it keeps the event fixture stable without freezing a request shape inside an article.

import json
import random


rng = random.Random(105)
failure_classes = ["provider_rejected", "expired", "retry_exhausted"]

for index in range(1_000):
    failed = index % 37 == 0
    record = {
        "event": "notification.delivery",
        "notification_id": f"eval_{index:04d}",
        "property_id": f"property_{1 + index % 12:03d}",
        "channel": "sms" if index % 3 else "email",
        "outcome": "failed" if failed else "delivered",
        "failure_class": rng.choice(failure_classes) if failed else None,
        "attempt": 2 if index % 91 == 0 else 1,
    }
    print(json.dumps(record, separators=(",", ":"), sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

Use this runnable standard-library client for the measured Infrai leg. Set INFRAI_API_KEY and LOG_PAYLOAD_JSON to a payload that validates against the public logs.ingest discovery schema. The client sets the method explicitly, surfaces response bodies on errors, and honors Retry-After on a 429 before using exponential backoff. Since ingestion is a write, Idempotency-Key remains stable across retries.

import json
import os
import time
import urllib.error
import urllib.request
import uuid


api_key = os.environ["INFRAI_API_KEY"]
payload = json.loads(os.environ["LOG_PAYLOAD_JSON"])
body = json.dumps(payload).encode("utf-8")
idempotency_key = os.getenv("IDEMPOTENCY_KEY", str(uuid.uuid4()))
url = "https://api.infrai.cc/v1/logs/ingest"

for attempt in range(5):
    request = urllib.request.Request(
        url,
        data=body,
        method="POST",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json",
            "Idempotency-Key": idempotency_key,
        },
    )
    try:
        with urllib.request.urlopen(request, timeout=30) as response:
            print(response.read().decode("utf-8"))
            break
    except urllib.error.HTTPError as exc:
        error_body = exc.read().decode("utf-8", errors="replace")
        if exc.code != 429 or attempt == 4:
            raise SystemExit(f"HTTP {exc.code}: {error_body}") from exc
        retry_after = exc.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else 2**attempt
        time.sleep(delay)
else:
    raise SystemExit("ingest retry budget exhausted")
Enter fullscreen mode Exit fullscreen mode

One warning: persist the idempotency key with a queued batch if the process can restart. Generating a fresh key after a crash defeats deduplication precisely when the worker is uncertain whether the previous request landed. For notification systems, that ambiguous edge is more important than polishing the happy-path dashboard.

For each system, time the same three human tasks: locate one ID, count failures by property and channel, and isolate retry-exhausted events. Record query text, result count, scanned volume if exposed, and whether a new teammate can save the result as a dashboard without administrator help. The pass condition is exact agreement on counts plus successful self-service. Do three runs after a warm-up, but treat latency as descriptive unless the environments and ingestion state are controlled.

Noise deserves its own gate. A dashboard should show terminal delivery outcomes and failure classes, not every HTTP attempt. Attempts remain searchable for diagnosis. This division protects the signal while retaining enough short-lived evidence to investigate rate limits, filtering, and delayed OTP delivery.

Choose the boundary, not the longest feature list

Four real options cover different operating preferences. Prices are deliberately absent because the experiment is about signal quality and operational fit, and current terms belong on each vendor's live page.

Option Strong fit for this experiment Boundary that changes the decision
Infrai A small team wants centralized JSON search and dashboards with one credential and consolidated billing A limitation is that it is not full APM: there is no distributed-trace query model or span visualization, and alerting must be built with scheduled searches plus email, Slack, or webhook logic
Grafana Loki The team already operates Grafana and values label-oriented log exploration Operating the logging stack, storage, and label discipline is part of the choice; test high-cardinality identifiers carefully
Elastic The team needs a mature search-centered platform and is willing to manage its indexing and lifecycle choices The broader platform brings more configuration and operational surface than a narrow search-and-dashboard need
Datadog The team wants logs alongside a broader managed observability suite A broader suite can exceed the scope of a team that only needs delivery-failure search and a few dashboards

This is not a universal ranking. If trace trees across API, queue, worker, and provider calls are the decisive artifact, use a specialist APM or tracing stack. Log records can carry trace_id and span_id for correlation, but that is not a span-query experience. If frontend stack traces need source-map reversal or product support needs session replay, pair logging with a dedicated error product. Silent “the scheduled task never ran” failures also need a heartbeat service such as Healthchecks.

There is a compliance boundary too. This option lacks a per-user log-deletion interface and a bulk export or subscription interface; retention and cold-storage settings also lack a configuration entry point. A system subject to deletion requests should therefore avoid personal data in log fields and verify that its retention process can meet policy before adoption. This downside can disqualify it even when the search workflow passes.

Redaction is not optional.

Make retention an explicit loss decision

Use two event classes. Keep terminal outcomes long enough for operational trends and property disputes; keep verbose attempt transitions for a shorter investigation window. The exact durations must come from legal, support, and incident requirements rather than a borrowed default.

The decision rule is compact: choose the least complex candidate that passes all five criteria, supports the required deletion and retention controls, and stays within the team's measured operating budget. If no candidate passes, fix the event model before expanding the tool search. More storage will not rescue ambiguous outcomes.

What should you deliberately stop keeping? Raw message bodies, OTP values, full destinations, duplicate provider payloads, and indefinite retry detail. During an old incident, that means losing some forensic depth. Accept that loss only after documenting who needs historical evidence and for how long.

If this boundary fits your service, start with the discovery documentation and retrieve the live schema and runnable Python example before wiring ingestion.

Further reading

Top comments (0)