DEV Community

YukiKobayashi880
YukiKobayashi880

Posted on

Notification Cost Attribution: Structured JSON API Logs for Small SaaS

For a small SaaS app, keep one compact, structured delivery outcome per attempt in the searchable logging service, and move verbose diagnostic context out of that tier before shortening retention. The dominant cost is usually determined by multiplication: attempts per day, bytes per event, indexed-field overhead, and retained days. A simple service is one whose bill can be assigned to a tenant, channel, and outcome without putting message bodies or exception dumps into every record.

TL;DR: start with a byte budget and an event budget, not a dashboard tour. Preserve the fields needed to count accepted, failed, retried, and terminal deliveries; sample repetitive diagnostics; keep correlation identifiers; and treat retention as the last lever. This protects the failure history while removing data that contributes storage and search work but does not change an operational decision.

What should a small SaaS app logging service index?

The useful first estimate is deliberately plain:

def retained_gib(events_per_day: int, bytes_per_event: int, days: int) -> float:
    return events_per_day * bytes_per_event * days / (1024 ** 3)


assumptions = {
    "delivery_attempts_per_day": 240_000,
    "compact_event_bytes": 900,
    "verbose_event_bytes": 6_000,
    "retention_days": 14,
}

for event_size in (assumptions["compact_event_bytes"], assumptions["verbose_event_bytes"]):
    print(round(retained_gib(
        assumptions["delivery_attempts_per_day"],
        event_size,
        assumptions["retention_days"],
    ), 2))
Enter fullscreen mode Exit fullscreen mode

Those values are a worked planning model, not a benchmark or a vendor quote. Under its assumptions, compact events retain about 2.82 GiB of raw JSON, while verbose events retain about 18.78 GiB. Indexes, replicas, compression, transport, and query processing can change the billed amount, so raw GiB can't forecast a price. It does expose the controlling term: changing 6,000 bytes to 900 bytes moves more data than shaving a day from an already short incident window.

That distinction matters because a delivery attempt tends to accumulate convenient debris: a rendered message, a provider response, request headers, an exception stack, and several copies of tenant metadata. Most of it is high-volume context. Only a smaller event spine is required to answer who incurred work, which channel failed, whether a retry occurred, and how the sequence ended.

OpenTelemetry's logs data model separates the log body from attributes and supports trace and span context. That boundary is useful even without adopting a particular backend: keep a bounded outcome in the body or event name, put stable dimensions in attributes, and retain correlation separately. Do not flatten an entire application object into indexed fields merely because the serializer permits it.

Attribute cost without turning tenants into metrics

Cost attribution for this developer-tools scenario needs three distinct questions: which tenant generated delivery work, which notification channel produced it, and which outcome caused extra attempts. A log record can carry the pseudonymous tenant key because investigation requires exact lookup. A metric label generally should not copy that unbounded key; aggregating by tenant can create a high-cardinality series set when the account population grows.

The event contract can stay small:

from datetime import datetime


REQUIRED_FIELDS = {
    "timestamp",
    "service",
    "region",
    "tenant_key",
    "notification_id",
    "channel",
    "attempt",
    "outcome",
}

ALLOWED_OUTCOMES = {"accepted", "temporary_failure", "permanent_failure", "delivered"}


def validate_delivery_event(event: dict) -> None:
    missing = REQUIRED_FIELDS - event.keys()
    if missing:
        raise ValueError(f"missing fields: {sorted(missing)}")

    datetime.fromisoformat(event["timestamp"].replace("Z", "+00:00"))
    if event["outcome"] not in ALLOWED_OUTCOMES:
        raise ValueError("unknown delivery outcome")
    if not isinstance(event["attempt"], int) or event["attempt"] < 1:
        raise ValueError("attempt must be a positive integer")

    forbidden = {"email", "phone", "message_body", "access_token"}
    exposed = forbidden & event.keys()
    if exposed:
        raise ValueError(f"sensitive fields present: {sorted(exposed)}")
Enter fullscreen mode Exit fullscreen mode

This schema supports exact per-tenant accounting in logs and bounded aggregation by region, channel, and outcome in metrics. It also avoids pretending that a region attribute proves where ingestion, indexing, replicas, support access, or backups occur. US and EU paths need their own documented processing boundaries; a JSON string is evidence about the event, not the storage architecture.

Keep attribution keys where exact investigation belongs. The tempting shortcut is to attach every customer and notification identifier to every telemetry type. That improves one query and quietly increases storage duplication, index width, and cardinality elsewhere.

Which bytes should remain searchable?

Retain the outcome spine for every attempt: occurrence time, service, deployment, region, pseudonymous tenant, notification correlation, channel, attempt number, and bounded outcome. Keep an ingestion timestamp when the platform exposes one, because occurrence order and observation order can differ. Preserve trace or request context when it exists; OpenTelemetry defines trace and span identifiers as part of log correlation.

Then split the bulky evidence by its operational value.

Data class Searchable-tier policy Cost and failure trade-off
Delivery outcome Keep each attempt for the incident window Supports retry and terminal-state reconstruction
Repeated stack trace Keep a fingerprint plus sampled examples Loses some instance-level diagnostic variation
Provider payload Extract bounded status; do not retain the raw body by default Reduces forensic detail and exposure
Rendered message Exclude Cannot inspect exact content from logs
Correlation context Keep stable identifiers Adds small per-event overhead but preserves joins
Tenant identifier Keep a pseudonymous key in logs Enables attribution while still requiring access controls

Sentry's documentation illustrates why a fingerprint is useful: events can be grouped automatically, and custom fingerprints can alter grouping. Grouping and delivery accounting are different jobs, though. A fingerprint can collapse repeated exceptions for diagnosis; it cannot replace the attempt records needed to determine that three sends were charged to one notification or that a temporary failure became a permanent one.

Do not sample terminal outcomes. Sample verbose repetitions after recording their compact outcome, and make the sampling decision explicit in a field or companion counter so an operator does not mistake sampled diagnostic volume for delivery volume.

Tiny records win.

Failure modes that distort both incidents and spend

Duplicate ingestion is the obvious trap. If the notification worker retries transmission after an ambiguous acknowledgment, the logging path may record the same logical attempt twice. Give each attempt a stable application identifier and define whether cost reports count unique attempts, received records, or provider submissions. These are not interchangeable totals.

Late and out-of-order arrival is subtler. Billing attribution by ingestion day can disagree with operational attribution by occurrence day near a reporting boundary. Preserve both timestamps where possible and state which one drives each report. Otherwise a delayed EU batch can appear as a fresh burst of notification failures and spend.

Backpressure also deserves a hard policy. Logging should sit outside the user request's critical path, with a bounded buffer and a measurable dropped-event signal. An unbounded queue transfers an indexing problem into application memory; silent dropping makes the cheap-looking system impossible to audit. Neither outcome is simple.

Schema drift creates a slower failure. If attempt changes from an integer to a string, or outcome starts carrying exception prose, filters fragment and aggregation becomes unreliable. Validate before emission, version semantic changes, and test the contract in continuous integration.

Set retention from the investigation window

Retention should cover the interval in which the team realistically discovers and investigates delivery failures, plus any reporting obligation established outside the logging system. No universal day count follows from the available evidence, so choose it from the notification workflow rather than copying a plan default. Measure event volume and size at the application boundary, then verify searchable and exported counts with synthetic data.

A practical decision order is:

  1. Remove secrets, destinations, rendered content, and raw payloads that should never have crossed the boundary.
  2. Replace repeated exception bodies with a stable fingerprint and retain sampled diagnostic examples.
  3. Stop indexing fields that are never filtered, grouped, or joined.
  4. Separate exact tenant investigation in logs from bounded aggregate dimensions in metrics.
  5. Shorten searchable retention only after the incident and reporting windows are explicit.

Compare candidate services with the same fixture and the same byte model. Verify that numeric fields remain numeric, exact filters compose across tenant and outcome, late records retain occurrence time, exports preserve structured fields, and access controls match the people who investigate failures. Pricing belongs in this final comparison, using the team's measured volume and the candidate's current terms; a transient advertised rate is a poor architecture constraint.

The deliberate loss is clear. Excluding message bodies means an incident responder cannot reconstruct exact rendered content from the log store. Sampling repeated stacks means rare variation may be absent. Shorter retention means an old tenant report may retain aggregate counts but no event-level path. Those costs are acceptable only when another governed system owns the evidence or the team has explicitly decided it is unnecessary. Cheap storage is not the goal. Accountable loss is.

Further reading

Top comments (0)