Keep a compact, queryable failure ledger in your own database and render customer-facing KPIs from pre-aggregated tables. That is the least complex way to show useful logistics checkout metrics without turning raw logs into an expensive analytics store. Use Metabase, Redash, Supabase Charts, or a managed metrics API as the presentation and query boundary; do not make any of them the only copy of the evidence.
TL;DR: the dominant cost is usually the number of retained events multiplied by their size and retention period, plus the work required to scan them. Record one small, structured outcome per checkout attempt, aggregate by tenant and time bucket, and retain raw failure evidence only long enough to investigate disputes and delivery gaps. This improves signal quality because expected declines, retries, and duplicate callbacks stop masquerading as separate incidents.
What are you actually paying to keep?
A dashboard looks like a charting problem. The bill starts earlier: collection, storage, indexing, repeated queries, egress, backups, and the engineering time spent controlling cardinality. A managed log service may price ingestion and indexed retention separately; Datadog's public pricing page is one concrete example of that model. The exact unit price changes, so it is a poor architecture anchor.
Measure the term you control instead. Suppose a checkout outcome contains 900 bytes after serialization. At 12 million attempts per month, one copy is about 10.8 GB before database overhead, indexes, replicas, and backups. Keeping every HTTP header, stack trace, carrier response, and retry as a 9 KB searchable event changes the base payload to roughly 108 GB. That tenfold change matters more than choosing a chart renderer.
The useful unit is not “a log line.” It is an outcome with a stable identity:
from dataclasses import dataclass
from datetime import datetime
@dataclass(frozen=True)
class CheckoutOutcome:
attempt_id: str
tenant_id: str
occurred_at: datetime
status: str # success, expected_decline, or system_failure
failure_class: str | None
retry_number: int
region: str
Do not put email addresses, phone numbers, street addresses, access tokens, or free-form exception messages in this record. Checkout observability crosses compliance boundaries quickly, and a KPI store rarely needs delivery credentials or personal data. Keep diagnostic detail behind narrower access and a shorter retention policy.
Reduce the event before choosing the chart
Signal quality comes from classification. A payment rejected for insufficient funds is an expected decline; a carrier quote timing out is a system failure; a customer clicking twice may produce a duplicate request but only one logical attempt. If all three increment checkout_failed, the dashboard is loud and operationally weak.
Start with an idempotent attempt identifier. Normalize retries under that identifier, then write one terminal outcome when possible. Late callbacks need an explicit correction path, because delivery systems reorder messages. The aggregation job can upsert a five-minute bucket keyed by tenant, region, and failure class.
from collections import Counter
def summarize(outcomes: list[CheckoutOutcome]) -> dict[str, int]:
latest: dict[str, CheckoutOutcome] = {}
for outcome in outcomes:
current = latest.get(outcome.attempt_id)
if current is None or outcome.retry_number >= current.retry_number:
latest[outcome.attempt_id] = outcome
counts = Counter(item.status for item in latest.values())
return {
"attempts": len(latest),
"successes": counts["success"],
"expected_declines": counts["expected_decline"],
"system_failures": counts["system_failure"],
}
This example deliberately has no currency amount and no customer identifier. It answers the operational question while keeping the dimensional surface small.
Small wins.
Should you use self-hosted Metabase, Redash, or Supabase Charts?
Metabase and Redash are self-hosted business-intelligence applications that can query a database and expose dashboards. Supabase Charts belongs in the same decision conversation when the metrics already live in a Supabase-backed data path. A managed metrics API moves more of the ingestion, query, and serving boundary outside your application. Those are deployment boundaries, not rankings.
| Boundary | Cost you operate | Best signal path | Main noise risk |
|---|---|---|---|
| Self-hosted BI over aggregates | Runtime, upgrades, database queries, access control | SQL over reviewed bucket tables | Ad hoc queries drifting onto raw events |
| Application-owned charts | API, cache, frontend, authorization | Narrow tenant-scoped KPI contract | Metric definitions diverging across screens |
| Managed metrics API | Sent series, retained dimensions, query volume, integration | Standardized counters and latency distributions | Unbounded labels and duplicate emissions |
The deciding question is where metric semantics live. If an analyst can silently redefine “failure” in a chart, an incident channel and a customer dashboard can show different truths. Put the definition in a reviewed transformation or application module, then let every presentation layer read the same aggregate.
Keep that boundary boring.
For an in-app view, tenant isolation deserves a separate check. A dashboard URL is not an authorization model. Enforce tenant scope before query execution, test it with two tenants, and cache using a key that includes tenant and metric version. This is also where self-hosting has a real labor cost: patching, backups, query limits, and authentication remain yours even when the software license has no usage charge.
Retention should follow the investigation clock
Use separate clocks for aggregates and evidence. Five-minute aggregates can remain useful for seasonal comparisons long after individual events have lost operational value. Raw outcomes should survive the longest realistic window for incident review, reconciliation, and support escalation. Detailed payloads should expire sooner still.
A practical policy might keep aggregate buckets for 13 months, normalized outcomes for 30 days, and restricted diagnostic payloads for 7 days. Those durations are examples, not universal compliance guidance. Legal obligations, chargeback windows, carrier reconciliation, and internal incident practice determine the real values. Document the owner and deletion mechanism for each tier. Retention also has a failure mode that teams discover late: deleting database rows while leaving replicas, exports, and backups untouched. Test expiration as a system behavior. Restore a backup in an isolated environment, verify what reappears, and record the maximum erasure delay. The change that moves the dominant term is dropping verbose evidence from the searchable path, not shaving pixels from a graph or chasing a temporarily lower service price. Sampling can help with repetitive diagnostics, but never sample the denominator needed to calculate a failure rate. A precise numerator over an estimated denominator is deceptive.
Test the dashboard like a delivery system
Dashboards fail quietly. Build fixtures for success, expected decline, timeout, duplicate callback, retry followed by success, and an event arriving after its bucket has closed. Then assert both the aggregate and the tenant-visible API response. A single end-to-end test should cross the write, aggregation, authorization, and rendering boundaries.
Watch freshness separately from checkout health. If the last completed bucket is old, show stale data rather than a comforting zero. Track aggregation lag, rejected events, correction count, and query latency as operational signals; keep them off the customer KPI panel unless customers need them to interpret the result.
Zero is ambiguous.
Web performance is part of the delivery path too. Core Web Vitals define Largest Contentful Paint, Cumulative Layout Shift, and Interaction to Next Paint, with guidance based on the 75th percentile. Reserve chart dimensions to prevent layout movement, defer noncritical series, and measure the actual embedded page. A fast SQL query does not guarantee a usable dashboard.
Deploy metric-definition changes with a version. During a transition, compute old and new definitions side by side, compare them, then switch readers deliberately. This catches classification drift before an apparent logistics outage is created by code rather than operations.
What I would deliberately stop retaining
I would stop keeping full request and response bodies in the general KPI path, along with arbitrary headers, raw personal details, duplicate retry events, and high-cardinality error strings. I would also reject dimensions that lack a named operational decision. Region can route an investigation; a randomly generated session label usually cannot.
The trade-off is concrete. After the restricted payload expires, an old incident can be explained by its attempt identity, class, timing, region, retry history, and aggregate context, but the exact remote response may be gone. That can make a rare historical dispute slower or impossible to reconstruct. Keeping a small audited sample or a separately governed case record may be justified, but indefinite searchable retention is not the default.
Choose the dashboard boundary only after making that loss explicit. The cheapest sustainable design is the one whose retained evidence answers real incident questions, whose aggregates stay trustworthy under retries, and whose operating work your team can actually own.
Top comments (0)