TL;DR: For a small fintech SaaS, start with one external probe in each customer region, one dead-man signal per scheduled job, and narrowly scoped application evidence that can reconstruct a customer incident. Attribute bytes and signal volume to a service, region, environment, and evidence class before choosing retention. The simplest workable design distinguishes an unreachable process, a broken dependency path, and a missing cron completion without retaining every log line.
This decision is often presented as a tool comparison. That skips the expensive question: what, exactly, will be retained? Probe history is rarely the dominant term. High-volume application events and high-cardinality metrics are stronger candidates. A sensible 2026 design begins with the evidence budget and works backward to collection, rather than turning on everything and discovering later that nobody can assign the bill or reconstruct a disputed payment.
What is the bill actually made of?
Model four quantities separately: items produced, bytes per item after encoding, retention duration, and the number of copies. Indexes, replicas, and cross-region copies are distinct multipliers even when an invoice collapses them into one line. Query work and alert evaluation may add another term, but retained volume is the first term worth making visible because policy controls it directly.
Use assumptions, not folklore. The following Python sketch is a calculator, not a benchmark. Its values are illustrative inputs: two regional probes at a 60-second interval, 25 metric series sampled every 15 seconds, 120 application events per minute, and 30 days of retention. Replace every value with measurements from the service before making a commitment.
from dataclasses import dataclass
SECONDS_PER_DAY = 86_400
@dataclass(frozen=True)
class EvidenceClass:
name: str
items_per_second: float
bytes_per_item: int
retention_days: int
copies: int = 1
def retained_bytes(self) -> float:
return (
self.items_per_second
* SECONDS_PER_DAY
* self.retention_days
* self.bytes_per_item
* self.copies
)
evidence = [
EvidenceClass("regional_probe", 2 / 60, 300, 30),
EvidenceClass("service_metrics", 25 / 15, 24, 30),
EvidenceClass("application_events", 120 / 60, 900, 30),
]
for item in evidence:
gib = item.retained_bytes() / (1024 ** 3)
print(f"{item.name}: {gib:.3f} GiB retained")
The arithmetic exposes the dominant term without pretending that an encoded metric sample or indexed event has a universal size. Measure those values after ingestion, including index and replication overhead. If application events dominate, changing a probe from 60 seconds to 30 seconds is probably a distraction; reducing duplicated payloads, cardinality, or retention moves the larger term.
Cost attribution needs stable dimensions, but restraint matters. Require service, environment, region, and evidence_class on stored telemetry, plus a team or cost-center mapping maintained outside the event payload. A customer identifier can help incident reconstruction, yet putting unbounded customer or transaction IDs into metric labels creates a series-growth failure mode. Keep those identifiers in controlled event records, then link them through a correlation key. Prometheus naming guidance recommends base units, meaningful suffixes, and labels whose combinations do not create excessive cardinality.
What should small SaaS Node uptime monitoring retain from health checks?
A public health response proves only that a particular request completed from a particular vantage point. It does not prove that every customer path works. Make the response cheap and bounded: process readiness may check local state, while a separate dependency-aware transaction can exercise the minimum path required to accept work. If the same endpoint writes against every dependency on every probe, monitoring becomes load generation and can amplify an outage.
Keep that boundary.
For an EU/US service, place external probes so each customer region is observed independently. Preserve the probe timestamp, region, target, outcome, latency, and a coarse failure class. One global success rate can hide a regional routing or dependency failure; two regional streams allow the incident timeline to say which path failed first. Do not interpret disagreement as noise until DNS, routing, certificate, and application causes have been separated.
Cron jobs need the opposite shape. A failed run can emit an event, but a process that never starts emits nothing. Record a completion heartbeat with the job identity, scheduled window, completion time, outcome, and an execution identifier. Alert on absence after the deadline plus an explicit grace period. Preserve retry attempts under the same execution lineage so a later success does not erase evidence of the first failure.
Silence is evidence.
The third signal is a compact application event around the customer-impacting boundary: request or job correlation ID, region, operation class, outcome, duration, and a redacted reason code. Syslog's severity scheme supplies a shared vocabulary, but severity alone cannot reconstruct a transaction. RFC 5424 defines severity numerically; the event still needs structured context.
| Evidence | Question it answers | Retention hazard | Attribution key |
|---|---|---|---|
| Regional probe | Could an outside client reach the service? | Frequent success records with little diagnostic value | Region and service |
| Cron heartbeat | Did scheduled work finish inside its window? | Retries counted as unrelated runs | Job and execution lineage |
| Application event | Which operation failed, where, and why? | Sensitive payloads and unbounded identifiers | Service and evidence class |
| Low-cardinality metric | Is the failure broad or isolated? | Series explosion from customer-level labels | Service, region, environment |
Retain a signal because it answers a reconstruction question, not because an agent can collect it.
How much evidence survives a customer dispute?
Start with a reconstruction worksheet before setting a duration. For a disputed transfer, an investigator may need to establish when the request entered the system, which region handled it, whether a scheduled settlement step ran, what outcome each boundary produced, and whether the external service was reachable at the same time. This is a chain of evidence. A dashboard screenshot is not the chain.
Assign each field one of three treatments: retain directly, derive and aggregate, or discard. Keep correlation and outcome fields long enough for the applicable investigation window. Aggregate routine successful probe results into coarse availability metrics after their short diagnostic window, if individual successes no longer add reconstructive value. Drop request bodies, credentials, and unrestricted error dumps from observability storage; they increase exposure and rarely improve the timeline enough to justify it. The exact duration is a policy decision tied to legal, contractual, and operational requirements. A generic monitoring checklist cannot supply it.
This is where logs and metrics stop competing. Metrics locate the interval and scope. Structured events explain a selected operation. Probe records establish an outside view. A cron heartbeat proves presence, while its absence triggers investigation. Trying to force all four questions into application logs yields expensive searches and fragile absence detection; forcing them all into metrics loses transaction detail.
There is also a consistency boundary. Telemetry delivery can lag or fail during the same incident it describes, so record source timestamps and ingestion timestamps separately, monitor the collector path, and avoid claiming a complete timeline until expected heartbeats have either arrived or exceeded a declared lateness bound. Duplicate delivery must not become duplicate evidence: use a stable event identifier and make ingestion idempotent where the storage interface permits it.
What breaks this design first?
Cardinality is the quiet failure. A label such as customer_id, payment_id, or raw URL can turn a compact metric family into an open-ended set of series. The repair is architectural: metrics carry bounded dimensions; event storage carries controlled correlation fields. Renaming the label does nothing. Clock ambiguity comes next. If regional probes, application nodes, and job runners disagree on time, the reconstructed sequence can reverse cause and effect. Preserve both occurrence and ingestion times, use UTC, and track clock synchronization as an operational dependency. A timestamp without its semantic origin is weak evidence. Then there is false health. A process can answer a shallow endpoint while its payment dependency is unusable, or a dependency-heavy endpoint can declare the entire service down during a noncritical subsystem failure. Define readiness around the work that instance can safely accept, and monitor important customer paths separately. Their alerts should identify the failed boundary rather than compress every condition into healthy: false.
Absence alerts fail around deployment pauses, daylight-saving assumptions, and retries. Store schedules in an explicit time zone, calculate deadlines deterministically, and represent maintenance suppression as data with an owner and expiry. Otherwise silence is indistinguishable from scheduler failure.
Cost attribution can become fiction when shared collectors stamp every byte onto the platform team. Attribute at ingestion from authenticated source metadata, reconcile unattributed volume, and reject dimensions that application code can freely rewrite. A monthly total with an unknown bucket is more honest than a precise allocation built from untrusted tags.
Measure first.
A retention decision that remains testable
Deploy the design as a contract. In staging, stop the HTTP listener, break one required dependency, suppress one cron completion, duplicate an event, and delay telemetry delivery. Each test should produce a different state and a reconstruction record with the expected service, region, and evidence class. Run the same checks after collector, schema, or routing changes; observability that is never failure-tested is an assumption.
Review volume by evidence class and owner at a fixed interval. A sudden rise should be explainable by traffic, a schema change, a retry storm, or a new dimension. Set a budget on volume and cardinality, but alert before enforcing destructive limits so the response does not erase the period under investigation. Producer error handling should be bounded: monitoring must not block a customer operation indefinitely when its sink is unavailable.
This design has explicit limitations. Regional probes cannot prove every transaction path. Aggregates cannot replay individual checks, and redacted events cannot recover an original submitted value. A very low-volume service with short investigations may accept detailed records for longer; a high-volume or sensitive service should discard more and invest in stronger correlation. No one retention tier fits both.
After the short diagnostic window, stop keeping individual successful probes and verbose success events when aggregates preserve the needed availability view. Do not retain raw payloads in telemetry. Keep narrowly structured failure and lineage records for the approved investigation period, then expire them under policy.
The cost is real. Once detailed successes expire, an investigator may prove that a region was broadly available but may no longer replay every probe. Storage saved by deletion is purchased with reduced reconstructive detail, and that trade-off should be approved before an incident.
References
- Prometheus, "Metric and label naming": https://prometheus.io/docs/practices/naming/
- IETF RFC 5424, "The Syslog Protocol": https://datatracker.ietf.org/doc/html/rfc5424
Top comments (0)