Short answer: for a Node SaaS app, simple uptime monitoring needs three separate contracts: whether US and EU users can reach a health endpoint, whether a nightly cron job has a missed run, and whether enough structured evidence remains to roll back a bad media-pipeline result. A heartbeat-only alternative leaves the first and third questions unanswered. The storage bill is driven by event volume multiplied by retained bytes and retention time, so reduce the last two terms deliberately; do not weaken the signals that page an operator.
For a concrete sizing model, let J be nightly jobs, E the structured events per job, B the average stored bytes per event after indexing overhead, and D the retained days. The hot evidence footprint is J * E * B * D. This is a planning identity, not a benchmark. If your own measurements show indexed event bytes dominate, changing a check interval will barely matter; shortening hot retention or storing fewer redundant success events will.
One illustrative policy uses a 48-hour rollback window and a 30-day compact completion ledger. Those numbers are design inputs to review against the pipeline's rollback promise, not universal defaults. I would stop keeping every successful step event after the hot window, while retaining immutable run summaries longer. The price of that choice is explicit: an investigation discovered after 48 hours can establish which run and artifact changed, but it may no longer reconstruct every successful intermediate step. This is the central trade-off, and measuring actual event size is the only honest way to know whether it moves a particular bill.
What should Node SaaS app uptime monitoring prove?
A single /health response proves very little. It can show that one HTTP path answered at one moment; it cannot prove that a nightly transcode, metadata extraction, or catalog publication completed. A cron-style heartbeat has the opposite blind spot: a worker can report completion while the public service is unreachable from one region. Keep those failure domains separate.
The external probes should exercise the same public boundary a client uses from at least the US and EU. The endpoint should remain cheap and deterministic, while its response represents only the dependencies required for serving traffic. The nightly pipeline needs a durable completion record containing a run identifier, scheduled window, completion timestamp, input snapshot identifier, output artifact identifier, status, and schema version. Freshness is then evaluated against the schedule plus an explicit grace period.
Silence wins.
A process that never starts cannot emit an exception, so the monitor must derive a missed run from the absence of a valid completion record. Generic health checks and dedicated dead-man's-switch services can both carry that signal, but neither changes the evidence contract. Event-grouping systems solve a different problem: they combine similar reported failures, and their grouping or fingerprint rules should not be mistaken for a liveness clock.
Model the evidence before choosing retention
Rollback safety depends on lineage, not on the sheer number of log lines. For each run, preserve enough information to answer four questions: what inputs were selected, what code or configuration produced the output, which artifact became current, and which previous artifact can be restored. A completion ledger with stable fields is more useful for that decision than thousands of free-form completed step messages. Consider a nightly catalog run that reads snapshot S42, writes artifact A43, and promotes a pointer from A42 to A43: the detailed logs explain individual transformations, but rollback needs the small, durable statement that S42 produced A43, that A42 was previously current, and that promotion occurred only after validation. If the pointer changes before that record is durable, an availability probe may stay green while operators lack a trustworthy reversal path. If the record is durable but verbose logs expire, rollback still works, though a later forensic reconstruction loses detail. That asymmetry is why the summary deserves a longer retention class.
Keep the chain small.
Metric names deserve the same discipline. Prometheus recommends a single-unit base unit and names whose aggregation remains meaningful. A counter such as media_pipeline_runs_total can be partitioned by a bounded status label; a run identifier should stay out of metric labels because it creates a new time series for every run. Put that identifier in the completion record and structured logs instead.
| Evidence | Retention role | Failure it exposes | Boundary |
|---|---|---|---|
| Regional HTTP probe result | Short operational history | Public path unavailable or region-specific failure | Does not prove batch completion |
| Last valid completion timestamp | Longer compact history | Missed or late nightly run | Does not explain the failed step |
| Run summary and artifact lineage | At least the rollback decision window | Wrong input, output, or promotion target | Needs immutable identifiers |
| Detailed per-step events | Short diagnostic window | Localizes a failure within a run | Highest volume and safest place to trim first |
This split changes the dominant storage term without discarding the rollback chain. Compact summaries live longer; verbose success-path events expire sooner. Failure events may justify a different retention class, but classify them with bounded fields rather than fragile message parsing. The limitation is equally plain: this approach is not suitable when policy requires complete event-level reconstruction for longer than the chosen hot window. In that case, retain the detailed records in a lower-cost archive and test retrieval time; do not pretend a compact ledger is an equivalent substitute.
Encode missed-run logic as data
The evaluator should be testable without a monitoring vendor or a running web service. The function below accepts UTC timestamps, treats equality at the deadline as on time, and rejects a completion recorded in the future. The caller remains responsible for reading the durable ledger and routing the resulting state. It is intentionally only the cron freshness core, not an uptime checker or a complete monitoring system.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
@dataclass(frozen=True)
class Freshness:
state: str
age_seconds: int
def evaluate_freshness(
now: datetime,
last_completed: datetime,
expected_interval: timedelta,
grace: timedelta,
) -> Freshness:
if now.tzinfo is None or last_completed.tzinfo is None:
raise ValueError("timestamps must be timezone-aware")
now_utc = now.astimezone(timezone.utc)
completed_utc = last_completed.astimezone(timezone.utc)
age = now_utc - completed_utc
if age < timedelta(0):
return Freshness("invalid_future_completion", int(age.total_seconds()))
deadline = expected_interval + grace
state = "on_time" if age <= deadline else "missed"
return Freshness(state, int(age.total_seconds()))
Test the boundary cases: exactly at the grace deadline, one second beyond it, a duplicated completion, an out-of-order write, a future timestamp, and a ledger read that times out. Also test a partial regional outage independently. Combining all of these into one red or green status makes rollback decisions harder because the operator cannot tell whether the suspect object is the public deployment, the scheduler, the worker, or the promoted media artifact.
Delivery semantics matter here. If two workers can complete the same logical run, use a deterministic run key and make the ledger write idempotent. If a late completion arrives after a retry has published a newer artifact, do not let arrival order redefine which artifact is current. Promotion should compare the intended run identity or generation, and the prior pointer should remain available until the rollback window closes.
Operate the rollback contract
A useful alert names the violated contract and carries stable correlation fields. pipeline_freshness=missed should include the scheduled window and last accepted run key; an external availability alert should include the probe region and endpoint class. Avoid embedding unbounded identifiers in metric labels. Link those identifiers at query time to logs or the completion ledger.
Deploy the monitor separately from the workload it watches. A health handler running inside the same process can still be valuable, but it shares process, host, network, and deployment failure modes. Regional probes add an outside view. The durable completion ledger adds a temporal view. Neither replaces the other.
Before changing retention, replay the rollback procedure against an older representative run. Confirm that the retained summary resolves both the current and previous immutable artifacts, that access controls permit restoration, and that the alert remains understandable after verbose events have expired. Then stage the policy change, observe stored bytes by evidence class, and verify deletion at the intended boundary.
The decision rule is strict: retain detailed events for the shortest period that covers normal diagnosis, retain compact lineage for the full rollback and audit obligation, and keep availability plus freshness as independent signals. This does not promise perfect reconstruction forever. It preserves the evidence required for the rollback you actually claim to support, while making the loss beyond that window visible before an incident forces the question.
Further reading
- Prometheus, metric and label naming: https://prometheus.io/docs/practices/naming/
- Sentry, event grouping and fingerprints: https://docs.sentry.io/concepts/data-management/event-grouping/
Top comments (0)