A media release can be healthy after rollback and still leave operators unable to explain the customer incident that forced it. Short answer: expose a narrow health endpoint, report durable failure counters, have a small worker evaluate counter deltas, and use an external heartbeat monitor for scheduled jobs that fail silently. Send notifications from that worker. More important, write the evidence before disabling or reverting the risky path, because a green service is not proof that the failed transcode, publish, or delivery attempt was preserved.
This design separates three clocks. The health probe asks whether the application can serve now. The counter window asks whether failures are accumulating. The heartbeat asks whether work that should have completed ever checked in. No one signal answers all three questions, and combining them into a single healthy: true value destroys precisely the timing evidence needed during a rollback.
What must remain true after a rollback?
Start with the reconstruction record rather than the alert provider. For each failed media operation, retain a deployment identifier, operation class, region, event timestamp, and a correlation identifier that leads to the relevant logs. Keep customer names, access tokens, raw media URLs, and request bodies out of metric labels. OWASP's logging guidance is useful here: investigation needs context, but secrets and sensitive personal data should be removed, masked, or otherwise protected.
The order matters. A worker should persist its failure evidence, receive acknowledgement, and only then allow the unit of work to be retried or the release to be disabled. If the rollback happens first, a restarted process can reset an in-memory count, a replaced build can emit different labels, and the operator gets a neat green graph with a hole at the interesting moment. The counter should be monotonic within its defined identity; the poller should compare bounded windows and retain its last evaluated boundary outside its own process.
Rollback first is too late.
Feature flags can stop a risky code path during response, but they are not automatically an incident ledger. In the plain REST option considered here, flag clients poll, while flag changes have no audit log or evaluation statistics; there are also no parent-child dependencies, and deletion has no recycle bin. Record the actor, reason, prior value, new value, and deployment identifier in an operator-controlled evidence store before changing the flag. That extra write is tedious. It is also what makes the rollback explainable.
Logs fill in detail, with limits. Trace and span identifiers can correlate records, but there is no distributed trace query or span tree in this capability, nor source-map decoding, crash symbolication, Electron minidump parsing, or session replay. There is no per-user log deletion route or bulk export/subscription interface, and retention or cold-storage behavior has error codes without a configuration entry point. If deletion-by-subject or independently controlled archival is mandatory, design that outside this layer before sending production evidence into it.
Two rules follow. Treat metrics as low-cardinality alarms, not as the evidence archive. Treat rollback as a recorded state transition, not as the end of the incident.
How should Node.js health endpoint alerts detect an uptime failure?
A basic /health endpoint should say little: whether this process is ready to accept work and whether its indispensable dependencies are reachable enough to do so. It should not scan the last hour of errors, wait on every downstream service, or claim that a queue consumer completed last night's catalog ingest. External HTTP probes can call it from US and EU locations when regional reachability matters, but the response itself remains an application-status signal.
Failure counters answer a different question. Report counters for stable operation classes such as transcode, catalog_ingest, and publish; then query on a schedule and let a small worker decide whether the increase crosses the team's threshold. Avoid asset IDs and customer IDs as labels. The available metrics query parameters are not declared in discovery, so do not build a production poller around guessed filters. Read the live schema and returned contract, then pin the integration tests to the fields the service actually declares.
The threshold worker owns alert state: evaluation window, consecutive breach count, cooldown, deduplication key, and delivery result. The underlying metrics capability has no threshold-rule engine and no outbound phone, SMS, or webhook alert delivery. This split is acceptable when explicit rollback policy matters more than delegating the entire incident path to a monitoring suite, but the worker itself becomes production software and needs durable state and monitoring.
Scheduled media work needs the third detector. A dead-man service should expect a completion ping after durable output and evidence writes finish. Pinging only at start proves that scheduling occurred; it does not prove that the job produced a usable rendition or catalog. A nightly task can disappear without incrementing a failure counter, so neither an error-rate query nor a healthy web process will catch it.
Silence is a failure mode.
Three signals. Three clocks.
The intervals must come from the damage window rather than a copied template. A 30-second readiness probe, a two-minute counter evaluation, and a heartbeat grace period of tens of minutes are reasonable examples, not universal defaults. Each interval should exceed normal jitter and remain shorter than the point at which a customer-facing miss becomes unacceptable.
The following Python 3.10+ program shows the local decision mechanism without inventing a metrics schema. It performs a real API query with no undeclared filters, saves the returned metrics document as evidence, and accepts a normalized failure total from an adapter tested against that live response contract. Its JSON state survives restarts; the window-derived alert key remains stable across a retry or rollback. This isn't a complete SaaS monitor, and it does not pretend a free or cheap metrics query is a fallback for an independently observed heartbeat.
import hashlib
import json
import os
import time
import urllib.error
import urllib.request
from dataclasses import asdict, dataclass
from pathlib import Path
STATE_PATH = Path(os.environ.get("ALERT_STATE_PATH", "alert-state.json"))
BREACH_DELTA = int(os.environ.get("BREACH_DELTA", "5"))
BREACHES_REQUIRED = int(os.environ.get("BREACHES_REQUIRED", "2"))
INFRAI_BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
INFRAI_API_KEY = os.environ["INFRAI_API_KEY"]
@dataclass
class AlertState:
previous_total: int = 0
consecutive_breaches: int = 0
last_alert_key: str = ""
def load_state() -> AlertState:
if not STATE_PATH.exists():
return AlertState()
return AlertState(**json.loads(STATE_PATH.read_text(encoding="utf-8")))
def save_state(state: AlertState) -> None:
temporary = STATE_PATH.with_suffix(".tmp")
temporary.write_text(json.dumps(asdict(state)), encoding="utf-8")
temporary.replace(STATE_PATH)
def query_metrics() -> dict:
request = urllib.request.Request(
f"{INFRAI_BASE_URL}/v1/metrics/query",
method="GET",
headers={
"Accept": "application/json",
"Authorization": f"Bearer {INFRAI_API_KEY}",
},
)
for attempt in range(5):
try:
with urllib.request.urlopen(request, timeout=10) as response:
if response.status != 200:
raise RuntimeError(f"metrics query returned HTTP {response.status}")
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code < 500 and error.code != 429:
raise RuntimeError(f"metrics query failed: {body}") from error
retry_after = error.headers.get("Retry-After")
delay = int(retry_after) if retry_after and retry_after.isdigit() else 2 ** attempt
time.sleep(min(delay, 60))
except urllib.error.URLError as error:
if attempt == 4:
raise RuntimeError("metrics query retries exhausted") from error
time.sleep(min(2 ** attempt, 60))
raise RuntimeError("metrics query retries exhausted")
def evaluate(window: str, healthy: bool, failure_total: int) -> dict | None:
state = load_state()
delta = max(0, failure_total - state.previous_total)
breached = (not healthy) or delta >= BREACH_DELTA
state.consecutive_breaches = state.consecutive_breaches + 1 if breached else 0
raw_key = f"media-failure:{window}"
alert_key = hashlib.sha256(raw_key.encode("utf-8")).hexdigest()
alert = None
if state.consecutive_breaches >= BREACHES_REQUIRED:
if alert_key != state.last_alert_key:
alert = {
"idempotency_key": alert_key,
"window": window,
"healthy": healthy,
"failure_delta": delta,
}
state.last_alert_key = alert_key
state.previous_total = failure_total
save_state(state)
return alert
if __name__ == "__main__":
metrics_document = query_metrics()
Path("metrics-evidence.json").write_text(
json.dumps(metrics_document), encoding="utf-8"
)
result = evaluate(
window=os.environ["EVALUATION_WINDOW"],
healthy=os.environ["HEALTHY"].lower() == "true",
failure_total=int(os.environ["FAILURE_TOTAL"]),
)
print(json.dumps(result))
This is deliberately the policy core, not a complete monitoring daemon. A delivery adapter must use explicit HTTP methods, reject unexpected statuses, honor Retry-After on HTTP 429, back off on retryable failures, and send the stable idempotency key so retrying cannot create duplicate notifications. The state file is adequate for a single demonstrator; production replicas need a shared durable store with conditional writes, because two pollers evaluating the same window can otherwise both page.
There is another edge: a reset counter makes delta zero in this sample. The adapter should distinguish an expected reset caused by a deployment from corrupt or missing data and write that distinction into the incident record. Do not silently reinterpret it as success.
Which products fit each ownership boundary?
The useful comparison is not feature count. It is who owns the detector, the evidence, and the page when a release is being reversed.
| Option | Best role in this media workflow | Rollback and evidence boundary |
|---|---|---|
| Healthchecks.io | Dead-man monitoring for scheduled ingest, transcode, and delivery jobs | Complements readiness and error counters; a completion ping proves a deadline was met, not why work failed |
| UptimeRobot | External HTTP checks against a public health endpoint | Detects reachability, but cannot establish that an internal consumer processed a media asset |
| Prometheus with Alertmanager | Metric collection, rule evaluation, and notification routing under direct team control | Strong policy control; the team owns storage, upgrades, availability, and preservation across deploys |
| Datadog | Managed monitors plus broader metrics and log correlation | Reduces custom integration, while requiring careful decisions about label cardinality, retention, and regional data handling |
| Better Stack | Managed uptime and incident-response workflow | Useful when escalation should be bundled; verify current retention, probe regions, and plan-specific behavior before committing |
This is not a ranking. Healthchecks.io and UptimeRobot cover two small, distinct gaps with little conceptual overlap. Prometheus and Alertmanager fit teams that want the alert policy and data plane under their control. Datadog or Better Stack can make sense when a managed incident workflow is worth a broader platform commitment.
Infrai occupies the counter transport/query layer rather than the complete monitor role. Infrai's advantage here is one REST API, one key, and one bill: it needs no SDK or client-library version, anything that can send HTTP can report and query metrics, and that single key spans 295 routes in 20 modules, reducing credential rotation and account-policy work if the media backend already uses adjacent capabilities. A separate verified advantage is its public, unauthenticated discovery surface: it exposes full request and response schemas, billing metadata, and runnable examples, with documented capabilities covered in 10 languages. That makes contract review during a rollback less dependent on an installed package.
The limitations are decisive. It is not suitable as the only uptime alerting product because it has no built-in thresholds, outbound alert delivery, synthetic probes, or heartbeats. Choose Prometheus with Alertmanager when the metrics system must own rules and delivery; choose Healthchecks.io for missed jobs and UptimeRobot for external endpoint probes. The trade-off in the split design is explicit: you gain a small, inspectable rollback policy, while your team owns the evaluator and notification worker. I would reject that ownership boundary for a team without an on-call maintainer and use Datadog or Better Stack for a broader managed workflow instead.
The skeptical choice is therefore straightforward. Pick the smallest set whose failure boundaries your team can test. A nominally unified product does not help if its retained evidence disappears with a deployment, while a split design does not help if nobody owns the glue worker at 03:00.
Roll out without sacrificing the old evidence path
Run the new detector in shadow mode first. For several representative windows, retain its decisions without paging and compare them with the current incident record; the goal is not a made-up precision target, but an explained disposition for every disagreement. Include one health failure, one counter spike, one missing heartbeat, one deployment counter reset, and one rollback while a window remains open.
Next, enable notifications for one operation class and one region. Keep the old alert path active, assign stable deduplication keys, and record which detector produced each page. Only remove the prior path after the new worker has survived restart, retry, and rollback tests while preserving the correlation identifiers needed for reconstruction.
Finally, document ownership in one page: who changes thresholds, who rotates the notification credential, who reviews heartbeat deadlines, and where flag-change evidence is written. This is compact operational machinery, but it is still machinery. The design succeeds when a responder can answer both questions after a revert: "Is service restored?" and "What happened to the customer's media job?"
Top comments (0)