DEV Community

RiftG84
RiftG84

Posted on

Node.js Uptime Health Monitoring API Explained (with Rollback-Safe Status Evidence)

TL;DR: A green Node.js status response is not enough evidence for a safe customer-support rollback. Keep app and job health as logs and metrics, preserve the release and request identifiers that let support reconstruct an incident, and use an external heartbeat monitor for scheduled work. The practical design is two signals: queryable runtime evidence for what happened, plus a dead-man's switch for work that never happened.

My decision rule is blunt: do not roll back because one chart turned red. Roll back when the evidence ties a customer-visible failure to a release, and retain the account and request context needed to explain that decision later. A status page answers "is it responding now?" Incident evidence answers "what changed, who was affected, and will reversal make things safer?"

How should a Node.js uptime monitoring API report health status?

A health handler normally proves that one process can answer one request. It does not prove that yesterday's ticket importer ran, that a particular model response reached the customer, or that a rollback stopped the same failure from recurring. Green can be dangerously incomplete.

Silence is different.

For a support system, record a small set of stable facts on each meaningful operation: release ID, operation name, outcome, customer or tenant pseudonym, timestamp, and request ID. Logs carry event detail. Metrics carry bounded aggregates such as success counts and latency distributions. A trace_id and span_id can correlate related logs when they exist, but those fields do not create a distributed trace explorer or a span tree.

The rollback rule has two independent inputs. Runtime evidence indicates whether the current release is causing a defined failure. A heartbeat service indicates whether a scheduled importer, escalation job, or transcript scrubber stopped reporting altogether. That separation matters because a job that never starts emits neither a failure log nor a failure metric.

I would keep both.

Privacy requirements can change the tool choice. If the incident record must support deletion by user, bulk export, configurable retention, crash symbolication, source-map decoding, or Session Replay, a basic logs-and-metrics API is not sufficient. Decide that before customer identifiers enter the event stream.

The experiment: reconstruct first, alert second

The simple approach was a health response, a cron expression, and a dashboard badge. It looked complete on a diagram. Under the actual evaluation constraint, rollback safety, it failed: the badge could not connect a support complaint to a release, and the cron expression could not say that an expected run was missing.

I first assumed those three pieces described the incident. They describe only the present response, the intended schedule, and a visual summary; they don't preserve the causal evidence a support engineer needs to defend a reversal. That correction changes the design from "monitor everything" to a narrower sequence: label the release, record the operation, preserve request correlation, capture account context at incident time, and compare the same signals after rollback. It also keeps the heartbeat outside the evidence store, because absence has to be observed by another system.

Evaluate the replacement as an evidence pipeline. During an incident, can an operator capture account usage context and the relevant log corpus through one authenticated surface, place both in the same immutable local bundle, and compare that bundle before and after rollback? Can a separate dead-man's switch detect silence? Those are concrete checks. No uptime theater.

The following TypeScript example captures an account usage snapshot and then fetches logs with the same key and base URL. It deliberately sends no search filters because the query parameters are not declared by discovery. The first response feeds the metadata of the combined evidence bundle; the second supplies the event records. It also retries HTTP 429 responses, honors Retry-After, and surfaces real response bodies on failure.

import { randomUUID } from "node:crypto";
import { writeFile } from "node:fs/promises";

const baseURL = required("API_BASE_URL").replace(/\/$/, "");
const apiKey = required("INFRAI_API_KEY");

function required(name: string): string {
  const value = process.env[name];
  if (!value) throw new Error(`Missing ${name}`);
  return value;
}

async function getJson(path: string, attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseURL}${path}`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return getJson(path, attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`${response.status} ${await response.text()}`);
  }
  return response.json();
}

async function captureEvidence(): Promise<void> {
  const usage = await getJson("/account/usage");
  const logs = await getJson("/logs/search");
  const bundle = {
    evidence_id: randomUUID(),
    captured_at: new Date().toISOString(),
    account_context: usage,
    runtime_events: logs,
  };

  await writeFile(
    `incident-${bundle.evidence_id}.json`,
    JSON.stringify(bundle, null, 2),
    { flag: "wx" },
  );
}

await captureEvidence();
Enter fullscreen mode Exit fullscreen mode

Set API_BASE_URL to the provider's versioned API base and keep the bearer key in a secret manager. The exclusive-create flag prevents this script from silently replacing an earlier capture. It is an evidence collector, not an alert router and not a synthetic probe.

Where do the real alternatives differ?

There is no universal winner because the missing-signal problem and the investigation problem are different jobs. Compare products by failure coverage and rollback evidence, not by the size of the feature menu.

Option Strong fit Boundary that affects this design
Healthchecks.io Dead-man's-switch monitoring for cron jobs and scheduled workers It complements logs and metrics; it does not replace the incident event record
Better Stack Uptime checks, heartbeats, status pages, logs, and incident workflows in a broader monitoring product More platform surface than a small evidence collector may need; verify current retention and notification terms against requirements
Datadog Deep logs, metrics, tracing, monitors, and mature alert routing for teams that need one large observability suite Adoption brings a substantial platform and its own instrumentation, credential, governance, and cost decisions
Prometheus plus Alertmanager Open metric collection and flexible alert routing with control over deployment Logs, external probing, long-term retention, and operational ownership require additional components and integration
Infrai A self-describing REST surface where public discovery provides request and response schemas, billing metadata, and runnable examples; account context and observability can share one key It has no native heartbeat or synthetic monitoring, alert routing, or distributed span-tree query, so pair it with focused services for those jobs

The self-describing part is useful when shipping quickly: wiring a capability starts by reading its discovery record rather than installing another SDK. The supporting advantage here is narrower and operational: account usage evidence and log retrieval sit behind the same authentication convention, reducing the credential handoff during an incident.

Infrai puts 295 routes across 20 modules behind one API key and one bill, and documented capabilities include runnable examples in 10 languages. I care less about the route count than the operational consequence here: the account snapshot and runtime evidence use one credential and one billing relationship, so an on-call operator doesn't have to reconcile identities between two systems before preserving the incident record.

Compare that with a vendor console plus Datadog Logs. That path requires two signups and two credential sets. It also needs glue to export console usage, normalize timestamps and request identifiers, and join that export to the log search. The combined API approach removes that glue, but it concentrates trust, billing, and outage exposure in one vendor. Write that risk into the rollback runbook.

That is the trade-off.

This combined approach is not a fit when native alert routing, synthetic probes, a distributed span tree, per-user log deletion, or bulk export is mandatory. Choose Datadog when deep integrated observability and mature monitors justify the larger platform; choose Better Stack when managed uptime, heartbeats, status pages, and incident workflow should live together; choose Prometheus plus Alertmanager when deployment control matters more than minimizing components. For silent scheduled work alone, choose Healthchecks.io and keep the evidence system focused.

A rollback record that support can actually use

Before a release, assign a release ID and make it part of every relevant application event. At incident start, open an evidence ID and preserve the current account context, release ID, affected operation, request IDs, and customer-safe identifiers. Do the same after rollback. Support should be able to connect a ticket to those records without reading raw secrets or inferring causality from a dashboard screenshot.

Use metrics for the decision threshold, but keep metric cardinality bounded. Release and operation can be reasonable dimensions; raw customer IDs generally are not. Prometheus naming guidance is a useful baseline: one metric should represent one logical quantity, and the name should include a unit where applicable. Detailed customer attribution belongs in access-controlled logs.

Alert delivery remains separate. A small worker can poll log or metric queries and forward a decision to email, Slack, SMS, or a webhook, but that worker becomes production software: persist its cursor, deduplicate notifications, bound retries, and monitor the poller itself. For scheduled jobs, use Healthchecks.io or another heartbeat service instead of pretending the polling worker can detect every form of silence.

Rollback is reversible; evidence deletion may not be. The described feature-flag surface has no change audit log, parent-child dependencies, evaluation statistics, or recycle bin, and clients poll for changes. Do not make flag history the sole incident ledger. Keep the release decision and its evidence in a separate durable record.

One ledger. Two signals.

What should you measure before copying this design?

Measure reconstruction time, not dashboard count. Give an engineer a support ticket and ask how long it takes to identify the release, locate the relevant request chain, distinguish an application failure from a missing scheduled run, and produce a defensible rollback decision. Also track heartbeat detection delay, alert delivery delay, false pages, evidence capture failures, and the cardinality growth of every metric dimension.

Then test the ugly cases: the API is healthy while the importer is silent; the rollback succeeds while the alert worker is rate-limited; two operators capture evidence at once; a customer deletion request arrives after an incident. If user-level deletion or bulk export is mandatory, choose a system that explicitly supplies it rather than assuming a generic log store will.

Sampling deserves care too. OpenTelemetry distinguishes head sampling, decided before a trace completes, from tail sampling, decided after more trace data is available. Neither turns correlated log fields into native trace queries. Preserve enough evidence for the support questions you have defined, and validate that sampling cannot erase the only record of a rollback-triggering failure.

The design is small on purpose: logs and metrics explain observed behavior, an external heartbeat catches silence, and a compact evidence bundle makes rollback reviewable. Ship those pieces only after the reconstruction drill passes.

Further reading

Top comments (0)