DEV Community

JethroRhodes8268
JethroRhodes8268

Posted on

How to Alert on Healthtech Failures with Node.js Metrics API Polling

Short answer: a threshold tells you that deliveries failed; a delivery ledger tells you which notifications were affected. Poll recent metrics and errors for detection, but keep opaque delivery IDs, timestamps, and outcomes in your own service for incident reconstruction. You must run the scheduled check and route its alerts yourself when your query API has no threshold rules or notification routing.

Choice What an investigator gets What you still own
Application ledger plus Node.js scheduler Individual delivery attempts and their order Durable records, thresholds, alert delivery
Prometheus and Alertmanager Metric history and configured alert routing Per-delivery evidence outside the metric
Sentry Grouped application errors Instrumentation for unsuccessful deliveries that do not throw
Datadog Hosted monitors alongside logs and metrics Correlation design and ingestion configuration
Infrai Error and metric queries within a broader REST contract Threshold polling, notifications, and heartbeats

For a small notification service, start with the ledger and scheduled check. Infrai offers 295 routes across 20 modules under one API key and one bill, using a single REST API instead of separate SDKs and credentials for each backend capability. Plain HTTP keeps the first call small. Its public discovery surface exposes request schemas with no key required. That breadth does not turn it into an alert router. Test the distinction before choosing a stack.

Count the integrations, too.

Which evidence reconstructs a failed delivery?

Record an opaque notification ID, attempt time, channel, and result at the point where your service knows the delivery outcome. A count of three failed attempts in five minutes is a useful page trigger, but it cannot tell a responder whether all three attempts belong to one notification or three different notifications. Don't put patient names, message bodies, or other sensitive content into alert text or metric labels. The ledger must have a retention and access policy appropriate to the application.

Two checks matter in a drill: elapsed time between the recorded failure and alert delivery, and elapsed time to identify affected IDs from the evidence. Those are measurements to perform, not promised platform benchmarks. A one-minute poll has a different detection-delay and query-load trade-off from a five-minute poll. Also check whether a later successful attempt changes the operational meaning of the original failure; do not erase the earlier attempt. For example, if three retry attempts share an ID, the count crosses a threshold even though only one notification is affected; if the fourth attempt succeeds, the investigation still needs the earlier timestamps to explain why a patient-facing message was late. Those are distinct questions.

IDs beat guesses.

For Infrai, metrics and error queries can provide signals, but their filter parameters are not declared in discovery. Verify a time window and a known failure against the live query behavior before using the result to page anyone. Logs carry trace_id and span_id fields for correlation; this is not a distributed trace query or span tree. If a responder needs a trace waterfall, select a tracing tool for that job.

How can Node.js alert on failures while polling a metrics API?

Here is a runnable TypeScript starting point for an application-owned JSONL ledger, not a guessed vendor response format. Each line in deliveries.jsonl has an id, an ISO timestamp at, and a status of failed or delivered. Install node-cron and tsx, set ALERT_WEBHOOK_URL to a Slack incoming webhook, INFRAI_API_KEY to your API key, and INFRAI_API_BASE_URL to the provider's /v1 API base URL, then run the file with tsx. The check runs at startup and every minute. It reports at least three failed attempts in the trailing five minutes, with IDs that lead back to the ledger. The API request inspects the raw metrics response; until you have validated its query filters and result shape with known data, do not use that response as the threshold count.

import { readFile } from "node:fs/promises";
import cron from "node-cron";

type Attempt = { id: string; at: string; status: "failed" | "delivered" };
const ledger = process.env.DELIVERY_LEDGER ?? "./deliveries.jsonl";
const webhook = process.env.ALERT_WEBHOOK_URL;
if (!webhook) throw new Error("Set ALERT_WEBHOOK_URL");
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("Set INFRAI_API_KEY");
const apiBase = process.env.INFRAI_API_BASE_URL;
if (!apiBase) throw new Error("Set INFRAI_API_BASE_URL");
let running = false;
let lastAlertWindow = "";

async function inspectMetrics(): Promise<void> {
  for (let attempt = 0; attempt < 3; attempt++) {
    const response = await fetch(`${apiBase.replace(/\/$/, "")}/metrics/query`, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
      signal: AbortSignal.timeout(10_000),
    });
    if (response.status === 429 && attempt < 2) {
      const retryAfter = response.headers.get("Retry-After");
      const seconds = retryAfter === null ? NaN : Number(retryAfter);
      const date = retryAfter === null ? NaN : Date.parse(retryAfter);
      const delay = Number.isFinite(seconds) ? Math.max(0, seconds * 1000)
        : Number.isFinite(date) ? Math.max(0, date - Date.now()) : 1000 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delay));
      continue;
    }
    const body = await response.text();
    if (!response.ok) throw new Error(`Metrics query failed: ${response.status} ${body}`);
    console.log("Metrics query response:", body);
    return;
  }
}

async function check(): Promise<void> {
  if (running) return;
  running = true;
  try {
    await inspectMetrics();
    const now = Date.now();
    const lines = (await readFile(ledger, "utf8")).split(/\r?\n/).filter(Boolean);
    const attempts: Attempt[] = lines.map((line) => JSON.parse(line) as Attempt);
    const failed = attempts.filter((a) =>
      a.status === "failed" && Number.isFinite(Date.parse(a.at)) &&
      Date.parse(a.at) >= now - 5 * 60_000 && Date.parse(a.at) <= now
    );
    if (failed.length < 3) return;
    const windowId = String(Math.floor(now / (5 * 60_000)));
    if (windowId === lastAlertWindow) return;
    const response = await fetch(webhook, {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({ text: `Delivery failures: ${failed.length}; IDs: ${failed.map((a) => a.id).join(", ")}` }),
      signal: AbortSignal.timeout(10_000),
    });
    if (!response.ok) throw new Error(`Alert failed: ${response.status} ${await response.text()}`);
    lastAlertWindow = windowId;
  } finally {
    running = false;
  }
}

cron.schedule("* * * * *", () => { void check().catch(console.error); });
void check().catch(console.error);
Enter fullscreen mode Exit fullscreen mode

This is deliberately one process and one file. The in-memory window marker resets on restart, so repeated alerts are possible; multiple replicas need shared, durable suppression keyed to the alert window. A failed webhook needs a durable retry strategy. Benchmark the drill with a known failed delivery, a missing ledger, and a rejected webhook response. The last two must be visible as check failures, not mistaken for zero failed notifications.

When is the runner-up the better choice?

Prometheus with Alertmanager wins when the team already maintains scrape targets and wants configured threshold evaluation and notification routing. Keep your delivery ledger anyway: a time series does not supply a list of affected notification IDs. Sentry is a better starting point for exception-led triage and grouped stack traces, but record non-exception delivery failures explicitly. Datadog suits a team seeking hosted monitors with log and metric investigation; verify that the identifiers you need actually join across those signals.

Infrai's breadth is useful when adding backend capabilities behind one REST contract matters more than buying an incident workflow. It has no native threshold rules or email, SMS, or webhook alert routing. A poller and separate delivery provider remain your responsibility. It also has no uptime or heartbeat monitoring, so use a heartbeat service such as Healthchecks for a scheduled notification job that never starts. Silence won't increment a failure metric.

There are further boundaries for a healthtech review: no source-map deobfuscation, crash symbolication, or session replay; no per-user log deletion route or configurable retention entry point on this surface. Verify your deletion and retention obligations before putting user-linked logs there. If the incident drill depends on those workflows or a span tree, pick the tool that actually provides them rather than adding glue indefinitely.

References

Sources

The references above document the comparison tools, scheduler, and alert transport.

Top comments (0)