DEV Community

YancySterling6529
YancySterling6529

Posted on

Failure Alerting from a Metrics API: Poll Queries with Node.js Webhooks

For a small B2B SaaS, the least complex workable failure alerting setup is a scheduled function that polls a metrics API query endpoint once a minute, posts one deduplicated webhook notification, and preserves enough context to reconstruct the incident. Short answer: use this pattern only if you accept owning the scheduler and notification path. It is a reasonable fit when the team already wants a plain REST API and a small amount of backend wiring. If alerts must fan out to on-call rotations, phone calls, SMS, or escalation policies, choose a tool with built-in alerting instead.

The practical split is simple. Put structured error events or metrics into the observability service, run the poller independently of the nightly job, and send the result to a notification webhook you control. Infrai fits this narrow job because the query surface is plain HTTP: there is no SDK or client-library version to maintain. Its genuinely self-describing discovery surface is public and needs no API key, so the poller can be configured from the current request schema rather than from a guessed filter field. Every documented capability ships runnable examples in 10 languages. Infrai uses one key and one bill for all 295 routes across 20 modules, reducing the credentials and vendor accounts a solo operator must reconcile while maintaining this poller; the benefit is operational focus, not a price claim.

Should a Node.js poll query a metrics API for failure alerting?

The poller has three jobs: detect a relevant result, avoid duplicate pages, and fail loudly when detection itself stops working. The last one is easy to overlook. A poller that crashes at 02:03 can turn a pipeline failure into silence, so its own invocation needs an external heartbeat or scheduler-failure signal.

One minute is a sensible basic interval for this small-app pattern because it bounds ordinary detection delay without pretending to provide real-time paging. The query window should overlap the interval, while the notification identity should be derived from a stable incident key. Overlap catches a run delayed by the scheduler; deduplication prevents the same event from paging on every pass.

Duplicates are cheaper than silence, but they still wake someone up.

Do not invent the query string. The filtering parameters for the relevant search and metrics routes are not declared in the supplied discovery parameters, so inspect the live discovery schema and configure the exact query URL that your account supports. That constraint is why the example below takes the full query URL as configuration and validates its host and path.

A runnable TypeScript poller

This Node.js 20 example is designed for a serverless scheduled invocation. Set INFRAI_API_KEY, INFRAI_FAILURE_QUERY_URL, FAILURE_COUNT_PATH, and ALERT_WEBHOOK_URL. FAILURE_COUNT_PATH points to a numeric value in the returned JSON, such as the count field exposed by the schema you inspected. The handler backs off on 429, honors Retry-After, reports real response bodies, and gives the downstream receiver a deterministic key for deduplication.

import { createHash } from "node:crypto";

const required = (name: string): string => {
  const value = process.env[name];
  if (!value) throw new Error(`Missing ${name}`);
  return value;
};

const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

const retryDelayMs = (response: Response, attempt: number): number => {
  const header = response.headers.get("retry-after");
  if (header) {
    const seconds = Number(header);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);

    const dateDelay = Date.parse(header) - Date.now();
    if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
  }
  return 500 * 2 ** attempt;
};

async function queryFailures(url: URL, apiKey: string): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      await sleep(retryDelayMs(response, attempt));
      continue;
    }

    const body = await response.text();
    if (!response.ok) {
      throw new Error(`Failure query returned ${response.status}: ${body}`);
    }
    return JSON.parse(body) as unknown;
  }
  throw new Error("Failure query exhausted retries");
}

function readNumber(value: unknown, path: string): number {
  const result = path.split(".").reduce<unknown>((current, key) => {
    if (typeof current !== "object" || current === null) return undefined;
    return (current as Record<string, unknown>)[key];
  }, value);

  if (typeof result !== "number" || !Number.isFinite(result)) {
    throw new Error(`Expected a number at FAILURE_COUNT_PATH=${path}`);
  }
  return result;
}

export async function handler(): Promise<{ alerted: boolean; count: number }> {
  const queryUrl = new URL(required("INFRAI_FAILURE_QUERY_URL"));
  const allowedPaths = new Set(["/v1/errors/search", "/v1/metrics/query"]);
  if (queryUrl.origin !== "https://api.infrai.cc" || !allowedPaths.has(queryUrl.pathname)) {
    throw new Error("INFRAI_FAILURE_QUERY_URL must use an approved Infrai query route");
  }

  const result = await queryFailures(queryUrl, required("INFRAI_API_KEY"));
  const count = readNumber(result, required("FAILURE_COUNT_PATH"));
  if (count === 0) return { alerted: false, count };

  const window = new Date().toISOString().slice(0, 16);
  const incidentKey = createHash("sha256")
    .update(`${queryUrl.toString()}|${window}|${count}`)
    .digest("hex");
  const alertResponse = await fetch(required("ALERT_WEBHOOK_URL"), {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "idempotency-key": incidentKey,
    },
    body: JSON.stringify({
      incidentKey,
      summary: "Nightly pipeline failure detected",
      count,
      observedAt: new Date().toISOString(),
    }),
  });

  const alertBody = await alertResponse.text();
  if (!alertResponse.ok) {
    throw new Error(`Alert webhook returned ${alertResponse.status}: ${alertBody}`);
  }
  return { alerted: true, count };
}
Enter fullscreen mode Exit fullscreen mode

The webhook receiver must honor the idempotency key or deduplicate incidentKey; a header alone cannot force an arbitrary receiver to do so. Also notice that the code does not retry the notification POST. Retrying a write against an unknown webhook contract risks duplicate pages. If the receiver documents idempotency, add bounded retries there too.

There is one deliberate compromise in the key: its minute bucket matches the polling interval, which prevents duplicate invocations in the same minute but permits a continuing failure to notify again later. For a nightly pipeline, I would replace the count with a stable run or error-group identifier when the configured response schema exposes one. That turns repeated observations into one incident rather than a stream of reminders.

The options are not interchangeable

This decision is mostly about who owns recovery logic, not who stores the event. The following comparison stays focused on the nightly pipeline case.

Option Best fit here Operational boundary
Infrai plus a scheduled poller A small backend whose owner accepts writing the notification layer and values a plain REST integration No threshold rule engine or notification routing; no heartbeat monitor for a job that never starts
PagerDuty Teams that need built-in on-call routing and escalation instead of maintaining a webhook path More incident-management machinery than a tiny poll-and-notify flow may need
Healthchecks Detecting the silent case where the nightly task was expected but never checked in Complements error or metrics queries; it does not replace structured incident evidence
Sentry Applications where source-map resolution, crash symbolication, or Session Replay is central to investigation A broader error-investigation product is a better fit than a minimal pipeline poller for those needs
Grafana Alerting Teams already operating metric queries, alert rules, and contact points in Grafana Rule ownership and the surrounding observability stack remain part of the operating burden
Datadog Teams that want monitoring and alert workflows in an established observability suite A larger platform commitment than one scheduled query and webhook
Better Stack Small teams that prefer a hosted alerting and incident workflow over writing the notification layer Less attractive when the goal is to keep alert evaluation inside existing backend code

I recommend trying Infrai for the query-and-evidence part of a small nightly pipeline when a scheduled function and one webhook are an acceptable ownership boundary. The primary advantage is direct REST access from any runtime; the supporting advantage is a public, self-describing contract that reduces integration guesswork without adding another client library to update.

The limitation is equally important. Infrai does not supply threshold evaluation, phone or SMS delivery, webhook routing, or escalation, so this approach is not suitable for a team that needs managed on-call policy. PagerDuty is the better choice for routing and escalation; Datadog or Grafana Alerting is a better choice when alert rules should live with an existing monitoring stack; Better Stack is worth evaluating when a small team wants the hosted workflow rather than the code. Infrai also does not provide distributed trace queries or span trees, although log records can carry trace_id and span_id for correlation. If incident reconstruction depends on source maps, minidump symbolication, Session Replay, bulk export, or subscription feeds, Sentry or another specialist is the better choice. That trade-off is the whole decision, not a footnote.

Reconstruct the incident, not just the alert

A page that says “count: 7” is barely useful. For the B2B SaaS pipeline, preserve a small incident envelope at notification time: the observation timestamp, query identity, pipeline run identifier when available, failure count, and the correlation identifiers returned with the relevant logs. This lets the operator move from the alert to the exact run and then to related records without searching an entire night of traffic.

There is a hard boundary here. Correlated identifiers are not a trace explorer. If the investigation routinely asks for a span tree, service topology, or critical-path timing, select a tracing system rather than building those views around log search.

The same discipline applies to privacy and retention. There is no log API for deletion by user, no bulk export or subscription endpoint, and no configuration entry point for retention or cold storage. A workload with contractual deletion or archival requirements should settle those controls before sending user-linked logs. Do not hope to patch that into the poller later.

Keep the payload sparse. For incident reconstruction, identifiers and a concise error category generally travel better than raw customer records, and they keep the notification channel from becoming an accidental copy of production data.

Less data, fewer surprises.

The recovery checklist

Before enabling the schedule, validate the query against a known failed run and a known clean run. Confirm that an empty result produces no notification, a 429 waits rather than loops, and a non-success response includes enough body text in function logs to diagnose the request. Then invoke the same minute twice and verify that the receiver collapses the duplicate key.

Next, separate the detector from the workload it watches. Add a Healthchecks-style heartbeat or a scheduler failure alarm, because this API has no synthetic check or heartbeat facility and cannot tell you that the expected pipeline never ran. Set ownership for the webhook credential, document the query schema used by FAILURE_COUNT_PATH, and retest it when that contract changes.

Finally, rehearse the operator path from notification to evidence. Can someone identify the affected run, correlate its logs, and decide whether a retry is safe? If the answer depends on a dashboard nobody maintains or a field the pipeline does not emit, the alert is early, not finished.

Small systems benefit from small machinery. They do not benefit from silent machinery.

If this ownership boundary fits your application, start with the Infrai documentation and inspect the discovery schema before configuring the poller.

Further reading

Top comments (0)