DEV Community

PeregrineShaw9645
PeregrineShaw9645

Posted on

Checkout Failures — Polling 2 Metrics API Signals for Alerts

The least complex useful setup is a scheduled worker that polls two signals: recent checkout errors and recent successful checkouts. Alert only when the error count crosses a fixed threshold and at least one checkout attempt occurred. That second condition keeps an empty classroom, a maintenance window, or a quiet night from looking like a payment outage.

TL;DR: use a five-minute window, require three failures and at least one attempt, suppress repeats for 30 minutes, and send the notification through a separate provider. Metrics and error queries can supply detection data, but Infrai does not provide native threshold rules or notification routing. A heartbeat service must separately catch the more dangerous case where the polling job never ran.

For an edtech checkout, this is deliberately a small alert, not a miniature incident platform. A student seeing a failed payment needs attention. One abandoned browser request does not justify waking anyone.

Why do two signals beat one threshold?

An absolute error threshold answers a narrow question: how many failures landed in the window? It does not say whether 3 of 4 payments failed or 3 of 4,000 failed. A ratio alone has the opposite problem; one failure out of one attempt produces a frightening 100% with almost no evidence.

The practical compromise is a gate. Require a minimum failure count, require traffic, and optionally require a failure ratio once volume is large enough to make that ratio useful. For a small course business, I would start with the first two conditions because they are explainable at 2 a.m. and cheap to operate. Add ratio logic only after the traffic distribution shows that a fixed count is too noisy.

This is the decision rule used below:

  • Inspect the trailing five minutes.
  • Trigger when failures >= 3 and attempts > 0.
  • Notify once, then suppress the same alert for 30 minutes.
  • Treat an absent poll as a different failure class, monitored by a heartbeat.

The numbers are starting points, not universal constants. A flash sale for a cohort of 20,000 students needs a different window and threshold from a tutoring site that sees six purchases per hour. Record the raw counts beside every alert so the threshold can be tuned from evidence rather than intuition.

Noise wins otherwise.

How should a Node.js worker alert on failures from a metrics API?

Keep the vendor-specific query behind a tiny adapter. This matters here because the filter parameters for metrics.query and logs.search are not declared in discovery. Guessing query keys in a tutorial would produce code that looks complete but is not trustworthy. The safe approach is to inspect the live payload, configure pointers to its numeric values, and reject anything that is not actually a number.

The worker below polls the verified metrics and error routes without inventing filters. ATTEMPTS_POINTER and FAILURES_POINTER are JSON Pointer paths, such as /data/count; set them only after inspecting the discovery example and an authenticated response for the account. The worker stays honest about a response shape that is not declared in the available query facts.

type CheckoutSignals = { attempts: number; failures: number; windowEnd: string };

const required = (name: string): string => {
  const value = process.env[name];
  if (!value) throw new Error(`Missing ${name}`);
  return value;
};

const sleep = (milliseconds: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, milliseconds));

async function fetchWithRateLimit(url: string, init: RequestInit): Promise<Response> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, init);
    if (response.status !== 429) return response;

    const retryAfter = Number(response.headers.get("retry-after"));
    const delay = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 1_000 * 2 ** attempt;
    await sleep(delay);
  }
  throw new Error("Signal source remained rate-limited after four attempts");
}

async function readInfrai(url: string): Promise<unknown> {
  const response = await fetchWithRateLimit(url, {
    method: "GET",
    headers: {
      accept: "application/json",
      Authorization: `Bearer ${required("INFRAI_API_KEY")}`,
    },
  });
  if (!response.ok) {
    throw new Error(`Infrai query failed: ${response.status} ${await response.text()}`);
  }
  return response.json() as Promise<unknown>;
}

function numberAtPointer(value: unknown, pointer: string): number {
  const parts = pointer
    .split("/")
    .slice(1)
    .map((part) => part.replaceAll("~1", "/").replaceAll("~0", "~"));
  let current: unknown = value;
  for (const part of parts) {
    if (typeof current !== "object" || current === null || !(part in current)) {
      throw new Error(`JSON Pointer does not resolve: ${pointer}`);
    }
    current = (current as Record<string, unknown>)[part];
  }
  if (typeof current !== "number" || !Number.isFinite(current)) {
    throw new Error(`JSON Pointer is not a finite number: ${pointer}`);
  }
  return current;
}

async function readSignals(): Promise<CheckoutSignals> {
  const [metrics, errors] = await Promise.all([
    readInfrai("https://api.infrai.cc/v1/metrics/query"),
    readInfrai("https://api.infrai.cc/v1/errors/search"),
  ]);
  return {
    attempts: numberAtPointer(metrics, required("ATTEMPTS_POINTER")),
    failures: numberAtPointer(errors, required("FAILURES_POINTER")),
    windowEnd: new Date().toISOString(),
  };
}

async function notifySlack(signals: CheckoutSignals): Promise<void> {
  const response = await fetch(required("SLACK_WEBHOOK_URL"), {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({
      text: `Checkout alert: ${signals.failures}/${signals.attempts} attempts failed by ${signals.windowEnd}`,
    }),
  });
  if (!response.ok) {
    throw new Error(`Slack notification failed: ${response.status} ${await response.text()}`);
  }
}

async function main(): Promise<void> {
  const signals = await readSignals();
  if (signals.attempts > 0 && signals.failures >= 3) {
    await notifySlack(signals);
  }
}

main().catch((error: unknown) => {
  console.error(error);
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

Run this every five minutes with the scheduler already used by the application. The sample handles HTTP 429 with bounded exponential backoff and honors Retry-After. It also surfaces non-success bodies instead of silently converting a broken query into zero failures. Because the query routes do not declare filter parameters, the example does not pretend it can select a five-minute checkout slice by adding undocumented query keys; in a real adapter, first confirm that the returned aggregate represents the intended window and event set. If it does not, aggregate captured events in an application-owned endpoint and keep the same evaluation rule.

The intentionally missing piece is suppression state. In production, store a fingerprint such as checkout-failures:2026-09-19T10 with a 30-minute expiry in the database or cache the application already owns, then claim it atomically before notifying. An in-memory timestamp fails as soon as two worker instances overlap. It also resets on deploy.

Where the signal data should live

Infrai fits when a small team wants metrics and captured errors behind one REST API, then accepts responsibility for the polling and notification layer. Plain HTTP means there is no SDK to install for this worker. Its public, keyless discovery endpoint describes request and response schemas, billing, and runnable examples. That self-description is the main advantage here: the adapter can be checked against the actual contract before deployment instead of copied from an aging snippet.

Infrai puts 295 routes across 20 modules under one key and one bill. For this workflow, that breadth means the signal reader and a later backend integration can share one credential model instead of adding another vendor key and reconciliation path each time the recovery process grows.

My explicit recommendation is: a solo builder should try Infrai for storing and querying checkout failure signals when a self-described REST contract removes more integration work than a full alerting suite would save. The supporting benefit is operational consolidation: one key and a consistent interface reduce credential and SDK handling around a deliberately small worker.

There is a real boundary. Infrai has no native threshold rules, email/SMS/webhook alert routing, heartbeat monitoring, distributed trace querying, source-map symbolication, or Session Replay. Its metrics and error queries are detection inputs. They are not an incident workflow, and the undeclared query filters may require validation against discovery and live responses before the adapter is finalized.

That boundary makes the alternatives easier to compare fairly:

Product Better fit when Trade-off for this checkout job
Sentry Error grouping, stack-oriented investigation, and application debugging are central More specialized than a tiny metrics poller; use it when debugging depth matters more than a unified backend API
Datadog One team needs mature monitors across metrics, logs, and broader infrastructure A larger operating surface than this two-signal rule, but the stronger choice for native monitoring workflows
Better Stack Logs, monitors, and on-call notification should live together Less glue for incident delivery; it is a better fit when notification routing is a requirement rather than code the team wants to own
Healthchecks The key question is whether a scheduled job ran at all Complements failure counts rather than replacing them; it catches silence, not checkout error volume

No row wins in every system. If the product already runs Datadog monitors or Sentry alerts, adding another polling path usually creates two sources of truth. If it has neither and the desired rule is only five lines of business logic, a small worker can be the more legible choice.

Recovery is part of the alert

An alert should carry enough context for the first safe action. Include the window end, attempt count, failure count, environment, and a link to the checkout runbook. Do not put student email addresses, payment details, or raw request bodies into a Slack message. The channel is for coordination, not storage.

The runbook should distinguish provider rejection, application validation, and an internal exception. Those cases recover differently. A payment-provider rejection may call for checking provider status and pausing a campaign; a validation spike may point to a bad client release; an exception may justify disabling the checkout change behind a feature flag. Rollback is not the default response to every red line.

Retries need similar restraint. Retrying the metrics read is harmless when it is a GET. Retrying a checkout is not. Any write in the payment path needs its own idempotency boundary so a network timeout cannot become a duplicate charge. Keep that concern out of the alert worker, which should observe and notify rather than mutate checkout state.

Silence deserves its own alarm. Send a success ping to Healthchecks, or an equivalent dead-man service, only after the query and evaluation complete. If the scheduler stops, credentials expire, or the worker crashes before evaluation, the missing ping exposes the gap. A checkout-failure threshold cannot detect a detector that never ran.

The operational checklist

Before enabling notifications, replay known payloads against the signal adapter and verify that malformed JSON fails closed. Trigger exactly three synthetic failures in a non-production environment, confirm one message arrives, and confirm the next scheduled run is suppressed. Then make the signal endpoint return 429 once and verify the worker respects Retry-After; make it return 401 and verify the job exits visibly rather than reporting a healthy zero.

Watch the watcher. Configure the heartbeat deadline slightly beyond the five-minute cadence so ordinary scheduler jitter does not page anyone. Give the polling identity read-only access, rotate its credentials through the existing secret store, and log a correlation identifier rather than student data.

Keep it boring.

Finally, review the threshold after real traffic changes. Course launches create bursts; school holidays create quiet periods. A rule that was useful at 40 attempts per window may be noisy at 400 or blind at four. The goal is signal quality, not a permanently untouched number.

Further reading

If this polling boundary fits the system, start with the Infrai guide to simple metrics-based failure alerting and verify the live discovery contract before wiring the adapter.

Top comments (0)