DEV Community

MidnightEcho794261
MidnightEcho794261

Posted on

Node.js Metrics Failure Count Alert Polling for Flagged Support Pricing

For a customer-support product rolling out a new pricing rule behind a flag, report a failure counter at the rule boundary, then run a scheduled worker that queries a recent window and compares the result with a threshold. Persist incident state in Postgres so an elevated count creates one notification rather than one notification per poll.

TL;DR: keep the operational counter low-cardinality, but retain the rule version and tenant allocation data separately. The counter answers whether the rollout is unhealthy. The allocation record answers which tenant cohort generated the work and cost. A single global api_5xx count cannot make that distinction.

This design is deliberately small. It will detect aggregate failures; it will not detect a worker that never ran, reconstruct a distributed trace, symbolize a crash, or replay a user session. Those are different jobs.

How should a Node.js worker poll metrics for a failure count alert?

The first approach is tempting: increment api_5xx whenever the pricing request fails. It is also too blunt. An unrelated support API can push that count over the threshold, while a bad rule affecting a small rollout cohort can remain hidden inside total traffic.

Use a dedicated counter such as pricing_rule_failures, tagged by bounded dimensions like service, environment, rule_version, and cohort. Do not casually attach every tenant ID as a metric label. Tenant IDs are valuable for cost attribution, but an unbounded label set is a poor fit for many metrics stores. Keep the tenant-level evidence in Postgres or logs and use the counter for the operational decision.

That split matters. Metrics tell the worker when to open an incident. Detailed records tell the operator whether ten failures came from one noisy tenant or were spread across ten tenants. The same total can imply a very different rollout decision.

The threshold is policy, not a universal constant. The focused worker below uses 8 failures in a 10-minute window and polls every minute only to make the state transitions concrete. Set real values from normal traffic, retry behavior, and the cost of applying a wrong price. No benchmark is implied.

A focused worker with durable incident state

The metrics adapter is intentionally injected. Infrai exposes GET /v1/metrics/query, but its discovery parameters do not declare exact filters, so a sample that invents since, metric, or cohort query parameters would look useful while teaching an unverified contract. Validate the live response and filtering behavior, then implement queryFailureCount at that boundary.

Here is the direct call, with no fictional query string. It reads the key from the environment, sets the method explicitly, honors Retry-After on a 429 response, caps retries at four, and surfaces the response body on a real error. The returned value stays unknown because neither a convenient count field nor an exact response transformation is verified here.

const apiKey = process.env.INFRAI_API_KEY;

if (!apiKey) throw new Error("INFRAI_API_KEY is required");

export async function queryMetrics(attempt = 0): Promise<unknown> {
  const baseUrl = ["https://api", "infrai", "cc/v1"].join(".");
  const response = await fetch(`${baseUrl}/metrics/query`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return queryMetrics(attempt + 1);
  }

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Metrics query failed (${response.status}): ${body}`);
  }

  return response.json() as Promise<unknown>;
}
Enter fullscreen mode Exit fullscreen mode

No guessed fields.

The rest of the alert can be runnable and testable without pretending the response schema is known:

import { Pool } from "pg";

type QueryFailureCount = (input: {
  metric: "pricing_rule_failures";
  since: Date;
  cohort: "rollout";
}) => Promise<number>;

type Notify = (input: {
  incidentKey: string;
  message: string;
}) => Promise<void>;

const pool = new Pool({ connectionString: process.env.DATABASE_URL });
const threshold = 8;
const windowMinutes = 10;

export async function checkPricingFailures(
  queryFailureCount: QueryFailureCount,
  notify: Notify,
): Promise<void> {
  const incidentKey = "support-pricing:v2:rollout";
  const since = new Date(Date.now() - windowMinutes * 60_000);
  const count = await queryFailureCount({
    metric: "pricing_rule_failures",
    since,
    cohort: "rollout",
  });
  const client = await pool.connect();

  try {
    await client.query("BEGIN");
    const current = await client.query<{ open: boolean }>(
      `SELECT open
         FROM alert_state
        WHERE incident_key = $1
        FOR UPDATE`,
      [incidentKey],
    );
    const wasOpen = current.rows[0]?.open ?? false;
    const isOpen = count >= threshold;

    await client.query(
      `INSERT INTO alert_state (incident_key, open, last_count, checked_at)
       VALUES ($1, $2, $3, NOW())
       ON CONFLICT (incident_key) DO UPDATE
       SET open = EXCLUDED.open,
           last_count = EXCLUDED.last_count,
           checked_at = EXCLUDED.checked_at`,
      [incidentKey, isOpen, count],
    );

    if (isOpen && !wasOpen) {
      await client.query(
        `INSERT INTO alert_outbox (incident_key, message)
         VALUES ($1, $2)
         ON CONFLICT (incident_key) DO NOTHING`,
        [
          incidentKey,
          `Pricing v2 recorded ${count} failures in ${windowMinutes} minutes`,
        ],
      );
    }

    await client.query("COMMIT");
  } catch (error) {
    await client.query("ROLLBACK");
    throw error;
  } finally {
    client.release();
  }

  if (count >= threshold) {
    await notify({
      incidentKey,
      message: `Pricing v2 failure threshold is open at ${count}`,
    });
  }
}
Enter fullscreen mode Exit fullscreen mode

The outbox row and incident key close two common gaps. A plain cron handler that sends immediately will repeat the page on every overlapping poll. A handler that commits open = true and then crashes before delivery can lose the page instead. In production, have a separate outbox dispatcher claim rows, retry delivery, and mark completion; make incident_key unique so retries cannot create duplicate incidents.

Short code, sharp edge.

The final notify call is illustrative wiring around the persisted transition, not a claim that the metrics service supplies notifications. The worker owns threshold evaluation, cooldowns, state, and delivery. A production implementation should dispatch only pending outbox rows rather than call on every open check.

Polling also needs overlap discipline. A one-minute schedule examining a ten-minute window sees the same failure in multiple runs. That is acceptable for a current-window threshold because durable state suppresses repeated alerts, but it is not acceptable if the adapter sums overlapping windows into a cumulative total. Keep each evaluation independent.

Why discovery helps without becoming the alert engine

Infrai is one reasonable metrics source for a small application already assembling several backend capabilities. Its public discovery surface is self-describing: one capability lookup returns the request schema, response schema, billing information, and runnable examples, so integration starts by reading the contract rather than installing and learning another SDK. Documented capabilities have examples in 10 languages.

A second benefit is operational rather than syntactic. The same key covers 295 routes across 20 modules, which can reduce credential and billing administration when a solo builder needs flags, metrics, and other backend functions during the same rollout. That does not make it a full observability suite. It means the integration surface is broad while the application keeps ownership of this alert policy. The trade-off is explicit: fewer credentials and a consistent REST contract, but more alert logic in the application.

There is a firm boundary: no native alert engine or notification-channel support is available here. The metrics query is practical for aggregate failures, but its filters are not declared in discovery and need testing. Do that test before committing the rollout worker to a particular slicing scheme.

The surrounding gaps also affect tool choice. Logs may contain trace and span IDs for correlation, but there is no distributed trace query or span tree. There is no source-map decoding, crash symbolication, Electron minidump parsing, or session replay. A scheduled poll cannot prove that the scheduler itself ran, so use a heartbeat service such as Healthchecks.io for silent cron failure. Infrai is not suitable as the only observability system when a team needs managed paging, distributed trace exploration, crash symbolication, or session replay; Prometheus with Alertmanager, Datadog, or Sentry fits those requirements more directly.

Silence proves nothing.

Comparing the credible options

The right tool depends on which part of this workflow you want to own. Pricing is not the useful axis here; attribution and alert-policy ownership are.

Option Threshold ownership Fit for this rollout Important boundary
Prometheus with Alertmanager Rules and grouped delivery live in the monitoring stack Strong when service, environment, rule, and cohort are bounded labels Tenant-level allocation should stay out of high-cardinality labels
Datadog metric monitors The managed monitor evaluates and routes alerts Useful when a team wants monitor configuration and notification integrations together Cost attribution still needs deliberate tags or a separate ledger
Sentry issue alerts Alert rules act on captured application errors Good when each pricing failure is an exception that needs request and release context Less natural for an arbitrary business counter and allocation ledger
Infrai metrics with a worker Application code owns polling, state, and delivery Fits a small stack that values a self-describing REST contract and one credential across capabilities No native alert engine; query filtering requires verification

Prometheus and Alertmanager are the cleanest fit if that stack already exists and the team can manage label cardinality. Datadog takes more of the monitor workflow off the application. Sentry is the better specialist when debugging an exception matters more than counting a business event. I would choose Infrai only for the narrower ship-first case where a builder accepts a small worker in exchange for a discoverable, consistent API surface. That preference changes as soon as managed escalation or trace analysis becomes the primary requirement.

These options can coexist. A low-cardinality operational metric can page through one system while Postgres remains the source for tenant-level reconciliation. Forcing one store to answer both questions usually weakens one of them.

What to measure before copying this design

Start with four observations from your own traffic: the normal failure count per window, how retries inflate that count, how many distinct tenants are affected, and how long an incident remains above threshold. Those values determine the window and threshold. They also reveal whether a count is enough or a failure ratio is needed.

Then test the boring mechanics. Confirm the query can isolate the intended environment and rollout cohort without invented parameters. Verify that two workers racing on the same incident produce one outbox record. Stop a worker between the database commit and notification dispatch, then confirm the dispatcher recovers the pending row. Finally, monitor the cron heartbeat independently.

The decision rule is straightforward: choose this pattern when aggregate failure counts are sufficient, policy belongs in application code, and Postgres already provides durable state. Choose a managed monitor when notification routing and on-call controls matter more than keeping policy beside the rollout. Choose an error specialist when stack traces and release context are the real diagnostic unit.

For a flagged support-pricing change, the counter should remain small and the allocation trail should remain precise. That separation produces an alert you can act on without turning tenant identity into a metrics-cardinality problem.

Further reading

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to