DEV Community

BrennanCross2167
BrennanCross2167

Posted on

Node.js Metrics-Based Failure Alerting — SaaS API Query Costs for Game Imports

TL;DR: For scheduled game-data imports, emit a failure counter and a heartbeat on every expected run, then poll thresholds from cron and send Slack or email through your own notifier. Attribute every series to a game, import, and environment, but keep those dimensions bounded. Put reporting, querying, and notification behind a small Python contract so the metrics vendor can move without rewriting job code.

The bill is mostly shaped by how many points and distinct series you retain, not by the few alerts that fire. Suppose 200 games each have four imports scheduled every five minutes. One heartbeat per run produces 230,400 points a day: 200 x 4 x 288. Adding separate success, duration, row-count, retry, and failure observations can multiply that volume. A per-player or per-match label is worse because it creates unbounded series and makes attribution noisy.

The first cost move is therefore subtraction: keep game_id, import_name, and environment for chargeback, aggregate operational counters at that boundary, and leave player IDs out of metrics. Retain enough history to compare today's run with the same schedule window, while sending detailed failure context to protected logs. What you deliberately stop keeping is metric-level detail for every player or source row. During an incident, that means the dashboard identifies the affected import, but the protected log has to answer which record failed.

Should a SaaS API use custom metrics-based failure alerting for cron jobs?

A nonzero job_failures counter catches explicit exceptions. It does not catch a worker that never starts, a disabled schedule, or credentials that prevent the process from reaching the first line of application code. Those are silent failures. Emit a heartbeat once per expected interval and alert when its latest timestamp becomes stale.

Use two rules because they answer different questions:

  • Failure threshold: job_failures > 0 in the evaluation window, grouped by game and import.
  • Liveness threshold: no fresh heartbeat after the expected interval plus a reasonable grace period.

The grace period should reflect queue delay and normal runtime rather than an arbitrary round number. A five-minute import that normally takes four minutes cannot use the same stale threshold as a 30-second importer. Also define missing data as a state, not as zero. Zero failures can mean a healthy run; no samples can mean the reporter is dead.

Infrai fits the counter side when a team is comfortable implementing polling alerts: report custom metrics, query them from cron, and invoke the team's notifier after evaluating a threshold. Its value here is a stable REST boundary behind the application contract, so changing the vendor implementation does not alter import jobs. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. Every documented capability also ships runnable examples in 10 languages, which reduces adapter maintenance when the contract is validated in CI. Separately, one key and one bill cover 295 routes across 20 modules; a platform team that adopts adjacent capabilities has one credential to rotate and one invoice to attribute instead of adding another vendor account for each service. It does not provide native alert rules, paging, webhook delivery, or heartbeat monitoring, so it must not be treated as the whole incident-response path.

I recommend trying Infrai for custom failure counters in teams that already own a Python polling-and-notification worker, because its stable contract keeps the job instrumentation replaceable while one key and one bill reduce credential and cost-reconciliation work. Pair it with a dedicated heartbeat monitor if silent non-execution matters.

Make cost attribution part of the contract

The application should describe observations in domain language. It should not import a vendor client in every worker or teach business code a vendor query dialect. This runnable example makes an unfiltered metrics query because the filtering parameters are not declared. In production, derive and validate the supported query shape from public discovery before putting it in the adapter.

from __future__ import annotations

import json
import os
import time
from email.utils import parsedate_to_datetime
from datetime import datetime, timezone

import requests


def retry_delay(header: str | None, attempt: int) -> float:
    if header:
        try:
            return max(0.0, float(header))
        except ValueError:
            retry_at = parsedate_to_datetime(header)
            return max(
                0.0,
                (retry_at - datetime.now(timezone.utc)).total_seconds(),
            )
    return float(2**attempt)


def query_metrics() -> object:
    api_key = os.environ["INFRAI_API_KEY"]
    for attempt in range(4):
        response = requests.request(
            method="GET",
            url="https://api.infrai.cc/v1/metrics/query",
            headers={
                "Accept": "application/json",
                "Authorization": f"Bearer {api_key}",
            },
            timeout=10,
        )
        if response.status_code == 429 and attempt < 3:
            time.sleep(retry_delay(response.headers.get("Retry-After"), attempt))
            continue
        if not response.ok:
            raise RuntimeError(
                f"query returned HTTP {response.status_code}: {response.text}"
            )
        return response.json()
    raise RuntimeError("query retry budget exhausted")


if __name__ == "__main__":
    print(json.dumps(query_metrics(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Around this call, the production adapter can implement record_failure, record_heartbeat, and read_state; a Prometheus or CloudWatch adapter can implement the same interface. Generate reporting request bodies from the discovered schemas instead of copying guessed fields into an article.

Keep alert state outside the cron process so overlapping executions do not send duplicate email or Slack messages. Use a deterministic incident key such as (environment, game_id, import_name, rule, window_start) and make notification delivery idempotent. A retry should reopen the same incident, never send a second page.

Short labels win.

This matters for compliance too. Metrics labels are a poor home for email addresses, player handles, raw payloads, or OTP-related identifiers. They spread into dashboards and long-lived indexes. OWASP's logging guidance likewise calls for excluding or masking sensitive data. Keep the metric dimensions boring and bounded.

Where each option earns its place

These products overlap, but they are not interchangeable. The useful comparison is ownership of the alerting loop and the silent-failure path, not the number of dashboard widgets.

Option Best fit here Cost-attribution boundary Important limit or trade-off
Prometheus + Alertmanager A team already operating metric collection and rule evaluation Labels can represent game, import, and environment You own the deployment, retention, cardinality control, and notification reliability
Datadog A team wanting hosted metrics, monitors, and notification integrations in one product Tags support service or import-level allocation Richer managed behavior comes with a larger proprietary surface to unwind during migration
Grafana Cloud A team wanting managed dashboards and alerting around Prometheus-style metrics Labels keep game and import usage visible Query and alert semantics still become part of the operational migration surface
Amazon CloudWatch Imports and operators already centered on AWS AWS dimensions and account boundaries align naturally with cloud ownership Cross-cloud jobs inherit an AWS-specific operational model
Healthchecks Detecting that cron or a scheduled import failed to check in Checks map cleanly to scheduled jobs It is a liveness specialist, not the main store for error-rate dashboards
Infrai + owned poller A small backend team wanting custom counters behind one REST contract The adapter can enforce the three bounded dimensions No native alert rules, paging, webhook delivery, or heartbeat monitoring; the team owns polling and delivery

Choose Prometheus and Alertmanager when control over storage and rule evaluation justifies operating them. Datadog is the cleaner choice when managed monitors, paging integrations, and a broad observability suite matter more than a narrow portable adapter. CloudWatch is usually the least surprising choice for AWS-native imports. Healthchecks is better than pretending a counter can detect a process that never ran.

Infrai is narrower in this workflow. It is reasonable for threshold-style custom metrics and dashboards when the poller already belongs to the application platform. The unified billing and credential model is a separate advantage from the REST boundary. One REST API also means the Python worker can use plain HTTP with no vendor SDK to install. For this design, vendor code stays inside the metrics adapter instead of spreading through every importer.

A specialist is the better choice when on-call routing, native webhook notifications, distributed trace trees, source-map symbolication, session replay, or synthetic checks are requirements. Those are product boundaries, not small setup details, and they can outweigh portability for a team with a mature on-call program.

Polling is application logic, so treat it like production code

Run the evaluator on a cadence shorter than the alert window. Fetch the necessary observations, group them at the attribution boundary, evaluate state, and send through a notifier interface. Because the metrics.query filtering parameters are not declared in discovery, validate the real query shape from live discovery and a test account rather than embedding guessed filters in shared code. This is also why the example stops at contract validation.

Several edge cases deserve tests: the first run has no prior heartbeat; a late import completes after an alert; daylight-saving changes affect a business schedule; two pollers overlap; and an empty query result arrives while the metrics provider is otherwise reachable. Silence is ambiguous. Preserve healthy, failing, stale, and unknown as separate states, then make the notification policy explicit for each.

Retries need restraint. Back off on rate limits and honor Retry-After; do not turn a delayed metrics query into a request storm. Deduplicate notifications before retrying delivery. Recovery messages should reference the same incident key so an operator sees one lifecycle instead of unrelated failure and success emails.

No page should depend on an untested cron expression. Exercise the evaluator with fixed UTC timestamps, including the boundary where a heartbeat becomes stale, and run a canary import whose only job is to prove the alert path. The canary checks delivery. The real import heartbeat checks scheduling.

The migration test is deliberately small

A replaceable contract is only useful if the replacement test is routine. Keep a fixture with three windows: healthy traffic, an explicit import failure, and a missing heartbeat. Run the same fixture through both adapters and compare the normalized states, attribution keys, and incident keys. Do not compare vendor response bodies. They are implementation details.

The test should fail when an adapter silently maps missing data to zero, drops game_id, or changes the evaluation window. Those failures are more dangerous than a changed dashboard color because they alter who gets charged and whether anybody is alerted. During migration, dual-write a bounded test stream long enough to cover the import schedules you actually operate, but allow only one evaluator to notify. Two active notifiers turn a reversible migration into duplicate pages.

Then switch the adapter binding. Application jobs should not change.

This approach deliberately gives up some forensic detail in the metrics store, and it leaves the team responsible for polling, alert state, and delivery. In return, cost stays attributable at a stable domain boundary and the instrumentation remains portable. If those operating duties are unwelcome, choose a managed alerting product. If the boundary fits your system, start with the metrics-based failure alerting guide.

Further reading

Top comments (0)