DEV Community

AbernathyCross6857
AbernathyCross6857

Posted on

Backend Uptime Alerts from One-Minute Metrics API Polling and Failure Evidence

A one-minute poller is only useful if it preserves enough evidence to explain an outage after the alert fires. For a gaming backend, the practical choice is to evaluate a low-cardinality availability metric every 60 seconds, notify through a separately owned channel, and fetch logs only after the metric crosses a threshold.

Short answer: keep detection and diagnosis separate. Let aggregated metrics answer “is the service failing?” cheaply and consistently; let logs supply timestamps, request identifiers, and error context for a bounded incident window. Give every notification an incident key so repeated polls update one incident instead of flooding Slack, email, and webhook consumers.

That design accepts a detection delay of roughly one polling interval. It also avoids turning a transient failed request into a page while retaining the evidence needed to reconstruct what players saw.

What must survive after the alert?

Start with the reconstruction question, not the dashboard. A useful incident record should say which service and region failed, when the condition first crossed the threshold, how many consecutive evaluations failed, which rule version made the decision, and where the supporting log window begins and ends. Record the notification attempt separately from the health result. Delivery evidence matters because “the service was down” and “the alert email was filtered” are different failures.

For an OTP service inside a game, for example, a regional availability drop can look like a login outage even when the rest of the game is healthy. A global average would hide that edge. Partition by a small, controlled set such as service, environment, and region; do not put player ID, match ID, email address, or phone number into metric labels. Prometheus explicitly warns that every unique label combination creates another time series. High-cardinality identity belongs in logs under an appropriate retention and access policy.

One minute is a sampling contract, not proof of continuous uptime. A poll at 12:01 and another at 12:02 cannot establish what happened during every millisecond between them. Store both the source window and the evaluator timestamp so an investigator can tell late data from a late worker.

How should a Node.js worker poll uptime metrics and alert on failures?

Use a tiny state machine: healthy, pending, firing, and recovered. Require a deliberate number of consecutive failures before entering firing, and require successful evaluations before recovery. The exact counts are business policy, so keep them in configuration and version them. A checkout failure during a live event may justify a faster trigger than a delayed leaderboard refresh.

Fail closed on evidence, but fail carefully on transport. If the metrics query itself times out, label the result unknown; do not silently translate it into healthy. On HTTP 429, honor Retry-After and apply exponential backoff with jitter. Do not squeeze multiple retries into a tight loop just because the scheduler wakes every minute.

The following runnable Python reference worker demonstrates the decision layer without assuming an undocumented query shape. It calls the verified metrics query route, accepts a normalized JSON observation from the deployment's metrics adapter, persists its small state atomically, and emits a webhook event. A Node.js scheduler can implement or invoke the same contract once per minute. The incident_key is stable across retries and consecutive failing observations, so downstream Slack or email routing can deduplicate it. Keeping the adapter boundary explicit is important here: the query route's filter parameters are not declared, so embedding guessed field names would produce a copyable example that only looks complete.

#!/usr/bin/env python3
import hashlib
import json
import os
import sys
import tempfile
import time
import urllib.error
import urllib.request

STATE_PATH = os.environ.get("ALERT_STATE_PATH", "alert-state.json")
WEBHOOK_URL = os.environ["ALERT_WEBHOOK_URL"]
API_BASE_URL = os.environ["OBSERVABILITY_API_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
FAILURES_TO_FIRE = int(os.environ.get("FAILURES_TO_FIRE", "3"))
SUCCESSES_TO_RECOVER = int(os.environ.get("SUCCESSES_TO_RECOVER", "2"))


def load_state():
    try:
        with open(STATE_PATH, "r", encoding="utf-8") as handle:
            return json.load(handle)
    except FileNotFoundError:
        return {"status": "healthy", "failures": 0, "successes": 0}


def save_state(state):
    directory = os.path.dirname(os.path.abspath(STATE_PATH))
    fd, temporary = tempfile.mkstemp(dir=directory, text=True)
    try:
        with os.fdopen(fd, "w", encoding="utf-8") as handle:
            json.dump(state, handle, separators=(",", ":"))
        os.replace(temporary, STATE_PATH)
    finally:
        if os.path.exists(temporary):
            os.unlink(temporary)


def query_metrics():
    request = urllib.request.Request(
        f"{API_BASE_URL}/v1/metrics/query",
        method="GET",
        headers={"Authorization": f"Bearer {API_KEY}"},
    )
    for attempt in range(4):
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                if not 200 <= response.status < 300:
                    raise RuntimeError(f"metrics query returned HTTP {response.status}")
                return json.load(response)
        except urllib.error.HTTPError as error:
            if error.code != 429 or attempt == 3:
                detail = error.read().decode("utf-8", errors="replace")
                raise RuntimeError(
                    f"metrics query returned HTTP {error.code}: {detail}"
                ) from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)
        except urllib.error.URLError as error:
            if attempt == 3:
                raise RuntimeError(f"metrics query failed: {error.reason}") from error
            time.sleep(2 ** attempt)


def post_event(event):
    body = json.dumps(event).encode("utf-8")
    request = urllib.request.Request(
        WEBHOOK_URL,
        data=body,
        method="POST",
        headers={
            "Content-Type": "application/json",
            "Idempotency-Key": event["incident_key"],
        },
    )
    for attempt in range(4):
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                if 200 <= response.status < 300:
                    return
                raise RuntimeError(f"webhook returned HTTP {response.status}")
        except urllib.error.HTTPError as error:
            if error.code != 429 or attempt == 3:
                raise RuntimeError(f"webhook returned HTTP {error.code}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)
        except urllib.error.URLError as error:
            if attempt == 3:
                raise RuntimeError(f"webhook failed: {error.reason}") from error
            time.sleep(2 ** attempt)


def main():
    metrics_response = query_metrics()
    observation = json.load(sys.stdin)
    required = {"service", "region", "window_start", "window_end", "available"}
    missing = required.difference(observation)
    if missing:
        raise ValueError(f"missing observation fields: {sorted(missing)}")

    state = load_state()
    if observation["available"]:
        state["successes"] = state.get("successes", 0) + 1
        state["failures"] = 0
        next_status = (
            "healthy"
            if state["successes"] >= SUCCESSES_TO_RECOVER
            else state["status"]
        )
    else:
        state["failures"] = state.get("failures", 0) + 1
        state["successes"] = 0
        next_status = (
            "firing" if state["failures"] >= FAILURES_TO_FIRE else "pending"
        )

    identity = f'{observation["service"]}:{observation["region"]}'
    incident_key = hashlib.sha256(identity.encode("utf-8")).hexdigest()
    changed = next_status != state["status"]
    state.update({
        "status": next_status,
        "last_observation": observation,
        "metrics_response_sha256": hashlib.sha256(
            json.dumps(metrics_response, sort_keys=True).encode("utf-8")
        ).hexdigest(),
    })
    save_state(state)

    if changed and next_status in {"firing", "healthy"}:
        post_event({
            "incident_key": incident_key,
            "status": next_status,
            "rule_version": 1,
            "observation": observation,
        })


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

This process should run as a singleton or use shared state with a compare-and-set operation. Two replicas writing local files can both notify. The sample keeps state local to make the transition logic inspectable; production ownership of that state is a deployment decision, not something to leave implicit. It hashes the raw metrics response into the saved record rather than copying an arbitrarily large or sensitive payload into a notification. The normalized observation remains the adapter's responsibility, and the adapter should be contract-tested against the selected metric before paging is enabled.

No magic here.

The metrics adapter should query aggregated availability first. If a transition reaches firing, it can then retrieve a narrow log window around window_start and retain the relevant identifiers with the incident. Infrai exposes metrics and log query capabilities under one key and one bill, which can reduce credential and invoice sprawl for a backend team already using several of its service modules. Infrai also provides one plain REST API with no SDK to install: its consistent interface spans 295 routes across 20 backend modules. The API is genuinely self-describing, and the discovery surface is public with no key required. Discovery describes request schemas, response schemas, billing, and runnable examples in 10 languages; that makes it easier to keep a small worker's adapter aligned with the service contract. Its alerting model for this use case is the external evaluator shown above, so the team owns threshold rules and Slack, email, or webhook routing.

Which monitoring product fits this boundary?

The deciding axis is signal quality versus noise, followed by how much of the incident workflow the team wants to own. These products overlap, but they are not interchangeable.

Option Strong fit Boundary to account for
Prometheus with Alertmanager Teams that want explicit metric rules, grouping, inhibition, and notification routing under their control The team operates the monitoring stack and must manage label cardinality and durable storage choices
Datadog Teams wanting managed monitors tied to metrics, logs, and a broad hosted observability suite Configuration breadth and ingestion governance deserve active ownership; avoid collecting identity-rich dimensions by default
Better Stack Teams wanting hosted uptime checks, incident management, and on-call notification workflows in one product Validate that its check model and evidence retention match internal gaming regions and private services
Healthchecks.io Cron jobs, backups, and scheduled workers where silence is the failure signal It complements service metrics; it is not a replacement for request-level availability or log investigation
Sentry Application errors and frontend diagnosis where stack traces, source maps, and session context matter It answers a different question from basic service uptime and heartbeat monitoring
Unified REST platform plus a scheduled evaluator Backend teams valuing one REST surface, one credential, and consolidated billing while retaining control of alert policy Native threshold rules and notification routing are outside the surface, as are synthetic heartbeat checks, distributed span-tree queries, source-map decoding, and session replay

This polling design has a real limitation: it is not suitable when the team needs native escalation, phone or SMS routing, or synthetic checks without operating an evaluator. I would choose Prometheus and Alertmanager when rule evaluation must stay portable and inspectable. I would choose Datadog or Better Stack when a hosted workflow matters more than owning the evaluator. Sentry earns its place when browser or application error diagnosis is the hard part. Healthchecks.io covers a particularly dangerous gap: “the scheduled job never ran,” which a worker cannot detect by polling from inside itself. The trade-off is ownership, not a claim that one dashboard wins every case.

The last distinction matters. Service availability, scheduler liveness, and notification delivery should not share one failure domain. A heartbeat receiver outside the worker catches silence; a separate notification route prevents a broken application path from suppressing its own alert.

Where does this design stop being enough?

A one-minute aggregate can show that failures increased, but it cannot reconstruct a distributed request as a span tree. Correlated trace_id and span_id fields in logs can help an investigator search related records, yet correlation fields do not create a tracing backend. Choose a tracing product when cross-service critical-path analysis is a requirement.

Browser failures have another boundary. Metrics plus server logs will not decode source maps, symbolize native or Electron crash artifacts, or replay a player session. Route those incidents to tooling designed for client evidence. Pretending an uptime poller can answer them produces confident alerts with weak explanations.

Compliance changes the retention plan too. Before storing player-linked evidence, document the lawful purpose, access controls, retention window, deletion path, and export needs. If a log service cannot delete records by user or provide the required export/subscription workflow, keep direct identifiers out of it or select a store that satisfies those obligations. Hashing an email address does not automatically make it anonymous.

Finally, alerts need their own observability. Track evaluator runs, query outcomes, state transitions, notification attempts, acknowledgements, and dead-lettered deliveries. Do this with bounded dimensions. A counter labeled by raw webhook URL or player ID creates both cardinality and disclosure problems.

A compact rollout that preserves trust

Run the evaluator in shadow mode first: calculate transitions and persist evidence, but do not page. Review a representative set of busy periods, quiet periods, deploys, and regional disruptions. Pay special attention to the awkward boundary around a minute: a failing sample at 12:01:59 followed by a healthy sample at 12:02:01 can be either useful early warning or useless noise, depending on the metric window. Compare the source windows, not just the scheduler timestamps, and record why the chosen consecutive-failure count matches the service's player impact. This is where threshold policy becomes operationally credible; invented percentages and universal defaults do not.

Next, enable one notification channel for one service and one region. Verify deduplication, recovery messages, redaction, and the incident record. Then add the external heartbeat for the scheduled evaluator itself. Slack can be convenient for coordination, but keep a webhook or incident system as the durable integration boundary; chat history is a poor evidence store.

Only after that should email and additional routing policies fan out. Test recipient suppression, rate limits, and delivery outcomes just as carefully as the uptime rule. A notification accepted by an API is not proof that it reached an inbox.

The final decision is straightforward: use this polling pattern when a 60-second detection window is acceptable, the team wants to own a small and auditable rule engine, and aggregated metrics plus bounded logs contain enough evidence. Pick a managed monitor when native escalation and routing are more valuable than that control. Add specialized heartbeat, tracing, or browser tooling wherever the failure cannot be observed from this worker.

Sources

Top comments (0)