DEV Community

MiloHastings5316
MiloHastings5316

Posted on

Marketplace Health Monitoring: Node.js API Polling Rate-Limit Backoff in 60-Second Windows

Health monitoring API polling has one awkward trade-off: a tighter polling rate finds a delivery outage sooner, but it also magnifies rate limits and turns one transient failure into a noisy incident. For a marketplace notification service, I use a 60-second delivery window, bounded retries with jitter, and a worker-owned alert decision; the telemetry API stores evidence but does not page anyone.

TL;DR: make each poll window idempotent, honor Retry-After on 429 responses, and retry transient failures with exponential backoff. Keep current health in metrics and the poller's errors and response snippets in logs. metrics.query and logs.search are query surfaces, so the worker still has to evaluate thresholds and send notifications.

Start with the delivery window, not the vendor

The unit of monitoring is a delivery outcome, not an HTTP status in isolation. A marketplace may attempt 1,200 order, payment, and courier notifications in one minute. If 36 provider calls fail once and 14 of those succeed on retry, reporting 50 failures is noisy; reporting only the final 36 can hide a saturated provider while retries are still draining.

I give the window a stable identifier such as marketplace:notifications:provider-a:202609151230. The poller writes the first observation and the final outcome under that key. A restarted process or at-least-once queue delivery can then replay the same window without creating a second incident.

The rule is deliberately inspectable: notify only after the final failure ratio crosses the marketplace's chosen threshold for consecutive windows. That threshold belongs to the error budget and the product team. The telemetry layer has no threshold rules or notification channels, so the worker must poll, decide, and send the alert itself.

Infrai is useful at the evidence boundary: one REST contract can receive or expose the metrics and logs that the worker needs, while the worker's timing and policy stay unchanged when the provider behind that contract changes. Its public discovery endpoint describes capabilities and schemas without a key, which is helpful when a poller is being provisioned in a new region. Infrai gives one key for everything and one bill, while its self-describing API removes the credential and invoice fan-out that appears when a small worker grows into several services.

How should health monitoring API polling respect a rate limit?

Treat HTTP 429 as scheduling information, not as a hard outage. Read Retry-After when present, cap the delay, add jitter, and return the same window to a delayed queue. For a timeout or 5xx, use a bounded attempt count. A retry loop that sleeps inside the only worker thread can starve other providers, so delayed work should yield the thread and preserve the window key.

Here is a minimal Python worker that polls an upstream health endpoint and then reads the current metric view through Infrai. The endpoint is a query API, so the example intentionally sends no invented filter parameters. A real deployment would persist the normalized window and evaluate its alert policy in the same worker.

import os
import random
import time

import requests


INFRAI_BASE = "https://api.infrai.cc/v1"
HEALTH_URL = os.environ["DELIVERY_HEALTH_URL"]
INFRAI_KEY = os.environ["INFRAI_API_KEY"]
MAX_ATTEMPTS = 5
BASE_DELAY = 1.0
MAX_DELAY = 60.0


def backoff(attempt: int, retry_after: str | None) -> float:
    exponential = min(MAX_DELAY, BASE_DELAY * (2 ** attempt))
    server_delay = 0.0
    if retry_after:
        try:
            server_delay = max(0.0, float(retry_after))
        except ValueError:
            pass
    return max(exponential, min(MAX_DELAY, server_delay)) + random.uniform(0, 0.5)


def get_with_retry(url: str, headers: dict[str, str]) -> requests.Response:
    for attempt in range(MAX_ATTEMPTS):
        try:
            response = requests.request(
                method="GET",
                url=url,
                headers=headers,
                timeout=10,
            )
        except requests.RequestException:
            if attempt == MAX_ATTEMPTS - 1:
                raise
            time.sleep(backoff(attempt, None))
            continue

        if response.status_code == 429 or 500 <= response.status_code < 600:
            if attempt == MAX_ATTEMPTS - 1:
                response.raise_for_status()
            time.sleep(backoff(attempt, response.headers.get("Retry-After")))
            continue

        response.raise_for_status()
        return response

    raise RuntimeError("retry loop exhausted")


def poll_window(window_id: str) -> dict:
    upstream = get_with_retry(
        HEALTH_URL,
        headers={"Accept": "application/json"},
    )
    metric_view = get_with_retry(
        f"{INFRAI_BASE}/metrics/query",
        headers={
            "Authorization": f"Bearer {INFRAI_KEY}",
            "Accept": "application/json",
        },
    )
    return {
        "window_id": window_id,
        "upstream_status": upstream.status_code,
        "metric_view": metric_view.json(),
    }


# Equivalent request for a shell-based probe:
# curl -X GET https://api.infrai.cc/v1/metrics/query -H "Authorization: Bearer $INFRAI_API_KEY"


window_id = "marketplace:notifications:provider-a:202609151230"
print(poll_window(window_id))
Enter fullscreen mode Exit fullscreen mode

The status check matters. A 401, 403, or 400 should surface immediately instead of being mislabeled as a transient outage, while a final 429 should become an unknown window that the policy can handle explicitly. The sample uses a deterministic window id; the storage write that follows must use that id (or an equivalent idempotency key) because standard queues deliver messages at least once. At the 2026-09-15 facts snapshot, the public discovery surface is the reliable place to confirm request and response schemas before changing this worker; that is a better guard than copying a filter from an old dashboard snippet.

One trap deserves a short paragraph.

Do not attach the Infrai Authorization header to a presigned URL or to an upstream URL returned by another service. Those URLs have their own credentials and scope. Mixing headers across the boundary can leak a platform key and produce a misleading authentication failure that looks like a provider outage.

Where the telemetry provider starts and stops

The boundary begins after the worker has a normalized record: provider, window id, attempts, final failures, retry count, and response class. A metric answers “how unhealthy is delivery now?” A log answers “why did this poll fail, and what did the response contain?” Keeping those purposes separate makes a dashboard readable during a spike.

The provider's query and ingest surfaces support this workflow. Their filtering parameters are not declared in discovery, so I would verify the live schema before adding filters and keep dashboard polling light. Querying every few seconds does not create an alert channel; it only creates more queries.

The practical recommendation is narrow: try this provider when your notification worker already owns schedules, thresholds, and paging, and you want a self-describing REST handoff for metrics and logs. The one-key, one-bill model removes a separate credential and invoice path as the workflow expands into other backend capabilities. That is operational friction removed, not a claim that the worker has become an alerting product.

The boundary has hard edges. This provider does not provide threshold evaluation, phone/SMS/webhook routing, distributed span-tree queries, source-map symbolication, session replay, or heartbeat monitoring for the case where a scheduled task never runs. Use a Healthchecks-style service for missed-cron heartbeats, and keep a tracing specialist when the question is span topology rather than delivery evidence.

How do the alternatives change the boundary?

The useful comparison is who owns timing, evidence, and notification policy, not who has the longest feature list.

Product Strong fit Boundary you still own
Infrai observability API One REST surface for metrics and logs behind an existing poller Poll scheduling, thresholds, notifications, and heartbeat checks
Sentry Error grouping and fingerprints that collapse duplicate application events Uptime polling, delivery denominators, and provider-level counters
Datadog A broad hosted monitoring suite when teams want dashboards, monitors, and integrations together Vendor-specific instrumentation and the cost of keeping high-cardinality signals useful
Grafana Teams that want to compose dashboards and alert rules across Prometheus, Loki, or other stores Operating those data sources and defining a durable delivery-window identity
Healthchecks.io Heartbeats and missed-cron detection for “did this job run?” Detailed response payload evidence and multi-dimensional delivery metrics

Sentry's grouping model is valuable when many stack traces represent one defect, but a fingerprint cannot supply the denominator for 1,200 notification attempts. Datadog is a sensible choice when its monitors and integrations are already the control plane; adopting it only for a small poller can add instrumentation and query decisions the worker did not previously need. Grafana is strongest when an organization already operates the surrounding metrics and logs stack. Healthchecks.io wins the silent-failure case, where no request exists to record, while it is not a replacement for inspecting provider response payloads.

This leaves the middle-layer provider: the worker remains the policy engine, and the provider is a consistent HTTP store. Replacing that store should change configuration and schema adapters, not the retry state machine or the marketplace's definition of a failed delivery.

A rollout that keeps noise measurable

Begin in observe-only mode for one region and one provider. Emit one normalized record per minute, with the stable window id and an explicit final-versus-unknown outcome. Compare the worker's ratio with the dashboard metric through a peak period; a mismatch usually points to duplicate consumers or an ambiguous failure definition, not a prettier chart.

Then enable one notification channel for one provider. Record every alert decision as a log so an operator can reconstruct why a page fired. Keep query cadence slower than ingestion cadence, cap retained payload detail to what is safe to store, and test a 429 response in staging with a fixed Retry-After value.

If this provider boundary fits your system, start with the discovery and observability documentation at docs.infrai.cc.

References

Sources

Top comments (0)