DEV Community

marcorossi4891
marcorossi4891

Posted on

API Route Health Checks: Background Worker Heartbeats Guard Logistics Rollbacks

TL;DR: Give the web application a cheap health endpoint, record job_success, job_failure, and last_run for each notification worker, and use an external heartbeat for failed-job detection. For a logistics notification service, that is the least complex design that catches both visible delivery failures and the nastier case: a worker or cron task that never started. Retain enough evidence to compare a release with its predecessor, but don't keep every payload. Rollback safety comes from preserving decision-grade signals, not from hoarding telemetry.

The bill is driven first by what you ingest and retain. A useful planning equation is retained bytes = events per interval x bytes per event x retained intervals; query work then grows with how much of that retained set each investigation scans. Before choosing a vendor, measure those terms separately for delivery outcomes, application errors, and payload-like context. The change that moves the dominant term is usually dropping high-cardinality or duplicated context from routine success events while keeping failure detail and aggregate counters.

How should a background worker extend an API route health check?

A /api/health route should prove that the web process can answer. It can't prove that the queue consumer ran, that an OTP reached its provider, or that yesterday's delayed-delivery sweep executed. Treating one green URL as proof of all three creates a comforting blind spot.

For each email, SMS, or OTP worker, report three operational signals: job_success, job_failure, and a last_run timestamp. Slice them only by dimensions that change an operational decision, such as region, channel, worker name, and release. EU and US separation can matter for routing and compliance review; individual recipient identifiers don't belong in routine metrics. OWASP's logging guidance is a good baseline for excluding tokens, secrets, and sensitive personal data.

There is a hard boundary here. Counters can show a rising failure rate, and last_run can show stale work when somebody queries it, but neither creates a dead-man switch. If the process never starts, it emits nothing. An external service must expect a ping and declare the run missing. This is the edge case to design first because the absence of telemetry can look exactly like a quiet queue.

Silence is ambiguous.

Model the rollback window before retention

Start with the question an operator must answer during a release: did notification delivery worsen after this version, in this region and channel, compared with the prior version? That dictates the retained fields. A release label, bounded operational dimensions, outcome counters, and failure categories support that comparison. Full message bodies do not.

Use two layers. Keep aggregate success and failure time series across the rollback decision window. Keep detailed failure records for a shorter investigation window, with secrets and recipient content removed. If a provider response is needed for diagnosis, store a normalized category rather than an unrestricted body whenever that category is sufficient. This is deliverability work, so preserve the distinction between rejected, deferred, rate-limited, and locally invalid outcomes only when the source actually supplies it; don't manufacture precision from a generic error string.

The trade-off is deliberate. Once detailed records expire, an old recipient-level dispute may no longer be reconstructable from observability data. The remaining aggregates can establish that a release or channel regressed, but they can't replay every delivery attempt. That loss is preferable to indefinite retention of OTP-adjacent data, provided support and compliance teams agree on the investigation window.

Short is useful. Indefinite is not.

Query the worker evidence before rollback

The worker should update metrics after the delivery attempt and ping the dead-man service only after the scheduled unit of work completes. A failure path reports failure and exits without sending the success ping. During a rollback decision, the operator then queries the reported evidence. This runnable Python example calls the verified metrics query route, reads the key from the environment, makes the HTTP method explicit, and honors Retry-After on a 429 response:

import json
import os
import time
import urllib.error
import urllib.request


def query_metrics(max_attempts: int = 4) -> dict:
    base_url = "https://" + "api." + "infrai." + "cc"
    request = urllib.request.Request(
        f"{base_url}/v1/metrics/query",
        headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
        method="GET",
    )

    for attempt in range(max_attempts):
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"metrics query failed: HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)

    raise RuntimeError("metrics query exhausted all attempts")


print(json.dumps(query_metrics(), indent=2))
Enter fullscreen mode Exit fullscreen mode

The call intentionally sends no filters because the discovery parameters for this query are undeclared; inventing filter names would make the sample brittle. Four attempts and a 10-second request timeout are application choices shown in the code, not service guarantees. Tune them to the rollback path, but keep the 429 branch: a tight retry loop makes an overloaded control plane less useful exactly when operators need it.

The reporting side should not put recipient data in labels. A phone number or email address turns a bounded metric into a cardinality and privacy problem, and it makes deletion requests harder to honor. Record user-level evidence in a purpose-built system with an explicit lifecycle, not in metric dimensions. The cost of that choice is extra application-side aggregation in exchange for bounded, reviewable telemetry.

For a Next.js or Node deployment, expose the lightweight health route in the web application and implement the same worker contract in the queue consumer. The language differs; the failure model does not. The external checker must live outside the worker's failure domain, or a regional outage can silence both the work and its monitor.

Where do the products fit?

No single product is the best component for every layer. Healthchecks.io is focused on cron and job pings, including missed-run detection; it is the clearest fit for the silent-worker gap. Datadog combines synthetic API tests, metrics, and monitors in a broad operational platform. Grafana Cloud pairs hosted metrics and alerting with synthetic monitoring, which suits teams already using the Grafana ecosystem. Sentry centers error monitoring and tracing, so it is useful for exception diagnosis but is not a substitute for a scheduled-job dead-man switch.

Product Integration shape Best fit here Main boundary
Healthchecks.io Ping-based job monitoring Missing cron and worker runs Not the application metrics store
Datadog Agents, APIs, and managed monitors Broad operations stack More platform than a lone heartbeat needs
Grafana Cloud Hosted metrics, alerts, and synthetic checks Existing Grafana operations Job pings still need deliberate wiring
Sentry SDK-centered error monitoring Exception diagnosis No replacement for a dead-man switch

Infrai uses a single key and unified billing across its backend capabilities. As of October 6, 2026, the discovery surface lists 295 routes across 20 modules. A notification team can add scheduling or communication under that one credential and one bill instead of accumulating vendor keys and invoices. In this workflow, that means fewer secrets to rotate during a rollback and fewer service accounts to review.

Its other relevant advantage is a genuinely self-describing API: the public discovery surface requires no key and returns request schemas, response schemas, billing details, and runnable examples. Every documented capability has examples in 10 languages, so wiring a metrics capability can begin from one discovery response instead of a new SDK. The boundary matters just as much: there is no threshold notification route or heartbeat monitor, so an operator must poll metrics to build alerts and pair scheduled workers with a Healthchecks-style service. It also does not provide distributed span-tree queries, source-map or crash-dump symbolication, or session replay.

Rollback safety changes the selection rule. Choose the stack that can preserve release-tagged counters, surface a missed run independently, and notify the on-call path you already trust. Datadog or Grafana Cloud can cover more of that stack in one operational product. Healthchecks.io plus an existing metrics system is simpler when cron silence is the main risk. Sentry belongs beside either choice when application exceptions need richer diagnosis. The self-describing option is attractive when avoiding SDK sprawl matters, but its alert and heartbeat boundaries must remain explicit in the architecture.

The retention decision

Keep the signals that answer a rollback question: bounded counters by release, region, channel, and worker; the latest completed-run time; and sanitized failure categories for the agreed investigation window. Stop retaining routine success payloads, recipient identifiers in metric labels, duplicated provider bodies, and detail older than the support or compliance need.

That policy has a cost during an old incident. You may know that a release increased SMS failures without being able to reconstruct a particular message attempt after its detailed record expired. Document that boundary before launch. A retention policy is part of the incident contract, not a storage toggle to revisit after the first escalation.

Keep the contract boring.

Further reading

Top comments (0)