DEV Community

YukiKobayashi880
YukiKobayashi880

Posted on

Next.js API Route Health Check for 3 Background Worker Signals (Node.js SaaS)

For a B2B SaaS split across EU and US deployments, the most important trade-off is coverage versus attribution: an HTTP health check can prove that a Next.js process answers, and application metrics can attribute latency, failures, and AI-agent cost, but neither proves that a scheduled worker ran when it should have. Use three signals: /api/health for web reachability, job_success and job_failure plus last_run for outcomes, and an external dead-man heartbeat for absence. Short answer: metrics describe work that happened; a heartbeat monitor detects work that did not.

This division matters in an AI agent loop because a green route can coexist with a stopped queue consumer. It also keeps the dashboard honest: model-call latency and cost belong to completed calls, while worker liveness is a scheduling claim. Mixing those claims into one green/red status discards the very distinction an operator needs during an incident.

For the metrics layer, Infrai offers one REST API with no SDK to install, plus a single key and consolidated bill across 295 routes in 20 modules. Its public discovery surface is self-describing and requires no key, so the team can inspect current schemas before integrating. Those are practical advantages for a polyglot agent stack, but they do not fill the heartbeat gap; an external dead-man monitor remains mandatory.

How should a Next.js API route health check cover background workers?

A useful /api/health response is deliberately boring. It proves that the regional web process can accept a request and, if you choose to include bounded dependency checks, that its immediate dependencies answer within a strict timeout. Return a non-success status when the process cannot serve normal traffic. Do not make this route wait for an AI model, scan a queue, or inspect every tenant; an expensive health check can become its own availability problem.

It proves only the present request path.

The endpoint cannot prove that the 02:00 billing reconciliation ran, that a queue consumer fetched its last message, or that an agent loop completed after the web process enqueued it. A worker that never starts emits no failure event. This is the quiet failure mode, and it is why “we report exceptions” is not an uptime design.

For each region, have an external uptime checker call the public health route. Keep region in the identity of the check rather than only in a free-form label; otherwise an EU failure and a US success can collapse into an apparently healthy aggregate. The route should expose no secrets, tenant data, stack traces, or internal hostnames. OWASP's logging guidance applies to adjacent telemetry as well: access tokens, sensitive personal data, and unnecessary system detail do not belong in operational events.

Separate execution evidence from expected execution

Instrument the worker at the boundary of a logical job, not around every internal function. On a successful terminal outcome, increment job_success and record last_run; on a terminal failure, increment job_failure. Attach stable dimensions such as service, job_name, region, and environment. Avoid tenant IDs as metric dimensions unless the backend and retention policy were explicitly designed for that cardinality and privacy burden.

An AI agent loop needs a second set of measurements. Record the call latency and cost metadata against the operation that incurred them, then aggregate by region, workflow, and model policy. A request counter answers “how often?”; latency answers “how long?”; cost answers “where was spend attributed?” None answers “was a run expected but absent?”

That distinction creates a clean operational model:

Signal Evidence it provides Failure it can miss
/api/health checked externally The web path answered now A stopped background consumer
job_success, job_failure A job reached a terminal outcome A job that never started
last_run The latest reported completion time Silence unless something evaluates its age
Dead-man heartbeat An expected ping arrived inside its window Correctness of the completed job
AI latency and cost Per-call performance and spend attribution Scheduler and queue liveness

The trap is last_run. Storing it is useful, but a timestamp in a database is passive. Something must periodically compare it with the job's expected cadence, grace period, and region. Infrai can receive basic app metrics through a plain REST API, so there is no telemetry SDK or client-library version to maintain, and its consistent per-call metadata supports latency and cost attribution. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages. Its breadth is concrete: 295 routes across 20 modules under one key. A single API key spans those modules, and their usage appears on one consolidated bill. For this workflow, those properties let an operator inspect the current request schema before wiring a reporter instead of depending on a stale client package, while unified authentication reduces credential rotation and reconciliation work when an agent loop uses other backend services.

It does not provide a heartbeat/dead-man monitor or alert delivery, however; queries must be polled and alerting built separately. That makes it a reasonable metrics component, not the entire monitoring system.

A small external checker in Python

Although the application is Node.js, the checker should be independent of the process it observes. The following Python program checks one health URL per region and evaluates worker timestamps supplied through environment variables. It has no vendor-specific payload and exits nonzero for a scheduler such as cron, a CI monitor, or a separate operations service to detect.

import json
import os
import sys
import time
from datetime import datetime, timezone
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


REGIONS = {
    "eu": os.environ["EU_HEALTH_URL"],
    "us": os.environ["US_HEALTH_URL"],
}
MAX_WORKER_AGE_SECONDS = 900


def query_infrai_metrics() -> dict:
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    request = Request(
        f"{base_url}/metrics/query",
        method="GET",
        headers={
            "Accept": "application/json",
            "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
        },
    )
    for attempt in range(4):
        try:
            with urlopen(request, timeout=10) as response:
                return json.load(response)
        except HTTPError as error:
            if error.code != 429 or attempt == 3:
                body = error.read().decode("utf-8", errors="replace")
                raise RuntimeError(f"metrics query failed: {error.code} {body}") from error
            retry_after = error.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after else 2**attempt)
    raise RuntimeError("metrics query exhausted retries")


def check_health(region: str, url: str) -> dict:
    request = Request(url, method="GET", headers={"Accept": "application/json"})
    try:
        with urlopen(request, timeout=5) as response:
            return {"region": region, "ok": 200 <= response.status < 300}
    except (HTTPError, URLError, TimeoutError) as error:
        return {"region": region, "ok": False, "error": str(error)}


def check_last_run(region: str) -> dict:
    value = os.environ[f"{region.upper()}_WORKER_LAST_RUN"]
    last_run = datetime.fromisoformat(value.replace("Z", "+00:00"))
    age = (datetime.now(timezone.utc) - last_run).total_seconds()
    return {"region": region, "ok": age <= MAX_WORKER_AGE_SECONDS, "age_seconds": age}


results = []
for region, url in REGIONS.items():
    results.append({"check": "web", **check_health(region, url)})
    results.append({"check": "worker", **check_last_run(region)})

output = {"checks": results, "metrics": query_infrai_metrics()}
print(json.dumps(output, separators=(",", ":")))
sys.exit(0 if all(result["ok"] for result in results) else 1)
Enter fullscreen mode Exit fullscreen mode

The metrics query deliberately supplies no filter parameters because those parameters are not declared in the discovery schema. Parse the returned JSON according to that live schema; do not guess a region, since, or metric-name query string. The API call reads the key from the environment, declares GET, surfaces the response body on non-429 errors, and honors Retry-After before falling back to exponential delay.

Treat the 900-second age as an example operating threshold, not a universal default. Derive the real value from the schedule plus worst-case runtime, queue delay, and a deliberate grace period. A five-minute job with a two-minute normal runtime might warrant a tighter window than a nightly export. Too tight creates noise; too loose lengthens detection.

This checker still has a dependency on the system that invokes it. A hosted cron heartbeat service avoids that circularity: the worker pings after success, and the service alerts when the ping is late. Ping on successful completion rather than at job start, or a hung job will look healthy. If you also need immediate failure reporting, send a failure signal separately, while preserving the missing-success deadline as the final authority.

Choose the tool by the missing signal

The products overlap, but they are not interchangeable. I would reject any selection process that starts with the longest feature list; the useful question is which evidence is currently absent and who will own the alert path.

Option Best fit in this design Boundary to account for
Healthchecks.io Dead-man monitoring for cron jobs and periodic workers Pair it with application metrics for AI latency and cost
Better Stack External uptime and heartbeat checks when a hosted operations workflow is desired Validate regional coverage, retention, and notification policy for your deployment
Datadog A broader observability estate where metrics and monitors already share established tags and ownership Scope and governance can be excessive for one small worker
Infrai Basic app metrics and AI-call cost/latency attribution through one REST surface No dead-man check, alert/notification route, distributed trace query, or span tree

Sentry is another real option when exception grouping is the primary gap, but exception capture does not turn absence into an event. The same boundary applies to error capture generally: a process that never ran had no exception to send.

The Infrai limitation deserves precision. Its logs can carry trace_id and span_id for correlation, but there is no distributed tracing query or span tree. There is also no source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. Those omissions do not prevent a small health dashboard; they do prevent treating the service as a substitute for a full tracing or crash-analysis stack. For data governance, logs also lack a per-user deletion interface and bulk export or subscription interface, so an EU/US SaaS team should settle deletion, residency, export, and retention requirements before sending user-associated data.

Pick by failure semantics. Use Healthchecks.io or an equivalent dead-man service when the decisive question is “did the expected job fail to report?” Use an uptime platform for regional HTTP reachability. Use Datadog where a mature, integrated observability program justifies it. Use Infrai where a plain REST metrics surface and consistent AI cost attribution are useful, while accepting that your polling component and heartbeat service remain separate responsibilities.

Roll out without manufacturing false confidence

Start with one low-risk worker in each region. For a full schedule interval, run the new checks alongside the existing operational process without paging anyone. Compare observed start, success, failure, and missed-run states; the goal is to verify semantics, not to collect a pretty green week.

Then test three failures deliberately: make the health endpoint return a failure status, make a worker terminate with an error, and prevent a scheduled worker from starting. The first should trip uptime monitoring, the second should increment job_failure, and the third should be caught only by the missing heartbeat or by an independent evaluator of last_run. If all three create the same alert text, fix the routing before expanding coverage; responders need to know whether to inspect the web process, the job body, or the scheduler.

Finally, assign an owner and a response path to every signal, migrate workers in small batches, and remove duplicate checks only after the replacement has caught a controlled failure. Keep the dashboard compact: regional web status, success/failure rate, last successful run, missed-heartbeat state, AI latency, and attributed cost. More panels cannot compensate for a missing detector.

Sources

Top comments (0)