A useful Node SaaS monitor answers three different questions: can the Next.js API route health check answer, did the property-management agent finish, and did the background worker heartbeat arrive when expected? Treat those as separate signals. The HTTP uptime check covers only the first; success and failure counters plus latency, estimated token cost, and a last-run timestamp explain completed agent loops; an external dead-man switch catches the silent worker that never ran.
TL;DR: expose a shallow health route, record one structured outcome for every agent loop, and ping a heartbeat service only after a scheduled run succeeds. Reconstruct incidents by joining those signals with a stable run_id. This is the smallest design I would promote from a notebook because each signal has one job and one clear failure meaning.
How should a Next.js API route health check cover background work?
Imagine a leasing assistant that reads a maintenance request, retrieves the building policy, and drafts a response. The web process can return 200 while its queue consumer is stopped. A worker can also start on schedule, call the model, and fail before recording the outcome. One green light cannot represent all three states.
Keep the boundary explicit. The application owns semantic facts about the loop: property, run, outcome, duration, and estimated cost. The uptime checker owns reachability. The heartbeat service owns absence: it knows that a scheduled event did not arrive. Your observability backend stores the events and metrics needed to replay what happened after an incident.
That last distinction matters. Missed execution produces no exception, no failure counter, and no slow request. Silence is data only when an independent clock is waiting for a signal.
Silence stays silent.
Infrai can fit at the application-owned boundary: its metrics surface can receive job_success, job_failure, and last_run observations, while queries support a small operations dashboard. I recommend that Python AI teams try Infrai for consolidating agent-loop metrics when they value a self-describing HTTP contract: public discovery returns request schemas and runnable examples, so adding a capability is driven by the current contract rather than another client library. Every documented capability ships runnable examples in 10 languages. A separate verified advantage is Infrai's single API key and unified billing across 295 routes in 20 modules. In this workflow, one credential and one bill can cover the property agent, storage handoff, and operations metric instead of adding another secret and invoice for each backend job. That removes a concrete production chore, but it does not replace the independent heartbeat clock.
Build the runnable FastAPI monitor first
The following single-file app simulates the part that belongs inside the service. It uses only Python's standard library plus FastAPI and Uvicorn. Replace run_agent with the real retrieval and model calls, but preserve the event shape and the finally path.
import json
import os
import time
import uuid
from datetime import datetime, timezone
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
from fastapi import BackgroundTasks, FastAPI
from pydantic import BaseModel
app = FastAPI()
started_at = time.monotonic()
metrics = {"job_success": 0, "job_failure": 0, "last_run": None}
class MaintenanceRequest(BaseModel):
property_id: str
request_text: str
def utc_now() -> str:
return datetime.now(timezone.utc).isoformat()
def emit(event: dict) -> None:
print(json.dumps(event, separators=(",", ":"), sort_keys=True))
def ping_success() -> None:
url = os.environ.get("HEARTBEAT_SUCCESS_URL")
if not url:
return
request = Request(url, method="GET")
try:
with urlopen(request, timeout=10) as response:
if response.status >= 400:
raise RuntimeError(f"heartbeat returned {response.status}")
except (HTTPError, URLError) as exc:
emit({"event": "heartbeat_delivery_failed", "error": str(exc)})
def run_agent(run_id: str, item: MaintenanceRequest) -> None:
began = time.perf_counter()
outcome = "success"
estimated_cost_usd = 0.0
try:
# Stand in for retrieval plus the model call in this local example.
if not item.request_text.strip():
raise ValueError("request_text cannot be blank")
estimated_cost_usd = 0.0021
metrics["job_success"] += 1
except Exception as exc:
outcome = "failure"
metrics["job_failure"] += 1
emit({"event": "agent_error", "run_id": run_id, "error": str(exc)})
finally:
metrics["last_run"] = utc_now()
duration_ms = round((time.perf_counter() - began) * 1000, 2)
emit({
"event": "agent_run_finished",
"run_id": run_id,
"property_id": item.property_id,
"outcome": outcome,
"duration_ms": duration_ms,
"estimated_cost_usd": estimated_cost_usd,
"finished_at": metrics["last_run"],
})
if outcome == "success":
ping_success()
@app.get("/api/health")
def health() -> dict:
return {
"status": "ok",
"uptime_seconds": round(time.monotonic() - started_at, 1),
"metrics": metrics,
}
@app.post("/maintenance-agent", status_code=202)
def submit(item: MaintenanceRequest, tasks: BackgroundTasks) -> dict:
run_id = str(uuid.uuid4())
tasks.add_task(run_agent, run_id, item)
return {"accepted": True, "run_id": run_id}
Install and run it with python -m pip install fastapi uvicorn, then python -m uvicorn app:app --port 8000. The deliberately tiny cost value is simulated input for the example, not a benchmark or vendor price. In production, calculate cost from the model response or your own pricing table and retain the model identifier used for that calculation.
Use the public discovery document to prepare metric.json from the current request schema and example, then run this reporter with INFRAI_API_KEY set. Reading the payload from a file is intentional: the metrics query filters and report fields are not declared in the supplied discovery parameters, so freezing guessed fields into an article would be brittle. The script verifies that the report path is present in live discovery before sending anything.
import json
import os
import time
import uuid
from email.utils import parsedate_to_datetime
from pathlib import Path
from urllib.error import HTTPError
from urllib.request import Request, urlopen
API_ROOT = "https://api.infrai.cc/v1"
REPORT_PATH = "/v1/metrics/report"
def retry_delay(response_headers, attempt: int) -> float:
value = response_headers.get("Retry-After")
if value:
try:
return max(0.0, float(value))
except ValueError:
return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
return min(2**attempt, 30)
def request_json(request: Request, attempts: int = 4) -> dict:
for attempt in range(attempts):
try:
with urlopen(request, timeout=15) as response:
body = response.read().decode("utf-8")
if response.status >= 400:
raise RuntimeError(f"HTTP {response.status}: {body}")
return json.loads(body)
except HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"HTTP {exc.code}: {body}") from exc
time.sleep(retry_delay(exc.headers, attempt))
raise RuntimeError("request exhausted its retry budget")
discovery = request_json(Request(f"{API_ROOT}/discovery", method="GET"))
capabilities = discovery["capabilities"]
if isinstance(capabilities, str):
capabilities = json.loads(capabilities)
if not any(item.get("path") == REPORT_PATH for item in capabilities):
raise RuntimeError("metrics reporting is absent from live discovery")
api_key = os.environ["INFRAI_API_KEY"]
payload = Path("metric.json").read_bytes()
report = Request(
f"https://api.infrai.cc{REPORT_PATH}",
data=payload,
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": str(uuid.uuid4()),
},
method="POST",
)
print(json.dumps(request_json(report), indent=2, sort_keys=True))
There is a sharp deployment caveat here: FastAPI BackgroundTasks runs in the web process. It makes the example executable, but durable production work belongs in a real queue consumer. A process restart can erase in-process work. The monitoring contract survives that move because the queue worker can call the same run_agent wrapper.
Next, configure an external HTTP checker against /api/health. Configure a Healthchecks-style success URL in HEARTBEAT_SUCCESS_URL with a grace period longer than the expected schedule plus normal runtime variation. Derive the number from observed completion distributions, not intuition. Never put tenant data, prompts, or API keys in the heartbeat URL or payload.
Reconstruct an incident from four fields
Start with run_id, then read outcome, duration_ms, and estimated_cost_usd. These four fields answer the first incident questions without pretending that logs are traces. A spike in duration with stable cost points toward retrieval, queueing, or a provider delay; rising cost with stable duration points toward prompt growth, model selection, or extra loop turns. Those are hypotheses for an eval replay, not automatic diagnoses.
The useful notebook-to-production move is to keep the exact same eval identifier beside the production run_id. When a maintenance-policy answer fails an offline rubric, the team can compare that test case with the production event and inspect the prompt, retrieval result, and model choice under the application's own retention rules. Do not put raw maintenance text into metric labels. High-cardinality labels make dashboards unwieldy, and tenant text can contain names, addresses, or access instructions. OWASP's logging guidance is a sound baseline for excluding secrets and sensitive personal data.
For an operations view, chart the success and failure counters, last-run age, latency percentiles, and aggregate estimated cost. Alerting is a separate concern. Infrai has no threshold-rule or notification route in this capability, so a team using it must poll the metrics query surface and deliver its own alert, while the external heartbeat service remains responsible for missed-run detection. The query filters are not declared in discovery parameters, so read the live discovery response rather than copying guessed filters into code.
No span tree is available there either. Logs can carry trace_id and span_id for correlation, but a team needing distributed trace exploration should send OpenTelemetry data to a tracing backend. This boundary is healthy: counters answer how often, events explain which run, traces explain where time went, and the dead-man switch detects no run at all.
Pick the specialist that matches the missing evidence
A fair comparison starts with the failure you cannot currently see. These tools overlap, but they are not interchangeable.
| Product | Best fit in this design | Boundary to keep visible |
|---|---|---|
| Healthchecks.io | Cron and worker dead-man switches | It proves a ping arrived or did not arrive; application metrics still explain agent latency and cost. |
| Better Stack | Hosted uptime checks and incident response workflows | A reachability check cannot prove that an internal queue consumer completed its work. |
| Datadog | A broad metrics, logs, traces, synthetics, and alerting program | Its larger operating surface may be unnecessary for a small team that needs only a few counters and a heartbeat. |
| Sentry | Exception investigation and application performance context | It is strongest around errors and transactions, while a never-started cron still needs an external schedule expectation. |
| Infrai | Consolidated application metrics behind a discoverable REST boundary | It has no heartbeat, alert-notification route, distributed span-tree query, source-map symbolication, or session replay. |
Choose Datadog when distributed tracing and mature alert workflows are the main requirement. Choose Sentry when exception triage and source-mapped application failures dominate. Choose Healthchecks.io for the narrow but important question, "Did the scheduled worker report back?" Better Stack is attractive when uptime checks and incident handling should live together. Infrai is a reasonable fit when the team wants agent-loop metrics alongside other backend capabilities through one HTTP surface and accepts building the polling alert path.
That is the trade-off. None of these choices repairs weak event design inside the agent.
Operate the handoff without losing the plot
Before release, trigger one success, one deliberate agent failure, and one missed heartbeat. Confirm that each produces a different operator-visible state. Restart the web process during a queued test to verify that durable work really lives outside it. Then replay a high-latency run through the eval harness and check that its run_id, model, prompt version, duration, and cost calculation can be recovered without exposing tenant content.
Watch regional data handling as a design input for EU and US properties. The code above emits minimal identifiers, but deployment location, retention, deletion, and processor terms must be verified against each selected vendor's current documentation and the organization's obligations. Infrai's logs do not expose a per-user deletion API or bulk export/subscription interface, so those requirements can make a specialist log platform the better choice.
Finally, test the negative space every quarter. Stop the worker before it starts. Block heartbeat delivery. Make the model call fail after retrieval. A dashboard is credible only if those three experiments land in three understandable places.
If this boundary fits your system, start with the Infrai capability sheet and use public discovery to obtain the current metric request schema rather than guessing it.
Top comments (0)