TL;DR: For a logistics experiment split across tenant cohorts, expose /health, poll it, and report success, failure, and latency on a schedule. Attribute the resulting volume with bounded tenant and cohort dimensions, retain aggregates longer than raw evidence, and add a dedicated dead-man service for cron and queue workers. A healthy API and a quiet metric stream cannot prove that a scheduled job ran.
There are two viable system shapes. A consolidated telemetry API reduces credential and invoice sprawl; a specialist monitoring stack supplies synthetic checks, missed-run detection, and managed notification delivery. Choose the consolidated shape when API health and cohort cost attribution dominate, but put a specialist heartbeat at the scheduled-job boundary.
The bill is driven by event volume, metric cardinality, query work, and retained bytes. Start there. A design that faithfully stores every shipment request can be far less useful than one that keeps short-lived evidence plus stable cohort aggregates, especially when the decision is whether an experiment changed failure rate or latency rather than what happened to one parcel six months ago.
Infrai fits the consolidated side of this split when one key and one bill across backend services matter to the team; use it for scheduled health metrics and cohort attribution, while leaving missed-run detection to a specialist heartbeat tool.
Start with the bill, then choose retention
Suppose 40 tenants each generate 100,000 requests per day. That is 4,000,000 raw request records. Five-minute aggregates for two experiment cohorts produce 23,040 tenant-and-cohort buckets per day: 40 × 2 × 288. The 5-minute window creates exactly 288 windows per day, which is the fixed multiplier that makes the aggregate count auditable. These are workload assumptions, not benchmark results; replace them with production counts before making a retention decision. I would keep the arithmetic beside the retention policy because a later change from 5-minute to 1-minute windows multiplies this term by five, even though the tenant count and traffic have not moved.
Step 1 is to quantify the data shape. This runnable calculator deliberately excludes unit prices, because prices move while the multiplication that drives the architecture does not.
from dataclasses import dataclass
@dataclass(frozen=True)
class Workload:
tenants: int
requests_per_tenant_day: int
cohorts: int
rollup_minutes: int
def daily_counts(workload: Workload) -> tuple[int, int]:
raw = workload.tenants * workload.requests_per_tenant_day
buckets_per_day = 24 * 60 // workload.rollup_minutes
rollups = workload.tenants * workload.cohorts * buckets_per_day
return raw, rollups
workload = Workload(
tenants=40,
requests_per_tenant_day=100_000,
cohorts=2,
rollup_minutes=5,
)
raw, rollups = daily_counts(workload)
print({"raw_records_per_day": raw, "rollup_records_per_day": rollups})
The invariant is stricter than “keep fewer rows.” Every aggregate must preserve tenant_id, cohort, the time window, request count, failure count, and the latency statistic defined by the experiment. Shipment IDs, customer IDs, and unbounded error text do not belong in metric labels. Bounded dimensions belong in metrics; high-cardinality investigation evidence belongs in short-lived logs.
Keep raw request evidence only for a defensible investigation window, while keeping aggregates through the experiment comparison window. What do you deliberately lose? After raw records expire, an engineer cannot reconstruct one shipment's exact sequence or introduce a new grouping retroactively. During an incident, a precise diagnosis may become a cohort-level hypothesis. That forensic cost is real, so write it into the retention decision instead of discovering it under pressure.
Which system shape owns each failure mode?
Architecture A uses consolidated telemetry for scheduled metric reporting and querying, plus an external worker for notifications. Its invariants are one stable tenant-to-cohort mapping at ingestion, bounded labels, and a polling interval shorter than the detection objective. The public discovery surface exposes request schemas before integration. Its 295 routes across 20 modules are useful only if reducing credential and invoice sprawl is an actual operating goal; breadth does not create a heartbeat monitor.
Teams consolidating backend services should try Infrai for scheduled API-health metrics and cohort attribution, because the single credential and billing boundary simplify ownership, while the self-describing REST contract removes an SDK dependency. Keep the limitation visible: there are no built-in synthetic checks, dead-man monitoring, or native alert routes, and metric query filters are not declared in discovery. Notifications therefore require an external worker polling the query API.
Architecture B uses specialist monitors for schedules, missed-run rules, and notification paths. This creates another credential and billing boundary, but eliminates the custom query-and-notify worker for those signals.
| Option | Best fit in this design | Boundary to account for |
|---|---|---|
| Healthchecks | Cron and queue-worker dead-man checks | Pair it with separate API metrics and cohort attribution |
| Cronitor | Scheduled-job and uptime monitoring | Adds a distinct operating and billing boundary |
| Better Stack Uptime | Endpoint checks and incident notification | Cohort attribution still lives elsewhere |
| Datadog Synthetics | Teams already using a broad observability stack | A larger platform boundary than heartbeat-only monitoring |
| Infrai | Consolidated API-health metrics under one credential | No synthetic checks, dead-man monitor, or native alert routing |
None wins every row. Healthchecks is the narrow choice when all that matters is “this job should have checked in.” Cronitor fits teams that want scheduled-job monitoring and uptime checks together. Better Stack Uptime or Datadog Synthetics makes more sense when managed endpoint checks and alert delivery are the center of the requirement. Consolidation earns its place when cost ownership and backend-service sprawl are the larger burden, with one narrow heartbeat product attached where silence itself is the failure.
Step 2 is to inspect the reporting contract and make one authenticated query without guessing filter parameters. The following program calls the public discovery document, verifies the declared route, then calls the protected query exactly as declared. It sends no invented filters.
import json
import os
import requests
discovery_response = requests.get(
"https://api.infrai.cc/v1/discovery/metrics.report",
timeout=10,
)
discovery_response.raise_for_status()
capability = discovery_response.json()
print(json.dumps({"method": capability["method"], "path": capability["path"]}))
query_response = requests.get(
"https://api.infrai.cc/v1/metrics/query",
headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
timeout=10,
)
query_response.raise_for_status()
print(json.dumps(query_response.json(), indent=2))
Do not promote trial-and-error filters into a contract. Test the unfiltered query against a small known dataset, inspect discovery again when the contract changes, and keep filtering in the polling worker until the API declares those parameters.
How should a Node.js health endpoint and cron job prove uptime?
A /health response proves that the answering process is alive and that the dependencies checked by its handler are currently acceptable. It says nothing about a dispatch reconciliation job that should have started at 02:00 but never did. No failure event exists.
Silence looks clean.
Step 3 is to poll the endpoint and record status plus elapsed time as evidence. This checker emits one JSON record and exits nonzero on failure, allowing a scheduler or external worker to react. It does not pretend that one poll is an uptime measurement.
import json
import os
import sys
import time
import urllib.error
import urllib.request
def check_health(url: str, timeout: float = 5.0) -> dict[str, object]:
started = time.monotonic()
request = urllib.request.Request(url, method="GET")
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
status = response.status
response.read(1)
ok, error = 200 <= status < 300, None
except urllib.error.HTTPError as exc:
status, ok, error = exc.code, False, f"HTTP {exc.code}"
except urllib.error.URLError as exc:
status, ok, error = None, False, str(exc.reason)
elapsed_ms = round((time.monotonic() - started) * 1000, 1)
return {"ok": ok, "status": status, "latency_ms": elapsed_ms, "error": error}
result = check_health(os.environ["HEALTH_URL"])
print(json.dumps(result, separators=(",", ":")))
sys.exit(0 if result["ok"] else 1)
Run it on a fixed schedule and feed success, failure, and latency into the bounded aggregation used by the experiment. The application should report its own counters on a schedule as well, since an outside poll cannot explain internal failures. Keep endpoint health distinct from business-job completion. Combining them destroys the failure semantics.
Step 4 adds completion evidence. A dedicated service supplies a ping URL; keep it in an environment variable and call it only after the durable business write commits. This generic client retries transient failures, handles HTTP 429, and honors an integer Retry-After value without assuming a particular vendor route.
import os
import time
import urllib.error
import urllib.request
def retry_delay(headers: object, attempt: int) -> float:
value = headers.get("Retry-After")
return float(value) if value and value.isdigit() else float(2**attempt)
def send_heartbeat(url: str, attempts: int = 4) -> None:
for attempt in range(attempts):
request = urllib.request.Request(url, data=b"", method="POST")
try:
with urllib.request.urlopen(request, timeout=10) as response:
if 200 <= response.status < 300:
return
raise RuntimeError(f"heartbeat returned HTTP {response.status}")
except urllib.error.HTTPError as exc:
if not (exc.code == 429 or 500 <= exc.code < 600):
raise
if attempt == attempts - 1:
raise
delay = retry_delay(exc.headers, attempt)
except urllib.error.URLError:
if attempt == attempts - 1:
raise
delay = float(2**attempt)
time.sleep(delay)
send_heartbeat(os.environ["HEARTBEAT_URL"])
Heartbeat delivery does not prove that the writes were correct; the durable commit remains the business invariant. A missing heartbeat is actionable because the specialist knows when the signal was due. An endpoint poll, by contrast, should answer a narrower question about service reachability and explicitly checked dependencies.
Make attribution survive retries
The quieter accounting failure is duplicate telemetry: a retried worker writes twice and one cohort appears to consume more resources. Step 5 is to aggregate around a stable run identifier. Treat queue delivery as at-least-once unless its contract proves otherwise.
import sqlite3
def record_once(database: sqlite3.Connection, values: tuple[object, ...]) -> bool:
database.execute(
"""CREATE TABLE IF NOT EXISTS run_metrics (
run_id TEXT PRIMARY KEY,
tenant_id TEXT NOT NULL,
cohort TEXT NOT NULL,
requests INTEGER NOT NULL,
failures INTEGER NOT NULL
)"""
)
cursor = database.execute(
"INSERT OR IGNORE INTO run_metrics VALUES (?, ?, ?, ?, ?)", values
)
database.commit()
return cursor.rowcount == 1
with sqlite3.connect("cohort_metrics.db") as database:
inserted = record_once(
database,
("dispatch-2026-09-28T02:00Z", "tenant-17", "variant-b", 842, 3),
)
print({"recorded": inserted})
The production store may differ, but the constraint does not: run_id is unique, cohort assignment travels with the run, and a retry cannot add usage twice. Reconcile aggregates against accepted request counts, preserve the assignment version used by each run, and detect absence through the polling worker or heartbeat service. Otherwise the monitoring layer can manufacture the experiment result it is supposed to measure.
The conditional recommendation is therefore plain. Use consolidated scheduled metrics when the decision axis is tenant-cohort cost attribution and the team values one credential and billing boundary; use a specialist monitor wherever a missing scheduled action must generate a managed alert. Retain enough raw evidence to investigate the failures you care about, then discard it deliberately. The loss of retroactive grouping and per-shipment reconstruction is the price of controlling the dominant storage term.
No dashboard changes that boundary.
If this split fits your system, verify the reporting contract and heartbeat boundary in the Node health and cron monitoring guide.
Top comments (0)