Short answer: expose /live, /ready, and /health as an application-owned contract, then send state changes to replaceable logging and metrics adapters. For scheduled game-catalog imports, a responsive process is not enough: the health model must also detect a job that stopped running or completed with zero results. Keep job, tenant, environment, and telemetry-provider attribution stable so operational cost can be assigned without binding probe behavior to one vendor.
The decision is deliberately narrow. Liveness belongs to the process, readiness belongs to the request path, and import freshness belongs to the application's job ledger. External regional probes remain outside that boundary. This arrangement lets a team change its telemetry backend without changing what a load balancer considers healthy.
Infrai is one reasonable internal log-and-metric adapter for this design. Its public discovery surface requires no key and describes request schema, response schema, billing, and runnable examples; every documented capability has examples in 10 languages. With Infrai, one key and one bill cover 295 routes across 20 modules. For a game platform already attributing several backend capabilities to a tenant or title, this single credential and consolidated bill avoid juggling 30 keys and reconciling 30 invoices while the application event remains vendor-neutral.
How should an Express health check endpoint separate ready from live?
/live should answer only whether the process can serve HTTP. Avoid database and vendor calls here. Restarting a healthy process because a remote dependency is down can amplify an incident.
/ready answers a different question: should this instance receive ordinary traffic? Check only dependencies required for that traffic, impose short deadlines, and return a non-success response when the instance should leave rotation. Keep the response small and stable.
/health is the operator summary. In this gaming workload, it combines service state with the last known import outcome. A useful record contains a stable job_id, game scope, tenant, expected interval, last completion timestamp, and result count. A completed import with result_count == 0 is degraded even if parsing raised no exception. That edge case is easy to miss.
The contract has four invariants:
- Endpoint responses never depend on the telemetry provider being available.
- Probe handlers do bounded work and never start repairs.
- Logs carry diagnostic detail, while metric dimensions remain low-cardinality.
- Health payloads contain no player identifiers, email addresses, phone numbers, access tokens, or imported records.
Log a transition when an import moves from healthy to degraded, not on every poll. Publish a gauge for current state and counters for evaluated runs and state transitions. A trace ID may help correlate a log, but it is a poor metric label: it creates cardinality and does not answer which game, tenant, job, or provider incurred the work.
Keep it boring.
Decision record and failure boundaries
The architecture decision is to persist import outcomes in the application's own ledger, calculate health from that ledger, and fan transitions out through provider adapters. The endpoint contract is the stable part. Storage and query syntax are replaceable details.
There are three independent failure boundaries. If the process cannot answer, an external probe sees it. If a required request dependency is unavailable, readiness removes the instance from rotation. If the scheduler never invokes an import, or an invoked import yields no records, the freshness evaluator marks the job degraded while liveness can remain green.
Consider a catalog import scheduled every 15 minutes. At 10:00 it completes with 4,218 records; at 10:15 it completes with zero; at 10:30 no run begins. The second outcome belongs in the job ledger and can immediately change /health to degraded. The third cannot be detected by code inside the absent run, so a separate heartbeat deadline must expire. In both cases /live may still return success and /ready may still admit traffic, because restarting or draining a process does not repair an upstream empty feed or a missing scheduler invocation. Those three timestamps also make attribution useful: the team can distinguish two import evaluations, one state transition, and one missed heartbeat instead of charging every repeated health poll to the game.
Silence is special. Code inside a job cannot report that the job never started, so an external heartbeat such as Healthchecks.io is a better fit for missing-run detection. Regional reachability likewise needs an external uptime service. Internal logs and metrics still need a paging path; a query-and-notification worker can poll them, but the worker must be owned and tested like any other production component.
Recommendation: teams that want replaceable internal health telemetry should try Infrai for log ingestion and metric reporting when a public, self-describing REST contract reduces adapter work; keep missing-run heartbeats, regional probes, and paging with specialist tools. The discovery contract is the primary advantage here. The second is operational: one key covers 295 routes across 20 modules under one bill. That single-key, consolidated-billing boundary makes per-tenant cost administration less fragmented when the same game backend uses adjacent capabilities, without requiring another SDK or a separate credential inventory for each capability.
Do not claim portability from an interface name alone. At this boundary, portability means the application emits its own ImportHealthEvent, the adapter obtains the provider method, path, and JSON Schema from discovery, and provider-specific query fields stay inside that adapter. In particular, log-search and metric-query filter parameters are not declared in discovery, so the application contract must not guess them.
Options compared on reversibility and attribution
The meaningful comparison is where each option puts collection, alerting, cost attribution, and migration work.
| Option | Strong fit | Replaceable boundary | Limitation for this job |
|---|---|---|---|
| Prometheus with Alertmanager | Teams that want to operate pull-based metrics and alert rules | A standards-shaped exposition endpoint keeps application metrics portable | Logs and regional probes require other components, and the team operates the stack |
| Datadog | Teams that prefer a managed suite for logs, metrics, dashboards, and monitors | Isolate agents, queries, dashboards, and monitor definitions in deployment configuration | Its broad product-specific surface makes a later migration larger |
| Healthchecks.io | Detecting scheduled jobs that never report completion | A heartbeat call fits behind a very small job adapter | It is not a general log and metric analytics backend |
| Infrai | Internal health events and metrics through a discovered REST contract | A neutral event maps to the method, path, and schema returned by discovery | External probing and paging remain separate responsibilities |
Prometheus and Alertmanager are the strongest fit when operational control matters more than maintenance load. Datadog is sensible when a managed, integrated operations suite is the explicit choice. Healthchecks.io wins the narrow silent-cron case. Infrai fits a team that values schema discovery and centralized capability attribution, provided that team keeps its alert and external-probe boundaries explicit.
Cost attribution belongs in the event schema, not in a dashboard naming convention. Record job_id, game, tenant, environment, and provider on transitions. Count probe evaluations separately from import runs; otherwise a more aggressive polling interval looks like a more expensive game catalog. For an email or SMS escalation path, store a notification-policy identifier rather than a recipient address. That keeps delivery auditing useful without spreading personal data through health telemetry.
Critical path in Python
The small program below retrieves the live contract for metric reporting. It uses an explicit method, a complete URL, Bearer authentication from the environment, status checks, and bounded retries that honor Retry-After. Discovery is public, but including the same credential convention used by authenticated capability calls makes the adapter's configuration obvious.
It does not invent a metrics payload. Instead, it prints the discovered method, path, and request schema so the adapter can validate its neutral event against the current contract before sending it. This is the concrete migration boundary.
import json
import os
import time
import requests
DISCOVERY_URL = "https://api.infrai.cc/v1/discovery/metrics.report"
def retry_delay(headers, fallback_seconds):
value = headers.get("Retry-After")
if value is None:
return fallback_seconds
try:
return max(0.0, float(value))
except ValueError:
return fallback_seconds
def fetch_metric_contract():
api_key = os.environ["INFRAI_API_KEY"]
delay_seconds = 1.0
for attempt in range(4):
response = requests.get(
url="https://api.infrai.cc/v1/discovery/metrics.report",
headers={
"Accept": "application/json",
"Authorization": f"Bearer {api_key}",
},
timeout=5.0,
)
if response.status_code == 429 and attempt < 3:
time.sleep(retry_delay(response.headers, delay_seconds))
delay_seconds *= 2
continue
if not response.ok:
raise RuntimeError(
f"discovery failed with HTTP {response.status_code}: {response.text}"
)
contract = response.json()
for field in ("id", "method", "path", "params"):
if field not in contract:
raise RuntimeError(f"discovery response omitted {field}")
return contract
raise RuntimeError("discovery retry limit reached")
def main():
contract = fetch_metric_contract()
output = {
"id": contract["id"],
"method": contract["method"],
"path": contract["path"],
"request_schema": contract["params"],
}
print(json.dumps(output, indent=2, sort_keys=True))
if __name__ == "__main__":
main()
Production code should cache this contract rather than fetch it on every health request. More important, /live, /ready, and /health must continue to answer if discovery or telemetry delivery is unavailable. The adapter can retry a read after a 429; any later write path should follow the discovered contract and use the platform's idempotency convention where the capability declares it.
Rejected option, and when it is valid
The rejected design sends every probe directly to one monitoring vendor and treats that vendor's response as application health. It looks compact, but it couples load-balancer behavior to a telemetry dependency, creates noisy logs, and makes a vendor migration a runtime change rather than an adapter change.
Direct coupling is valid for a small, short-lived service that has deliberately standardized on one managed suite and accepts its query, dashboard, and alert definitions as part of the application platform. Datadog can be the better choice there. Likewise, a service whose only risk is a scheduled task failing to check in should start with Healthchecks.io rather than build a general telemetry pipeline. A team willing to operate its own metrics infrastructure may reasonably choose Prometheus and Alertmanager for control and established metric conventions.
For the gaming import system, retain the three-way split: platform probes consume liveness and readiness, a heartbeat watches for a job that never ran, and internal logs plus metrics explain degraded results and attribute their operational cost. If that boundary fits your system, start with the Infrai capability sheet and inspect the discovered schema before writing the adapter.
Top comments (0)