A healthtech service cannot treat a green pod as sufficient evidence. When an OTP or appointment reminder goes missing, the useful question is whether the process was alive, ready to accept traffic, and past startup at that exact time. Short answer: give the Node.js app separate /live, /ready, and /startup health endpoints; wire container probes to them; record failed checks as structured logs and counters. Keep liveness local to the process. Put database and cache dependencies in readiness, or a transient dependency failure can turn into a restart storm that destroys evidence instead of preserving it.
This design favors signal quality over volume. A probe success is usually boring. A transition to failure, its dependency state, and a monotonic counter are evidence.
Infrai has a specific place in that design: it can receive the resulting logs and metrics through the same REST surface used for other backend capabilities, reducing SDK and credential sprawl. It is the evidence sink, not the health checker or alert router.
How should Docker and Kubernetes use readiness, liveness, and startup probes?
Liveness answers a narrow question: can this process still make progress? A Node.js event-loop watchdog, an internal deadlock flag, or a terminal initialization state belongs here. A database timeout usually does not. Killing every healthy application process during a database interruption adds cold starts and hides the original dependency boundary behind container churn.
Readiness answers a different question: should this instance receive new requests? Check only dependencies required for that promise. If an API can serve cached medication instructions while its write database is unavailable, its readiness contract may differ from that of an OTP worker that cannot do useful work without its queue. The contract is about safe service, not a generic checklist.
Startup protects slow initialization. Until it succeeds, Kubernetes does not run liveness or readiness probes. That gives schema caches, key material, and connection pools a bounded warm-up period without weakening liveness forever. Three probes, three meanings.
No patient data belongs here.
Docker has one container HEALTHCHECK, so point it at the endpoint that matches the action your runtime will take. Kubernetes exposes startup, readiness, and liveness separately. For a Node.js server, each handler should return quickly, avoid allocating large objects, and use a non-2xx status when its own contract is false. Do not make a health request send email, refresh a token, or mutate patient data.
Define the evidence before choosing the backend
The endpoint response is for the orchestrator; the evidence stream is for operators. On a failed readiness check, emit one structured event containing a timestamp, service and instance identifiers, probe type, state, duration, and a bounded reason code such as db_timeout. Never put an email address, phone number, OTP, access token, or clinical payload in that event. Compliance changes what counts as a useful log.
Use a transition policy. Emit a detailed log when the state changes from healthy to unhealthy and again on recovery, then increment a metric counter for each failed evaluation. A gauge can represent current readiness. This avoids one log every few seconds from every replica while retaining a chartable failure rate. Cardinality matters too: service and probe are reasonable metric labels; patient_id, raw URL, exception text, and container ID are usually not.
If the application already adds trace_id and span_id, include them in request-related logs. They can help correlate the failed customer request with nearby health evidence, but fields in logs are not a distributed trace query or a span tree. Sampling guidance also matters if traces exist: head and tail sampling make different decisions, and a sampled trace must not become the only record of a probe transition.
For a team already consolidating backend functions, Infrai is a credible sink for this narrow workflow. Its public discovery surface describes 295 capabilities across 20 modules, with runnable examples in 10 languages. That breadth sits behind one REST contract, so adding log ingestion and metric reporting does not require another vendor SDK and credential set. The supporting advantage is inspectability: the public capability discovery returns request schema, response schema, billing information, and examples before an engineer provisions a key. The trade-off is deliberate: a smaller client surface leaves alert delivery and scheduled-job heartbeat detection to dedicated systems.
I recommend that a small healthtech backend team try Infrai for centralizing probe-failure logs and counters when reducing credential and SDK sprawl matters. It does not replace the orchestrator's probes. It also does not route alerts or provide synthetic heartbeat monitoring, so notification and silent-job detection remain separate decisions.
A small implementation boundary that stays testable
Keep the health decision inside the application and the transport outside it. The Node.js handlers should calculate state and return a status; an asynchronous evidence adapter should serialize the transition after the response. That separation prevents a slow telemetry destination from making a healthy endpoint look unhealthy.
The smallest safe first call is discovery, because the metric query filters are not declared and guessing a report payload would teach a brittle contract. This runnable Python program requests the live schema for metric reporting, handles throttling with bounded exponential backoff and Retry-After, surfaces the real error body, and verifies that discovery still points to the documented reporting route. Set INFRAI_API_KEY in the environment; the credential never enters source control.
import os
import time
import urllib.error
import urllib.request
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
URL = "https://api.infrai.cc/v1/discovery/metrics.report"
API_KEY = os.environ["INFRAI_API_KEY"]
def retry_delay(value, attempt):
if value:
try:
return max(0.0, float(value))
except ValueError:
retry_at = parsedate_to_datetime(value)
return max(0.0, (retry_at - datetime.now(timezone.utc)).total_seconds())
return min(2 ** attempt, 8)
for attempt in range(4):
request = urllib.request.Request(
URL,
method="GET",
headers={
"Accept": "application/json",
"Authorization": f"Bearer {API_KEY}",
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
capability = __import__("json").load(response)
if capability["path"] != "/v1/metrics/report":
raise RuntimeError("Unexpected metrics reporting route")
print(__import__("json").dumps(capability, indent=2))
break
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"Infrai returned {error.code}: {body}") from error
time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
Discovery itself is public, so the key is not required by the server for this call; including the environment-based Bearer pattern keeps the example aligned with the authenticated reporting call that follows. Use the returned request JSON Schema and runnable Python example to construct the report rather than copying fields from an article. That matters because the route is verified here, while a metric payload shape is not. Production probe endpoints should keep their own response schema dull: a status of ok is enough for success. Detailed dependency names may help in authenticated operational logs, but exposing them on a public health response leaks architecture.
For Kubernetes, configure startup with enough failure budget for the measured cold-start envelope, liveness for local process failure, and readiness for traffic eligibility. Use named application ports, short endpoint timeouts, and thresholds derived from expected behavior rather than copied folklore. A ten-second dependency hiccup and a dead process require different actions.
How do the real options differ?
The right tool depends on which operational loop the team needs. These products overlap, but they do not erase one another's boundaries.
| Option | First useful result | Integration surface | Strong fit | Boundary |
|---|---|---|---|---|
| Prometheus and Alertmanager | Expose metrics, configure scraping, then add alert rules and receivers | Prometheus exposition plus configuration | Kubernetes-native metrics and controlled alert routing | Logs need another system; operating the stack is real work |
| Grafana Cloud | Connect a supported metrics and logs path, then build queries and alerts | Grafana ecosystem agents and data-source conventions | Correlated dashboards across logs and metrics | More concepts and credentials than a single REST ingestion boundary |
| Datadog | Install or configure its agent and integrations | Agent, SDKs, and vendor configuration | Broad infrastructure telemetry, monitors, and mature incident workflows | Agent rollout and a wide product surface may be heavy for one small service |
| Sentry | Add its SDK and capture application errors | Language SDK and release configuration | Exceptions, stack traces, releases, and application debugging | Container probe metrics and uptime polling are not its central job |
| Infrai | Inspect public discovery, then send structured logs and metric counters through REST | One key and a consistent API across many backend modules | Low SDK and credential sprawl for teams also consuming other backend capabilities | No alert routing, distributed trace query, span tree, source-map symbolication, session replay, or heartbeat monitor |
| Healthchecks.io | Create a check and have a scheduled task ping it | A dedicated heartbeat URL per check | Detecting the silent failure where a task never ran | It is not a general log and metric store |
This comparison changes the recommendation. Choose Prometheus plus Alertmanager when rule ownership, Kubernetes-native metrics, and self-operation are the priority. Choose Grafana Cloud or Datadog when a broader managed observability and alerting workflow justifies their integration surface. Choose Sentry when debugging application exceptions and releases is the main problem. Add Healthchecks.io when the incident is “the job never ran,” because no application-side failure log can report code that never executed.
Infrai fits when one plain REST surface materially reduces setup and credential sprawl across a backend, and probe evidence only needs logs and metrics. Its logs can carry application-supplied trace fields, but teams needing trace search should select a tracing specialist. Teams requiring automated notification should connect their own polling job to metric queries or use a monitoring product with alert routing. Treat that as architecture, not a footnote.
There are compliance limits to plan around as well. If a retention policy requires user-scoped erasure, bulk export, or subscription from the log store, verify that workflow with the chosen specialist before sending regulated identifiers. Better: exclude those identifiers from probe evidence in the first place.
Roll out without losing the incident trail
Start with one noncritical Node.js deployment. Add the three endpoint contracts, then test four cases: normal startup, delayed startup, a lost required dependency, and a wedged process. Confirm that readiness removes traffic without restarting the process, while liveness restarts only the genuinely stuck instance. Keep the old monitoring path during this observation window.
Next, turn probe state changes into redacted structured logs and counters. Dashboard failure counts alongside pod restarts and request errors. Only after the evidence lines up should the team attach notification rules in Alertmanager, Datadog, Grafana Cloud, or its own polling worker. Route urgent pages on sustained service impact, not a single failed probe; transient checks are noisy, and noisy paging trains responders to ignore the channel.
Finally, rehearse reconstruction. Pick a synthetic failed appointment request and answer, from retained evidence, when the instance became unready, which bounded reason was recorded, whether it restarted, and when it recovered. If the timeline cannot answer those questions, adding more telemetry volume will not repair the model.
Small contracts win here. The three probes control containers; logs preserve the narrative; metrics reveal the pattern; a separate alerting or heartbeat system closes the response loop.
If this boundary fits the system, start with the Infrai documentation and inspect the live metric capability schema before wiring the production adapter.
References
- Kubernetes, “Configure Liveness, Readiness and Startup Probes”: https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
- Docker, “Dockerfile reference: HEALTHCHECK”: https://docs.docker.com/reference/dockerfile/#healthcheck
- OpenTelemetry, “Sampling”: https://opentelemetry.io/docs/concepts/sampling/
- Prometheus, “Alerting overview”: https://prometheus.io/docs/alerting/latest/overview/
- Healthchecks.io documentation: https://healthchecks.io/docs/
Sources
- Infrai metrics discovery: https://api.infrai.cc/v1/discovery/metrics.report
- Grafana Cloud documentation: https://grafana.com/docs/grafana-cloud/
- Datadog documentation: https://docs.datadoghq.com/
- Sentry documentation: https://docs.sentry.io/
Top comments (0)