A page saying "checkout API unavailable" is already late. The on-call needs to see which pods stopped accepting traffic, whether they were restarted, and what changed immediately before the customer-visible failure. For a containerized Node.js service, the practical answer is to expose separate liveness, readiness, and startup health endpoints; let Kubernetes or Docker call the appropriate endpoint; and record every failed check in both logs and metrics.
Short answer: keep liveness restricted to process health, put database and cache dependencies in readiness, and give slow startup its own budget. That split preserves evidence and makes rollback safer because a dependency outage drains traffic instead of creating a restart storm. Logs explain an individual failure, while a counter shows whether failures are becoming a trend. Neither signal sends a page by itself, so use a polling job or an uptime service for notification delivery.
1. Reconstruct the failure from the page backward
Start at the action, not at the dashboard. A useful page identifies the service and environment, reports that readiness failures crossed an SLO-derived threshold, and links to the relevant logs and deployment record. A weak page says only that an endpoint returned a non-200 response. That forces the responder to reconstruct scope while customers wait.
Imagine the checkout service has ten replicas. A new release leaves four replicas alive but unable to reach their cache. Their readiness endpoint should fail, removing those replicas from service, while liveness remains successful and preserves the processes for inspection. The first investigation then asks whether remaining ready capacity can carry the load and whether rollback is lower risk than waiting for the dependency. Restarting those four processes adds churn without repairing the cache.
This is the capacity-planning reflex that health-check tutorials often omit: a readiness threshold is also a capacity threshold. If losing four of ten replicas pushes the remaining six beyond the latency SLO, the alert must fire before the customer error-rate objective is exhausted. The exact threshold belongs to your traffic and error-budget data; a copied percentage is guesswork.
Fast diagnosis needs one event per state transition, not one event per probe. Log the endpoint name, previous and current state, reason, pod or container identity, release identifier, and timestamp. If the application already creates trace_id and span_id fields, include them so request logs can be correlated. Those fields do not create a distributed trace query or a span tree.
2. How do Docker and Kubernetes readiness, liveness, and startup probes differ?
The three checks answer different operational questions. Liveness asks, "Can this process still make progress?" Readiness asks, "Should this instance receive new traffic now?" Startup asks, "Has initialization had enough time to finish before liveness judgments begin?" Collapsing them into one deep dependency check turns an ordinary database interruption into repeated process replacement.
Keep the liveness handler local and cheap. It can verify that the event loop is responsive and that the process has not entered an unrecoverable internal state; it should not fail merely because PostgreSQL, Redis, or a third-party API is unavailable. The readiness handler may test dependencies required to serve the next request, but each dependency needs a tight timeout and a reason code that avoids leaking credentials or customer data. The startup probe should cover known initialization work and then get out of the way.
For Kubernetes, configure startupProbe, readinessProbe, and livenessProbe against three distinct HTTP paths on the same application port. Set failureThreshold and periodSeconds from measured startup and recovery behavior, then verify that the total startup allowance covers a slow but valid deployment. Docker's HEALTHCHECK can call the readiness-style endpoint when Docker itself is responsible for container health, although it does not replace Kubernetes traffic routing.
One warning matters more than another page of configuration: do not make readiness depend on every downstream system. If a product-catalog timeout still permits checkout to serve a degraded response, marking the whole pod unready throws away useful capacity. Probe only dependencies whose failure makes this instance unable to perform its required job. For example, suppose 4 of 10 checkout replicas lose Redis while the other 6 sit at 70% CPU during the normal peak; draining the affected replicas may be correct, but the remaining capacity is now the constraint, so the rollout controller needs to stop before it removes another healthy replica. A shallow liveness response keeps the failed instances available for log inspection, a readiness failure protects new requests, and the deployment's capacity guard protects the SLO. Those are three separate controls, and folding them into one endpoint makes rollback behavior much harder to predict.
Keep them separate.
3. Turn a state transition into evidence
A probe runs frequently. Logging every success creates volume, hides the transition that matters, and couples observability cost to probe frequency. Emit a structured log when a check changes state, increment a failure counter on each failed evaluation, and expose a current-state gauge for dashboards. Keep the metric labels bounded: service, environment, check, and reason can be reasonable; request IDs, pod UIDs, and raw error text are not.
The following Go program is intentionally outside the Node.js process. It checks the readiness URL, logs a state transition, and reports a caller-supplied metric document through the observability API when the check fails. The metric JSON must match the current public discovery schema; keeping it in configuration avoids freezing an undeclared field shape into application code. The example uses one write route, an environment variable for the key, an explicit method, bounded exponential backoff for HTTP 429, and error-body handling.
package main
import (
"bytes"
"context"
"fmt"
"io"
"log"
"net/http"
"os"
"time"
)
func main() {
target := os.Getenv("READINESS_URL")
baseURL := os.Getenv("INFRAI_BASE_URL")
apiKey := os.Getenv("INFRAI_API_KEY")
metricJSON := os.Getenv("METRIC_REPORT_JSON")
if target == "" || baseURL == "" || apiKey == "" || metricJSON == "" {
log.Fatal("READINESS_URL, INFRAI_BASE_URL, INFRAI_API_KEY, and METRIC_REPORT_JSON are required")
}
client := &http.Client{Timeout: 2 * time.Second}
wasReady := true
ticker := time.NewTicker(10 * time.Second)
defer ticker.Stop()
for range ticker.C {
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
req, err := http.NewRequestWithContext(ctx, http.MethodGet, target, nil)
ready := err == nil
if ready {
resp, requestErr := client.Do(req)
ready = requestErr == nil && resp.StatusCode >= 200 && resp.StatusCode < 300
if resp != nil {
resp.Body.Close()
}
}
cancel()
if ready != wasReady {
log.Printf("service=checkout check=readiness ready=%t", ready)
wasReady = ready
}
if !ready {
if err := reportMetric(client, baseURL, apiKey, []byte(metricJSON)); err != nil {
log.Printf("metric_report_error=%q", err)
}
}
}
}
func reportMetric(client *http.Client, baseURL, apiKey string, body []byte) error {
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequest(http.MethodPost, baseURL+"/metrics/report", bytes.NewReader(body))
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
req.Header.Set("Content-Type", "application/json")
resp, err := client.Do(req)
if err != nil {
return err
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("metric report status %d: %s", resp.StatusCode, responseBody)
}
delay := time.Duration(1<<attempt) * time.Second
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if parsed, err := time.ParseDuration(retryAfter + "s"); err == nil {
delay = parsed
}
}
time.Sleep(delay)
}
return fmt.Errorf("metric report remained rate limited after 3 attempts")
}
Run this as one highly available component, not one copy beside every application pod, or duplicate reports become part of the incident. It still does not page anyone. Notification delivery needs a separate polling job or uptime service, plus deduplication and a durable record of attempts.
That boundary is easy to miss.
4. Move telemetry without moving the health contract
Keep the application contract stable: structured transition logs, a failure counter, a current-state gauge, and consistent service and environment attributes. Then the backend behind that contract can move without rewriting health handlers. OpenTelemetry helps at the instrumentation boundary, but sampling deserves care: health failures are sparse and operationally important, so a trace sampling policy should not be mistaken for log retention.
The buy-versus-build decision is mostly about on-call ownership and rollback friction, not feature count.
| Option | Operating model | Useful fit | Boundary to plan around |
|---|---|---|---|
| Prometheus plus Alertmanager | Self-hosted or managed components | Teams wanting direct control of metric collection and alert routing | The team owns sizing, retention choices, upgrades, and alert-path reliability when self-hosted |
| Grafana Cloud | Managed metrics, logs, dashboards, and alerting | Teams wanting a managed observability stack with Grafana workflows | Validate ingestion, retention, and portability against the intended rollback plan |
| Datadog | Managed monitoring with integrated infrastructure views and alerting | Teams prioritizing a broad hosted operations workflow | Agent rollout and vendor-specific monitors increase migration work |
| Healthchecks | Dead-man's-switch monitoring for jobs and heartbeats | Detecting "the task never ran" failures | It complements service probes; it is not a general logs-and-metrics backend |
| Infrai | One REST contract across backend capabilities | Teams that value swapping the provider behind a capability while keeping application code stable | It stores and queries logs and metrics, but alert routing, synthetic checks, trace queries, and span trees require separate components |
Infrai's practical advantage here is one REST API and one API key across backend capabilities, while its public discovery surface provides schemas and runnable examples for contract validation during a provider change. The limitation is equally concrete: its logs can carry application-supplied trace and span IDs for correlation, but they do not provide distributed-trace navigation, alert routing, or synthetic checks. Query filters are also not declared for its log and metric query operations, so it is not a fit when the rollback runbook depends on those filters or on an integrated pager; choose Datadog or Grafana Cloud for a hosted alerting workflow, or Prometheus with Alertmanager when the team wants to own that path.
Choose the smallest arrangement that meets the recovery objective. A platform team already operating Prometheus may reasonably keep it. A small team trying to reduce on-call surface may prefer Datadog or Grafana Cloud. A scheduled build or backup needs a dead-man's-switch service such as Healthchecks because a silent non-run cannot fail an application probe. My decision rule is blunt: if one backend failure can suppress both the signal and the page, the design has not earned the rollback window.
This is a trade-off, not a leaderboard.
5. Rehearse the failed rollback before release
Before release, test four states: healthy startup, slow but valid startup, loss of a required dependency, and recovery. Confirm that startup delay does not trigger liveness, dependency loss removes the instance from traffic without restarting it, transitions appear once in logs, and the failure counter moves as expected. Then deploy a canary and verify that enough ready replicas remain to satisfy the latency and availability objectives during rollback.
A health check is a control-plane input. Treat changing its semantics like changing a load balancer rule. Version the behavior with the application, review it, and make rollback restore the previous probe contract as well as the previous binary.
False positives have a concrete cost: they wake people, drain healthy capacity, and train responders to distrust pages. False negatives spend the error budget silently. Reduce both by alerting on sustained state and customer impact, keeping liveness shallow, and setting startup allowances from observed initialization rather than optimism.
No single threshold is safe everywhere. Revisit it after traffic shape, dependency behavior, or replica count changes; the arithmetic that protected ten replicas may be wrong at three. The page should arrive early enough to act, but every earlier page consumes attention. That is the trade.
Further reading
- Kubernetes: Configure Liveness, Readiness and Startup Probes: https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
- Dockerfile HEALTHCHECK reference: https://docs.docker.com/reference/dockerfile/#healthcheck
- Prometheus alerting overview: https://prometheus.io/docs/alerting/latest/overview/
- Grafana Cloud alerting documentation: https://grafana.com/docs/grafana-cloud/alerting-and-irm/alerting/
- Datadog monitors documentation: https://docs.datadoghq.com/monitors/
- Healthchecks documentation: https://healthchecks.io/docs/
- OpenTelemetry sampling concepts: https://opentelemetry.io/docs/concepts/sampling/
Top comments (0)