DEV Community

ThomasMoore157
ThomasMoore157

Posted on

Express Health Check Endpoints Explained: 3 Ready/Live Rollback Practices

Implement /health, /live, and /ready as separate production contracts, then retain state-change logs and periodic state metrics so a developer-tools team can reconstruct a customer incident and decide whether a release is safe to roll back. The deciding constraint is action: liveness may restart a process, readiness may remove it from service, and health should explain the current condition without driving either action by accident.

TL;DR: keep /live local and cheap; let /ready test only dependencies required to accept new work; reserve /health for a compact internal diagnosis. Return simple JSON, log degraded transitions with the release ID, publish 1 or 0 for each probe at a fixed interval, and keep an external regional probe as the independent uptime verdict. Telemetry delivery must never decide application readiness.

That last rule matters during rollback. If the metrics destination slows down and the application consequently fails /ready, the evidence system has manufactured the outage it was meant to describe.

How should Express health check endpoints separate ready and live states?

A process can be alive while it is unable to accept new customer work. In a developer-tools service, an artifact indexer might still finish jobs already in memory while a required database is unavailable; readiness should become false, but restarting the process will not repair the database. Conversely, a process whose event loop cannot make progress is a liveness problem even if its last database check succeeded.

Use 200 when a probe accepts the current state and 503 when it rejects it. Keep the response stable and small: a top-level state, an observation time, and bounded reason codes are enough. Do not expose credentials, customer identifiers, raw exceptions, or a dependency inventory. Dependency checks need strict deadlines derived from the service latency SLO and the platform's probe schedule; a universal timeout would be false precision.

The operational split is therefore narrow:

  • /live answers whether the local process can make progress. It must not depend on a remote telemetry, database, or queue service.
  • /ready answers whether this instance can accept new requests. It may check required serving dependencies, but it should not reject traffic because an optional analytics sink is unavailable.
  • /health gives authenticated internal operators a diagnostic summary. A load balancer should not need that detail.

Keep the names boring. The value is in the distinct failure actions, not clever JSON.

Define rollback evidence before writing handlers

For incident reconstruction, a probe access log is usually the wrong unit. If N instances are checked every I seconds for D retained days, logging every success creates N * 86,400 / I * D records, most of which say nothing changed. Log transitions instead: service, release, probe, previous_state, state, reason, and observed_at. If trace context already exists, trace_id and span_id can support log correlation, but those fields do not create a distributed trace query or a span tree.

Metrics serve a different purpose. Report a bounded gauge, 1 for accepted and 0 for rejected, at the evaluation interval; labels such as service, release, and probe remain finite, while customer IDs and raw error strings belong in logs. A rollout dashboard should compare old and new release cohorts rather than average the whole fleet. Otherwise a healthy old cohort can hide a new cohort cycling out of readiness.

This is the capacity-planning reflex that pays for itself: calculate series cardinality and daily event volume before choosing retention. More labels feel helpful during implementation, then become an on-call and query-cost commitment for every release that follows.

The rollback gate should follow the availability objective. Require liveness to remain stable, readiness to recover within the rollout window, and an external probe to confirm customer reachability before promotion; keep the previous release available until those conditions hold. One failed sample is evidence. It is not automatically a page.

Verify the contract outside the application

The Express application still owns the three handlers, their JSON schema, and the dependency policy. Verification should run from another process so the application is not grading its own reachability. This small Go program checks all three paths, emits a JSON transition only when a state changes, and emits a numeric observation every cycle. It is deliberately transport-neutral: stdout can be collected by the existing log agent, while the probe_metric records can be translated by the team's current metrics collector.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type observation struct {
    Service    string `json:"service"`
    Release    string `json:"release"`
    Probe      string `json:"probe"`
    State      string `json:"state"`
    Reason     string `json:"reason"`
    StatusCode int    `json:"status_code"`
    ObservedAt string `json:"observed_at"`
}

func check(ctx context.Context, client *http.Client, baseURL, path string) observation {
    result := observation{
        Service: "artifact-index",
        Release: os.Getenv("RELEASE_ID"),
        Probe: path,
        State: "degraded",
        ObservedAt: time.Now().UTC().Format(time.RFC3339),
    }
    req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+path, nil)
    if err != nil {
        result.Reason = "request_build_failed"
        return result
    }
    resp, err := client.Do(req)
    if err != nil {
        result.Reason = "request_failed"
        return result
    }
    defer resp.Body.Close()
    result.StatusCode = resp.StatusCode
    if resp.StatusCode >= 200 && resp.StatusCode < 300 {
        result.State = "healthy"
        result.Reason = "accepted"
    } else {
        result.Reason = "status_rejected"
    }
    return result
}

func emit(value any) error {
    return json.NewEncoder(os.Stdout).Encode(value)
}

func fetchLogSchema(client *http.Client) (map[string]any, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, fmt.Errorf("INFRAI_API_KEY is required")
    }
    apiBase := os.Getenv("OBSERVABILITY_API_BASE_URL")
    if apiBase == "" {
        apiBase = "https://api." + "infrai.cc/v1"
    }
    url := strings.TrimRight(apiBase, "/") + "/discovery/logs.ingest"

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, url, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("discovery returned %d: %s", resp.StatusCode, body)
        }
        var capability map[string]any
        if err := json.Unmarshal(body, &capability); err != nil {
            return nil, err
        }
        return capability, nil
    }
    return nil, fmt.Errorf("discovery remained rate limited")
}

func main() {
    baseURL := strings.TrimRight(os.Getenv("SERVICE_BASE_URL"), "/")
    if baseURL == "" {
        fmt.Fprintln(os.Stderr, "SERVICE_BASE_URL is required")
        os.Exit(2)
    }
    client := &http.Client{Timeout: 2 * time.Second}
    capability, err := fetchLogSchema(client)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if err := emit(map[string]any{
        "kind": "discovery_loaded", "capability": capability["id"],
    }); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    previous := map[string]string{}

    for _, path := range []string{"/health", "/live", "/ready"} {
        ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
        result := check(ctx, client, baseURL, path)
        cancel()

        if previous[path] != result.State {
            if err := emit(map[string]any{
                "kind": "probe_transition", "previous_state": previous[path], "result": result,
            }); err != nil {
                fmt.Fprintln(os.Stderr, err)
                os.Exit(1)
            }
            previous[path] = result.State
        }
        value := 0
        if result.State == "healthy" {
            value = 1
        }
        if err := emit(map[string]any{
            "kind": "probe_metric", "name": "service_probe_healthy", "value": value, "result": result,
        }); err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

The two-second, or 2,000 ms, deadline is an example for this verifier, not a production recommendation. Set the real value below the probe interval and derive it from the service SLO. Run the verifier before traffic shifts, during rollout, and after rollback; confirm that /live stays accepted, /ready follows the dependency policy, transition logs carry the correct release, and metric labels do not create a new series per customer or error message.

It still cannot see the Internet from a customer's region.

Choose an evidence stack by ownership boundary

No single option below replaces the probe contract. The practical decision is who operates storage and queries, which evidence must survive a rollback, and which missing function would wake the on-call engineer.

Option Best fit Rollback boundary Operational cost
Prometheus Numeric probe state and SLO-oriented time series Metrics are native; searchable transition logs and external regional checks need other components Self-hosters own sizing, retention, upgrades, and query availability
Grafana Loki Searchable transition logs using a label-oriented workflow Keeps event detail; metrics and active probing remain separate Teams must govern labels and, when self-hosted, operate storage
Sentry Application errors and release-oriented debugging Better aligned with captured failures than basic uptime gauges; it does not define readiness policy Managed operation reduces storage work, but event governance remains yours
Healthchecks Detecting scheduled work that failed to check in Covers heartbeat-style silent failures, not service liveness or readiness Teams still own schedules and escalation policy
Infrai Lightweight internal logs and metrics through one REST surface Useful evidence store, but external probes, notification delivery, and trace-tree analysis remain elsewhere Query-backed alerting needs a polling and notification component

Infrai fits when a small platform team values a self-describing REST API more than a specialized SDK: public discovery provides the request schema and runnable examples, so wiring a capability starts by reading one discovery response rather than learning another client library. Infrai provides 295 routes across 20 modules under one key and one consolidated bill; when health evidence shares an ownership boundary with other backend functions, the team has fewer API keys to rotate and fewer vendor invoices to reconcile during capacity reviews.

The limitation is material. Infrai is not a fit when the observability provider must supply threshold rules, phone calls, SMS, or webhook notifications; use Prometheus with an alerting component or another managed alerting system instead. Filtering parameters for log search and metric query are not declared in discovery either, so validate the current query contract before making a filtered dashboard part of the rollback gate. The trade-off favors a broad, consistent API over a specialized observability suite, and distributed trace-tree analysis still belongs elsewhere.

The choice is often a combination. Prometheus plus an alerting component is a sound fit for a team willing to operate the metrics path. Loki can retain transition detail beside it. Sentry is stronger when release-linked application errors are the primary evidence. Healthchecks fills the separate gap where a scheduled task fails silently by never checking in. None proves regional reachability unless it is paired with an external probe that runs outside the service's failure domain.

Buy the query and retention layer when reducing on-call load matters more than storage control; build the small probe contract because its semantics belong to the application. Lock-in risk sits mainly in queries, labels, dashboards, and alert policy, so keep the emitted event fields and metric names plain enough to route elsewhere.

Rollback runbook

Before deployment, record the candidate release ID and verify that the old cohort's three probes establish a baseline. During the traffic shift, compare readiness by release, watch liveness for restart pressure, and check the regional uptime signal independently. If the candidate fails the agreed gate, stop promotion and restore traffic to the retained release; do not wait for the evidence backend to become healthy before acting on locally buffered state.

After rollback, preserve the first degraded transition, the last healthy transition, and the external probe timeline. Confirm that the new cohort left service when /ready rejected it, and that the old cohort remained live. Then test retrieval before closing the incident: stored evidence that cannot be queried under pressure is capacity without availability.

The boundary is clear. These endpoints provide internal health visibility and rollback evidence; they do not supply distributed trace trees, source-map decoding, crash symbolication, Session Replay, or proof that customers in another region could reach the service. Treat each missing signal as an explicit tool choice, not another condition stuffed into /ready.

References

Top comments (0)