For a Node.js app running in Docker on Kubernetes, readiness, liveness, and startup probes should protect the process lifecycle; app metrics and logs should decide whether a tenant cohort experiment is safe. The page says the experiment cohort is unhealthy, yet on-call sees an aggregate error-rate increase, a deployment timestamp, and hundreds of healthy containers. None of that answers the urgent question: should we roll back the application, disable the experiment for one tenant cohort, or leave both alone because a dependency is briefly slow?
TL;DR: treat startup, liveness, readiness, and cohort outcome as four separate signals. Startup protects slow initialization, liveness detects a stuck process, readiness controls whether an instance should receive traffic, and the cohort signal decides whether an experiment is safe. A rollback page should name the affected cohort, compare it with a control, show sample volume, and link the result to the exact flag revision. Container health cannot substitute for experiment health.
For a gaming SaaS, this distinction is operationally decisive. A server can answer its probe while players in one tenant cohort fail to complete matchmaking. Conversely, a cohort can look worse during a low-volume interval while every application instance is behaving correctly. The first case calls for a flag decision; the second calls for patience, not a restart storm.
Healthy isn't harmless.
How should Node.js readiness, liveness, and startup work on Kubernetes?
The first screen should present the decision boundary rather than a wall of telemetry. A useful page says that cohort treatment-b crossed its rollback SLO, that control did not, that the comparison covers a stated number of attempts, and that the current experiment revision is the common change. It should also say whether readiness changed at the same time. Those facts separate a bad treatment from a bad release.
That's the distinction.
The page should not fire merely because one pod became unready. Instances leave and enter service for ordinary reasons, and a single lifecycle transition says nothing about player outcomes. Nor should it wait for the whole service to fail. The actionable condition is a sustained breach in a sufficiently populated cohort, with a control that is still inside its objective and enough healthy serving capacity to make the comparison credible.
Use four signals, each with one job:
| Signal | Question it answers | Safe automated action | Dangerous interpretation |
|---|---|---|---|
| Startup | Has initialization completed? | Delay the other probe decisions | Treat slow startup as a dead process |
| Liveness | Is the process irrecoverably stuck? | Restart that instance | Restart because a dependency is slow |
| Readiness | Should this instance receive new traffic? | Remove it from service | Declare the experiment unhealthy |
| Cohort outcome | Is the treatment violating its rollback objective? | Disable or roll back that flag revision | Restart healthy instances |
This is an ownership boundary as much as an instrumentation boundary. The platform owns the first three contracts and their capacity effects. The experiment owner defines the player outcome and the acceptable rollback threshold. On-call needs both on one page, but combining their meanings creates brittle automation.
Who owns rollback authority?
Suppose the page arrives after the treatment cohort has accumulated enough failed matchmaking attempts to breach its rollback objective. Work backward. The alert should have been preceded by a rising treatment-to-control error ratio; before that, each request should have carried a stable experiment revision and a privacy-safe tenant cohort identifier; before that, the flag evaluation should have produced a reason that could be correlated with the request result.
The application metrics therefore need two dimensions that generic container telemetry does not provide: cohort and flag revision. Keep both bounded. A raw tenant ID or player ID turns a useful metric into an unbounded cardinality problem, and player-level identifiers in logs create a separate retention and erasure obligation. GDPR Article 17 is a useful forcing function here: if an identifier may need erasure, do not casually copy it into every operational stream.
Logs carry the diagnostic detail that metrics deliberately omit. A result record can include a request correlation ID, experiment revision, coarse tenant cohort, outcome class, and latency bucket. The corresponding metric should aggregate the same outcome classes. Do not put account names, player handles, or unconstrained error messages into labels.
Feature flags also need explicit operational categories. Martin Fowler distinguishes release, experiment, ops, and permissioning toggles because they have different lifetimes and decision dynamics. An experiment toggle is a routing decision tied to a cohort; treating it like a release toggle makes rollback attribution ambiguous. Record the evaluated revision, not merely the flag's human-readable name, so a changed allocation can be separated from unchanged application code.
Four signals are enough. More signals may help diagnosis, but they should not acquire restart authority merely because they are measurable.
No shared kill switch.
Trace the alert backward to its earliest signal
The Node.js service can expose three small endpoints for orchestration and emit application metrics and structured logs through the same internal health model. The handler implementation is less important than the contract: startup remains false until mandatory initialization finishes; liveness checks only local progress; readiness checks whether the process can accept new work without exceeding its own queue or dependency budget.
The following Go model is deliberately framework-neutral. It shows the state machine that the Node.js handlers should mirror and, more importantly, keeps cohort rollback outside the process-health result.
package health
import (
"encoding/json"
"net/http"
"sync/atomic"
"time"
)
type State struct {
started atomic.Bool
lastProgress atomic.Int64
queueDepth atomic.Int64
queueCapacity int64
}
type response struct {
Status string `json:"status"`
}
func (s *State) Startup(w http.ResponseWriter, _ *http.Request) {
s.write(w, s.started.Load())
}
func (s *State) Liveness(w http.ResponseWriter, _ *http.Request) {
last := time.Unix(0, s.lastProgress.Load())
s.write(w, time.Since(last) < 30*time.Second)
}
func (s *State) Readiness(w http.ResponseWriter, _ *http.Request) {
ready := s.started.Load() && s.queueDepth.Load() < s.queueCapacity
s.write(w, ready)
}
func (s *State) write(w http.ResponseWriter, healthy bool) {
w.Header().Set("Content-Type", "application/json")
if !healthy {
w.WriteHeader(http.StatusServiceUnavailable)
_ = json.NewEncoder(w).Encode(response{Status: "unhealthy"})
return
}
_ = json.NewEncoder(w).Encode(response{Status: "ok"})
}
The 30s progress window and queue capacity are examples, not universal defaults. Derive both from the service's latency SLO, event-loop work profile, expected initialization distribution, and replica capacity. A process doing valid work must not be killed because an arbitrary interval looked tidy in configuration.
There is another trap in dependency checks. If readiness requires every downstream system to answer immediately, one dependency outage can remove every otherwise useful replica at once. If readiness ignores all dependencies, instances may accept work they cannot finish. Classify dependencies by the request paths they gate, preserve degraded paths where the product permits it, and reserve global unready status for loss of the service's minimum viable function.
This four-signal approach has a real limitation: it is not suitable when the application cannot identify a stable cohort at request time, and it cannot establish causality when control and treatment receive materially different traffic. In those cases, use the probes only for lifecycle control, pause automatic cohort rollback, and fix allocation or event attribution before trusting the comparison. The trade-off is slower action, but acting quickly on incomparable populations is false precision.
Enforce the experiment decision separately
Rollback logic needs a numerator, denominator, comparison window, minimum sample volume, and recovery rule. Without all five, “treatment errors are high” is an impression, not an operational policy.
For matchmaking, the numerator might be terminal failures and the denominator completed attempts. Compare treatment with control over the same time window and region mix. Require a minimum volume before paging, because one failure out of two attempts produces a dramatic percentage and almost no evidence. Then require consecutive evaluation windows or another explicit persistence rule so the flag does not flap at the boundary.
The rollback path should disable the specific experiment revision for the affected cohort while leaving unrelated cohorts and healthy application instances alone. Feature toggles make that separation possible, but they also add inventory: every rollback must preserve an audit record, every temporary experiment needs an owner, and expired flag paths need removal. A toggle without lifecycle ownership is deferred operational risk.
My decision rule is skeptical by design: page for human action when treatment breaches its player-outcome objective, control remains within its objective, sample volume is adequate, and serving capacity is healthy. If readiness is also collapsing across treatment and control, investigate the release or dependency path first. If only liveness is failing, repair the process failure; do not infer an experiment result from it.
Budget cardinality before choosing the control plane
The question is not which dashboard looks best. It is where rollback authority lives, how much on-call load the system creates, and whether the evidence remains portable.
| Approach | Rollback safety | On-call burden | Lock-in boundary | Capacity-planning concern |
|---|---|---|---|---|
| Build the full path | Maximum control over evaluation and actuation | Team owns storage, alerting, retries, audit, and recovery | Data and rules remain internal | Must budget ingestion, cardinality, retention, and failover |
| Managed telemetry, internal rollback | Raw signals leave the cluster; actuation stays controlled | Less storage operations, continued policy ownership | Query and alert semantics may be provider-specific | Export volume and label growth still need forecasts |
| Managed evaluation and action | Shortest integration path when contracts align | Provider handles more machinery; team validates failure modes | Policy and history may be harder to move | External control-plane availability enters the rollback budget |
I would keep the actuation boundary narrow regardless of acquisition choice. The component allowed to change a cohort assignment should accept a revision-scoped command, reject stale revisions, record who or what initiated the change, and support a tested recovery path. Telemetry readers do not automatically deserve write access to flags.
Capacity planning belongs in the initial design. Estimate active cohorts, flag revisions retained at once, outcome classes, regions, and evaluation frequency before choosing labels. Multiply them. Then include headroom for a launch, when cohort traffic and diagnostic logging usually rise together. If the estimated series count already strains the team's operational budget, aggregate earlier or move detail to sampled logs; hoping cardinality stays small is not a plan.
Test the failure paths in staging and during controlled production exercises. Hold startup open and verify that liveness does not interfere. Saturate the local queue and verify that readiness drains the instance without restarting it. Stop local progress and verify that liveness acts only after the chosen budget. Finally, inject treatment failures with a healthy control and confirm that the page names the cohort and revision while the containers remain untouched.
The false-positive bill arrives on-call
An aggressive threshold looks safe on paper because it reacts quickly. In practice, a low-volume cohort can cross it by chance, trigger a rollback, erase the evidence needed to compare the experiment, and wake someone who cannot make a better decision from the page. Repeated false pages train responders to distrust the signal; repeated automatic reversals make the experiment state harder to reason about.
The opposite error is costly too. Long windows and oversized sample requirements can preserve statistical calm while players experience a bad treatment. Set the alert window from the rollback SLO and expected cohort volume, then rehearse both sparse and busy periods. Track pages that led to no action as part of the alert's quality, and review the threshold when cohort allocation changes.
The clean design is boring: probes govern process lifecycle, bounded metrics reveal cohort divergence, structured logs explain individual outcomes, and a revision-aware flag action limits blast radius. Keep those contracts separate, and the page can answer the only question that matters at 3 a.m.: what is safe to change now?
Top comments (0)