DEV Community

Davi Orlandi
Davi Orlandi

Posted on

Consul Service Discovery That Does Not Lie: Health Checks First

Service discovery without honest health checks is a catalog of corpses. I learned that the expensive way: a payments instance that answered 200 on / while its worker pool was deadlocked. Consul kept advertising it. Customers timed out. Dashboards looked fine. The personal stake was simple and humiliating. Discovery had done exactly what we configured it to do. We had configured a lie.

This essay is about making Consul return instances that can do useful work. The quality of payments.service.consul is not determined by how pretty the catalog looks. It is determined by how you define healthy.

Registration is a readiness story

Define the service and nested check blocks (or a checks array). Register with the local agent. Reload if the service was already registered. Confirm clients filter to the statuses you intend. Initial check status defaults to critical until the first success. That prevents newborns from receiving traffic before they are ready. Override with status = "passing" only when you understand the trade, because starting green is how you advertise an unfinished boot.

Prefer HTTP and TCP checks over shell folklore

If you are curling an endpoint from a script check, use an HTTP check. If you are probing a port with netcat, use a TCP check. Native checks avoid shell sprawl and reduce remote-execution risk compared with broadly enabled script checks.

service {
  name = "payments"
  port = 8080

  check {
    id       = "payments-http"
    name     = "HTTP /health/ready"
    http     = "http://127.0.0.1:8080/health/ready"
    method   = "GET"
    interval = "10s"
    timeout  = "2s"
    failures_before_critical = 3
  }
}
Enter fullscreen mode Exit fullscreen mode

Probe 127.0.0.1 so you evaluate the local process, not a load balancer illusion. Keep timeouts short. Use failure thresholds to avoid removing capacity on a single blip. Consul's HTTP mapping is worth memorizing: 2xx is passing, 429 is warning, anything else is critical.

check {
  id       = "payments-tcp"
  name     = "TCP 8080"
  tcp      = "127.0.0.1:8080"
  interval = "10s"
  timeout  = "1s"
}
Enter fullscreen mode Exit fullscreen mode

Design /health on purpose

Separate concerns before you wire Consul to them.

Endpoint Purpose Typical Consul use
/health/live Process is up Orchestrator restarts
/health/ready Can accept traffic Consul check target

Avoid marking every instance critical when a shared database hiccups unless evacuating everyone is truly safer than serving degraded. Mass critical status amplifies outages. Readiness should prove the instance can accept work: for example a cheap in-process slot acquire with a timeout that returns 503 when saturated. Liveness should answer a smaller question.

Other check types and when they earn their complexity

TTL checks let the application heartbeat Consul. Useful when only the app knows readiness. Missed heartbeat becomes critical. gRPC checks fit services that expose the standard health protocol. Alias checks mirror another service or node health for sidecars. Script checks are powerful and dangerous; prefer enable_local_script_checks over broadly enabling remote script checks.

You may attach several checks to one service. Keep the set small. Three flaky checks create a jittery pool that looks like network chaos and trains operators to ignore status changes.

Client discovery only becomes trustworthy after checks are honest

After checks tell the truth, clients can resolve Consul DNS with local caching, use health HTTP APIs with blocking queries, or use Connect when you need mTLS and intentions. DNS detail matters: by default warning statuses may still appear in DNS answers. Set dns_config.only_passing = true (or filter explicitly in the HTTP API) when only passing instances should receive traffic. I have seen teams "debug networking" for hours when the real bug was DNS still answering for warning.

Registration hygiene that keeps the catalog adult

Give each instance a unique id (payments-i-0abc). Tag for version or canary. Deregister on shutdown, or keep intervals short enough that dead boxes disappear. Enable maintenance mode before disruptive node work so discovery stops advertising without pretending the process crashed. Pair Consul readiness with orchestrator readiness gates during deploys so both systems tell the same story.

Closing

Consul discovery is only as trustworthy as your checks. Prefer native HTTP, TCP, or gRPC probes. Separate liveness from readiness. Configure DNS and API clients to honor the statuses you actually want. Healthy discovery is less about the catalog and more about refusing to advertise lies. The day I stopped pointing readiness at a decorative / endpoint was the day discovery started feeling like infrastructure again instead of theater.

Top comments (0)