DEV Community

Cover image for Your cluster says 3/3. It will still go dark with one zone.
Sai Pisey
Sai Pisey

Posted on Originally published at saipisey.com

Your cluster says 3/3. It will still go dark with one zone.

Every deployment in this three-zone cluster reports ready:

$ kubectl get deployments
NAME            READY   UP-TO-DATE   AVAILABLE
checkout-api    3/3     3            3
session-store   1/1     1            1
web             3/3     3            3
Enter fullscreen mode Exit fullscreen mode

Now take us-east-1a away:

$ kubectl survive-zone
Losing us-east-1a  ->  2 lost, 0 degraded, 1 impaired by a dependency
  WORKLOAD       PLACEMENT                               VERDICT
  checkout-api   us-east-1a:3                            LOST
  session-store  us-east-1a:1                            LOST
  web            us-east-1a:1 us-east-1b:1 us-east-1c:1  IMPAIRED
Enter fullscreen mode Exit fullscreen mode

Two workloads are simply gone. The third still has two healthy replicas, and stops serving anyway.

Five of the seven pods sit in us-east-1a. Losing that zone takes checkout-api and session-store down outright and leaves web without its dependency.

Nobody made a mistake

checkout-api was scheduled while us-east-1a was the only zone. The scheduler put all three replicas there, correctly. Two more zones were added later, and nothing moved them.

That is not a bug. Topology spread constraints are applied when a pod is scheduled, and never checked again. The Kubernetes docs say it plainly: there is no guarantee they stay satisfied once pods move.

So the manifests say multi-AZ, and the cluster quietly says something else.

The row a linter would pass

web is the interesting one. It is spread one replica per zone, which is exactly what a spread linter wants to see. It still goes down, because the only session-store pod lives in the zone that died.

All three web replicas depend on session-store, whose only replica is in us-east-1a. Two healthy web replicas still cannot serve.

Zone survivability is a property of the dependency graph, not of each workload on its own. kubectl survive-zone follows each Service to its EndpointSlices to build that graph, so a workload can be marked impaired even when its own pods look perfect.

It will not hand you a fix that causes an outage

Most zone fixes can make pods unschedulable. Tighten a spread constraint from ScheduleAnyway to DoNotSchedule and you can end up with replicas stuck in Pending, which is the outage you were trying to avoid.

So kubectl survive-zone fix only prints a patch after proving two things: that it is schedulable on your real nodes, using the upstream kube-scheduler's own Filter plugins, and that it actually survives the zone loss.

A fix is printed only if it is schedulable and survives the zone loss. A PodDisruptionBudget is shown, and marked as not a fix.

And it keeps watching

Placement decays with nothing changed in git. A node drains at night, pods reschedule, and a workload that was spread now sits in one zone. Run as a Prometheus exporter, the tool turns that into an alert while the zone is still up.

survive_workload_survives drops from 1 to 0 after two zones are cordoned, with no deploy and no change in git.

Two things the full post covers that surprised me

A node drain that hangs forever, behind a disruption budget that looks completely fine. Budget arithmetic says the drain is safe. It is not, and the reason only shows up when you simulate the replacement pod against the real scheduler.

An ordinary rolling update that quietly broke an enforced maxSkew: 1. I found it by accident while recording the demo, and the fix is a one-line field most spread constraints do not set.

The full write-up also covers how every verdict is checked nightly against a real control plane in 200 randomised scenarios, and what the harness caught that no unit test did.

Read the full post on saipisey.com →

Try it in under a minute. It is read-only and never writes to your cluster:

$ kubectl krew install survive-zone
$ kubectl survive-zone
Enter fullscreen mode Exit fullscreen mode

Top comments (3)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to

Collapse
 
dev_in_the_fog profile image
Jason Y. (dev_in_the_fog) •

Really thoughtful post! Documenting real-world engineering hurdles and actionable solutions like this brings immense value to the community.

Collapse
 
sai_pisey_02 profile image
Sai Pisey •

Thank you so much!