<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sai Pisey</title>
    <description>The latest articles on DEV Community by Sai Pisey (@sai_pisey_02).</description>
    <link>https://dev.to/sai_pisey_02</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174831%2Fec21c07e-4849-4365-816e-6f52a37e2b60.png</url>
      <title>DEV Community: Sai Pisey</title>
      <link>https://dev.to/sai_pisey_02</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sai_pisey_02"/>
    <language>en</language>
    <item>
      <title>Your cluster says 3/3. It will still go dark with one zone.</title>
      <dc:creator>Sai Pisey</dc:creator>
      <pubDate>Sat, 10 Oct 2026 10:36:48 +0000</pubDate>
      <link>https://dev.to/sai_pisey_02/your-cluster-says-33-it-will-still-go-dark-with-one-zone-3l3k</link>
      <guid>https://dev.to/sai_pisey_02/your-cluster-says-33-it-will-still-go-dark-with-one-zone-3l3k</guid>
      <description>&lt;p&gt;Every deployment in this three-zone cluster reports ready:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl get deployments
&lt;span class="go"&gt;NAME            READY   UP-TO-DATE   AVAILABLE
checkout-api    3/3     3            3
session-store   1/1     1            1
web             3/3     3            3
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now take &lt;code&gt;us-east-1a&lt;/code&gt; away:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl survive-zone
&lt;span class="gp"&gt;Losing us-east-1a  -&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;2 lost, 0 degraded, 1 impaired by a dependency
&lt;span class="go"&gt;  WORKLOAD       PLACEMENT                               VERDICT
  checkout-api   us-east-1a:3                            LOST
  session-store  us-east-1a:1                            LOST
  web            us-east-1a:1 us-east-1b:1 us-east-1c:1  IMPAIRED
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two workloads are simply gone. The third still has two healthy replicas, and stops serving anyway.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpw04hszix8hzo5rcb1ah.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpw04hszix8hzo5rcb1ah.png" alt="Five of the seven pods sit in us-east-1a. Losing that zone takes checkout-api and session-store down outright and leaves web without its dependency." width="800" height="352"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody made a mistake
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;checkout-api&lt;/code&gt; was scheduled while &lt;code&gt;us-east-1a&lt;/code&gt; was the only zone. The scheduler put all three replicas there, correctly. Two more zones were added later, and nothing moved them.&lt;/p&gt;

&lt;p&gt;That is not a bug. Topology spread constraints are applied when a pod is scheduled, and never checked again. The &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/topology-spread-constraints/#known-limitations" rel="noopener noreferrer"&gt;Kubernetes docs&lt;/a&gt; say it plainly: there is no guarantee they stay satisfied once pods move.&lt;/p&gt;

&lt;p&gt;So the manifests say multi-AZ, and the cluster quietly says something else.&lt;/p&gt;

&lt;h2&gt;
  
  
  The row a linter would pass
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;web&lt;/code&gt; is the interesting one. It is spread one replica per zone, which is exactly what a spread linter wants to see. It still goes down, because the only &lt;code&gt;session-store&lt;/code&gt; pod lives in the zone that died.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7or2b822ftg412keyrsf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7or2b822ftg412keyrsf.png" alt="All three web replicas depend on session-store, whose only replica is in us-east-1a. Two healthy web replicas still cannot serve." width="799" height="311"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Zone survivability is a property of the dependency graph, not of each workload on its own. &lt;code&gt;kubectl survive-zone&lt;/code&gt; follows each Service to its EndpointSlices to build that graph, so a workload can be marked impaired even when its own pods look perfect.&lt;/p&gt;

&lt;h2&gt;
  
  
  It will not hand you a fix that causes an outage
&lt;/h2&gt;

&lt;p&gt;Most zone fixes can make pods unschedulable. Tighten a spread constraint from &lt;code&gt;ScheduleAnyway&lt;/code&gt; to &lt;code&gt;DoNotSchedule&lt;/code&gt; and you can end up with replicas stuck in Pending, which is the outage you were trying to avoid.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;kubectl survive-zone fix&lt;/code&gt; only prints a patch after proving two things: that it is schedulable on your real nodes, using the upstream kube-scheduler's own Filter plugins, and that it actually survives the zone loss.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmzigewydzqha8aedjrw2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmzigewydzqha8aedjrw2.png" alt="A fix is printed only if it is schedulable and survives the zone loss. A PodDisruptionBudget is shown, and marked as not a fix." width="800" height="289"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  And it keeps watching
&lt;/h2&gt;

&lt;p&gt;Placement decays with nothing changed in git. A node drains at night, pods reschedule, and a workload that was spread now sits in one zone. Run as a Prometheus exporter, the tool turns that into an alert while the zone is still up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb1y3y5jjke98xc6cb05m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb1y3y5jjke98xc6cb05m.png" alt="survive_workload_survives drops from 1 to 0 after two zones are cordoned, with no deploy and no change in git." width="799" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things the full post covers that surprised me
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A node drain that hangs forever, behind a disruption budget that looks completely fine.&lt;/strong&gt; Budget arithmetic says the drain is safe. It is not, and the reason only shows up when you simulate the replacement pod against the real scheduler.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An ordinary rolling update that quietly broke an enforced &lt;code&gt;maxSkew: 1&lt;/code&gt;.&lt;/strong&gt; I found it by accident while recording the demo, and the fix is a one-line field most spread constraints do not set.&lt;/p&gt;

&lt;p&gt;The full write-up also covers how every verdict is checked nightly against a real control plane in 200 randomised scenarios, and what the harness caught that no unit test did.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://saipisey.com/blog/kubectl-survive-zone" rel="noopener noreferrer"&gt;Read the full post on saipisey.com →&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Try it in under a minute. It is read-only and never writes to your cluster:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl krew &lt;span class="nb"&gt;install &lt;/span&gt;survive-zone
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl survive-zone
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>sre</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
