Your AKS cluster upgrade is stuck.
The control plane is waiting for a node to drain. The node is waiting for a pod to terminate. The pod is refusing to move.
Look at the logs. The workload has a PodDisruptionBudget that allows exactly zero disruptions.
It is doing exactly what it was told to do. Protect the application at all costs. Even if it means halting the cluster upgrade.
To upgrade a node, AKS has to drain it. It cordons the node and evicts the pods.
But the eviction process respects disruption budgets.
If your deployment has three replicas, and your budget dictates that three must remain available, the eviction fails. The node stays tainted. The upgrade stalls.
You didn't just configure high availability. You configured a deadlock.
Replica headroom isn't just a best practice for scaling. It is the physical space required for the control plane to do maintenance.
Before your next maintenance window, look at your math.
Check your minAvailable or maxUnavailable settings against the actual number of running replicas.
If you use minAvailable with a percentage, remember that Kubernetes rounds up. A setting of 100% means zero disruptions allowed. A setting of 90% with three replicas rounds up to three.
You have to leave room for the scheduler to breathe. Configure these values deliberately, not just to pass a compliance check.
High availability means surviving a failure. It shouldn't mean blocking your own operations team.
Your availability policy guarantees the app survives a zone failure. But has anyone tested whether it survives routine maintenance?
If your zero-disruption configuration blocks your upgrade pipeline, you haven't built a resilient system.
You have built a hostage situation.
Top comments (0)