DEV Community

Cover image for How one stuck PDB doubled our EKS autoscaler bill in six days
Muhammad Hassaan Javed for Infraforge

Posted on Edited on Originally published at infraforge.agency

How one stuck PDB doubled our EKS autoscaler bill in six days

The EC2 bill for our EKS cluster went from $4,200 to about $8,600 over six days, against the same six days the month before. Same cluster, no launch, no traffic event, no team asking for headroom. kubectl get nodes returned 60 m5.4xlarge instances where the working set is normally 20, and forty of them were sitting under 8% CPU. The cluster autoscaler had scaled up overnight for a batch burst that finished by 6 am; it had never scaled back down.

Problem signals:

  • The EC2 spend for your EKS node groups jumps without a matching traffic or deploy event
  • kubectl get nodes returns dozens more nodes than the working set, most at low CPU utilization
  • Cluster autoscaler logs repeat 'cannot be removed: not enough pod disruption budget to move' for pods of the same workload
  • A single PodDisruptionBudget in the cluster shows ALLOWED DISRUPTIONS = 0 while others are 1 or higher
  • A workload somewhere in the cluster has a pod stuck in ImagePullBackOff that nobody is watching

Six days, $8,600, and 40 idle nodes

The Friday morning spike

The bill spike was six days old by the time anyone looked. The AWS Budgets alert at $12,000 month-to-date on this cluster's spend, which normally fires around the 18th, fired on the 13th. The six days since the batch had cost about $8,600, against $4,200 for the same six days the month before. As far as anyone knew nothing had shipped that week, CI had been quiet since Wednesday afternoon, and no product team was asking for extra capacity.

Compute Optimizer had already flagged the cluster's managed node group as significantly overprovisioned. kubectl get nodes returned 60 m5.4xlarge nodes; the normal working set for this cluster is 20. Forty of the extras were sitting under 8% CPU, hosting kube-system DaemonSet pods and one pod each of a single Deployment nobody had looked at in months.

The math is straightforward. m5.4xlarge on-demand in us-east-1 is $0.768 an hour. Forty extra nodes for six days is about $4,400 of surprise burn, which is the whole jump from the $4,200 baseline. That baseline is the cluster's full EC2 spend, volumes and data transfer included, not just its 20 nodes. The overnight autoscale had happened on the previous Saturday to run a monthly batch job; the batch finished by 6 am, and then the cluster autoscaler just stopped scaling down. Six days of that pattern is how a bill doubles.

Why HPA runaway and traffic burst were both wrong

Two theories that died fast

Two theories came up on the bridge in the first five minutes. The obvious one was HPA runaway: a Horizontal Pod Autoscaler misreading its metric and scaling a Deployment to hundreds of replicas, forcing the cluster autoscaler to hold capacity to place them. The second was a traffic burst: some overnight event pushing real request volume up, driving replicas up, then subsiding but leaving the nodes.

Both died fast. kubectl get hpa --all-namespaces showed every HPA sitting comfortably below its ceiling, none within striking distance of a scale trigger. Prometheus request-rate graphs for every ingress-fronted service were flat across the whole week. When we summed pod counts across the cluster, we got 342, roughly what a normal Friday looks like, and nowhere near the 900+ pods it would take to justify 60 m5.4xlarge nodes at typical density.

The demand side was fine. Something on the supply side was refusing to shrink.

The cluster-autoscaler log that led to exactly one PDB

One PDB behind every blocked node

We went to the cluster autoscaler's own logs, which is where CA tells you plainly why it will not do the thing you want. kubectl -n kube-system logs deploy/cluster-autoscaler --since=10m | grep 'cannot be removed' (the line is logged at verbosity 2, which the --v=4 in the AWS example manifest includes) returned the same line forty times per pass, each naming a different node and a pod of the same Deployment.

I0313 09:32:04.892417       1 cluster.go:169] Node ip-10-42-14-88.ec2.internal cannot be removed: not enough pod disruption budget to move data-platform/nightly-rollup-7f8b9c5d4-4kx2m
I0313 09:32:04.895102       1 cluster.go:169] Node ip-10-42-15-19.ec2.internal cannot be removed: not enough pod disruption budget to move data-platform/nightly-rollup-7f8b9c5d4-9h4tn
I0313 09:32:04.897733       1 cluster.go:169] Node ip-10-42-15-33.ec2.internal cannot be removed: not enough pod disruption budget to move data-platform/nightly-rollup-7f8b9c5d4-b7q9r
... (37 more lines in this pass, each naming a pod of the nightly-rollup Deployment) ...
Enter fullscreen mode Exit fullscreen mode

Every blocked node in every iteration named a pod of the same Deployment. Cluster autoscaler names the pod, not the PDB, so the next step is finding which PDB covers it.

The log names pods, not PDBs, so we listed every PDB to find the one covering nightly-rollup and to confirm no other was involved.

$ kubectl get poddisruptionbudgets --all-namespaces
NAMESPACE       NAME                     MIN AVAILABLE   MAX UNAVAILABLE   ALLOWED DISRUPTIONS   AGE
api-gateway     api-gateway-pdb          2               N/A               3                     47d
auth            auth-service-pdb         1               N/A               2                     47d
data-platform   nightly-rollup-pdb       39              N/A               0                     213d
data-platform   warehouse-writer-pdb     1               N/A               2                     47d
payments        payments-api-pdb         2               N/A               1                     47d
search          search-indexer-pdb       1               N/A               2                     47d
Enter fullscreen mode Exit fullscreen mode

Six PodDisruptionBudgets. Exactly one at zero allowed disruptions. Every other PDB had one to three to spare, which is the healthy state.

The outlier was data-platform/nightly-rollup-pdb, and it was 213 days old, while the other five had all been re-created seven weeks earlier as part of a cluster upgrade. Something about this specific PDB had been quietly wrong for a long time.

We looked at the underlying workload. The 'nightly-rollup' service started life as a CronJob and had been reimplemented at some point as a 40-replica Deployment with soft pod anti-affinity, running continuously to feed a downstream aggregation service, and it rolled out with maxSurge: 0 and maxUnavailable: 1. Its PDB carried minAvailable: 39, the usual 'allow one voluntary disruption at a time' setting for node drains; PDBs do not limit a Deployment's own rolling updates. Fine when the workload is healthy.

kubectl get pods -n data-platform -l app=nightly-rollup told the rest of the story. Thirty-nine pods Running, one pod stuck in ImagePullBackOff. Six days earlier, an old pipeline had triggered a rolling update with a container image reference that pointed at an ECR registry we had migrated away from during a project the previous quarter. Every other workload had been re-pointed at the new registry during that migration; this one Deployment had been missed, and nobody had noticed because it 'just worked' for months. The overnight burst had left exactly one nightly-rollup pod on each of the forty extra nodes. The broken rollout went out at 6:10 am, minutes after the batch finished and before cluster autoscaler had drained a single burst node.

The instant one pod dropped out of Ready, healthy went from 40 to 39 against a minAvailable floor of 39, and allowedDisruptions clamped to zero. From that moment on, no node hosting a nightly-rollup pod could be drained: evicting one would take healthy below the floor, and cluster autoscaler is careful about that. Forty pods, forty nodes, one stuck PDB. Each nightly-rollup pod requested about 400m CPU and 1.5 GiB of memory, which is a rounding error on an m5.4xlarge, but the pod's presence plus the PDB at zero meant the node was unremovable regardless of how much headroom the rest of the box had. Forty nodes at 5-8% CPU, all of them technically pinned by one broken image reference.

Why we deleted the PDB instead of pushing the fixed image

Delete the PDB, or fix the image

We saw two paths out. Fix the underlying pod: patch the Deployment's image reference, let the rolling update roll forward, healthy count returns to 40, allowedDisruptions goes positive, CA is free. Or delete the PDB directly: no PDB, no eviction block, CA can drain nodes.

We almost went with the image fix, because that is the 'correct' answer in an abstract sense; the PDB is doing what it was written to do. The problem was who owned the Deployment. The team that had originally built the nightly-rollup service had been reorganized eighteen months earlier, and the current owning team had inherited it in a spreadsheet handoff and had never reviewed it; the rollout six days earlier came from an old pipeline that still referenced the old registry, where the new tag had never been pushed. To fix the image correctly we needed either to push the correct image to the old ECR registry (which required someone with write access on an account we were sunsetting) or to patch the Deployment to point at the new registry (which required knowing whether the image tag existed there and had feature-parity, which we did not, at 9:30 am).

The PDB, on the other hand, was clearly orphaned. Its author was gone. Its minAvailable: 39 was arithmetically fine for a 40-replica Deployment during healthy operation but it meant one broken pod would freeze the entire fleet, which is exactly what had happened. Deleting the PDB did not remove any replicas or affect the currently-running pods; it just removed the eviction guardrail.

kubectl delete poddisruptionbudget -n data-platform nightly-rollup-pdb
# poddisruptionbudget.policy "nightly-rollup-pdb" deleted
Enter fullscreen mode Exit fullscreen mode

One command. The gate is the PDB; remove the gate.

Twelve minutes later, cluster autoscaler started removing nodes. It took about ninety minutes to work through the fleet. The nightly-rollup pods that got evicted rescheduled onto the remaining nodes without issue because the anti-affinity was preferredDuringSchedulingIgnoredDuringExecution (soft), which meant CA could consolidate the fleet onto fewer nodes when memory allowed. The pending pod stayed pending, because the underlying image reference was still wrong, but a single pending pod is a monitoring problem, not a bill problem.

By 11:30 am the cluster was back at 24 nodes. We paged the data team's on-call to fix the image reference at their leisure the following Monday, and when they did, the PDB went back in as maxUnavailable: 2. That Friday's instance spend for the node group still came to about $740, because the fleet ran at 60 nodes until mid-morning; Saturday, still at 24 nodes, came to about $440.

The instinct in the first thirty minutes had been to kubectl delete node on the idle boxes. That would have made things worse. kubectl delete node drops the Node object from the API server, and the kubelet registers only once, at startup, so the node does not come back. Cluster autoscaler then treats the instance as unregistered and may remove it after 15 minutes, and the pod garbage collector deletes the pods that were bound to the vanished node without going through eviction, so the PDB is bypassed for every pod on it. The PDB is the gate. Remove the gate or fix the gated pod; deleting Node objects skips both and takes the pods with it.

In hindsight there was a third path, and it should have come first: roll the Deployment back. The rollout had stalled with the previous ReplicaSet still serving 39 pods on an image every node could pull, so undoing it replaces the stuck pod with a working one, puts the healthy count back at 40 and lets scale-down start again, though slower than after deleting the PDB: cluster autoscaler drains one node at a time either way, and with minAvailable: 39 standing, each drain also waits for the evicted pod's replacement to become Ready. Then widen the budget, for example to maxUnavailable: 2, so one stuck pod still leaves room for a drain. At 40 replicas, maxUnavailable: 1 is the same budget as minAvailable: 39 and would freeze the fleet the same way. The rule is that the budget has to allow more disruptions than a stalled rollout can hold unavailable. Nobody on the bridge raised it.

kubectl rollout history deployment/nightly-rollup -n data-platform
kubectl rollout history deployment/nightly-rollup -n data-platform --revision=<N>
kubectl rollout undo deployment/nightly-rollup -n data-platform --to-revision=<last good revision>
Enter fullscreen mode Exit fullscreen mode

Check the history first: the plain list shows only revision numbers and change causes, so read each candidate's pod template and image with --revision before you undo. A bare undo goes to the previous revision, which is only the good one if nothing else rolled out since. If Argo CD or Flux manages the Deployment, roll back in Git instead, or the next sync reapplies the broken revision; with Helm, use helm rollback so the release record matches. And treat it as a stopgap: the old image still lives in the registry you are retiring.

The Kyverno policy and PDB alert that closed the gap

The rules we shipped the next week

The postmortem produced three changes we shipped inside the week.

The first was a cluster-level admission policy that rejects new PDBs unless they carry an owner annotation and an expires-at annotation. If a PDB is going to have the power to freeze half a cluster, someone needs to own it and someone needs to argue for renewing it. Kyverno was already installed for other reasons, so the rule was about thirty lines.

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: pdb-lifecycle-required
spec:
  rules:
    - name: require-owner-and-expiry
      match:
        any:
          - resources:
              kinds:
                - PodDisruptionBudget
      validate:
        failureAction: Enforce
        message: "PDBs must carry infraforge.io/owner and infraforge.io/expires-at annotations. See runbook: https://internal.infraforge.io/rb/pdb-lifecycle"
        pattern:
          metadata:
            annotations:
              infraforge.io/owner: "?*"
              infraforge.io/expires-at: "?*"
Enter fullscreen mode Exit fullscreen mode

Kyverno ClusterPolicy blocking any PDB that arrives without an owner or an expiry. failureAction sits on the validate rule (Kyverno 1.13 or later); the older spec-level validationFailureAction is deprecated.

The exemption path for genuine long-lived PDBs (data-plane primary databases, for example) is to set infraforge.io/expires-at to a far-future date and put the owning team alias in infraforge.io/owner. That does not prevent the specific failure we hit, but it does mean that in eighteen months when the next team reorganization happens, we can grep for the PDBs whose owners no longer exist before they become orphaned. Add-on charts that create their own PDBs need the annotations too, or an exclude block for their namespaces.

The second change was a Prometheus alert that fires when any PDB has allowedDisruptions == 0 for more than thirty minutes. Not five, not sixty. Thirty is long enough that a healthy rolling update has finished but short enough that a stuck one still gets caught inside the same business day. It only sees stuck rollouts that drive a budget to zero; a widened budget, like nightly-rollup's now, hides a single stuck pod from it.

The third was a workload audit. Over the following week we pulled seven days of CPU usage from Prometheus for every HPA with minReplicas > 1, and dropped the floor on five workloads whose real utilization was under 30%. Two went from minReplicas: 3 to 1, and for those two we changed the PDB from minAvailable: 1 to maxUnavailable: 1 in the same change, because minAvailable: 1 on a single replica allows zero disruptions, which is this incident again. That work did not fix the incident (the PDB was the actual bug), but it lowered the cluster's headroom baseline enough to knock another two nodes off the normal working set.

On the observability side we added AWS Cost Anomaly Detection on the cluster's cost allocation tag. A flat month-to-date Budgets threshold is not a spike detector: ours normally fires around the 18th, and this incident only moved it to the 13th. Anomaly detection works from Cost Explorer data, which can lag by up to 24 hours, so it would have flagged the jump within about a day.

When your EKS bill doubles without a traffic event

When this is happening to you

Bill spikes without a matching traffic event almost always trace to a controller doing exactly what it was configured to do, in a way the config never anticipated. Cluster autoscaler is the most common source of this specific flavor, because its downscale logic is deliberately conservative around PodDisruptionBudgets; a single PDB with zero allowed disruptions can pin nodes indefinitely, and if the PDB is stuck on a pod nobody currently owns, nobody notices until the bill does. HPA misconfiguration, Karpenter provisioner drift, and stuck node-termination lifecycle hooks are the other three variants we see most often in the same shape.

We run recovery engagements with this exact shape most quarters. The fastest path to the answer is usually the cluster autoscaler logs, not a Prometheus dashboard, because CA will just tell you what it is refusing to do. If your EKS bill doubled inside a week and Compute Optimizer is flagging the node group as overprovisioned, we can be on a bridge with your platform team the same day. Book an infrastructure review and we will start with a 30-minute diagnostic call this week; the more artifact you can share ahead of the call (bill line item, kubectl get nodes output, and a kubectl -n kube-system logs deploy/cluster-autoscaler --since=1h grab), the faster we can name the specific blocker.

If you are earlier in the shape (rising EKS bill, no obvious cause yet, no acute page), the Kubernetes and CI/CD stabilization work we do covers the workload-audit and admission-policy side of what this article described. We have written more on the general pattern of cost spikes driven by controller behavior in the cloud cost spikes problem write-up.


Originally published at https://infraforge.agency/insights/eks-autoscaler-stuck-pdb-cost-spike/.

If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — see /review.

Top comments (0)