DEV Community

polasamy-eng
polasamy-eng

Posted on

Why Your Kubernetes HPA Won't Scale Down (It's Probably Not Stuck)

You scaled up under load, traffic dropped ten minutes ago, and your HPA is still sitting at 8 replicas instead of 2. First instinct: something's broken. Second instinct, after kubectl describe hpa shows nothing wrong: confusion.

It's probably not stuck. It's doing exactly what it's configured to do — you just never configured it.

The default nobody sets on purpose

If you don't define a behavior.scaleDown block on your HPA, Kubernetes falls back to a 300-second stabilization window. That's not a bug, it's a deliberate anti-flapping guard. The catch: that window resets every time a metric sample comes in above your target, not just once. On a noisy CPU metric, "scale down after load drops" can feel like it never happens, because the window keeps getting pushed back by brief spikes.

The fix is one explicit block

behavior:
  scaleDown:
    stabilizationWindowSeconds: 120
    policies:
      - type: Percent
        value: 50
        periodSeconds: 60
Enter fullscreen mode Exit fullscreen mode

Setting this explicitly — rather than trusting the default — gives you a scale-down window you actually chose, and a policy that caps how fast replicas get removed per step. No more guessing whether 300 seconds means 300 seconds or "300 seconds from whenever the metric last blinked."

Try it yourself

I put together a minimal, runnable reproduction — a demo Deployment, an HPA with the behavior block set explicitly, and a load generator script so you can watch kubectl get hpa -w actually behave the way you'd expect:

github.com/polasamy-eng/devsaas-devops-examples — see kubernetes-hpa-scaling-demo/

Full breakdown of how to tell a genuinely stuck HPA apart from one that's just conservatively (or accidentally) configured:

Kubernetes HPA Not Scaling Down →

If you've hit a scaling edge case that doesn't fit this — multiple metrics fighting each other, custom metrics via Prometheus adapter, whatever — I'd like to hear it. Most write-ups on this stop at "here's the YAML" without touching the cases that actually bite in production.

Top comments (0)