DEV Community

Cover image for Autoscaling doesn't save you at 30,000 requests per second — your dependencies do
Kennedy Njoroge
Kennedy Njoroge

Posted on

Autoscaling doesn't save you at 30,000 requests per second — your dependencies do

There's a comfortable story about peak traffic: load goes up, the autoscaler notices, pods multiply, everyone keeps their evening. I've watched that story hold, and I've watched it fail in a specific and instructive way — the autoscaler worked perfectly and the platform degraded anyway.

Running transaction paths at sustained peaks in the tens of thousands of requests per second teaches you that horizontal scaling is the easy half of the problem. The hard half is that everything your new pods talk to did not scale, and now there are more of them asking.

The failure that taught me this

Scale-up triggers. Pods go from 12 to 40 in about ninety seconds. Each one opens its configured pool of 20 database connections on startup.

The database's connection limit is 500.

Twelve pods needed 240 connections and everything was fine. Forty pods want 800. The database starts refusing connections, the pods that can't connect fail their readiness probes, the orchestrator restarts them, they try to connect again, and you have built a very efficient machine for denying yourself service. The autoscaler, meanwhile, sees rising latency and scales up further.

The dashboards looked bizarre: CPU low, memory low, replica count healthy, error rate vertical.

Autoscaling converts a capacity problem into a concurrency problem on whatever you depend on. If that dependency is a fixed-size resource — a connection limit, a license seat count, a rate-limited upstream, a legacy system with a thread pool — scaling out makes things worse, not better, and it does so fast.

Budget connections globally, not per pod

The fix is to stop thinking about per-pod configuration and start thinking about a global budget:

pool_size_per_pod × max_replicas ≤ dependency_limit × safety_factor
Enter fullscreen mode Exit fullscreen mode

With a 500-connection database, 0.8 safety factor (leave headroom for migrations, admin sessions, the other service nobody told you about), and a 40-replica ceiling:

pool_size ≤ (500 × 0.8) / 40 = 10
Enter fullscreen mode Exit fullscreen mode

Then pin both halves so neither drifts:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
  minReplicas: 12
  maxReplicas: 40          # ← this is a dependency constraint, not a cost one
  metrics:
    - type: Pods
      pods:
        metric:
          name: http_inflight_requests
        target:
          type: AverageValue
          averageValue: "35"
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 30
      policies:
        - type: Percent
          value: 50
          periodSeconds: 60
    scaleDown:
      stabilizationWindowSeconds: 600
Enter fullscreen mode Exit fullscreen mode

Three things worth calling out.

maxReplicas is load-bearing. Document why it's 40, next to the number, or someone will raise it during an incident to "give it more capacity" and cause the exact outage they're trying to stop.

Scale on in-flight requests, not CPU. A service waiting on a slow downstream has low CPU and is in serious trouble. In-flight count — or queue depth — tracks saturation. CPU tracks work, and at peak the problem is usually waiting, not working.

Asymmetric windows. Scale up fast, scale down slowly. Flapping replicas means constantly re-establishing connections, which is precisely the pressure you're trying to avoid.

Retry budgets, or how three services become twenty-seven

Service A calls B calls C. Each retries three times on failure, which is the default in a lot of client libraries and sounds reasonable in isolation.

C has a bad minute. B's single call becomes 3. A's single call becomes 9. Your users' single click becomes 27 requests to a service that is already struggling. Retries turn a brownout into an outage, reliably, and the blame lands on whoever's service fell over last rather than whoever configured the retries.

Two rules that have held up for me:

Retry at one layer only. Usually the outermost one that can still do something useful. Everything beneath it fails fast and propagates.

Cap retries as a fraction of total traffic, not per request. This is what service meshes mean by a retry budget — "retries may not exceed 20% of requests to this destination". Under normal conditions nobody notices. Under failure, the budget exhausts and retries simply stop, which is exactly the behaviour you want and exactly the behaviour per-request configuration cannot express.

Pair it with a circuit breaker so that a hard-down dependency fails in microseconds instead of consuming a connection for a 30-second timeout:

outlierDetection:
  consecutive5xxErrors: 5
  interval: 10s
  baseEjectionTime: 30s
  maxEjectionPercent: 50
Enter fullscreen mode Exit fullscreen mode

maxEjectionPercent: 50 is the detail people skip. Without it, a bad deploy or a network partition can get every endpoint ejected and you've turned a partial failure into a total one.

The failover you haven't run is not a failover

Every platform I've worked on had a documented DR procedure. The useful question is never whether it exists, it's: when did a human last execute it, and how long did it take?

If the answer is "during the audit, and about forty minutes", you don't have failover. You have a document. Forty minutes of manual steps under pressure is where people typo a hostname into a production config.

Making it real means making it a script, and making the script run on a normal Tuesday:

#!/usr/bin/env bash
set -euo pipefail

TARGET_SITE="${1:?usage: failover.sh <dr|primary>}"

# 1. verify the target is actually healthy before sending traffic to it
./healthcheck.sh "$TARGET_SITE" || { echo "target unhealthy, aborting"; exit 1; }

# 2. pre-warm — a cold site that receives 100% of peak traffic will fall over
ansible-playbook -i "inventory/${TARGET_SITE}" scale-up.yml \
  --extra-vars "target_replicas=${PEAK_REPLICAS}"
./wait-for-ready.sh "$TARGET_SITE" --timeout 180

# 3. shift traffic in steps, checking error rate between each
for weight in 10 25 50 100; do
  ./set-traffic-weight.sh "$TARGET_SITE" "$weight"
  sleep 30
  ./assert-error-rate-below.sh 0.01 || { ./set-traffic-weight.sh "$TARGET_SITE" 0; exit 1; }
done
Enter fullscreen mode Exit fullscreen mode

The pre-warm step is the one most often missing. Shifting full peak traffic onto a site running at minimum replicas gives you a cold start, an empty cache, and a thundering herd against the database, all at once — and the resulting outage gets attributed to the failover rather than to the lack of pre-warm.

The traffic ramp matters for the same reason: it gives you a cheap, reversible answer to "is the DR site actually serving correctly" before you're fully committed.

Telemetry that triggers things, not just dashboards

Most observability setups stop at displaying. The step after that is letting signals drive action — scale-up, pre-warm, failover — with clear, bounded authority.

What I'd automate, in order of how confidently I'd do it:

  1. Scale-up on saturation. Safe, reversible, happens constantly.
  2. Pre-warm the DR site on sustained primary degradation. Also safe — you're just spending some compute on standby capacity. Do this well before you'd consider failing over.
  3. Traffic shift. Automate the steps, gate the decision on a human, at least until you've watched it work a dozen times. Automated failover with a flaky health signal gives you a system that fails over and back repeatedly, which is worse than either site being down.

And whatever you automate, the signal needs to be multi-source. A single probe failing is a probe problem until something else agrees with it. Requiring agreement between an error-rate signal and a synthetic transaction before acting cuts false positives dramatically, at the cost of a little latency on genuine failures. That's a good trade for anything that moves traffic between sites.

What to check this week

If peak is coming and you've got an afternoon:

  • Multiply your pool size by max replicas. Compare that number against every fixed-limit dependency. Do it for database connections, and then do it again for anything with a rate limit.
  • Trace a request through your stack and count the retry layers. If it's more than one, fix that before peak.
  • Ask when the failover was last executed by a human. If it's more than six months, schedule a rehearsal rather than a review.
  • Check what your HPA actually scales on. If it's CPU, ask what happens when the service is slow but idle.

None of this is exotic. It's just the half of scaling that doesn't show up in the autoscaler's dashboard, which is why it's the half that's usually missing.


Writing up what I learn about reliability, backend systems and data infrastructure. If you've had an autoscaler make an outage worse, I'd like to hear how.

Top comments (0)