DEV Community

Jeroen Dirks
Jeroen Dirks

Posted on AI-assisted

Your Fargate Autoscaler Doesn't Understand Your JVM (by Default): Lessons Learned the Hard Way

Originally published on the AWS Builder Center.

Your autoscaler doesn't understand your JVM

Autoscaling a stateless web service feels like a solved problem: pick a CPU
target, set min and max, walk away. That works right up until the service is a
JVM — and I learned the hard way that it doesn't just underperform, it
fails in ways that look nothing like a capacity problem.

Out of the box, your autoscaler watches two things: container CPU and container
memory. Neither understands the runtime inside the box. A JVM has its own
memory manager, its own thread model, and its own pathological failure modes,
and the container metrics describe the box, not the runtime — so when the
two views disagree, a policy wired to the default signal does something worse
than nothing. The textbook case: it scales out on a signal that never recovers,
and the fleet ratchets up and never comes back down.

The default signals also go quiet at exactly the wrong moments — when the
unusual happens
. A downstream slows down and requests pile up waiting on I/O;
CPU and memory stay calm while the service silently runs out of threads. One
giant request inflates the heap toward an OOM while container memory, in an
oversized task, barely moves. A sudden spike arrives faster than any task can
launch. Each of these is a real load profile that naive, container-metric
scaling gets wrong — and each has a JVM-truthful signal that sees it, if you
wire the autoscaler to understand the runtime.

This post walks through five load profiles for one Java-on-Fargate service, and
for each asks the only question that matters: which emitted signal actually
reflects this pressure, which one lies — and how do you make the fleet scale
right when the unusual happens?

The task envelope

Every Fargate task is a fixed box. The reference task here is 4 vCPU /
10 GiB memory / 40 GiB ephemeral disk
, running one JVM with -Xmx6144M
under G1GC.

That sizing is deliberate. A 6 GiB heap in a 10 GiB task leaves realistic room
for non-heap (metaspace, thread stacks, direct buffers) and the OS — which
means container-memory utilisation actually tracks JVM pressure. Put the same
6 GiB heap in an oversized 24 GiB task and the heap could never move the
container-memory needle; the metric would read ~25% while the JVM was one GC
away from death. Hold on to that distinction — it is the crux of the whole
post.

Autoscaling never changes the box. It changes how many boxes run. To decide
that well you have to know what fills each box and which fills the autoscaler
can actually see.

Each diagram below shows four stacked bars for one load state: memory
(with the fixed -Xmx heap band and, nested inside it, the live-set that
survives GC), worker-thread pool occupancy, CPU, and ephemeral
disk
. We take the load types roughly in the order a healthy service meets
them — concurrency first, because on a well-tuned system it is the signal that
should drive scaling most of the time.

1 · Idle

It is the middle of the night. Traffic is a trickle — health checks and the odd
request — and the fleet has already scaled all the way in to its minimum,
say 3 tasks (you keep a small floor for availability across AZs, not
because the load needs it). The JVM has warmed up and settled: the post-GC heap
live-set is small, only a handful of worker threads are ever busy at once, CPU
idles in the single digits, and disk sits at its baseline (image + build
objects + a little log).

This is the reference state every other scenario is measured against, and
it is the state a right-sized fleet should return to overnight — sitting
quietly at minCapacity until demand returns. If a fleet can't get back here
— if task count stays pinned high while traffic falls — that is almost always a
scale-in bug (a policy stuck "active"), not real demand. Hold that thought;
it comes back in the memory section.

Verdict: no scaling action — the fleet is correctly resting at its
minCapacity floor (~3 tasks). The thing to verify here is the opposite of
scaling: that nothing is wedging the fleet above this floor overnight.

When the spike beats the autoscaler

Before we walk the steady-state load types, deal with the ugly one. Suppose
from that quiet 3-task floor traffic suddenly jumps far faster, and far higher,
than anything forecast — a flash sale, a retry storm, a dependency recovering
and dumping its backlog on you. Autoscaling cannot save you here: a new
Fargate task takes minutes to launch and pass health checks, and the spike
arrives in seconds. For that window the current tasks are all there is.

So the behaviour that matters is not "scale fast" — it's what each individual
task does when it is handed more than it can serve.
The rule:

Each task accepts only as many concurrent requests as its worker-thread
pool
has slots, and fast-fails everything beyond that — an immediate,
cheap rejection rather than an accepted request left to time out.

That gives you a brownout, not a blackout: the requests inside the pool's
capacity are served at normal latency, and the overflow is shed instantly so
the client can retry or degrade. The task's throughput of successful work
stays flat and predictable right through the surge, instead of collapsing as an
unbounded queue drags every request past its deadline and pegs CPU and GC. As
the autoscaler then brings new tasks online, total pool capacity across the
fleet rises and the shed fraction shrinks until the service is fully available
again. A fixed amount of good work per task, multiplied by a growing task
count
— that is the whole recovery curve.

One carve-out is non-negotiable: a task in overload must keep passing its
load-balancer health checks.
If an overloaded task starts failing health
checks, the load balancer marks it unhealthy and pulls it out of rotation — so
its share of traffic piles onto the remaining tasks, pushing them into
overload, and the fleet unravels one task at a time exactly when you need every
task serving. Health-check handling must therefore be independent of the
request-serving thread pool
(a dedicated path or a reserved slot), so a task
can be shedding customer load and still truthfully answer "yes, I'm alive and
in service." Shedding load is healthy behaviour; it must not look like death.

This is why the size of the worker-thread pool is the single most important
tuning knob
, and why it should be set so the pool saturates before CPU or
memory does
under normal conditions. The pool is then a clean admission gate
at the front door — full pool means instant, cheap rejection — rather than
letting load leak in until CPU or GC melts down with no clean gate at all. How
to actually find that number is a stress-testing problem; we come back to it in
Sizing the pool once the load
types are on the table.

2 · Concurrency-loaded

tart here, because on a healthy, well-tuned service this is the normal case
— the one that should drive most of your scaling decisions. A well-behaved
request-serving JVM spends much of its time waiting on downstreams (a database,
a dependency service, a cache), so the resource that tracks real demand most
faithfully is not CPU or memory but how many requests are in flight at once.

Picture a downstream slowing down — not failing, just answering in 300 ms
instead of 50. TPS hasn't changed, but each request now holds its worker
thread far longer
while it waits on I/O. The threads aren't doing work;
they're parked. So CPU stays moderate and memory barely moves — the two metrics
you would normally scale on both look calm — while the worker-thread pool
quietly fills toward its ceiling. When the last free thread is taken, the next
request is queued or rejected, and latency spikes off a cliff even though
the box looks half-idle.

This is the case raw CPUUtilization / MemoryUtilization completely miss. The
signal that sees it is thread-pool occupancy % (in-flight requests ÷ pool
size). Most Java service frameworks expose an equivalent — a busy-threads
gauge, an active-connection count, an in-flight-request counter — that you can
publish to CloudWatch. Scale out on it so you add capacity while threads are
filling, not after they're exhausted. And because occupancy falls cleanly as
load drops, it is also a trustworthy scale-in signal — which is why, on a
well-tuned system, concurrency ends up being the metric the fleet breathes on
day to day.

Verdict: scale out (and in) on thread-pool occupancy — the primary
signal on a healthy service, and the only one that sees thread exhaustion.

3 · CPU-loaded

Now a contrasting case. Some workloads — or some operations within a workload —
are genuinely CPU-bound: each request spends its time on the box (parsing,
resolving, serialising, compressing) rather than waiting on a downstream. For
those, CPU utilisation tracks demand almost linearly and the thread pool may
never be the thing that saturates first.

That is fine until CPU gets too high. Once the vCPUs saturate, runnable
threads queue in the OS run-queue waiting for a core. Nothing is broken, but
every request now spends extra time simply waiting to be scheduled, and
that shows up directly as rising service latency.

The goal is to add tasks before that happens: scale out on CPU utilisation at
a target (e.g. 40%) that sits well below the point where run-queue waiting
begins, so new tasks are in service and taking traffic before latency is ever
touched. Because new Fargate tasks take minutes to launch and pass health
checks, the target has to leave enough headroom to cover that lag on a rising
ramp.

Verdict: scale out on CPU utilisation — and CPU is a trustworthy
scale-in signal too. Treat it as the complement to concurrency: whichever of
the two is the binding constraint for your workload leads, and both are safe
to drive scale-in.

4 · Memory-loaded

A few unusually large requests arrive — a client with a huge payload, or a
caching layer that retains objects for the life of a request. These cost little
CPU and few threads, but they inflate the heap live-set. As live data
climbs toward -Xmx, G1 has less room to work in, so it collects more often
and for longer
— and the post-GC floor (the memory that survives every
collection) keeps rising instead of dropping back.

That is the danger sign: the JVM is spending an increasing share of its time in
GC (burning CPU to no useful end), and if the live-set reaches -Xmx the task
takes an OutOfMemoryError and dies. The honest early-warning signal is
heap-used-after-GC as a percentage of -Xmx.

Here is where sizing decides everything. Because this box is well-sized (6 GiB
heap in a 10 GiB task), container memory util also reads high (~80%) — the two
signals agree, which is exactly what a right-sized task buys you. Contrast:
put this same 6 GiB heap in a 24 GiB task and container util would read only
~35% at the same OOM-imminent moment. Container-RSS memory% would be blind
to the pressure. That is why heap-after-GC, not container memory, is the
metric to scale on.

Two non-obvious rules make a heap-based policy safe:

  • Scale-out only. A heap-based scale-in policy stays "active" whenever warm caches or a sticky live-set keep heap high. Under the AWS rule that a target the fleet scales in only when all policies are inactive, a heap scale-in policy jams scale-in fleet-wide — the exact "scales up, never down" bug from the idle section. Drive scale-in from CPU / concurrency only.
  • Use p90, not Maximum. A single large request pins one task's heap high. Maximum reads the single hottest task and would chase it toward maxCapacity even though adding tasks can't cool a lone hot task. p90 fires only when the top ~10% of tasks are hot — real fleet-wide pressure. Keep the alarm on Maximum; use p90 only for the scaling policy.

Verdict: scale out on heap-used-after-GC p90, scale-out only, at a
threshold below the heap alarm so capacity arrives before a page.

5 · Disk-loaded

This one doesn't follow traffic at all. Over days, ephemeral disk creeps
upward on every task in the fleet at the same slow rate, regardless of load.

The usual cause is log files that never get cleaned up — rotated logs that
should be deleted are piling up because whatever is supposed to reap them isn't
running — or, less often, repeated core dumps from crashing tasks. CPU,
memory and threads all read completely healthy; only disk is in trouble.

This is the one scenario where the answer is not autoscaling. Adding tasks
just starts more tasks with the same broken reaper that will fill at the same
rate; removing tasks does nothing. Treat disk as an operational alarm that
pages a human to fix the reaper or raise ephemeral storage — never as a scaling
axis.

Verdict: do not autoscale on disk. Alarm, page, fix the reaper.

A note on the diagrams: they hold the OS/agent and non-heap bands constant
across scenarios for clarity. In reality those tick up a little under load
(the ECS agent, log shipping, OS bookkeeping, and thread stacks all do
slightly more work as volume rises), but the movement is small relative to
the JVM's own footprint and well within the fixed headroom. Scale on the
JVM-truthful signals; treat the non-JVM drift as background noise.

What the system tells you

Fargate and the JVM emit two different classes of signal. The whole art of
scaling a JVM service is knowing which is truthful for which load.

Metric Source What it truly measures Blind spot
CPUUtilization ECS container Aggregate vCPU use across the task Sees I/O-blocked concurrency as "calm"; GC inflates it
MemoryUtilization ECS container (RSS) Whole-container resident memory Only as good as the sizing — tracks pressure in a well-sized box, goes blind in an oversized one. Prefer heap-after-GC.
Thread-pool occupancy % App framework (in-flight ÷ pool size) Real concurrency saturation Must be emitted by the app; must stay below the pool ceiling
In-flight request count App framework Raw count of concurrent requests Raw count is meaningless across a large fleet — prefer the normalised occupancy %
HeapUsedAfterGC % JVM via JMX → CloudWatch Post-GC live-set as % of -Xmx — the truthful memory-pressure signal Not emitted by default; sawtooth, so publish the post-collection value and use p90
Ephemeral disk used % Host/system metrics Disk fill An ops failure, not a capacity signal — alarm, don't scale

When to add or shed tasks

Scale OUT if any of:

  • Thread-pool occupancy high (approaching the pool ceiling) — the primary signal on a healthy service, and the one CPU/memory miss.
  • CPUUtilization > target (e.g. 40%) — the CPU-bound case.
  • HeapUsedAfterGC p90 > ~60% then ~80% (stepped) — the memory case, scale-out only, thresholds set below the heap alarm.

Scale IN only when all scale-in-eligible signals are low — driven by
thread-pool occupancy and/or CPUUtilization falling, with conservative
cooldowns (e.g. one hour). Never let heap or container-memory drive
scale-in.

Never autoscale on: container MemoryUtilization as a JVM-heap proxy (drop
it, or set its target well above resting RSS), or disk (alarm + page + fix the
reaper).

Emitting heap-after-GC

HeapUsedAfterGC is the one signal in that table that isn't free — the JVM
doesn't publish it to CloudWatch for you. The standard, framework-agnostic way
to capture it uses only java.lang.management / com.sun.management from the
JDK:

  1. Get the GC beans with ManagementFactory.getGarbageCollectorMXBeans(), cast each to a NotificationEmitter, and register a listener.
  2. On each notification, convert the payload with GarbageCollectionNotificationInfo.from(cd) and read GcInfo.getMemoryUsageAfterGc(), which returns a Map<String, MemoryUsage> of each pool's usage after the collection.
  3. Sum the heap pools (e.g. the old generation and any survivor/eden pools — skip the non-heap pools like Metaspace) and divide by -Xmx to get the post-GC heap-used percentage. Publish that number as a custom CloudWatch metric on your preferred emission interval.

That GarbageCollectionNotificationInfo + GcInfo pattern is the canonical
JDK recipe for "how full was the heap after the last collection"; it is the
same mechanism tools like Micrometer's JVM GC metrics build on, so if you
already run Micrometer you likely have the raw signal and only need to shape and
export the percentage. Many Java service frameworks ship a built-in equivalent.

Two cheap safeguards are worth adding so the policy never silently no-ops:

  1. Verify the metric exists before pointing a policy at it. Run aws cloudwatch list-metrics against the target account; a policy aimed at an absent metric sits in INSUFFICIENT_DATA forever — worse than no policy, because it looks configured.
  2. Alarm on the metric's absence. A deploy that drops the JMX listener should page, not quietly disable your memory scaling.

A minimal CDK sketch

The shape of a safe policy in aws-cdk-lib (aws_applicationautoscaling /
aws_ecs), with the two rules baked in — heap is scale-out only on p90,
and requests own scale-in:

const scaling = fargateService.autoScaleTaskCount({
  minCapacity: 10,
  maxCapacity: 200,
});

// Concurrency: the primary signal — thread-pool occupancy scales out AND in.
scaling.scaleToTrackCustomMetric('PoolOccupancy', {
  metric: poolOccupancyPercent, // in-flight / pool-size * 100, from the app
  targetValue: 60,
});

// CPU: the complement for CPU-bound work — also scales both out and in.
scaling.scaleOnCpuUtilization('Cpu', {
  targetUtilizationPercent: 40,
  scaleInCooldown: Duration.minutes(5),
  scaleOutCooldown: Duration.minutes(1),
});

// Heap: SCALE-OUT ONLY, p90, stepped, thresholds below the heap alarm.
scaling.scaleOnMetric('HeapOut', {
  metric: heapUsedAfterGcPercent.with({ statistic: 'p90' }),
  adjustmentType: AdjustmentType.CHANGE_IN_CAPACITY,
  scalingSteps: [
    { lower: 60, change: +5 },
    { lower: 80, change: +10 },
  ],
  cooldown: Duration.minutes(3),
  // no negative step → this policy can never scale in
});
Enter fullscreen mode Exit fullscreen mode

Sizing the pool: what the stress test is for

The spike section set the rule: the
thread pool is the admission gate, and it should saturate before CPU or
memory
under normal load so overload turns into a clean shed rather than a
meltdown. The open question that leaves is a number — how big should the pool
actually be?
That is what a stress test answers.

Drive rising concurrency at a single task and find the "knee" where p99
starts climbing steeply. Set the pool to fill at or just before that knee —
large enough to use the box well (don't fast-fail while CPU sits at 30%), small
enough that a full pool still means healthy latency with CPU and memory
headroom to spare. Too small wastes the hardware; too large removes the clean
admission gate and lets overload leak through to a CPU/GC meltdown.

The honest caveat: a stress test is a starting point, not the truth. The
lab number rarely matches production, for two structural reasons:

  • Different dependencies. Test-environment downstreams have different latency, throttling and failure behaviour than prod — and since thread occupancy is dominated by time spent waiting on downstreams, the right pool size in the lab can be wrong in prod.
  • Different traffic mix. Synthetic data is rarely the real blend of request shapes. A generator hammering one cheap operation looks nothing like the prod mix of small and large payloads, cache hits and misses, cheap and expensive code paths.

Treat the stress-test number as an initial pool size, then validate against
real traffic: watch production p99 versus pool occupancy, confirm the pool
saturates before CPU/memory under real load, and adjust. Whenever you change
the pool size, move the occupancy scale-out threshold with it so the two stay
consistent.

The core lesson

The container sees the box (RSS, aggregate CPU). The JVM sees the truth
(heap-after-GC, thread occupancy). How far the two diverge depends on sizing:
in a well-sized task container memory tracks heap pressure closely, but in an
oversized task a 6 GiB heap can leave the box looking ~35% idle while the JVM is
one GC from an OOM kill.

Either way:

  • Scale out on whichever signal genuinely reflects the active load — concurrency first on a healthy service, then CPU for CPU-bound work, and heap-after-GC for memory pressure.
  • Scale in only on the load-shedding signals (CPU / concurrency), never on heap or container memory.
  • Treat disk as an operational alarm — never a scaling axis.

Get the signal right and the fleet breathes with demand: out on the morning
ramp, back in overnight. Get it wrong and you get a fleet that only knows how to
grow.

Further reading

Found this useful, or scaling a JVM fleet differently? I'd love to hear how in
the comments.

Top comments (0)