<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jeroen Dirks</title>
    <description>The latest articles on DEV Community by Jeroen Dirks (@jeroen_dirks_29c3ab2f43a7).</description>
    <link>https://dev.to/jeroen_dirks_29c3ab2f43a7</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4173937%2Fe88a398a-3b07-406d-a7e9-7a9540f7f3fd.png</url>
      <title>DEV Community: Jeroen Dirks</title>
      <link>https://dev.to/jeroen_dirks_29c3ab2f43a7</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jeroen_dirks_29c3ab2f43a7"/>
    <language>en</language>
    <item>
      <title>Your Fargate Autoscaler Doesn't Understand Your JVM (by Default): Lessons Learned the Hard Way</title>
      <dc:creator>Jeroen Dirks</dc:creator>
      <pubDate>Fri, 09 Oct 2026 17:35:10 +0000</pubDate>
      <link>https://dev.to/jeroen_dirks_29c3ab2f43a7/your-fargate-autoscaler-doesnt-understand-your-jvm-by-default-lessons-learned-the-hard-way-2f57</link>
      <guid>https://dev.to/jeroen_dirks_29c3ab2f43a7/your-fargate-autoscaler-doesnt-understand-your-jvm-by-default-lessons-learned-the-hard-way-2f57</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://builder.aws.com/content/3KSu7HdOFfv7hWgVE2P4MDmLuAE/your-fargate-autoscaler-doesnt-understand-your-jvm-by-default-lessons-learned-the-hard-way" rel="noopener noreferrer"&gt;AWS Builder Center&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Your autoscaler doesn't understand your JVM
&lt;/h2&gt;

&lt;p&gt;Autoscaling a stateless web service feels like a solved problem: pick a CPU&lt;br&gt;
target, set min and max, walk away. That works right up until the service is a&lt;br&gt;
&lt;strong&gt;JVM&lt;/strong&gt; — and I learned the hard way that it doesn't just underperform, it&lt;br&gt;
fails in ways that look nothing like a capacity problem.&lt;/p&gt;

&lt;p&gt;Out of the box, your autoscaler watches two things: container CPU and container&lt;br&gt;
memory. Neither understands the runtime inside the box. A JVM has its own&lt;br&gt;
memory manager, its own thread model, and its own pathological failure modes,&lt;br&gt;
and the container metrics describe the &lt;em&gt;box&lt;/em&gt;, not the &lt;em&gt;runtime&lt;/em&gt; — so when the&lt;br&gt;
two views disagree, a policy wired to the default signal does something worse&lt;br&gt;
than nothing. The textbook case: it scales out on a signal that never recovers,&lt;br&gt;
and the fleet ratchets up and never comes back down.&lt;/p&gt;

&lt;p&gt;The default signals also go quiet at exactly the wrong moments — &lt;strong&gt;when the&lt;br&gt;
unusual happens&lt;/strong&gt;. A downstream slows down and requests pile up waiting on I/O;&lt;br&gt;
CPU and memory stay calm while the service silently runs out of threads. One&lt;br&gt;
giant request inflates the heap toward an OOM while container memory, in an&lt;br&gt;
oversized task, barely moves. A sudden spike arrives faster than any task can&lt;br&gt;
launch. Each of these is a real load profile that naive, container-metric&lt;br&gt;
scaling gets wrong — and each has a JVM-truthful signal that sees it, if you&lt;br&gt;
wire the autoscaler to understand the runtime.&lt;/p&gt;

&lt;p&gt;This post walks through five load profiles for one Java-on-Fargate service, and&lt;br&gt;
for each asks the only question that matters: &lt;strong&gt;which emitted signal actually&lt;br&gt;
reflects this pressure, which one lies — and how do you make the fleet scale&lt;br&gt;
right when the unusual happens?&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  The task envelope
&lt;/h2&gt;

&lt;p&gt;Every Fargate task is a fixed box. The reference task here is &lt;strong&gt;4 vCPU /&lt;br&gt;
10 GiB memory / 40 GiB ephemeral disk&lt;/strong&gt;, running one JVM with &lt;code&gt;-Xmx6144M&lt;/code&gt;&lt;br&gt;
under G1GC.&lt;/p&gt;

&lt;p&gt;That sizing is deliberate. A 6 GiB heap in a 10 GiB task leaves realistic room&lt;br&gt;
for non-heap (metaspace, thread stacks, direct buffers) and the OS — which&lt;br&gt;
means container-memory utilisation &lt;em&gt;actually tracks&lt;/em&gt; JVM pressure. Put the same&lt;br&gt;
6 GiB heap in an oversized 24 GiB task and the heap could never move the&lt;br&gt;
container-memory needle; the metric would read ~25% while the JVM was one GC&lt;br&gt;
away from death. Hold on to that distinction — it is the crux of the whole&lt;br&gt;
post.&lt;/p&gt;

&lt;p&gt;Autoscaling never changes the box. It changes &lt;strong&gt;how many boxes&lt;/strong&gt; run. To decide&lt;br&gt;
that well you have to know what fills each box and which fills the autoscaler&lt;br&gt;
can actually see.&lt;/p&gt;

&lt;p&gt;Each diagram below shows four stacked bars for one load state: &lt;strong&gt;memory&lt;/strong&gt;&lt;br&gt;
(with the fixed &lt;code&gt;-Xmx&lt;/code&gt; heap band and, nested inside it, the live-set that&lt;br&gt;
survives GC), &lt;strong&gt;worker-thread pool occupancy&lt;/strong&gt;, &lt;strong&gt;CPU&lt;/strong&gt;, and &lt;strong&gt;ephemeral&lt;br&gt;
disk&lt;/strong&gt;. We take the load types roughly in the order a healthy service meets&lt;br&gt;
them — concurrency first, because on a well-tuned system it is the signal that&lt;br&gt;
should drive scaling most of the time.&lt;/p&gt;
&lt;h2&gt;
  
  
  1 · Idle
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foyvzlvr4vbh1snj4essl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foyvzlvr4vbh1snj4essl.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is the middle of the night. Traffic is a trickle — health checks and the odd&lt;br&gt;
request — and the fleet has already &lt;strong&gt;scaled all the way in to its minimum&lt;/strong&gt;,&lt;br&gt;
say &lt;strong&gt;3 tasks&lt;/strong&gt; (you keep a small floor for availability across AZs, not&lt;br&gt;
because the load needs it). The JVM has warmed up and settled: the post-GC heap&lt;br&gt;
live-set is small, only a handful of worker threads are ever busy at once, CPU&lt;br&gt;
idles in the single digits, and disk sits at its baseline (image + build&lt;br&gt;
objects + a little log).&lt;/p&gt;

&lt;p&gt;This is the &lt;strong&gt;reference state&lt;/strong&gt; every other scenario is measured against, and&lt;br&gt;
it is the state a right-sized fleet should return to overnight — sitting&lt;br&gt;
quietly at &lt;code&gt;minCapacity&lt;/code&gt; until demand returns. If a fleet &lt;em&gt;can't&lt;/em&gt; get back here&lt;br&gt;
— if task count stays pinned high while traffic falls — that is almost always a&lt;br&gt;
&lt;strong&gt;scale-in bug&lt;/strong&gt; (a policy stuck "active"), not real demand. Hold that thought;&lt;br&gt;
it comes back in the memory section.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; no scaling action — the fleet is correctly resting at its&lt;br&gt;
&lt;code&gt;minCapacity&lt;/code&gt; floor (~3 tasks). The thing to verify here is the &lt;em&gt;opposite&lt;/em&gt; of&lt;br&gt;
scaling: that nothing is wedging the fleet above this floor overnight.&lt;/p&gt;
&lt;h2&gt;
  
  
  When the spike beats the autoscaler
&lt;/h2&gt;

&lt;p&gt;Before we walk the steady-state load types, deal with the ugly one. Suppose&lt;br&gt;
from that quiet 3-task floor traffic suddenly jumps far faster, and far higher,&lt;br&gt;
than anything forecast — a flash sale, a retry storm, a dependency recovering&lt;br&gt;
and dumping its backlog on you. Autoscaling &lt;em&gt;cannot&lt;/em&gt; save you here: a new&lt;br&gt;
Fargate task takes minutes to launch and pass health checks, and the spike&lt;br&gt;
arrives in seconds. For that window the current tasks are all there is.&lt;/p&gt;

&lt;p&gt;So the behaviour that matters is not "scale fast" — it's &lt;strong&gt;what each individual&lt;br&gt;
task does when it is handed more than it can serve.&lt;/strong&gt; The rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Each task accepts only as many concurrent requests as its &lt;strong&gt;worker-thread&lt;br&gt;
pool&lt;/strong&gt; has slots, and &lt;strong&gt;fast-fails everything beyond that&lt;/strong&gt; — an immediate,&lt;br&gt;
cheap rejection rather than an accepted request left to time out.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That gives you a &lt;strong&gt;brownout, not a blackout&lt;/strong&gt;: the requests inside the pool's&lt;br&gt;
capacity are served at &lt;em&gt;normal latency&lt;/em&gt;, and the overflow is shed instantly so&lt;br&gt;
the client can retry or degrade. The task's throughput of &lt;em&gt;successful&lt;/em&gt; work&lt;br&gt;
stays flat and predictable right through the surge, instead of collapsing as an&lt;br&gt;
unbounded queue drags every request past its deadline and pegs CPU and GC. As&lt;br&gt;
the autoscaler then brings new tasks online, total pool capacity across the&lt;br&gt;
fleet rises and the shed fraction shrinks until the service is fully available&lt;br&gt;
again. &lt;strong&gt;A fixed amount of good work per task, multiplied by a growing task&lt;br&gt;
count&lt;/strong&gt; — that is the whole recovery curve.&lt;/p&gt;

&lt;p&gt;One carve-out is non-negotiable: &lt;strong&gt;a task in overload must keep passing its&lt;br&gt;
load-balancer health checks.&lt;/strong&gt; If an overloaded task starts failing health&lt;br&gt;
checks, the load balancer marks it unhealthy and pulls it out of rotation — so&lt;br&gt;
its share of traffic piles onto the remaining tasks, pushing &lt;em&gt;them&lt;/em&gt; into&lt;br&gt;
overload, and the fleet unravels one task at a time exactly when you need every&lt;br&gt;
task serving. Health-check handling must therefore be &lt;strong&gt;independent of the&lt;br&gt;
request-serving thread pool&lt;/strong&gt; (a dedicated path or a reserved slot), so a task&lt;br&gt;
can be shedding customer load and still truthfully answer "yes, I'm alive and&lt;br&gt;
in service." Shedding load is healthy behaviour; it must not look like death.&lt;/p&gt;

&lt;p&gt;This is why the &lt;strong&gt;size of the worker-thread pool is the single most important&lt;br&gt;
tuning knob&lt;/strong&gt;, and why it should be set so the pool &lt;strong&gt;saturates before CPU or&lt;br&gt;
memory does&lt;/strong&gt; under normal conditions. The pool is then a clean admission gate&lt;br&gt;
at the front door — full pool means instant, cheap rejection — rather than&lt;br&gt;
letting load leak in until CPU or GC melts down with no clean gate at all. How&lt;br&gt;
to actually find that number is a stress-testing problem; we come back to it in&lt;br&gt;
Sizing the pool once the load&lt;br&gt;
types are on the table.&lt;/p&gt;
&lt;h2&gt;
  
  
  2 · Concurrency-loaded
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feypcck82m6wx9uhx89a8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feypcck82m6wx9uhx89a8.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;tart here, because on a healthy, well-tuned service &lt;strong&gt;this is the normal case&lt;/strong&gt;&lt;br&gt;
— the one that should drive most of your scaling decisions. A well-behaved&lt;br&gt;
request-serving JVM spends much of its time waiting on downstreams (a database,&lt;br&gt;
a dependency service, a cache), so the resource that tracks real demand most&lt;br&gt;
faithfully is not CPU or memory but &lt;strong&gt;how many requests are in flight at once&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Picture a downstream slowing down — not failing, just answering in 300 ms&lt;br&gt;
instead of 50. TPS hasn't changed, but each request now &lt;strong&gt;holds its worker&lt;br&gt;
thread far longer&lt;/strong&gt; while it waits on I/O. The threads aren't doing work;&lt;br&gt;
they're parked. So CPU stays moderate and memory barely moves — the two metrics&lt;br&gt;
you would normally scale on both look calm — while the worker-thread pool&lt;br&gt;
quietly fills toward its ceiling. When the last free thread is taken, the next&lt;br&gt;
request is &lt;strong&gt;queued or rejected&lt;/strong&gt;, and latency spikes off a cliff even though&lt;br&gt;
the box looks half-idle.&lt;/p&gt;

&lt;p&gt;This is the case raw &lt;code&gt;CPUUtilization&lt;/code&gt; / &lt;code&gt;MemoryUtilization&lt;/code&gt; completely miss. The&lt;br&gt;
signal that sees it is &lt;strong&gt;thread-pool occupancy %&lt;/strong&gt; (in-flight requests ÷ pool&lt;br&gt;
size). Most Java service frameworks expose an equivalent — a &lt;code&gt;busy-threads&lt;/code&gt;&lt;br&gt;
gauge, an active-connection count, an in-flight-request counter — that you can&lt;br&gt;
publish to CloudWatch. Scale out on it so you add capacity &lt;em&gt;while&lt;/em&gt; threads are&lt;br&gt;
filling, not after they're exhausted. And because occupancy falls cleanly as&lt;br&gt;
load drops, it is also a trustworthy &lt;strong&gt;scale-in&lt;/strong&gt; signal — which is why, on a&lt;br&gt;
well-tuned system, concurrency ends up being the metric the fleet breathes on&lt;br&gt;
day to day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; scale &lt;em&gt;out&lt;/em&gt; (and &lt;em&gt;in&lt;/em&gt;) on thread-pool occupancy — the primary&lt;br&gt;
signal on a healthy service, and the only one that sees thread exhaustion.&lt;/p&gt;
&lt;h2&gt;
  
  
  3 · CPU-loaded
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz00qwbo5v86dud4lmym.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz00qwbo5v86dud4lmym.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now a contrasting case. Some workloads — or some operations within a workload —&lt;br&gt;
are genuinely &lt;strong&gt;CPU-bound&lt;/strong&gt;: each request spends its time on the box (parsing,&lt;br&gt;
resolving, serialising, compressing) rather than waiting on a downstream. For&lt;br&gt;
those, CPU utilisation tracks demand almost linearly and the thread pool may&lt;br&gt;
&lt;em&gt;never&lt;/em&gt; be the thing that saturates first.&lt;/p&gt;

&lt;p&gt;That is fine until CPU gets &lt;em&gt;too&lt;/em&gt; high. Once the vCPUs saturate, runnable&lt;br&gt;
threads queue in the OS run-queue waiting for a core. Nothing is broken, but&lt;br&gt;
every request now spends extra time simply &lt;strong&gt;waiting to be scheduled&lt;/strong&gt;, and&lt;br&gt;
that shows up directly as rising &lt;strong&gt;service latency&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal is to add tasks &lt;em&gt;before&lt;/em&gt; that happens: scale out on CPU utilisation at&lt;br&gt;
a target (e.g. 40%) that sits well below the point where run-queue waiting&lt;br&gt;
begins, so new tasks are in service and taking traffic before latency is ever&lt;br&gt;
touched. Because new Fargate tasks take minutes to launch and pass health&lt;br&gt;
checks, the target has to leave enough headroom to cover that lag on a rising&lt;br&gt;
ramp.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; scale &lt;em&gt;out&lt;/em&gt; on CPU utilisation — and CPU is a trustworthy&lt;br&gt;
scale-&lt;em&gt;in&lt;/em&gt; signal too. Treat it as the complement to concurrency: whichever of&lt;br&gt;
the two is the binding constraint for &lt;em&gt;your&lt;/em&gt; workload leads, and both are safe&lt;br&gt;
to drive scale-in.&lt;/p&gt;
&lt;h2&gt;
  
  
  4 · Memory-loaded
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F16vfknxx7w37z4istwum.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F16vfknxx7w37z4istwum.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A few unusually &lt;strong&gt;large requests&lt;/strong&gt; arrive — a client with a huge payload, or a&lt;br&gt;
caching layer that retains objects for the life of a request. These cost little&lt;br&gt;
CPU and few threads, but they inflate the &lt;strong&gt;heap live-set&lt;/strong&gt;. As live data&lt;br&gt;
climbs toward &lt;code&gt;-Xmx&lt;/code&gt;, G1 has less room to work in, so it collects &lt;strong&gt;more often&lt;br&gt;
and for longer&lt;/strong&gt; — and the post-GC floor (the memory that survives every&lt;br&gt;
collection) keeps rising instead of dropping back.&lt;/p&gt;

&lt;p&gt;That is the danger sign: the JVM is spending an increasing share of its time in&lt;br&gt;
GC (burning CPU to no useful end), and if the live-set reaches &lt;code&gt;-Xmx&lt;/code&gt; the task&lt;br&gt;
takes an &lt;code&gt;OutOfMemoryError&lt;/code&gt; and dies. The honest early-warning signal is&lt;br&gt;
&lt;strong&gt;heap-used-after-GC&lt;/strong&gt; as a percentage of &lt;code&gt;-Xmx&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here is where sizing decides everything. Because this box is &lt;em&gt;well-sized&lt;/em&gt; (6 GiB&lt;br&gt;
heap in a 10 GiB task), container memory util also reads high (~80%) — the two&lt;br&gt;
signals agree, which is exactly what a right-sized task buys you. &lt;strong&gt;Contrast:&lt;/strong&gt;&lt;br&gt;
put this same 6 GiB heap in a 24 GiB task and container util would read only&lt;br&gt;
~35% at the same OOM-imminent moment. Container-RSS memory% would be &lt;strong&gt;blind&lt;/strong&gt;&lt;br&gt;
to the pressure. &lt;em&gt;That&lt;/em&gt; is why heap-after-GC, not container memory, is the&lt;br&gt;
metric to scale on.&lt;/p&gt;

&lt;p&gt;Two non-obvious rules make a heap-based policy safe:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scale-out only.&lt;/strong&gt; A heap-based &lt;em&gt;scale-in&lt;/em&gt; policy stays "active" whenever
warm caches or a sticky live-set keep heap high. Under the AWS rule that a
target the fleet scales &lt;strong&gt;in only when &lt;em&gt;all&lt;/em&gt; policies are inactive&lt;/strong&gt;, a heap
scale-in policy jams scale-in fleet-wide — the exact "scales up, never down"
bug from the idle section. Drive scale-in from CPU / concurrency only.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use p90, not Maximum.&lt;/strong&gt; A single large request pins &lt;em&gt;one&lt;/em&gt; task's heap high.
&lt;code&gt;Maximum&lt;/code&gt; reads the single hottest task and would chase it toward
&lt;code&gt;maxCapacity&lt;/code&gt; even though adding tasks can't cool a lone hot task. &lt;code&gt;p90&lt;/code&gt;
fires only when the top ~10% of tasks are hot — real fleet-wide pressure.
Keep the &lt;em&gt;alarm&lt;/em&gt; on &lt;code&gt;Maximum&lt;/code&gt;; use &lt;code&gt;p90&lt;/code&gt; only for the scaling &lt;em&gt;policy&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; scale &lt;em&gt;out&lt;/em&gt; on heap-used-after-GC p90, scale-out only, at a&lt;br&gt;
threshold below the heap alarm so capacity arrives before a page.&lt;/p&gt;
&lt;h2&gt;
  
  
  5 · Disk-loaded
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjkyijlsoi1pfu0hrnx6f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjkyijlsoi1pfu0hrnx6f.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This one doesn't follow traffic at all. Over &lt;strong&gt;days&lt;/strong&gt;, ephemeral disk creeps&lt;br&gt;
upward on every task in the fleet at the same slow rate, regardless of load.&lt;/p&gt;

&lt;p&gt;The usual cause is &lt;strong&gt;log files that never get cleaned up&lt;/strong&gt; — rotated logs that&lt;br&gt;
should be deleted are piling up because whatever is supposed to reap them isn't&lt;br&gt;
running — or, less often, repeated &lt;strong&gt;core dumps&lt;/strong&gt; from crashing tasks. CPU,&lt;br&gt;
memory and threads all read completely healthy; only disk is in trouble.&lt;/p&gt;

&lt;p&gt;This is the one scenario where &lt;strong&gt;the answer is not autoscaling&lt;/strong&gt;. Adding tasks&lt;br&gt;
just starts more tasks with the same broken reaper that will fill at the same&lt;br&gt;
rate; removing tasks does nothing. Treat disk as an &lt;strong&gt;operational alarm&lt;/strong&gt; that&lt;br&gt;
pages a human to fix the reaper or raise ephemeral storage — never as a scaling&lt;br&gt;
axis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; do &lt;em&gt;not&lt;/em&gt; autoscale on disk. Alarm, page, fix the reaper.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A note on the diagrams: they hold the OS/agent and non-heap bands constant&lt;br&gt;
across scenarios for clarity. In reality those tick up a little under load&lt;br&gt;
(the ECS agent, log shipping, OS bookkeeping, and thread stacks all do&lt;br&gt;
slightly more work as volume rises), but the movement is small relative to&lt;br&gt;
the JVM's own footprint and well within the fixed headroom. Scale on the&lt;br&gt;
JVM-truthful signals; treat the non-JVM drift as background noise.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  What the system tells you
&lt;/h2&gt;

&lt;p&gt;Fargate and the JVM emit two different classes of signal. The whole art of&lt;br&gt;
scaling a JVM service is knowing which is truthful for which load.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What it truly measures&lt;/th&gt;
&lt;th&gt;Blind spot&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CPUUtilization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;ECS container&lt;/td&gt;
&lt;td&gt;Aggregate vCPU use across the task&lt;/td&gt;
&lt;td&gt;Sees I/O-blocked concurrency as "calm"; GC inflates it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MemoryUtilization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;ECS container (RSS)&lt;/td&gt;
&lt;td&gt;Whole-container resident memory&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;Only as good as the sizing&lt;/em&gt; — tracks pressure in a well-sized box, goes blind in an oversized one. Prefer heap-after-GC.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Thread-pool occupancy %&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;App framework (in-flight ÷ pool size)&lt;/td&gt;
&lt;td&gt;Real concurrency saturation&lt;/td&gt;
&lt;td&gt;Must be emitted by the app; must stay below the pool ceiling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;In-flight request count&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;App framework&lt;/td&gt;
&lt;td&gt;Raw count of concurrent requests&lt;/td&gt;
&lt;td&gt;Raw count is meaningless across a large fleet — prefer the normalised occupancy %&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;HeapUsedAfterGC %&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;JVM via JMX → CloudWatch&lt;/td&gt;
&lt;td&gt;Post-GC live-set as % of &lt;code&gt;-Xmx&lt;/code&gt; — the truthful memory-pressure signal&lt;/td&gt;
&lt;td&gt;Not emitted by default; sawtooth, so publish the post-collection value and use p90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ephemeral disk used %&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Host/system metrics&lt;/td&gt;
&lt;td&gt;Disk fill&lt;/td&gt;
&lt;td&gt;An ops failure, not a capacity signal — alarm, don't scale&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  When to add or shed tasks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Scale OUT&lt;/strong&gt; if &lt;em&gt;any&lt;/em&gt; of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Thread-pool occupancy high (approaching the pool ceiling) — the primary
signal on a healthy service, and the one CPU/memory miss.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;CPUUtilization&lt;/code&gt; &amp;gt; target (e.g. 40%) — the CPU-bound case.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;HeapUsedAfterGC&lt;/code&gt; p90 &amp;gt; ~60% then ~80% (stepped) — the memory case,
&lt;strong&gt;scale-out only&lt;/strong&gt;, thresholds set &lt;em&gt;below&lt;/em&gt; the heap alarm.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Scale IN&lt;/strong&gt; only when &lt;em&gt;all&lt;/em&gt; scale-in-eligible signals are low — driven by&lt;br&gt;
thread-pool occupancy and/or &lt;code&gt;CPUUtilization&lt;/code&gt; falling, with conservative&lt;br&gt;
cooldowns (e.g. one hour). &lt;strong&gt;Never&lt;/strong&gt; let heap or container-memory drive&lt;br&gt;
scale-in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never autoscale on:&lt;/strong&gt; container &lt;code&gt;MemoryUtilization&lt;/code&gt; as a JVM-heap proxy (drop&lt;br&gt;
it, or set its target well above resting RSS), or disk (alarm + page + fix the&lt;br&gt;
reaper).&lt;/p&gt;
&lt;h2&gt;
  
  
  Emitting heap-after-GC
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;HeapUsedAfterGC&lt;/code&gt; is the one signal in that table that isn't free — the JVM&lt;br&gt;
doesn't publish it to CloudWatch for you. The standard, framework-agnostic way&lt;br&gt;
to capture it uses only &lt;code&gt;java.lang.management&lt;/code&gt; / &lt;code&gt;com.sun.management&lt;/code&gt; from the&lt;br&gt;
JDK:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Get the GC beans with
&lt;a href="https://docs.oracle.com/en/java/javase/11/docs/api/java.management/java/lang/management/ManagementFactory.html#getGarbageCollectorMXBeans()" rel="noopener noreferrer"&gt;&lt;code&gt;ManagementFactory.getGarbageCollectorMXBeans()&lt;/code&gt;&lt;/a&gt;,
cast each to a
&lt;a href="https://docs.oracle.com/en/java/javase/11/docs/api/java.management/javax/management/NotificationEmitter.html" rel="noopener noreferrer"&gt;&lt;code&gt;NotificationEmitter&lt;/code&gt;&lt;/a&gt;,
and register a listener.&lt;/li&gt;
&lt;li&gt;On each notification, convert the payload with
&lt;a href="https://docs.oracle.com/en/java/javase/11/docs/api/jdk.management/com/sun/management/GarbageCollectionNotificationInfo.html" rel="noopener noreferrer"&gt;&lt;code&gt;GarbageCollectionNotificationInfo.from(cd)&lt;/code&gt;&lt;/a&gt;
and read
&lt;a href="https://docs.oracle.com/en/java/javase/11/docs/api/jdk.management/com/sun/management/GcInfo.html#getMemoryUsageAfterGc()" rel="noopener noreferrer"&gt;&lt;code&gt;GcInfo.getMemoryUsageAfterGc()&lt;/code&gt;&lt;/a&gt;,
which returns a &lt;code&gt;Map&amp;lt;String, MemoryUsage&amp;gt;&lt;/code&gt; of each pool's usage &lt;em&gt;after&lt;/em&gt; the
collection.&lt;/li&gt;
&lt;li&gt;Sum the heap pools (e.g. the old generation and any survivor/eden pools —
skip the non-heap pools like &lt;code&gt;Metaspace&lt;/code&gt;) and divide by &lt;code&gt;-Xmx&lt;/code&gt; to get the
post-GC heap-used percentage. Publish that number as a custom CloudWatch
metric on your preferred emission interval.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That &lt;code&gt;GarbageCollectionNotificationInfo&lt;/code&gt; + &lt;code&gt;GcInfo&lt;/code&gt; pattern is the canonical&lt;br&gt;
JDK recipe for "how full was the heap &lt;em&gt;after&lt;/em&gt; the last collection"; it is the&lt;br&gt;
same mechanism tools like Micrometer's JVM GC metrics build on, so if you&lt;br&gt;
already run Micrometer you likely have the raw signal and only need to shape and&lt;br&gt;
export the percentage. Many Java service frameworks ship a built-in equivalent.&lt;/p&gt;

&lt;p&gt;Two cheap safeguards are worth adding so the policy never silently no-ops:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Verify the metric exists before pointing a policy at it.&lt;/strong&gt; Run
&lt;code&gt;aws cloudwatch list-metrics&lt;/code&gt; against the target account; a policy aimed at
an absent metric sits in &lt;code&gt;INSUFFICIENT_DATA&lt;/code&gt; forever — worse than no policy,
because it looks configured.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alarm on the metric's absence.&lt;/strong&gt; A deploy that drops the JMX listener
should page, not quietly disable your memory scaling.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  A minimal CDK sketch
&lt;/h2&gt;

&lt;p&gt;The shape of a safe policy in &lt;code&gt;aws-cdk-lib&lt;/code&gt; (&lt;code&gt;aws_applicationautoscaling&lt;/code&gt; /&lt;br&gt;
&lt;code&gt;aws_ecs&lt;/code&gt;), with the two rules baked in — heap is &lt;strong&gt;scale-out only&lt;/strong&gt; on &lt;strong&gt;p90&lt;/strong&gt;,&lt;br&gt;
and requests own scale-in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;scaling&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;fargateService&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;autoScaleTaskCount&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;minCapacity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;maxCapacity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Concurrency: the primary signal — thread-pool occupancy scales out AND in.&lt;/span&gt;
&lt;span class="nx"&gt;scaling&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scaleToTrackCustomMetric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PoolOccupancy&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;poolOccupancyPercent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// in-flight / pool-size * 100, from the app&lt;/span&gt;
  &lt;span class="na"&gt;targetValue&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// CPU: the complement for CPU-bound work — also scales both out and in.&lt;/span&gt;
&lt;span class="nx"&gt;scaling&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scaleOnCpuUtilization&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Cpu&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;targetUtilizationPercent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;scaleInCooldown&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Duration&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;minutes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;scaleOutCooldown&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Duration&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;minutes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Heap: SCALE-OUT ONLY, p90, stepped, thresholds below the heap alarm.&lt;/span&gt;
&lt;span class="nx"&gt;scaling&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scaleOnMetric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;HeapOut&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;metric&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;heapUsedAfterGcPercent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;with&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;statistic&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;p90&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="na"&gt;adjustmentType&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;AdjustmentType&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CHANGE_IN_CAPACITY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;scalingSteps&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;change&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;change&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;cooldown&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Duration&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;minutes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="c1"&gt;// no negative step → this policy can never scale in&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Sizing the pool: what the stress test is for
&lt;/h2&gt;

&lt;p&gt;The spike section set the rule: the&lt;br&gt;
thread pool is the admission gate, and it should &lt;strong&gt;saturate before CPU or&lt;br&gt;
memory&lt;/strong&gt; under normal load so overload turns into a clean shed rather than a&lt;br&gt;
meltdown. The open question that leaves is a number — &lt;em&gt;how big should the pool&lt;br&gt;
actually be?&lt;/em&gt; That is what a stress test answers.&lt;/p&gt;

&lt;p&gt;Drive rising concurrency at a single task and find the &lt;strong&gt;"knee"&lt;/strong&gt; where p99&lt;br&gt;
starts climbing steeply. Set the pool to fill at or just before that knee —&lt;br&gt;
large enough to use the box well (don't fast-fail while CPU sits at 30%), small&lt;br&gt;
enough that a full pool still means healthy latency with CPU and memory&lt;br&gt;
headroom to spare. Too small wastes the hardware; too large removes the clean&lt;br&gt;
admission gate and lets overload leak through to a CPU/GC meltdown.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The honest caveat: a stress test is a starting point, not the truth.&lt;/strong&gt; The&lt;br&gt;
lab number rarely matches production, for two structural reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Different dependencies.&lt;/strong&gt; Test-environment downstreams have different
latency, throttling and failure behaviour than prod — and since thread
occupancy is dominated by &lt;em&gt;time spent waiting on downstreams&lt;/em&gt;, the right pool
size in the lab can be wrong in prod.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Different traffic mix.&lt;/strong&gt; Synthetic data is rarely the real blend of request
shapes. A generator hammering one cheap operation looks nothing like the prod
mix of small and large payloads, cache hits and misses, cheap and expensive
code paths.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Treat the stress-test number as an &lt;em&gt;initial&lt;/em&gt; pool size, then validate against&lt;br&gt;
real traffic: watch production p99 versus pool occupancy, confirm the pool&lt;br&gt;
saturates before CPU/memory under real load, and adjust. Whenever you change&lt;br&gt;
the pool size, move the occupancy scale-out threshold with it so the two stay&lt;br&gt;
consistent.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core lesson
&lt;/h2&gt;

&lt;p&gt;The container sees &lt;em&gt;the box&lt;/em&gt; (RSS, aggregate CPU). The JVM sees &lt;em&gt;the truth&lt;/em&gt;&lt;br&gt;
(heap-after-GC, thread occupancy). How far the two diverge depends on sizing:&lt;br&gt;
in a well-sized task container memory tracks heap pressure closely, but in an&lt;br&gt;
oversized task a 6 GiB heap can leave the box looking ~35% idle while the JVM is&lt;br&gt;
one GC from an OOM kill.&lt;/p&gt;

&lt;p&gt;Either way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Scale out&lt;/strong&gt; on whichever signal genuinely reflects the active load —
concurrency first on a healthy service, then CPU for CPU-bound work, and
heap-after-GC for memory pressure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scale in&lt;/strong&gt; only on the load-shedding signals (CPU / concurrency), never on
heap or container memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat disk as an operational alarm&lt;/strong&gt; — never a scaling axis.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Get the signal right and the fleet breathes with demand: out on the morning&lt;br&gt;
ramp, back in overnight. Get it wrong and you get a fleet that only knows how to&lt;br&gt;
grow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.aws.amazon.com/AmazonECS/latest/developerguide/capacity-autoscaling-best-practice.html" rel="noopener noreferrer"&gt;Optimizing Amazon ECS service auto scaling&lt;/a&gt; — the official "scale out if any policy is active, scale in only when all are inactive" rule this post leans on.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://aws.amazon.com/blogs/containers/jvm-memory-cpu-and-classpath-best-practices-for-java-containers-on-aws/" rel="noopener noreferrer"&gt;JVM memory, CPU, and classpath best practices for Java containers on AWS&lt;/a&gt; — the AWS Containers Blog companion on JVM-in-cgroup sizing.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/npayyappilly/request-rate-based-autoscaling-why-cpu-metrics-lie-and-how-to-fix-them-pie"&gt;Request-Rate-Based Autoscaling: Why CPU Metrics Lie&lt;/a&gt; — a complementary take on the concurrency case.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://cspinetta.substack.com/p/load-shedding-autoscaling-overload" rel="noopener noreferrer"&gt;Autoscaling Won't Save You from Overload&lt;/a&gt; — the load-shedding / brownout argument in depth.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Found this useful, or scaling a JVM fleet differently? I'd love to hear how in&lt;br&gt;
the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>fargate</category>
      <category>autoscaling</category>
      <category>java</category>
      <category>jvm</category>
    </item>
  </channel>
</rss>
