DEV Community

Cover image for CPU Quota Does Not Look at Your Average
Mustafa ERBAY
Mustafa ERBAY

Posted on Originally published at mustafaerbay.com.tr

CPU Quota Does Not Look at Your Average

It is very easy to look at a container's CPU graph and conclude the thing is loafing. Average
consumption: a fifth of its quota. Flat line. No alerts.

That same container was stopped 509,912 times over the past ten days.

These two sentences do not contradict each other. The contradiction is in what I was looking at.
Linux's CFS quota mechanism has no interest whatsoever in your average; it cares about 100
millisecond windows. And for years, every time I typed --cpus on a container, I was not saying
"half a core per second" — I was saying "50 ms in every 100 ms, no matter what." Those are not
the same thing. In the resource-hog container guide
I wrote back in May, I handed out the --cpus "0.5" example without a second thought; the advice
in that guide is not wrong, but this measurement showed me it was incomplete.

I swept the fleet and found 16 quota'd cgroups

On VPS3 (Ubuntu, 6.8.0-142-generic, 18 cores) I walked every cgroup that had a quota and read
its cpu.stat. In all 16 cgroups where I had set a quota, cpu.max.burst is 0 and the
nr_bursts counter is 0. I will get to why that surprised me shortly.

The resulting table gave me pause, because the two most heavily throttled services were the two
with the least tolerance for latency. Top five by throttle count:

cgroup quota periods throttled throttle rate throttled time avg. quota use
vpsman-etcd 0.5 CPU 9,297,557 509,912 5.48% 14.24 hours 17.55%
licman-etcd 0.5 CPU 5,195,009 308,177 5.93% 8.45 hours 19.82%
messageman-api-1 1 CPU 3,336,709 7,270 0.22% 18 minutes 3.46%
burncpu-app 2 CPU 125,783 1,725 1.37% 22 minutes 12.24%
messageman-pg-1 1 CPU 201,165 813 0.40% 41 seconds 5.16%

etcd. The consensus ledger. The thing that triggers a leader election when heartbeats are late,
the thing every write waits on. Because its average load is 17% of quota, I had assumed for years
that "half a core is more than enough." Technically true. In practice it stopped etcd half a
million times over ten days.

Cumulative counters tell you about the past, not about today. So I read cpu.stat twice, 120
seconds apart:

vpsman-etcd, 120-second window
  periods     : 1191   (= 119.1 s — the timer almost never goes idle)
  throttled   : 107 times (8.98% of periods)
  throttled time : 12.42 s
  CPU used       : 13.60 s = 22.84% of its quota
Enter fullscreen mode Exit fullscreen mode

One hundred and seven times in two minutes. While using a quarter of its quota. No average-based
dashboard will show you this, because the average genuinely is low — the problem is that the
average is the wrong question.

The window is 100 ms, and there is no average in there

The mechanism is stated plainly in the kernel docs: within each period the group is allocated
quota microseconds of CPU, and once that is exhausted its threads are throttled until the
period refreshes. The default period is 100 ms, the default quota is unconstrained, the minimum
quota and period are 1 ms, and the maximum period is 1 second.

Wondering whether quota carries across periods, I went to the "Caveats" section of the docs —
because I was about to write "it never carries," and I would have been wrong. What it says is
that once a slice is assigned to a CPU it does not expire; if every thread on that CPU
becomes unrunnable, all but 1 ms of the slice may be returned to the global pool. So there is a
small carryover — "typically at most 1 ms per cpu or as defined by min_cfs_rq_runtime." In the
kernel source that constant sits there as 1 * NSEC_PER_MSEC.

The nuance matters, and so does its scale: on an 18-core machine that is a few milliseconds of
residue at best. Rounding error next to an 80 ms peak demand. The same paragraph also names what
this carryover is meant to eliminate: "the propensity to throttle these applications while
simultaneously using less than quota amounts of cpu." The exact phenomenon these counters record,
named in the kernel's own documentation — and still happening on my server.

It is worth doing the arithmetic once. A half-core quota grants 50 ms per period. If your job
needs 80 ms of CPU in one go, it must be split no matter what your average consumption is —
because what is missing is not the budget but its distribution. vpsman-etcd uses 17% of its
quota on average; across 10.76 days of periods it never spent 82% of its entitlement, close to
nine days' worth.

Which is exactly what happened to etcd. For a read-only range request — and for a request this
small, at that:

{"level":"warn","caller":"txn/util.go:93","msg":"apply request took too long",
 "took":"220.804807ms","expected-duration":"100ms",
 "prefix":"read-only range ","request":"key:\"health\" "}
Enter fullscreen mode Exit fullscreen mode

Reading the health key took 220 ms. The work itself takes microseconds. The rest is waiting.

Four ways to misread cpu.stat

I tripped over these counters four times. I could not be confident until I opened the kernel
source (kernel/sched/fair.c, v6.10) and found the lines — and all four run against intuition.

nr_periods is not wall-clock time. The period timer is deactivated when the group goes
idle; the comment at the head of do_sched_cfs_period_timer says so explicitly. So
nr_periods × 100 ms is not elapsed time, it is "time during which there was work." In a lab run
that lasted 124 seconds, the counter showed 750 periods (75 seconds). Inverted, it makes a nice
diagnostic: for vpsman-etcd, 9,297,557 periods works out to 10.76 days against an uptime of
10.80 days — across ten days there was only about an hour in which it had no runnable work.
licman-etcd passes the same test independently: 6.01 days against 6.04.

throttled_usec can exceed wall-clock time. The counter is kept per cgroup but accumulated
separately as each CPU's queue is released: cfs_b->throttled_time += rq_clock(rq) -
cfs_rq->throttled_clock
. If the group is throttled on four cores simultaneously, 100 ms of real
waiting adds 400 ms to the counter. The measurement itself shows this: dividing 51,261 seconds by
509,912 throttle events gives 100.5 ms per event, which is above the period length — and since no
single period can hold more waiting than that, the excess comes from summing across cores. So the
"14.24 hours" in the table is not real waiting, it is an upper bound.

The counters are not hierarchical. The docs cover this in a single sentence, but the
consequence is heavy: those five CFS fields account only for throttling caused by that cgroup's
own bandwidth limit. If the quota lives on a parent slice (kubepods.slice under Kubernetes, a
parent unit under systemd), the container's own cpu.stat looks spotless. In September, when
I wrote 0 to cgroup.pressure
the counter went quiet; this is the same blindness wearing a different file. There is a remedy,
and it was news to me: cpu.stat.local also reports throttling inherited from ancestors, and it
has been there since Linux 6.6. I found the file on VPS3 (6.8); for my containers the two files
hold identical values, meaning the quota really does sit at the leaf. But in a setup that puts
the quota on a parent unit, that is the file to read. Note that cpu.stat.local is also summed
across CPUs, so it does not solve the second trap.

The threshold in an application log is the application's, not yours. etcd's 24-hour log holds
891 "apply request took too long" lines for vpsman-etcd and 12,395 for licman-etcd. For a
moment I got excited that the distribution never dipped below 100 ms — "look, the exact period
boundary!" No. etcd's expected-duration threshold is 100 ms; it never logs anything below that.
The floor belongs to the logger, not the scheduler. For licman-etcd the actual distribution is
a smooth tail from 100 ms out to a second: p50 141 ms, p90 278 ms, p99 601 ms, worst case 996 ms.

The fix has been in the kernel for four years: cpu.max.burst

Banking unused quota and spending it later is not a new idea. Commit f4183717b370
("sched/fair: Introduce the burstable CFS controller") is dated 21 June 2021 and landed in Linux
5.14; the cpu.max.burst interface arrived in the same release. In v5.13 there is no such thing
as cfs_burst.

The logic is ten lines inside __refill_cfs_bandwidth_runtime:

Diagram

(The nr_burst and burst_time in the diagram are the kernel's internal field names; in the
cgroup v2 interface you see the same things as nr_bursts and burst_usec.)

The critical line is the ceiling at the bottom: runtime = min(runtime, quota + burst). The most
you can spend in a single period is quota plus burst. The docs give the permitted range for
cpu.max.burst as [0, $MAX], default 0. So if you set burst equal to quota, you can spend at
most twice the quota in one period, and not a microsecond more.

That ceiling helped me design the experiment — because it tells you in advance what should happen.

The experiment: half a burst buys you nothing

In Docker Desktop's linuxkit VM (6.10.14-linuxkit, cgroup v2) I created my own cgroup and wrote
a deliberately spiky workload: each round, idle for 400 ms, then burn 80 ms of CPU, 250 rounds.
Quota 50000 100000. Average demand is only 33% of the quota. The quantity under measurement is
the wall-clock duration of the work — since the burn loop spends exactly 80 ms of CPU as counted
by CLOCK_PROCESS_CPUTIME_ID, the difference is waiting.

I ran three arms:

quota 50000/100000, work: 400ms idle + 80ms CPU x 250 rounds

burst=0       p50  97.3 ms   throttled 250/250 rounds   throt 4.298 s   nr_bursts   0
burst=25000   p50  98.0 ms   throttled 250/250 rounds   throt 4.416 s   nr_bursts 250  burst 6.250 s
burst=50000   p50  80.0 ms   throttled   0/250 rounds   throt 0     s   nr_bursts 208  burst 3.765 s
Enter fullscreen mode Exit fullscreen mode

The third row was what I expected: with burst equal to quota there is not a single throttle event
and latency settles at 80.0 ms — the pure CPU cost of the work. A 22% latency tax, gone.

The second row is what stopped me. With burst set to half the quota, nr_bursts came out at 250:
the mechanism fired on every single round. burst_usec is 6.250 seconds, that is exactly
25,000 µs per round — the budget spent down to the last drop. And latency did not improve by one
millisecond.

The ceiling formula explains why: 50 ms quota + 25 ms burst = 75 ms, and the job needs 80 ms.
Five milliseconds short. Waiting for a period boundary over five milliseconds produces the same
delay as waiting over fifty. Here a half measure arrives at full cost and zero benefit — while
the counters report back that "burst is working!" Green on the dashboard, waiting for the user.

If it were me, I would take exactly one rule away from this: there is no such thing as enabling
burst "a little." If you do not know the peak your workload demands in one go, you do not know
whether the burst value you picked will do anything at all.

The dial nobody turns: period

Here is what gets lost in the burst conversation: cpu.max is two numbers. Everyone tunes the
first one and leaves the second at its default for four years. The same 0.5 CPU ratio, run at
three different periods with burst disabled:

same work (400ms idle + 80ms CPU), 0.5 CPU ratio throughout, burst=0

5000/10000      (10 ms period)   p50 151.0 ms   throttled 2892/3414 periods
50000/100000    (100 ms period)  p50  96.4 ms   throttled  197/600  periods
250000/500000   (500 ms period)  p50  80.0 ms   throttled    0/194  periods
Enter fullscreen mode Exit fullscreen mode

Shortening the period made latency 57% worse. That may read as counterintuitive — I had
assumed "refreshing more often means waiting less" — but the opposite happens: as the window
narrows, the slice you can take in one go shrinks, and an 80 ms job gets cut into sixteen pieces.
Every wait is short, but there are a lot of waits.

Lengthening the period gave the same result as burst: zero throttling, 80.0 ms. And this dial
works today — you set it with docker run --cpu-period --cpu-quota, it is part of the
container configuration so it survives a restart, and it needs no new tooling. The cost is
symmetric: once you genuinely exhaust the quota, a single stall can now last up to 500 ms, and
the group holds the machine for longer while it does. The documented ceiling is 1 second.

So a solution that needs no burst at all was sitting there, and for two weeks I never looked at
the second number in cpu.max.

Where the chain breaks

As for burst: the reason it is off in all 16 of my cgroups is that nothing in my stack turns it
on. I checked the chain link by link:

  • Kernel: present since v5.14. ✓
  • OCI runtime-spec: a burst field is defined, with the constraint that it must not exceed a positive quota. ✓
  • runc: specconv maps r.CPU.Burst to c.Resources.CpuBurst; the actual write happens in fs2/cpu.go inside opencontainers/cgroups, to the cpu.max.burst file. ✓
  • Podman: settable via --cgroup-conf=cpu.max.burst=50000 — documented behavior, writing to an arbitrary cgroup v2 file as part of the container configuration. (I am taking this one from the docs; I did not measure it, as podman is not installed on my machine.) ✓
  • Docker CLI: absent. On my 27.4.0 install, the word "burst" appears zero times in docker run --help; the current docker run reference lists --cpu-period, --cpu-quota and --cpu-shares, and lists neither burst nor any --cgroup-conf-style passthrough. ✗
  • Compose spec: zero hits in deploy.md. ✗
  • systemd: systemd.resource-control has CPUQuota and CPUQuotaPeriodSec, with no burst equivalent (VPS3 runs systemd 255). ✗
  • Kubernetes: no burst field in the Pod spec. On the CRI side there is a generic cgroup v2 passthrough map called LinuxContainerResources.unified — so the pipe is not technically closed, there is just no user-facing surface for the kubelet to fill it from. ✗

The chain is not broken end to end; it breaks at the Docker CLI and the kubelet's Pod
interface
. If you run podman, the switch is within reach today.

"I will just write it by hand on Docker," I said, and tried. I started a container with
docker run --cpus=0.5, wrote 50000 into its cgroup, verified it, then ran docker restart:

docker --cpus=0.5  -> cpu.max=[50000 100000]  burst=[0]
written by hand    -> burst=[50000]
after restart      -> cpu.max=[50000 100000]  burst=[0]   (same container id)
Enter fullscreen mode Exit fullscreen mode

Docker rewrites cpu.max because that is its configuration; it does not rewrite burst because
it knows no such thing exists. Any value you write by hand evaporates on the first restart. This
is not configuration, it is a poke.

What to ask about your own setup

  1. First ask why the quota is there at all. My etcd instances are squeezed into half a core on an 18-core machine. The quota was put there to stop them starving neighboring projects, but etcd is not the profile that starves anyone. The right answer here may be to drop the quota and give them a share of the contention via cpu.weight — but that is not free: a weight only divides the spoils during contention, it sets no ceiling. A runaway compaction can still take all 18 cores. A quota is a hard wall; a weight is a negotiation.
  2. Consider the period; you have probably never touched it. The measurement above shows that lengthening the period at the same ratio achieves what burst does, and it works in Docker today. Shortening it backfires. Limits: the documented maximum period is 1 second, and a longer period means a longer single stall.
  3. Raising the quota is boring, and it works. The difference from burst is this: doubling the quota also doubles your entitlement on average, whereas burst only gives back entitlement you previously did not spend.
  4. Burst is not free — the docs say so outright. The burst section of the kernel documentation introduces the mechanism and its price in the same sentence: borrowing against a future underrun happens "at the cost of increased interference against the other system users." The consolation is that the tardiness is bounded, and that per the docs the interference stays limited when there are many cgroups or the CPU is under-utilized. Still, do not read it as "protects the neighbors": it protects the average and raises the instantaneous interference.
  5. If you do choose burst, measure peak demand rather than guessing it — and work out up front how to make it stick. If it does not come through Docker, you need something that reapplies it on every restart.
  6. At minimum, put the counter on a dashboard. Even if you do none of the above, reading nr_throttled and throttled_usec is free — or cpu.stat.local if the quota lives on a parent unit. An average CPU graph structurally cannot show you this event.

The last one is the cheapest and the most useful. A pause you do not measure has not been
cancelled, only made invisible. And if you have an appetite for changing the scheduler more
fundamentally, the sched_ext side
is worth a look — though the answer to the quota problem is not there.

What I did not prove

I need to draw a line here, because the story looks too clean and for a while I believed it too
much myself.

In the lab there is causation: same work, same quota, burst the only variable, result 250 throttle
events down to zero. But the lab workload is single-threaded. The docs state that burst is not
transferred between cores; in a multi-threaded group like etcd, per-CPU slice distribution enters
the picture and the behavior may not be identical. What I proved is the single-threaded form of
the mechanism.

In production, all I have is a strong coincidence. I know etcd is slow, I know it is being
throttled, I know the requests that slow down are things like key:"health" with no CPU cost at
all, and I know storage is not the culprit (against 891 and 12,395 apply warnings in 24 hours,
slow fdatasync appears only 1 and 3 times). I wanted to test it with a measurement: I took eight
one-minute readings on licman-etcd and lined up throttle counts against slow-request counts
across the seven deltas.

The result did not support me: Pearson correlation 0.437, n=7. That does not mean "no
relationship" — it is moderately positive but not significant at this sample size, so the
measurement neither supports nor refutes; it is underpowered. The reason is visible inside it:
against an average of 61.2 throttle events per minute there are only 11.3 slow-request warnings.
Most throttles never turn into a warning on the request path — they land on background threads,
and one slow request can span several throttles besides. At one-minute resolution the signal
drowns in noise, and the way to fix that is to enable burst in production and run an A/B, which I
am not going to do on a live consensus ledger.

So what I hold is this: exposure measured, mechanism proven in its single-threaded form,
production causation unproven. Keeping those three apart matters as much as the measurement itself.

What is certain, though, is that the question I had been asking since the day I set that quota was
the wrong one. I was asking "how much of its quota is this service using?" and for ten days I
proudly collected the answer "17%." What I should have asked was "how many times has this service
hit its quota?" The first question measures an average. The second measures what the user is
waiting for.

Official Sources

Top comments (0)