DEV Community

Cover image for cgroup.pressure = 0: 86 ms of Stall Went Unrecorded
Mustafa ERBAY
Mustafa ERBAY

Posted on Originally published at mustafaerbay.com.tr

cgroup.pressure = 0: 86 ms of Stall Went Unrecorded

PSI — Pressure Stall Information — is my favourite kernel interface of recent years. Unlike load average, it does not ask "how many jobs are queued" but "how long did somebody wait for a resource". Read memory.pressure and you can see a container genuinely struggling, at microsecond resolution. It went onto my Grafana boards, into my alerts, and the systemd-oomd runbook on this blog is written around these files.

It turns out this interface also has an off switch. It has been sitting in every cgroup directory since Linux 6.1, it is called cgroup.pressure, and its default is 1.

So I set up a lab and turned it off. Then, inside a cgroup capped at 64 MiB, a 512 MiB file got read twelve times — three runs, four reads each; a workload that forces the kernel to fight page reclaim, and that produces pressure by definition. During that period the parent cgroup recorded 86,399 microseconds of stall. When I opened the ledger of the leaf cgroup where the stall actually happened, the counter had not budged: the difference was exactly 0.

Not zero pressure. No pressure. Both are read from the same file, as the same number.

This article chases that gap. At the end of it there are three things: files that disappear, a counter that never lies but goes quiet, and a poll() flag that cannot tell an alarm apart from a teardown.

Where the switch came from

cgroup.pressure arrived in Linux 6.1. I confirmed this against the tags rather than trusting my memory: it does not appear at all in Documentation/admin-guide/cgroup-v2.rst at v6.0, and has a full section at v6.1. Likewise kernel/cgroup/cgroup.c gains the psi_files[] array and the show/hide code in v6.1.

The rationale sentence in the documentation is unambiguous:

PSI accounts stalls for each cgroup separately and aggregates it at each level of the hierarchy. This may cause non-negligible overhead for some workloads when under deep level of the hierarchy.

So the switch was designed as a performance escape hatch: turn accounting off in non-leaf cgroups, escape the cost of the hierarchy walk. The documentation stresses one more thing, and it drives most of the rest of this article:

This control attribute is not hierarchical, so disable or enable PSI accounting in a cgroup does not affect PSI accounting in descendants and doesn't need pass enablement via ancestors from root.

That sentence says enablement does not propagate. It says nothing about what happens to the recording. That gap turned out to be the most surprising part of this article.

The lab

The measurements come from Docker Desktop's own linuxkit virtual machine, writing to the host cgroup tree from inside a --privileged container. Kernel 6.10.14-linuxkit, pure cgroup v2. My reason for picking that VM is simple: when poking at global kernel switches I want a disposable machine. Running the same experiment on a live server would make that server's monitoring part of the experiment.

To generate pressure, a cgroup with memory.max set to 64 MiB reads a 512 MiB file. My first attempt produced nothing, and the reason was instructive: the file had been written with dd in the container's own cgroup, so the page cache was charged there. Reading it from the leaf was all cache hits — pgscan zero, pressure zero. Calling drop_caches before the load changed the picture: pgscan 377,152, workingset_refault_file 189, and memory.pressure started talking for the first time.

Prove that the thing you are trying to measure actually happens. When you see a zero, the first suspect is not the subject but the experiment.

Finding 1: the file is not zeroed, it disappears

My expectation on writing 0 to cgroup.pressure was that the files would stay put and the values would freeze. Here is what happens instead:

$ ls /sys/fs/cgroup/psilab | grep pressure
cgroup.pressure
cpu.pressure
io.pressure
memory.pressure

$ echo 0 > /sys/fs/cgroup/psilab/cgroup.pressure
$ ls /sys/fs/cgroup/psilab | grep pressure
cgroup.pressure

$ cat /sys/fs/cgroup/psilab/memory.pressure
cat: can't open '...': No such file or directory
Enter fullscreen mode Exit fullscreen mode

Three files are removed from the directory. (Four if your kernel also has irq.pressure; the kernel loop runs over NR_PSI_RESOURCES. This machine does not have it.) There is a one-line counterpart to this in the kernel source, in cgroup_pressure_write() in kernel/cgroup/cgroup.c:

    psi = cgroup_psi(cgrp);
    if (psi->enabled != enable) {
        int i;

        /* show or hide {cpu,memory,io,irq}.pressure files */
        for (i = 0; i < NR_PSI_RESOURCES; i++)
            cgroup_file_show(&cgrp->psi_files[i], enable);

        psi->enabled = enable;
        if (enable)
            psi_cgroup_restart(psi);
    }
Enter fullscreen mode Exit fullscreen mode

cgroup_file_show() makes the kernfs node visible or invisible. That has three separate consequences on the monitoring side, and each one got its own test:

  • A fresh open() while disabled → ENOENT. An agent that restarts gets a clean "no such file" error. Honest behaviour.
  • Reading from a previously open fd → ENODEV (errno 19). A collector holding the file open does not quietly read zeros; it gets a hard error. Also honest behaviour.
  • A registered PSI trigger → a different story, coming up shortly.

Validation matches the source exactly: the kernel rejects writing 2 or -1 (enable < 0 || enable > 1 → -ERANGE), and x fails at the kstrtoint() level. The file is fastidious about the two values it accepts. The real question is what those two values mean.

Finding 2: not hierarchical — but not in the direction you expect

When the documentation says "not hierarchical", it means the setting does not propagate downwards, and that is true: turning a cgroup off leaves its child's cgroup.pressure at 1, with all three pressure files still in place.

What I actually wanted to know was the opposite. If a node is turned off, do the stalls that happen there keep being written into its ancestors' ledgers? Three configurations, same workload, three runs each. The numbers are one run's increase in some ... total=, in microseconds:

Configuration leaf intermediate root cgroup /proc/pressure*
A — all enabled 27,253 27,249 26,156 26,156
B — intermediate disabled 22,416 no file 15,724 15,724
C — leaf disabled no file 23,832 22,485 22,485

* The last two columns are not two independent measurements: the root cgroup and /proc/pressure read the same psi_group object — which is what the next section is about. The numbers are one run's increase; across three runs the leaf-to-root gap in row B moved between 4% and 30%, which is within-run noise.

Row C is the heart of it. Accounting is off in the cgroup where the stall physically happens, yet both the intermediate node and the root see the event perfectly well. "Disabled" does not mean "these stalls are no longer recorded"; it means "this cgroup no longer keeps its own ledger". The recording continues throughout the rest of the hierarchy.

The same switch at the root is a different switch

The first draft of this article had a checklist item saying "disabling the root does not affect /proc, they are two separate ledgers." It was wrong. When writing it I had put a 0 into /sys/fs/cgroup/cgroup.pressure, then read /proc/pressure/memory once, seen a total greater than zero, and moved on thinking "still counting". What I had in front of me was a frozen number; it proved the file was still readable, not that it was still counting.

Measuring the delta flipped the result:

== root PSI ENABLED ==
    /proc delta=32614 us | leaf delta=24102 us
    /proc delta=35606 us | leaf delta=36182 us
== root PSI DISABLED ==
    /proc/pressure/memory still readable: some ... total=2441532
    /proc delta=0 us     | leaf delta=39605 us
    /proc delta=0 us     | leaf delta=33862 us
== root PSI RE-ENABLED ==
    /proc delta=25133 us | leaf delta=34557 us
Enter fullscreen mode Exit fullscreen mode

While the leaf recorded 33,000–39,000 microseconds of stall, /proc/pressure/memory advanced by exactly zero. The reason is a single line in include/linux/psi.h:

static inline struct psi_group *cgroup_psi(struct cgroup *cgrp)
{
    return cgroup_ino(cgrp) == 1 ? &psi_system : cgrp->psi;
}
Enter fullscreen mode Exit fullscreen mode

The root cgroup's inode number is 1. So writing to cgroup.pressure at the root does not touch one cgroup's ledger: it touches psi_system, the very object the /proc/pressure/* files read. This is not a cgroup switch but a machine-wide PSI kill switch — the runtime counterpart of the psi=0 boot parameter.

And that makes it the purest example of this article's thesis. The three files under /sys/fs/cgroup do disappear at the root too, so that noise is still there. But /proc/pressure/* stays put: it opens, it reads, and it hands you a frozen total with zero averages. No error code, no missing file. That is exactly why the first draft had it wrong: one reading, and a file that looked like it was doing its job.

Diagram

That is good news for anyone using the switch the way the documentation suggests: disabling intermediate nodes does not corrupt the table above. If your monitoring reads the leaves' own files, you are fine too.

The trouble starts when somebody is looking at the node you turned off.

Finding 3: the counter is not lying, it is silent

The canonical run went like this. First a warm-up: the leaf's total reads 39,362 µs, the cgroup above it 39,346 µs. Then cgroup.pressure goes to 0 on the leaf and the same workload runs three more times. The parent climbs to 125,745 — recording 86,399 µs of stall in between. Re-enabling the leaf:

leaf total = 39362 us   (before disabling: 39362 us  -> delta: 0 us)
some avg10=0.18 avg60=0.03 avg300=0.00 total=39362
Enter fullscreen mode Exit fullscreen mode

The counter picks up exactly where it left off. The kernel source explains this too: psi_group_change() has an early-return branch when accounting is off, and the comment at the top of psi_cgroup_restart() states the intent plainly — on re-enable, groupc->state_start is restarted from now. The disabled interval is not billed retroactively.

I had an assumption here, and the measurement killed it. My guess was that the averages would freeze as well — that avg10, avg60 and avg300 would hold at their values from the moment of shutdown. So I ran a controlled A/B: same workload, then 40 seconds of waiting. Once with PSI left on, once with those 40 seconds spent disabled and the switch flipped back at the end.

right after the load 40 seconds later
Control (always on, idle) avg10=2.15 avg60=0.84 avg300=0.20 avg10=0.04 avg60=0.44 avg300=0.18
Experiment (40 s disabled) avg10=3.46 avg60=1.65 avg300=0.50 avg10=0.06 avg60=0.85 avg300=0.44

This is a single pair of runs, and the starting points do not match exactly (2.15 against 3.46), so what I am comparing is the rate of decay rather than the absolute numbers. The result is clear regardless: the averages do not freeze. They decay — along almost exactly the same curve as a healthy idle 40 seconds: avg10 falls to 0.04 in the control and 0.06 in the experiment. Yet the machine was not idle during the experiment window; the leaf was wrestling with page reclaim the whole time.

This is where I want to state the thesis, because to my mind it is the one real trap in this interface:

A disabled window is indistinguishable from a perfectly healthy one. The monotonic counter has not advanced, and the averages have decayed as though there had been no pressure at all. A collector taking deltas will say "no pressure in this interval". An alert watching the averages stays quiet. Neither of them returns a wrong answer — both return the right answer to the wrong question.

This is also where the noisiness of the disappearing file turns out to be a virtue. ENOENT and ENODEV wake you up. But if you are on a path that never opens the file — reading the output of an agent that collects these values for you, say — or if somebody touched the switch in the middle of your sampling interval, you get no warning at all. Black boxes spend their worst nights without telling anyone.

Finding 4: teardown looks just like an alarm

PSI's real strength is not polling but triggers. Write some 50000 1000000 to memory.pressure and the kernel says "wake me if there are 50 ms of stall in a 1-second window", then wakes you via poll() with POLLPRI.

A note in passing: systemd-oomd does not take this route. Its man page calls it a service that "uses cgroups-v2 and pressure stall information (PSI)", but there is no trigger registration in its source (src/oom/oomd-manager.c); it wakes on its own timer at MEM_PRESSURE_INTERVAL_USEC (one second) and reads avg10 — and it reads the full line's avg10, not the some line this article's lab works with. That drops oomd straight into the third finding: on a disabled cgroup, avg10 decays to a healthy value and oomd sees nothing out of the ordinary. The quietest victim of the off switch may be the very service written for pressure.

So I registered a trigger and then pulled PSI out from under it. The result:

poll(300ms) before disabling = []
cgroup.pressure=0 written
poll(1000ms) after disabling = [(3, 10)]
event flags: ['POLLPRI', 'POLLERR']
read from trigger fd: OSError errno=19 (ENODEV)
Enter fullscreen mode Exit fullscreen mode

revents is 10: POLLPRI (2) and POLLERR (8) set together. POLLPRI is precisely the flag that arrives when a threshold is crossed. To a loop that checks only for that flag, this reads as "pressure threshold exceeded". What actually happened is the opposite: the measurement source is gone.

The kernel documentation already spells out the correct behaviour — for POLLERR it says "event source is gone". But because the flag arrives alongside POLLPRI, a badly written loop has to ask about POLLERR specifically to notice.

The price tag got measured too. POLLERR is level-triggered, so it stays ready until the file is closed:

PSI enabled,  poll returns in 200 ms: 0
PSI disabled, poll returns in 200 ms: 444,247
Enter fullscreen mode Exit fullscreen mode

(Method: a single fd, poll() with a timeout of 0, spinning in an empty loop for 200 ms; the count is the number of non-empty poll() returns.)

Four hundred and forty-four thousand. In two hundred milliseconds. A monitoring agent that only looks at POLLPRI will either emit 444,000 phantom pressure events in that window or burn a core spinning for nothing. Probably both.

The right order is this: if revents & POLLERR is set, close the fd and try reopening the file; if that gives you ENOENT, PSI has been disabled on this cgroup — and you should report that as a loss of monitoring event, not a pressure event. The two must not land on the same graph.

So is there really any overhead

The switch exists for performance reasons. I could not confirm them.

A chain of 20 cgroups below the root ran 200,000 pipe round trips between two processes at the very bottom — each round trip a context switch, which is exactly the path where psi_group_change() walks the chain to the top. In the disabled rounds all 20 of those cgroups were turned off, leaving the root alone for the reason above. To cancel out drift the rounds alternated: seven rounds, each with one enabled and one disabled run, 14 runs in total.

Medians: 9,319 ms with PSI on, 9,197 ms with it off. A difference of 1.3%. But the enabled runs span 9,037–9,816 ms and the disabled ones 8,910–9,622 ms, so the within-condition spread is roughly six times the gap between the medians. On top of that, the disabled side was slower in three of the seven rounds. That is a result you cannot tell apart from a coin flip.

Do not read this as "there is no overhead"; read it as "I could not resolve it". There are two plausible reasons. First, this is a virtual machine on macOS, so the noise floor is high. Second, and more interesting: disabling does not remove the hierarchy walk. psi_group_change() still traverses the chain from end to end, hitting if (!group->enabled) at each node and returning early. What gets saved is test_states() and record_times(), not the loop itself. If the true cost of a deep hierarchy is the loop, this switch cannot remove it anyway.

If you are after a real gain, measure it on your own workload and your own kernel. The phrase "non-negligible" in the documentation is a warning, not a measurement result.

A checklist before you touch it

  • Count first. find /sys/fs/cgroup -name cgroup.pressure -exec grep -l '^0$' {} + gives you every cgroup that is disabled today. The default is 1, so a non-empty result means somebody, or some tool, wrote that.
  • Know who can write it. In an ordinary Docker container /sys/fs/cgroup is mounted ro and a write from inside fails with "Read-only file system". So whoever flips this switch is either root on the host, a privileged container, or a manager that delegated the cgroup to you. It does not change on its own; it can change if your automation is one of those three.
  • If you are going to disable, disable an intermediate node, not a leaf. That is the documentation's intent. Disabling leaves deletes the only local record at the place where the workload actually lives.
  • Put the node you disabled on the dashboard. If your monitoring query targets a cgroup by reading memory.pressure, collect that cgroup's cgroup.pressure value as a separate metric too. Only that distinction separates zero pressure from no measurement.
  • Always ask about POLLERR in your trigger loop. Code that watches only POLLPRI confuses a vanished measurement with an alarm — up to 444,000 times per 200 ms.
  • Do not swallow ENODEV and ENOENT silently. Most collectors skip an unreadable file and move on. What gets skipped here is a cgroup's entire pressure history.
  • Never disable at the root. Writing 0 to /sys/fs/cgroup/cgroup.pressure does not disable one cgroup, it disables psi_system: the /proc/pressure/* files stay in place but stop counting. Every PSI consumer on the machine goes blind at once.
  • Confirm PSI is on at all first. Many distribution kernels are built with CONFIG_PSI_DEFAULT_DISABLED=y; without the psi=1 boot parameter the pressure files are never created. An empty result from the "count first" item may not mean "nothing is disabled" — check cat /proc/pressure/memory to confirm PSI exists.
  • Before measuring, prove your workload really produces pressure. If pgscan and workingset_refault_file are zero, your experiment is broken, not the mechanism.
  • Reverting is free; the history does not come back. echo 1 restores the files instantly, but the disabled interval is permanently lost.

Conclusion

I am used to settings on Linux that quietly have no effect; hunting down who actually listens to ionice on this blog turned up something similar. cgroup.pressure sets the opposite trap: the setting works perfectly, does exactly what the documentation says, and even rejects invalid values properly. The flaw is not in the setting but in how its disabled state looks.

Because a measurement interface being switched off should not take on the same shape as the absence of the thing being measured. If total has not advanced there are two possibilities — no stall occurred, or nobody was watching — and this interface tells both stories with the same number. The disappearing file is actually good design: it is loud, noticeable, and gives you ENODEV. But that noise only reaches the code that opens the file directly. For anyone sampling that cgroup's total as a delta, alerting on its avg10, or unlucky enough that the switch moved mid-interval, a disabled window looks exactly like a perfectly healthy one. Whoever pulls the aggregate from above has a different problem: they still see the pressure, but they can no longer see which child is living it.

So I have widened my own rule by one clause. You know a metric is healthy not because its number looks good, but because you can separately confirm that the mechanism producing that number was switched on at the time. If the pressure is zero, ask this first: was it really zero, or was the ledger closed?

Official Sources

Top comments (0)