The sched(7) man page has said the same thing for years: when autogrouping is on, a process's nice value only matters within its own session, and the split between sessions is governed by the nice value you write into /proc/PID/autogroup. A tidy sentence. Easy to memorise, easy to repeat, hard to check.
This morning I started a CPU burner in each of two separate sessions on vps5 and wrote -20 and +19 into that file. Then I waited. Twenty seconds later both had taken exactly half the core: 998 ticks to 998 ticks. Precisely nothing had happened.
This piece chases that "nothing". It ends at a two-line condition in the kernel source and at a systemd unit file nobody ever opens. Here is the short answer up front: autogroup exists on all seven servers in my fleet, it shows up under /proc, its counter keeps climbing — and on none of them does it do anything.
I previously measured who actually listens to ionice: there the hint was correct and the layer meant to read it (the block scheduler) simply wasn't there. This story is the mirror image. The hint does get read, to the letter — just not in the box you think you're in. And you didn't pick that box.
Where nice actually lives
nice looks like a priority number, but to the scheduler it is a weight. The kernel's sched_prio_to_weight[] table does the conversion literally: 1024 for nice 0, 15 for nice 19, 88761 for nice -20. When two tasks compete in the same pool, CPU time is split in proportion to those weights. nice 0 against nice 19 means 1024 against 15 — roughly 68 to 1.
The load-bearing words are "the same pool". With group scheduling (CONFIG_FAIR_GROUP_SCHED) active, the kernel first weighs groups against each other, then divides a group's share among the tasks inside it. So nice is not absolute but local: it applies inside whichever group you happen to be in. The split between your group and some other group is decided by the group's weight, not by your nice.
This is where autogroup comes in. Every new session (setsid — a new terminal, a new SSH connection) gets its own automatic group; the point was to keep your browser responsive while make -j64 runs on the desktop. When a process forks, the child inherits its parent's group.
So who decides which group a task belongs to? In the kernel the entire decision lives in one small function in kernel/sched/autogroup.h:
static inline struct task_group *
autogroup_task_group(struct task_struct *p, struct task_group *tg)
{
extern unsigned int sysctl_sched_autogroup_enabled;
int enabled = READ_ONCE(sysctl_sched_autogroup_enabled);
if (enabled && task_wants_autogroup(p, tg))
return p->signal->autogroup->tg;
return tg;
}
The tg passed in is the group derived from the task's cgroup. So the call is really asking: "the cgroup handed me this group — should I substitute the autogroup instead?" The answer comes from task_wants_autogroup() in kernel/sched/autogroup.c, which makes up its mind on its very first line:
bool task_wants_autogroup(struct task_struct *p, struct task_group *tg)
{
if (tg != &root_task_group)
return false;
...
if (p->flags & PF_EXITING)
return false;
return true;
}
Those two lines are the whole decision. If the task is not in the root CPU cgroup, autogroup is switched off without ever being consulted. sysctl_sched_autogroup_enabled can be 1, /proc/PID/autogroup can be readable, it can contain a perfectly nice-looking number — none of that matters. If the group coming from the cgroup differs from the root, return tg runs and autogroup is never mentioned again.
There is a subtlety here that most summaries skip, and the rest of this article leans on it. tg is derived not from the task's cgroup path but from the cgroup's CPU controller css; the line in core.c reads tg = container_of(task_css_check(tsk, cpu_cgrp_id, true), struct task_group, css);. In cgroup v2, if the cpu controller is not enabled on a cgroup, the task inherits the css of the nearest ancestor where it is — which is usually the root. So "I am not in the root cgroup" is not enough on its own: what decides the matter is whether the cpu controller is enabled along the path. sched(7) is careful about this too, saying "root CPU cgroup" rather than merely "root cgroup".
This is not a bug; it is the design. If you deliberately built a cgroup CPU hierarchy, it would be absurd for the kernel to ignore it and impose its own per-session grouping. The catch is that most of us did not build that hierarchy deliberately.
A lab on a single core
I don't write claims I haven't measured, so I set up a small rig on vps5 (six cores, kernel 7.0.0-34-generic, getconf CLK_TCK = 100). The rules are simple: every burner is nailed to a single core with taskset -c 5, so the machine's other five cores carried on with their day. Each measurement runs 20 seconds — 2000 ticks on that core. (That 100 is USER_HZ, the unit the /proc counters report in; the kernel's own tick rate, CONFIG_HZ, may differ — and it is the kernel's that HZ/10 refers to further down.) I took each process's CPU consumption from the utime+stime delta in /proc/PID/stat, and I also recorded CPU 5's total busy ticks from /proc/stat so that "is something else eating the core?" never stays an open question.
The burner is that famous one-liner:
taskset -c 5 bash -c 'while :; do :; done'
The rig covers five scenarios, and the numbers below come from a run repeated three times; the first, single-pass round went in the bin, for reasons I get to at the end. The three rounds differed by at most a few ticks, so a single round is quoted below.
E1 — same session, nice 0 against nice 19. Same terminal, same autogroup, same cgroup. Expected behaviour: nice works.
A=1968 ticks (98%) B=29 ticks (1%) ratio=67.9:1
A+B=1997 CPU5 busy=2002 window=2000 ticks
67.9 to 1. The table's 1024/15 predicted 68.3; the measurement matched theory to almost a decimal place. The totals are clean too: the two processes shared the whole core, nothing leaked.
E2 — separate sessions, nice 0 against nice 19. This time I launched both burners via setsid; /proc/PID/autogroup confirms they are in different groups (197625 and 197626). On the man page's reading, nice should be powerless here.
A=1969 ticks (98%) B=29 ticks (1%) ratio=67.9:1
A+B=1998 CPU5 busy=2001 window=2000 ticks
Nothing changed. nice worked perfectly well across sessions. That was my first surprise, and I'll admit my first instinct was to go looking for a measurement error.
E3 — the autogroup nice. Same rig, but now both processes have their own nice at 0; I try to create the split by writing into /proc/PID/autogroup: -20 for A, +19 for B. The writes succeed, and reading back shows nice -20 and nice 19.
A=998 ticks (50%) B=998 ticks (50%) ratio=1.0:1
A+B=1996 CPU5 busy=2001 window=2000 ticks
That is where this article's title comes from. The kernel accepted the write, stored the value, showed it back to me — and the scheduler never read it.
Taken together the two results are consistent: if the tasks really were competing as separate autogroups, E2's nice gap would have vanished and E3's autogroup gap would have flattened everything. The exact opposite happened. So both are in the same group, and that group is not the autogroup. /proc/PID/cgroup had been saying so all along: 0::/user.slice/user-1000.slice/session-158.scope.
Why the write succeeds in silence
The most maddening part of E3 is that writing raises no error. On the kernel side the reason is plain: proc_sched_autogroup_set_nice() takes the written value entirely seriously. It range-checks it (-EINVAL outside MIN_NICE–MAX_NICE), asks the LSM, demands privilege for negative values (-EPERM if can_nice() says no), and without CAP_SYS_ADMIN it rate-limits consecutive writes to one per HZ/10 with -EAGAIN — the operation is annotated as heavy work that takes global locks. Having passed all those gates, it converts the nice value to a weight through sched_prio_to_weight[] and calls sched_group_set_shares(ag->tg, shares).
So the write is not fake. The group's share really is changed. There just aren't any tasks in that group: your processes are running in a different group, the one the cgroup handed them. You have adjusted the weight of an empty container. For the kernel to object, the write path would have to check whether the autogroup is currently in use — and that check lives on the scheduling path, inside autogroup_task_group(), where the /proc writer's news never arrives.
This detail also explains why the file looks so convincing: the nice -20 you read back really was written, into ag->nice.
Give the root back and everything inverts
The honest way to test a theory is to change one variable and expect both results to flip. My variable is the cgroup: I moved the processes into the root cgroup. In cgroup v2 the root is exempt from the "no internal processes" rule, so this is allowed:
sudo sh -c "echo $PID > /sys/fs/cgroup/cgroup.procs"
That write makes the kernel regroup the task; now that tg == root_task_group, task_wants_autogroup() returns true this time.
E4 — root cgroup, separate sessions, nice 0 against nice 19:
A=999 ticks (50%) B=1000 ticks (50%) ratio=1.0:1
A+B=1999 CPU5 busy=2002 window=2000 ticks
nice went quiet. The 68-to-1 gap from E2 collapsed to 1-to-1 with a single cgroup move.
E5 — root cgroup, equal nice values, autogroup nice A=-20 / B=+19:
A=1999 ticks (100%) B=1 tick (0%)
A+B=2000 CPU5 busy=2001 window=2000 ticks
And the file that had just been useless worked, brutally, on the same machine and the same kernel. B got a single tick in twenty seconds; the ratio is beyond tick resolution. Theory predicts 88761/15, i.e. 5917 to 1 — the measurement doesn't confirm that, it only says "larger than I can measure". That's the honest phrasing.
Four measurements together say the world the man page describes is real in the root cgroup, and not on my servers.
Who closed the box
One question remained: why isn't my SSH session in the root cgroup? I started hunting at system.slice and Docker, because vps5 has more than twenty container scopes and every one of them carries Delegate=yes.
Scanning the fleet showed that was the wrong trail. Six of the seven servers run Docker; one (vps7) does not. On the Docker machines the root's cgroup.subtree_control line is wide: cpuset cpu io memory hugetlb pids rdma misc dmem. On Docker-free vps7 it is short: cpu memory pids. The list shrinks, but cpu never leaves. So Docker widens the controller set; it is not what turns on cpu.
I found the culprit on vps7, in systemd's own unit file:
$ systemctl show -p Delegate -p DelegateControllers user@1000.service
Delegate=yes
DelegateControllers=cpu memory pids
Nor is this a distribution choice; upstream systemd's units/user@.service.in says it verbatim: Delegate=pids memory cpu. The CPU controller is delegated to the user's session manager so that users can put resource limits on their own units.
systemd's own documentation explains the rest: when a controller is enabled somewhere, then "because of how the cgroup hierarchy works" it is automatically enabled for all parent units and for any sibling units, starting with the lowest level at which it is enabled. The same document also states plainly that units for which a controller is enabled may be subject to resource control even if they have no explicit configuration at all.
The chain is complete:
- systemd delegates the
cpucontroller foruser@1000.service. - The controller is enabled at
user-1000.sliceand, per the rule, propagates to parents (including the root) and to siblings. - Every SSH session is born inside a
session-N.scope, i.e. a non-root cgroup. -
task_wants_autogroup()returnsfalse. - Autogroups keep getting created, their counter keeps climbing, they keep showing up under
/proc— and they are never used.
Across all seven servers in the survey, sched_autogroup_enabled was 1 and the login session sat in a non-root cgroup. So on these machines autogroup's practical value is zero. The kernel code isn't dead; it is unreachable in my setup. The distinction matters, because it is still alive in the root cgroup: for kernel threads and for processes that never entered that hierarchy, the mechanism works exactly as documented.
The wrong lesson to draw would be "systemd breaks things". systemd enabled something that gives you per-session resource control; in exchange it disabled a desktop heuristic from 2011. The trade is reasonable. And in fairness the trade is not undocumented either — though I only found that out after the measurements. In one of the closing sentences of its autogroup section, sched(7) says it plainly: "The use of the cgroups(7) CPU controller to place processes in cgroups other than the root CPU cgroup overrides the effect of autogrouping." systemd's own CPUWeight= documentation warns the reader too: the kernel may divide resources automatically by session id, and because that "is similar to the cpu controller with no explicit configuration", the two should not be mistaken for one another.
So the information is there; the problem is that nothing points you at the sentence. It sits at the tail of the autogroup section, just ahead of the "nice value and group scheduling" section I quoted at the top — by which point someone who opened the page to find out how to write /proc/PID/autogroup has already found their answer and stopped reading. And /proc/PID/autogroup still speaks the language of the old world: it shows you a group, accepts your write, stores the value, and gives not one hint that the group is unused.
So what does work
We are not ripping the mechanism out; we are using its replacement. If you genuinely want to apportion CPU between sessions or services, the right lever is the cgroup weight, i.e. CPUWeight. That one went through the same rig.
Same core, two separate systemd scopes, both with their own nice at 0:
sudo systemd-run --unit=aglab-a --scope -p CPUWeight=1000 \
taskset -c 5 bash -c 'while :; do :; done'
sudo systemd-run --unit=aglab-b --scope -p CPUWeight=100 \
taskset -c 5 bash -c 'while :; do :; done'
A=1817 ticks (90%) B=181 ticks (9%) ratio=10.04:1
A+B=1998 CPU5 busy=2002
10.04 to 1 — the requested 10:1, within 0.4%. And /proc/PID/autogroup reported the same group for both processes, so the measurement itself proves the split did not come from autogrouping.
The same rule applies to containers, and that's where most people never look. A process in a Docker container is born in a scope under system.slice, which by definition is not the root cgroup; so writing to /proc/PID/autogroup inside a container gives you the same silent result. To apportion CPU between workloads in containers the lever is again weight: cpu.weight written into the container's own cgroup. docker run --cpu-shares still works, but the name is a cgroup v1 leftover; under v2 Docker maps it onto cpu.weight through a lossy conversion, so the number you pass is not the weight you get. If you dislike that ambiguity, go straight to cpu.weight.
In practice this means: if your goal is "don't let that backup job crush the live service", writing nice -n 19 does nothing when the job runs in its own service unit — the unit is already a separate cgroup and nice only applies inside it. The fix is CPUWeight= on the unit. The converse holds too: when several processes share one unit, nice still behaves exactly as you expect (E1 shows this). Pick the right scale and both tools are honest.
What the lab charged me
I threw away the first round of these measurements, and leaving out why would be dishonest.
In my first attempt I launched the burners from a shell function as A=$(spawn ...) and collected the PIDs inside the function with PIDS="$PIDS $p". Because command substitution runs in a subshell, that variable never made it back to the parent shell (the burners' output went to /dev/null, which is why the substitution returned instead of hanging); my cleanup function walked an empty list every round. The result: each experiment's burners piled on top of the next one's. In the second experiment the pair I was measuring totalled 1000 ticks instead of 2000 — because two leftovers from the previous run were on the core beside my two, and the weights split the share right down the middle (1024+15 against 1024+15). By the third experiment the total had fallen to 992: four strays plus two new processes, so my share of a 4×1024 + 2×15 weight pool works out at 49.6%. By the time I re-ran the rig, twelve burners had accumulated on the core. What saved the numbers wasn't cleverness — it was having written a CPU5 busy column into the log as well. When the A+B total stopped matching the window, the measurement reported itself.
The second one is funnier. To clear the strays I ran the command remotely: ssh vps5 'pkill -9 -f "while :; do :; done"'. Run that way, the thing carrying the command on the server is a bash -c wrapper whose command line contains the pattern verbatim. pkill killed the burners, then its own shell; the channel dropped and ssh exited 255 — which is ssh's own connection-error code. At an interactive prompt the same command would not have killed me: a login shell's command line is -bash, and pkill skips its own pid. The cleanup was a success — I just got the news a little late.
Both carry the same lesson: put an independent witness into your rig, one that can tell you whether the measurement is sound. Mine was core busy-time read from /proc/stat, and for a two-line cost it stopped me from publishing something false.
Checklist
You can reach the same diagnosis on your own server in five minutes:
-
cat /proc/self/autogroup— seeing a group only tells you the group exists, not that it is used. -
cat /proc/self/cgroup— if the output isn't0::/, autogroup is probably off for you. A necessary clue, but not sufficient on its own. -
cat /sys/fs/cgroup/cgroup.subtree_control— this is the decisive test. Ifcpuis in the list and you are not at the root, autogroup is dead for you. Ifcpuis absent, autogroup keeps working even in a non-root cgroup — because the group is set by the cpu controller's css, not by the cgroup path. -
systemctl show -p DelegateControllers user@$(id -u).service— if you seecpu, you've found your reason. - Need a split between sessions or services? Use
CPUWeight=, notnice. Need a split between jobs inside one unit?niceis the right tool. -
CPUWeight=only compares siblings. Weighting a service insidesystem.slicedoes not protect it fromuser.sliceload; for the "backup job vs live service" case you have to set weights at every level of the path. The range is 1-10000 and the kernel default is 100 — soCPUWeight=1000means ten times any unweighted sibling. - Read back what you wrote:
systemctl show -p CPUWeight <unit>andcat /sys/fs/cgroup/<path>/cpu.weight. This whole article is about writes that succeed and are never read; apply the same suspicion to my own advice. - Rollback and monitoring:
systemctl set-property --runtime <unit> CPUWeight=...gives you a non-persistent trial, andsystemd-cgtopshows how the shares actually land. - Weight only shows up under contention. Don't test on an idle box and conclude this is as fake as
/proc/PID/autogroup; you have to fill the core to see it. - Thinking of setting
sched_autogroup_enabledto 0? Check the second item first: you are probably about to switch off a knob that is already inert.
One small naming note as well: the man page still explains the mechanism in terms of CFS, whereas the kernel moved to EEVDF in 6.6 and vps5 is already on the 7.0 series. The weight table and the group logic are unchanged, so the measurements here are unaffected — but keep the difference in mind while reading the documentation.
Conclusion
What actually bothers me isn't that autogroup doesn't work. It's that it never says so. /proc/PID/autogroup accepts the write, stores the value, reads it back; the nice you wrote is sitting right there. No error code, no warning, no dmesg line. The file behaves like an interface when it is nothing more than a ledger.
Much of Linux tuning is like this: you write a value, it is accepted, and somewhere else a condition you have never seen decides whether that value will ever be read. With ionice the condition was which block scheduler had been picked for the disk; here it is which cgroup the task was born into. In both cases the write succeeds, in both cases there is no effect, and in both cases the system doesn't tell you.
So I keep one rule, and this exercise confirmed it again: you know a setting works because you can measure the difference, not because you were allowed to write it. A one-line burner and two tick counters told me which sentence of a fifteen-year-old document I needed to read, before the document did.
Official Sources
- kernel/sched/autogroup.c — Linux source
- kernel/sched/autogroup.h — autogroup_task_group()
- kernel/sched/core.c — the sched_prio_to_weight table
- CFS Scheduler — group scheduling extensions
- EEVDF Scheduler
- Control Group v2
- systemd — units/user@.service.in
- systemd.resource-control — CPUWeight and controller enablement
- sched(7) — the autogroup section
Top comments (3)
Really very enjoyable 😃❤️😊
(っ.❛ ᴗ ❛.)っThanks for sharing this!!
Thank you so much! 😄❤️
I’m really glad you enjoyed it — especially since this one started with me staring at two CPU counters and wondering, “Why on earth did absolutely nothing happen?” 😂
Sometimes “nothing happened” turns out to be the beginning of the best debugging stories.
Thanks for reading and for the lovely comment! 😊🙏
Welcome ❤️😀