Nothing in Unix is more boring than a pipe. Two ends, one buffer, 64 kilobytes. Same thing for
forty years. Every | opens one, and nobody has ever once wondered how many kilobytes it
holds.
Today I started two containers. The first opened 1024 pipes and held on to them. The second — a
separate container, separate PID, mount and network namespaces — opened the first pipe of
its life, and that pipe came back at 8 kilobytes, not 64.
No error was returned. pipe() returned zero and the program carried on. Moving the same data
just started taking eight times longer.
The two containers shared exactly one thing: a uid.
Three files, three different jobs
The kernel has three knobs for pipe buffers, and all three do something different. They are
also the three that get confused with one another most often.
$ cat /proc/sys/fs/pipe-max-size # 1048576 -> ceiling for a single pipe, BYTES
$ cat /proc/sys/fs/pipe-user-pages-soft # 16384 -> soft quota per uid, PAGES
$ cat /proc/sys/fs/pipe-user-pages-hard # 0 -> hard quota per uid, PAGES (0 = off)
The first is bytes, the other two are pages. The first is about one pipe, the other two are
about all of that uid's pipes. The default soft quota is 16,384 pages: 64 MiB with 4 KiB
pages. The kernel computes that number as PIPE_DEF_BUFFERS * INR_OPEN_CUR, 16 × 1024 — or, as
the documentation puts it, enough "to allocate up to 1024 pipes at their default size".
My first surprise came while poking at pipe-max-size. F_SETPIPE_SZ does not give you the
size you ask for; it rounds up to the next power of two:
request 4096 -> 4096
request 4097 -> 8192
request 65536 -> 65536
request 65537 -> 131072
request 70000 -> 131072 <-- ask for 68 KiB, pay for 128 KiB
request 1048576 -> 1048576
request 1048577 -> EPERM
Asking for 70,000 bytes and paying for 131,072 sounds harmless. It stops sounding harmless once
you remember the quota is counted in pages: the 70,000 bytes you asked for would have fitted in
18 pages, and that pipe costs you 32. Nearly half of it wasted. The fcntl return value
reports the capacity actually allocated, so code that ignores the return value is wrong about
its own budget.
The one thousand and twenty-fifth pipe
Finding the edge of the quota only takes a patient program: open pipes one at a time, ask each
one for its capacity, print the point where the answer changes.
pipe #1 capacity=65536
pipe #1025 capacity=8192 (accumulated by then: 16384 pages)
Right at the bottom of the quota, in a single pipe, a division by eight. The kernel source
makes it obvious why this is silent — alloc_pipe_info() charges the request first, sees that
it crossed the quota, shrinks the request, fixes the accounting and carries on:
user_bufs = account_pipe_buffers(user, 0, pipe_bufs);
if (too_many_pipe_buffers_soft(user_bufs) && pipe_is_unprivileged_user()) {
user_bufs = account_pipe_buffers(user, pipe_bufs, PIPE_MIN_DEF_BUFFERS);
pipe_bufs = PIPE_MIN_DEF_BUFFERS;
}
There is no return -ENOSPC here, no pr_warn, no counter. There is an assignment. That is
exactly what makes the soft quota soft: it does not stop you, it quietly shrinks you.
The value it falls back to is two pages, 8192 bytes. That number has a small story of its own.
When the quota mechanism was added in 2016 the fallback was one page. In 2021 Alex Xu
raised it to two, and the reasoning sits in the comment that patch added to fs/pipe.c: with a
single-page pipe "a write to a non-empty pipe may block even if the pipe is not full", which
means deadlock for "GNU make jobserver or similar uses of pipes as semaphores". The patch's
reproducer starts from the same place as my count above — open 1025 pipes, look at the capacity
— but then goes one step further and demonstrates the hang; I only looked at the capacity.
As for the documentation: fs.rst kept calling the fallback "a single page" until July 2025.
For four years the official docs stated the wrong number, until Štěpán Němec came along and
fixed it. When people ask why I insist on reading source code, this is the example I give now.
The counter exists nowhere
The real problem here is not the 8 kilobytes. The real problem is that there is no supported
interface for asking where in the quota you are.
The counter lives inside struct user_struct:
struct user_struct {
...
atomic_long_t pipe_bufs; /* how many pages are allocated in pipe buffers */
...
};
It has no counterpart under /proc. None under /sys. vmstat does not show it, slabtop
does not show it. On a kernel with BTF you can walk into user_struct with drgn and read the
counter exactly, but that is a debugging trick, not an interface you build monitoring on. You
have a 64 MiB budget and nowhere to ask for your balance. It is like a bank with no ATM out
front, so you push your card into the machine and wait for "insufficient funds" — except this
machine does not say that either. It just hands you less money.
One detail makes the budget slipperier still: the counter charges reserved ring slots, not
resident memory. alloc_pipe_info books the full ring at creation
(pipe->nr_accounted = pipe_bufs), while the actual pages are only allocated on write, via
alloc_page(GFP_HIGHUSER | __GFP_ACCOUNT). So a thousand empty pipes can eat the entire quota
while using almost no RAM. The converse holds too: the data pages carry __GFP_ACCOUNT, so
your cgroup memory counter does see pipe contents — what it does not see is the quota itself.
In practice the only thing you can do is send in a canary: open a pipe and look at its
capacity. If you get the expected default you are under the quota; if you get two pages, you
crossed it a while ago.
python3 - <<'EOF'
import os, fcntl
page = os.sysconf('SC_PAGESIZE')
pmax = int(open('/proc/sys/fs/pipe-max-size').read())
r, w = os.pipe()
cap = fcntl.fcntl(w, fcntl.F_GETPIPE_SZ)
expected = min(16 * page, pmax) # PIPE_DEF_BUFFERS x page, clamped by pipe-max-size
print(cap, expected, 'QUOTA EXCEEDED' if cap <= 2 * page < expected else 'normal')
EOF
Do not hard-code 65536. The default capacity is PIPE_DEF_BUFFERS × page size, clamped by
pipe-max-size — on an arm64 machine with 64 KiB pages a fresh pipe comes back at 1 MiB, not
- An alarm with a fixed threshold will cry wolf there every single day.
There is one more asymmetry, and it helps when you are diagnosing: while you are over the
quota, growing a pipe returns EPERM, while shrinking one is always allowed.
F_SETPIPE_SZ(65536) over quota -> EPERM
F_SETPIPE_SZ(4096) over quota -> 4096
So if your application says "give me a 1 MiB pipe" at startup, that request comes back as
EPERM once you are over the quota. How the program handles that is entirely its own business
— whether it carries on with the default or dies at startup, you cannot tell from the outside.
If you know the answer for your own code, you are a step ahead.
The quota is on the uid, not the container
Now back to the measurement I opened with, because this is the part worth taking seriously.
Pipe pages are counted through user_struct, which means through the kuid. Containers,
cgroups, PID namespaces, mount namespaces — none of them enter this accounting. A controlled
experiment with three containers:
| Container | uid | State | Capacity of its first pipe |
|---|---|---|---|
| Clean start | 1000 | — | 65536 |
| Clean start | 1001 | — | 65536 |
| B | 1000 | A holds 1024 pipes on the same uid | 8192 |
| C | 1001 | A is still holding them | 65536 |
Only one variable changes: the uid. Container C runs on the same machine, at the same moment,
on the same kernel, and feels nothing because its uid is different. B is born penalized without
ever having opened a pipe.
This concerns anyone running containers, because in a default Docker container processes run as
uid 0, and unless you set up userns-remap that uid 0 is the same kuid as the host's root.
So if every container on the machine runs as root, all of them drink from one 64 MiB bucket.
"But root is privileged, it goes past the quota," is what I assumed too. Here is how the kernel
defines privileged:
bool pipe_is_unprivileged_user(void)
{
return !capable(CAP_SYS_RESOURCE) && !capable(CAP_SYS_ADMIN);
}
The capability mask of a default Docker container is 0x00000000a80425fb. Bit 21
(CAP_SYS_ADMIN) and bit 24 (CAP_SYS_RESOURCE) are not in that mask. As far as this
mechanism is concerned, root inside the container is an unprivileged user. It cannot pass
the quota; it gets clipped silently.
And the clipping point is not even fixed. Running the same experiment as root in a default
container, the edge showed up not at #1025 but at #736:
pipe #1 capacity=65536
pipe #736 capacity=8192
Because part of uid 0's quota had already been spent by the VM's own root processes — dockerd,
containerd and friends. How much? I can work it backwards from my own pipes: since #735 did not
trigger it and #736 did, the baseline has to sit between 4608 and 4624 pages, around 18 MiB. On
my first run the same arithmetic gave 4512–4528 pages. In other words the edge moves depending
on what else happens to be running as root at that moment. Not the most beloved way to run the
same image twice on the same machine and get different behavior.
That is also why running the canary as root on my own server proves nothing, a mistake I made
on my first attempt. In a privileged container I opened 1300 pipes — 20,800 pages, 1.27 times
the soft quota — and every one came back at 65536, with F_SETPIPE_SZ(65536) succeeding too.
Root holding CAP_SYS_RESOURCE never sees this ceiling. A canary you run as root on the host
tells you only that root is privileged.
The hard quota and a borrowed errno
pipe-user-pages-hard defaults to 0, which means off. Here is what happens when you give up on
that and write a number into it. With the hard quota set to 16,500 pages, the same program
gives:
pipe #1 capacity=65536
pipe #1025 capacity=8192
pipe #1083 FAILED errno=23 "Too many open files in system"
The arithmetic lands exactly: after the soft quota every pipe costs 2 pages, so 16384 + 58×2 =
16500, and 58 pipes later the door closed.
Now look at the errno. ENFILE — "too many open files in system". The reflex on seeing that is
fs.file-max, fs.nr_open and ulimit -n, and none of the three has anything to do with this
event. The error is not undocumented, though, which I learned while writing this: pipe(2) has
two separate ENFILE entries in its ERRORS section, and the second says exactly this —
"the user hard limit on memory that can be allocated for pipes has been reached and the caller
is not privileged; see pipe(7)". The documentation is right; the errno is overloaded and the
reflex sends you to the wrong files. I wrote about how those files get confused with each other
in who changed fs.file-max,
and now the pipe quota borrows the same errno on top of it.
So should you enable the hard quota? The first draft of this piece said "don't, the soft quota
already protects you". That was wrong. The soft quota limits pipe size, not the total:
once you cross it every new pipe arrives at two pages, but the number of them is unbounded. My
own measurement shows it — 58 more pipes opened after #1025, and without the hard quota it
would have continued until the file descriptors ran out. The patch that added the quota does
the arithmetic in its own commit message: 256 processes, 1024 fds each, 1024*64kB +. Where you imagine a 64 MiB bucket, one user can hold
(256*1024 - 1024) * 4kB = 1084 MB
gigabytes; the same patch notes that it mitigates a CVE (CVE-2013-4312). And it states why the
hard limit ships disabled: to avoid breaking applications that make heavy use of pipes,
splice in particular.
The honest version: if your machine has local users or tenants you do not trust, the hard quota
is the only thing bounding total memory. The price is diagnosis difficulty, and you avoid most
of that price simply by knowing the second ENFILE entry above. On a machine running only your
own services, leaving it off is reasonable — but telling a multi-tenant operator "don't enable
it" is bad advice.
What eight kilobytes cost
"Quietly shrinks you" ought to have a price tag. So here is the bench: a writer process pushes
1 GiB in 64 KiB chunks into a pipe while a reader consumes it. The only variable is the pipe's
capacity.
| Capacity | Time | Throughput | Writer's voluntary context switches |
|---|---|---|---|
| 65536 | 0.74 s | 1388 MiB/s | 12,090 |
| 16384 | 2.49 s | 412 MiB/s | 65,463 |
| 8192 | 5.02 s | 204 MiB/s | 130,965 |
| 4096 | 9.62 s | 106 MiB/s | 262,014 |
(Each row is the median of three runs.)
Close to a sevenfold slowdown (1388/204 = 6.8) and eleven times as many wakeups. The
context-switch count gives the mechanism away directly: pushing 1 GiB through 8 KiB windows
takes 131,072 rounds, and the measurement says 130,965. Every time the pipe fills the writer
sleeps, and when the reader drains it the writer wakes. It is not doing work; it is queueing.
The 12,090 on the 64 KiB pipe sits below the expected 16,384 because at that capacity the
reader sometimes keeps up and the writer never blocks at all.
What bothers me about this table is that in all three rows the code is the same, the image is
the same, the cgroup quota is the same. The only thing that changed is how many pipes the same
uid happens to be holding somewhere else. That is not a quantity I have ever put into a
capacity plan.
What this looks like on my own server
Labs are nice, but I wanted to know what the measurement says on a real machine, so I looked at
my own. On VPS3 I counted distinct pipe inodes whose descriptors show up under a given uid,
through /proc — a read-only scan, kernel 6.8.0-142-generic:
| uid | Distinct pipes | Assuming 16 pages | Share of soft quota |
|---|---|---|---|
| 0 | 606 | 9696 pages | 59.2% |
| 99 | 84 | 1344 pages | 8.2% |
| 1001 | 54 | 864 pages | 5.3% |
| 70 | 39 | 624 pages | 3.8% |
| 999 | 36 | 576 pages | 3.5% |
| 1000 | 36 | 576 pages | 3.5% |
Two caveats are needed here. First, "assuming 16 pages" is deliberate: /proc does not expose
a pipe's capacity, only its inode. Second — and I only noticed this after finishing the piece —
the kernel charges the quota to whoever created the pipe (pipe->user, fixed at creation),
while my scan looks at whoever is currently holding the descriptor. When a pipe crosses a
uid boundary, which is exactly what happens in the container scenario, the two diverge. So this
table is not the quota itself but a shadow near it. The canary returned the default capacity
for both unprivileged uids, so at least I know those two are under the ceiling.
So who is holding uid 0's pipes? Breaking the same scan down by process name:
| Process | Distinct pipes |
|---|---|
| containerd-shim | 278 |
| docker-proxy | 137 |
| node | 71 |
| next-server (two processes) | 26 |
| sshd | 13 |
The rows do not add up, and should not: a pipe has two ends and they can sit in two different
processes, so the same inode is counted on more than one row. The direction is clear all the
same — the biggest eater of the quota is the container runtime itself. Every container means a
shim, most of them also mean a docker-proxy, and both drink from root's bucket. The more
containers a machine runs, the fewer pipes are left for the root processes inside those
containers.
One more thing: a few minutes passed between the two scans, and uid 0's total fell from 606 to
- The bucket is constantly filling and draining. That is why the edge moved between runs in the lab — what you are measuring is not a fixed limit but who happens to be holding what.
Fifty-nine percent is not a fault. But it does make you think: on this machine the containers
run as root and all of them drink from the same bucket. The day one of them starts doing
something that holds long-lived pipes — a subprocess pool, a log shipper, make -j — the first
complaint will not come from that container, but from the next one slowing down. It is the same
problem as a limit I thought I had written per container but had not; it lands in exactly the
same place as the piece on how CPU quota does not look at your
average.
What to do
Four things, in order of importance:
1. Ask the canary as an unprivileged uid, and make the threshold page-size aware. Add a few
lines to your monitoring: open a pipe as the relevant service user, read F_GETPIPE_SZ, and
compare it against min(16 × page, pipe-max-size) rather than a hard-coded number. Two pages
means the quota has been crossed. Asking as root measures nothing — privileged root gets full
capacity even at 1300 pipes.
2. Run containers under distinct uids — but userns-remap alone does not do this. Here I
had to correct my own advice. Docker's userns-remap uses a single daemon-wide mapping:
every container's uid 0 lands on the same host kuid. So it moves the bucket away from host
root, but it does not separate containers from one another — the A and B experiment above
is a photograph of exactly that situation. What actually separates them is a distinct identity
per container: a different USER in the image, or a runtime that hands each container its own
uid range (podman's --userns=auto, for instance). Container C is the proof: its neighbor had
drained the quota and it felt nothing, because its uid was different.
3. Lowering pipe-max-size is a trade, not a free win. I set this up as a prediction and
then measured it. Since Linux 4.9 pipe-max-size also clamps the default capacity of a new
pipe; pull it down to 16384 and the default becomes 4 pages, so the edge should move past pipe
4096:
pipe-max-size=16384
pipe #1 capacity=16384
pipe #4097 capacity=8192
The prediction held exactly, and the fall becomes twofold instead of eightfold. But the bill
for that is sitting in the table above, and I left it out of my first draft: a 16 KiB pipe
carries 412 MiB/s where a 64 KiB pipe carries 1388. You are paying a permanent 3.4× on
every pipe to avoid a 6.8× penalty that happens rarely. On a machine where pipes carry real
data that is a bad trade; where pipes are only control and semaphores, a good one. On top of
that, an application wanting a big pipe will see EPERM, so you need to know everything that
uses splice and large buffers first.
4. Decide the hard quota by what kind of machine it is. Only your own services on it: leave
it off. Local users or tenants you do not trust: it is the only thing bounding total memory —
the soft quota bounds size, not the total.
And a recovery note: once the quota has been crossed, existing pipes do not grow back
on their own; only new ones are born small. Stopping the process that drained the quota will
not fix the penalized neighbor by itself — you have to restart that neighbor too, and only
after the hog has let go.
What I did not prove
There is no production incident in this piece. This mechanism has not bitten me; I got curious,
measured it, and then looked at how close my own server sits to it. Fifty-nine percent is a
measure of distance, not of an event.
The lab ran inside Docker Desktop's linuxkit VM (6.10.14-linuxkit, x86_64, 4 KiB pages). No
sysctl was written on the macOS host or on VPS3; the two settings I changed temporarily inside
the linuxkit VM were restored afterwards and verified from a separate container. On a machine
with a different page size — 64 KiB pages on arm64, say — every number changes: the quota is
counted in pages, and 16,384 pages is 1 GiB there.
The throughput measurement used a single writer and a single reader with 64 KiB chunks. A
workload doing small writes will show a completely different ratio; this table is not the
general answer to "how much worse does a small pipe get", it is that workload's answer.
The page counts in the VPS3 table are derived, not measured: I assumed 16 pages per pipe,
because /proc will not tell me the capacity. And as noted above, the scan looks at the uid
holding the descriptor while the kernel charges the pipe's creator; for pipes that cross a uid
boundary the two diverge, and the bias runs systematically towards undercounting root. In
short, 59.2% is an order of magnitude, not a balance.
The kernel version matters too: before 5.14 the fallback is one page rather than two, which
puts the documented GNU make jobserver deadlock back on the table. The patch was backported to
the stable branches, but if you run a kernel frozen before August 2021 or a vendor fork, you
inherit that risk. This is the one thing a reader can check on their own machine in two seconds
with uname -r.
Closing
The pipe quota is not new; it has been there since 2016 and on most machines it probably never
gets touched. What makes me think is not the mechanism itself but its shape.
Every limit has two identities: whose name the accounting is kept under, and where you believe
the isolation is. When those two land in the same place, the limit is honest — cross it and you
get an error, and the error belongs to you. When they land in different places, the limit goes
quiet: the one who pays is not the one who spent the quota, but the one standing next in line.
The pipe quota keeps its books under a uid while we imagine the isolation sits at the
container. In the gap between the two there is no error message, only slowness.
The question I ask when looking at container limits is no longer "how big is this limit". It is
"whose name is this limit's counter kept under". pipe_bufs is not the first place I have
asked that, but it turned out to be the place that hides the answer best: nobody can see the
counter at all.
Official Sources
- Linux kernel documentation —
fssysctls (pipe-user-pages-soft/pipe-user-pages-hard) - Kernel source —
fs/pipe.c(alloc_pipe_info,pipe_set_size,round_pipe_size) pipe(7)man page — pipe capacity and the/procinterfacespipe(2)man page — the two separateENFILEentriesF_GETPIPE_SZ/F_SETPIPE_SZman page — rounding,EPERMandEBUSY- Commit 759c0114 — the patch that added the quota (Linux 4.5)
- Commit 46c4c9d1 — raising the fallback to two pages (Linux 5.14)
- Commit 7069b529 — fixing the "a single page" error in the docs (2025)
Top comments (0)