I sat down to write this piece certain of one thing: the write brake on my servers never presses. Its target is set to 75 milliseconds and my reads are measured in microseconds — a thousandfold margin. A threshold like that doesn't get crossed, and a threshold that doesn't get crossed isn't a brake, it's decoration.
I was wrong. The kernel's trace showed the brake firing 23 times across 45 seconds of load, and clamping write depth all the way down.
The story starts on 23 September. While writing about how read_ahead_kb climbed to 8192, a suspicion stuck with me: if a disk's claim that it spins determines the read-ahead size, how many other settings read that same claim? I opened the list this week, wbt_lat_usec came up first — and six of my seven servers had 75000 in it.
Seven servers, one bit
Here's the raw table. The sched column is the disk's active I/O scheduler, wbt is the contents of /sys/block/<dev>/queue/wbt_lat_usec:
| Server | Ubuntu | Kernel | Device | rotational |
sched |
wbt_lat_usec |
|---|---|---|---|---|---|---|
| vps1 | 26.04 | 7.0.0-34 | sda | 1 | none | 75000 |
| vps2 | 26.04 | 7.0.0-34 | sda | 1 | none | 75000 |
| vps3 | 24.04 | 6.8.0-142 | sda | 0 | none | 2000 |
| vps4 | 24.04 | 6.8.0-142 | vda | 1 | none | 75000 |
| vps5 | 26.04 | 7.0.0-34 | sda | 1 | none | 75000 |
| vps6 | 26.04 | 7.0.0-34 | sda | 1 | none | 75000 |
| vps7 | 26.04 | 7.0.0-34 | sda | 1 | none | 75000 |
Exactly one column separates them: rotational. vps3 says zero and gets 2000; the other six say one and get 75000. The kernel version, the distribution release and the device name have nothing to do with that number. The decision lives in a single bit.
If you remember the table from five days ago: vps1 and vps2 showed 25.04 and 6.14 there. They really did change in between — /var/log/dist-upgrade/main.log on both is dated 25 September, so I moved them to 26.04 two days after that piece went out. This table is from the morning of 28 September.
Note the sched column too, we'll need it shortly: every one of them is none. So these disks have no I/O scheduler at all. Nobody is handing out fairness among queued requests, nobody is watching priorities. Under those conditions the only mechanism protecting read latency is wbt itself.
The ten lines that write 75000
Writeback throttling — wbt for short — is a module plugged into the block layer's rq_qos framework. Its job in one sentence: squeeze background write traffic the moment read latency starts to suffer. Jens Axboe wrote it in 2016, and it has been in place ever since; its most recent change landed this June, so nobody has retired it. The comment at the top of the file also confesses its inspiration: CoDel, the queue management scheme from networking. The only difference is that the block layer doesn't have the luxury of dropping packets.
The default target latency comes from this function in block/blk-wbt.c:
static u64 wbt_default_latency_nsec(struct request_queue *q)
{
/*
* We default to 2msec for non-rotational storage, and 75msec
* for rotational storage.
*/
if (blk_queue_rot(q))
return 75000000ULL;
return 2000000ULL;
}
The comment explains itself, and the reasoning is perfectly sound: on a platter-spinning disk, head seeks already cost milliseconds, a 2 ms target would be exceeded constantly and the brake would never let go. Giving a rotating disk 75 ms is a courtesy.
The question is what blk_queue_rot() actually looks at. That macro reads q->limits.features & BLK_FEAT_ROTATIONAL, i.e. the flag the driver set at registration. Drivers fill in that flag with wildly different levels of seriousness.
Three drivers, three different defaults. NVMe is the cleanest: drivers/nvme/host/core.c starts with the flag clear and only sets it when if (info->is_rotational) holds. So an NVMe device that says nothing counts as solid state.
SCSI does the exact opposite, and admits it plainly in a comment:
/*
* set the default to rotational. All non-rotational devices
* support the block characteristics VPD page, which will
* cause this to be updated correctly and any device which
* doesn't support it should be treated as rotational.
*/
lim->features |= (BLK_FEAT_ROTATIONAL | BLK_FEAT_ADD_RANDOM);
drivers/scsi/sd.c says "rotating" first, then tries to read the B1h VPD page; if the page exists and reports rot == 1, it clears the flag. If the page is missing, sd_read_block_characteristics() returns without doing anything and the default stands. So SCSI does ask — but assumes "rotating" when it gets no answer. Reasonable in 2014; less so in a world where most disks are virtual.
virtio-blk doesn't ask at all:
struct queue_limits lim = {
.features = BLK_FEAT_ROTATIONAL,
.logical_block_size = SECTOR_SIZE,
};
That's in drivers/block/virtio_blk.c, in the opening lines of the probe function. The flag arrives switched on and nowhere in the file is it cleared. So every virtio-blk disk, whatever sits underneath it, declares itself a spinning disk.
Back to my table: only vps4's vda goes down that path. The sda on the other five servers is SCSI, so it travels through sd.c and arrives with the model string QEMU HARDDISK — a virtual SCSI disk doesn't present the B1h page, so the default stands. Two different drivers, two different reasons, the same outcome. They share the result, not the route; that distinction matters, because the fix lives in a different place for each.
Blaming the driver here is easy but unfair. I looked at the protocol's header file: virtio_blk.h defines fifteen feature bits, covering size, topology, discard, zoned and more, and not one of them asks whether the disk rotates. Where would the driver get information that doesn't exist? The problem isn't the flag itself, it's how quietly the number of decisions resting on that flag has grown over the years. The same bit multiplies read-ahead by 970 in one place and loosens the write brake by thirty-seven and a half in another. Nobody planned for this bit to do that much work; the uses just accumulated.
A version note: the blk_queue_rot() macro and the queue_limits.features field above are new. On vps3 and vps4, running 6.8, the same decision goes through the QUEUE_FLAG_NONROT flag instead. The mechanism lands in the same place under different names — if you're reading the source of an older kernel, look for blk_queue_nonrot().
When does the brake actually press?
Once I understood the default, the real question became: does a 75 ms target do anything in practice? Read latencies on my servers are measured in microseconds. Seventy-five milliseconds is a thousand times a typical read. My expectation was that this threshold would effectively never be crossed and the brake would never press in its lifetime.
While forming that expectation I had read the latency_exceeded() function, and one detail caught my eye. Over the statistics it collects in 100 ms windows, wbt looks not at the average but at the smallest read latency:
/*
* If the 'min' latency exceeds our target, step down.
*/
if (stat[READ].min > rwb->min_lat_nsec) {
So it doesn't complain that "my p99 degraded"; it throttles only if even the fastest read in the window is above the target. A very conservative criterion. Combined with a 75 ms target, I thought, this brake will never engage.
Rather than assume it didn't, the honest move was to count. The kernel already offers four tracepoints: wbt_lat, wbt_step, wbt_stat and wbt_timer. wbt_step writes a line every time the brake changes a step, wbt_lat every time the target is exceeded. Counting is both easier and more honest than guessing.
Lab: same load, three settings
My lab machine is a Lima virtual machine on Apple silicon: Ubuntu 26.04, kernel 7.0.0-34-generic, 12 vCPUs, 14 GB of memory. The root disk is /dev/vda, i.e. virtio-blk — and predictably rotational is 1 and wbt_lat_usec is 75000. The active scheduler here is mq-deadline rather than none, because this disk has a single hardware queue. block/elevator.c gives mq-deadline to single-queue devices and none to multi-queue ones. My servers' disks are multi-queue, which is why they say none.
The load looks like this: four parallel buffered writers rewrite a 512 MiB file in a loop while a single reader does random 4 KiB O_DIRECT reads from another file. I dropped caches before every run.
There are two separate runs here and they shouldn't be conflated. The latency run measures each setting four times, 12 seconds each. The trace run switches on the wbt tracepoints and executes each setting three times, 15 seconds each. The three settings are the same in both: target 75000 (the default), 2000 (what the disk would get if it were honest) and 0 (brake fully off).
The trace run first. I saved the trace to a file at the end of each run and recounted the numbers from the raw file:
wbt_lat_usec |
Violation events | Smallest | Median | Largest | Step-downs |
|---|---|---|---|---|---|
| 75000 | 23 | 83.3 ms | 290.3 ms | 758.2 ms | 34 |
| 2000 | 32 | 4.9 ms | 247.5 ms | 815.1 ms | 42 |
| 0 | 0 | — | — | — | 0 |
My hypothesis collapsed. Even with the 75 ms target, the brake fired 23 times across 45 seconds of total load and stepped down 34 times. I found the reason back in the code: latency_exceeded() checks something else before it looks at the minimum latency. It measures how long a pending synchronous request has gone unanswered, and if that time is longer than the window, it declares a violation without consulting the statistics at all. The lines in the trace file are exactly that: 184 ms, 288 ms, 369 ms, 429 ms, 480 ms — one single pending read, growing.
So under a write storm, reads really do stall for hundreds of milliseconds. That's why the brake presses even at 75 ms. My assumption that "this threshold is never crossed" amounted to confusing the disk's behaviour on a healthy day with its behaviour under load. A classic mistake; this time it was mine.
But where does the difference land?
The two settings fire a similar number of times. The real difference is when they fire. The smallest violation recorded at 75000 was 83.3 milliseconds, with a median of 290 milliseconds. So the brake engages at best after reads have stalled for 83 ms, and typically only after they've passed a quarter of a second. At 2000 the smallest violation was 4.9 milliseconds.
Turn the number around and it gets sharper: of the 32 violations recorded at 2000, seven were below 75 ms. Roughly a fifth of the signal. Those seven were unlikely to show up at the default setting — I won't say "impossible", because the pending-request branch above looks at the window rather than the target, and the window shrinks to 50 ms, so a record in the 50-75 ms band is theoretically possible. But the shape is clear: the brake isn't broken, it's late.
You can also watch the steps move in the trace — and the ladder isn't one-directional. The resting maximum depth is RWB_DEF_DEPTH, i.e. 16. When only writes are in flight, the code temporarily raises the depth upward: 31, 61, 121 and finally 192. That last one is a ceiling, three quarters of nr_requests (256 × 3/4). Once reads arrive and latency degrades, the ladder comes down: 16, 8, 4, 2, 1.
A depth of 1 means background writeback gets a single request in flight — the brake is floored. Meanwhile the monitoring window shortens too: from 100 ms to 72.7, then 59.3, then 50. The harder the brake presses, the more often it looks.
Those numbers may look odd at first; the window comes out at 72.7 rather than 100/√2 = 70.7. The reason is that the code uses an integer square root: int_sqrt(512) returns 22 rather than 22.6, and the division runs against that rounded-down root. A small detail, but one that will cost half an hour to anyone verifying the number by hand.
The depth arithmetic lives in calc_wb_limits(): normal writes get half the maximum depth, background writeback a quarter, both rounded up. When the maximum depth drops to 2 or below, a special branch takes over and normal writes get the maximum depth itself while background is pinned to 1. The last two steps in my trace fall squarely inside that branch. Writing a target of 0 zeroes both, i.e. the limit disappears.
Now for what I could not measure, because not stating it would be a con: I found no reliable difference in the latency the reader actually saw across the three settings.
In my baseline runs without a writer, read throughput swung between 7,810 and 11,327 IOPS across four repetitions; in an earlier run that I eventually discarded, the same baseline reached 17,320. So the environment's own noise is 45% on the kindest reading, and more than twofold across every observation. The differences under load stayed inside that band. What's more, the single worst run in the whole set came at 2000, the setting that should in theory do better: 252 IOPS and a p99 of 26.6 ms. I don't count that as evidence against 2000, because I saw the reverse too — that's how wide the band is.
Stack a virtual machine, the Mac's own filesystem and the host cache on top of each other and the measurement can't hold that precision. The trigger counts are deterministic; the latency gain is not. I can't write up a result I don't have.
One more limit: my lab ran under mq-deadline, while every one of my servers uses none. Under none, nr_requests is pinned to the hardware depth, so both the ceiling and the floor of the ladder start somewhere else. Don't carry these numbers to the fleet one-for-one.
What BFQ quietly switches off
I saved the most annoying finding for last, because it hits my own advice directly.
In the piece about who actually listens to ionice, the conclusion was this: none and kyber don't see I/O priority at all, mq-deadline distinguishes only the class, and the one scheduler that takes it seriously down to the level is BFQ. For someone who wants to push a backup job into the background with ionice -c3, the natural move out of that piece is to switch to BFQ.
That move has a side effect. block/bfq-iosched.c runs these two lines while setting the queue up:
blk_queue_flag_set(QUEUE_FLAG_DISABLE_WBT_DEF, q);
wbt_disable_default(q->disk);
BFQ turns wbt off because it has its own congestion control. Defensible as a design decision. The problem is that it isn't written down anywhere. I tried it in order on a freshly booted machine:
1. boot (untouched) sched=mq-deadline wbt=75000 enabled=1
2. BFQ sched=bfq wbt=0 enabled=3
3. kyber sched=kyber wbt=75000 enabled=1
4. back to mq-deadline sched=mq-deadline wbt=75000 enabled=1
The instant you write echo bfq > scheduler, wbt_lat_usec drops to zero. No warning, no log line. Note that kyber does not do the same — kyber has its own latency targets but doesn't touch wbt, so that disk runs two separate brakes stacked on top of each other.
The enabled column comes from debugfs and maps one-to-one onto the counter in the code: 1 is on by default, 2 on manually, 3 off by default, 4 off manually. Switching to BFQ drops the state to 3 — "off by default".
Here's the subtle part. wbt_disable_default() only does anything while the state is 1. If you have written a value by hand at any point, the state becomes 2 and BFQ can no longer touch your setting:
5. manual 2000 sched=mq-deadline wbt=2000 enabled=2
6. BFQ (after manual write) sched=bfq wbt=2000 enabled=2
7. undo with echo -1 sched=mq-deadline wbt=75000 enabled=2
The last line hides a trap. echo -1 is documented as "reset to the default" and it does return the number to 75000 — but it leaves the state at 2. So you get the number back, not the state. Once you've touched it by hand, that queue counts as manually managed until the device is re-registered, which in practice means a reboot. That may sound trivial; it isn't for whoever switches to BFQ and burns an hour asking why the brake didn't switch off.
The writes the brake never touches
So far we've only talked about when the brake presses. The question people get wrong in production is a different one: what does it actually throttle? The answer is in wbt_should_throttle() and it's surprisingly narrow:
static inline bool wbt_should_throttle(struct bio *bio)
{
switch (bio_op(bio)) {
case REQ_OP_WRITE:
/*
* Don't throttle WRITE_ODIRECT
*/
if ((bio->bi_opf & (REQ_SYNC | REQ_IDLE)) ==
(REQ_SYNC | REQ_IDLE))
return false;
fallthrough;
case REQ_OP_DISCARD:
return true;
default:
return false;
}
}
Three lines of summary: wbt throttles buffered writes and discard, nothing else. O_DIRECT writes are exempted by name in the comment, and reads are exempt via the default branch. So if your database writes its data with O_DIRECT — PostgreSQL's WAL under the right settings, many databases' data files, VM images with cache=none — then whatever you write into wbt_lat_usec changes nothing. If you have a latency problem and most of your disk traffic is direct I/O, the hours you spend chasing this setting are wasted.
The read latency the brake protects is measured, but it is not the side being squeezed; wbt uses reads purely as a thermometer. All the throttling happens on the write side.
One more trap: changing nr_requests resets the ladder through wbt_queue_depth_changed() — the step counter returns to zero and the throttling lifts at that moment. The target you wrote by hand and its state survive. Worth knowing that the brake is quietly released in the middle of a tuning round that touches queue depth.
What I did
I'd love to end this piece with "go write 2000 right now", but my measurement doesn't earn that sentence. Here's the decision framework I'm applying to my own fleet instead.
Look first, decide second. One line is enough, and it changes nothing:
for d in /sys/block/[sv]d*; do
printf "%-8s rot=%s sched=%-12s wbt=%s\n" "$(basename $d)" \
"$(cat $d/queue/rotational)" \
"$(sed -n 's/.*\[\(.*\)\].*/\1/p' $d/queue/scheduler)" \
"$(cat $d/queue/wbt_lat_usec 2>/dev/null || echo none)"
done
If you see rot=1 next to wbt=75000 and nothing is actually spinning under that disk, your default rests on a false premise.
Know what the disk really is. Knowing that your cloud provider sells SSDs isn't enough; whatever lsblk -o NAME,ROTA,MODEL reports is what the kernel sees too. If there really is a rotating disk there, 75 ms is already the right value — leave it alone.
If you're going to change it, change it by measuring. That's what the trace is for, and it's cheap:
T=/sys/kernel/debug/tracing
echo 16384 > $T/buffer_size_kb # the default ring overflows on a busy disk
echo > $T/trace; echo 1 > $T/events/wbt/enable
# ... run your workload ...
echo 0 > $T/events/wbt/enable
grep -c 'scale down' $T/trace
awk '/overrun/{s+=$2} END{print "lost events:", s+0}' $T/per_cpu/cpu*/stats
Don't skip the last line. The trace file is the current contents of the ring buffer, not a cumulative counter; if the buffer overflows you undercount, and nobody tells you. If overrun isn't zero, grow the buffer and repeat. No step-downs at all means you don't need the brake; far too many means the target may genuinely be loose. Unlike my latency measurement, this number held up against the noise.
If you want to see the brake's current state, debugfs exposes all of it — the files are root-readable only:
ls /sys/kernel/debug/block/sda/rqos/wbt/
# curr_win_nsec enabled id inflight min_lat_nsec unknown_cnt wb_background wb_normal
If you make it permanent, use udev, not rc.local — the value is decided at device registration. Here's the rule I tested in the lab; after udevadm control --reload I watched udevadm trigger turn 75000 into 2000:
ACTION=="add|change", SUBSYSTEM=="block", KERNEL=="sd*|vd*", \
ATTR{queue/rotational}=="1", ATTR{queue/wbt_lat_usec}="2000"
Don't stop at vd*: six of the seven disks in my fleet are sda. Write the rule for virtio only and you'll miss exactly the machines you meant to fix — which is what I did on my first attempt.
If you're switching to BFQ, do it knowingly. If you genuinely use ionice levels, BFQ may be the right choice, but know that you're switching wbt off in the same breath. If you want both, order matters: write a value to wbt_lat_usec by hand first, then change the scheduler. The other way round does nothing.
This setting matters more on disks using none. With no scheduler there's nobody handing out fairness in the queue; wbt is the only thing protecting read latency. All seven of my servers use none, and on six of them that single protection is set to 75 ms.
I haven't made a permanent change to my own fleet yet. First I'll count the step-downs under production load; if the number comes out zero, there's nothing to change, because the brake never presses anyway.
The bill for one bit
Two weeks apart, I've now dug up two separate consequences of the same flag. The first multiplied read-ahead by 970; this time it loosens the write brake by thirty-seven and a half. In neither case did anyone make a conscious choice; a driver planted a fixed flag years ago because it had no information to fill in, and the kernel gradually hung more and more decisions on that flag.
The lesson I take from this isn't really about tuning. In our systems, the gap between "the device says so" and "the device is so" turns into policy the moment we forget it's there. Virtualisation separated those two permanently; the kernel still behaves in many places as though they were never separated. Virtual machines aren't the exception any more, they're the majority.
And there's a lesson about myself in it too: I started this piece believing that the brake never presses at 75 ms, and the trace showed it pressing 23 times. Had I written without measuring, I'd have published a thoroughly convincing falsehood. That isn't why I like reading the mechanism in the source on this blog; it's why I like reading the source and then measuring anyway.
Top comments (0)