DEV Community

Cover image for I Cut Requests 31x, the IOPS Ceiling Did Not Budge
Mustafa ERBAY
Mustafa ERBAY

Posted on Originally published at mustafaerbay.com.tr

I Cut Requests 31x, the IOPS Ceiling Did Not Budge

There is a pleasant-looking number in my server's /proc/diskstats line: since boot, 3,282,202
write requests have been merged. In other words the kernel caught a quarter of the 13,918,839
write fragments that applications submitted and packed each one into a neighbour's request. At
first glance that reads like free money: fewer requests, fewer trips, a happier disk.

Here is the whole line — my numbers are derived from it, and it is printed as-is so you can derive
them too (the server had been up for 41,242 seconds):

   8       0 sda 639705 234489 36923566 ... 10636637 3282202 135302774 ...
#            reads: completed/merged/sectors    writes: completed/merged/sectors
Enter fullscreen mode Exit fullscreen mode

The average write request reaching the device — 135,302,774 sectors over 10,636,637 requests — is
6.36 KiB.
So merging does work, but what it achieves is turning 4 KiB fragments into 6 KiB fragments. That
does not look much like the "fewer trips" story.

What I actually wanted to know was this: is that counter telling me something, or just keeping a
ledger? Five days ago a similar question came up on the CPU side —
CPU quota looks at the window, not the average.
I started on the block side with the same suspicion, and the answer came back harder: I improved
merging thirty-one fold and it changed nothing at all where the limit actually binds.

First, the vocabulary: the three levels of nomerges

The knob that disables merging is /sys/block/<disk>/queue/nomerges. The kernel's stable ABI
document has described it since January 2010 like this: setting it to 1 disables complex merge
checks but keeps the simple one-shot merge with the previous request; setting it to 2 disables all
merge attempts. The default is 0, meaning every kind of attempt is enabled.

On the source side the same taxonomy sits as a comment in block/elevator.c:

/*
 * Levels of merges:
 *      nomerges:  No merges at all attempted
 *      noxmerges: Only simple one-hit cache try
 *      merges:    All merge tries attempted
 */
Enter fullscreen mode Exit fullscreen mode

Clean so far. That was the expectation here too: 0 fully on, 1 half on, 2 off. Measuring showed
the first two are wrong: 0 is not on under every condition, and 1 disables nothing on this path.

The lab: turning the knob is not the same as changing how you submit

The experiment runs on my own machine: the Docker Desktop linuxkit VM on a Mac, kernel
6.10.14, cgroup v2, 14 vCPUs. The target device is vda — a virtio disk with mq-deadline as
the active scheduler, nr_requests=256, max_sectors_kb=1280. The write target is a Docker
volume on vda1, the workload is fio doing 4 KiB sequential O_DIRECT writes with iodepth=64,
256 MiB per run. Every arm ran three times and I took the counters as before/after deltas from
/proc/diskstats.

I varied two things: the nomerges value and fio's submission batch
(iodepth_batch_submit). The second one decides how many requests the application puts on the
table at once when it tells the kernel "take these".

Arm Device requests Merged bios Average request
nomerges=0, batched (32) 2,074 / 2,108 / 2,086 63,522 / 63,627 / 63,634 126.5 / 127.0 / 126.1 KiB
nomerges=1, batched (32) 2,067 / 2,142 / 3,971 63,510 / 63,699 / 63,844 126.9 / 126.5 / 242.0 KiB
nomerges=2, batched (32) 66,326 / 65,748 / 65,644 0 / 0 / 0 4.1 / 4.0 / 4.0 KiB
nomerges=0, one at a time (1) 65,628 / 65,592 / 65,554 219 / 166 / 26 4.2 / 4.0 / 4.0 KiB

The first three columns come from the counters, the last one from the sector delta; since vda sees
writes other than my fio, the two do not multiply out exactly (the starkest example is the third
nomerges=1 repeat, discussed below). I printed all three repeats, because the real information
is in the consistency. 65,536 writes of
4 KiB collapse into 2,074 device requests under batched submission — a factor of 31.6. Across the
three repeats: 31.6 / 31.1 / 31.4. A spread that tight says the measurement comes from the
mechanism, not from noise.

The last row is the whole point. nomerges=0 means merging is fully enabled; even so, the device
saw 65,628 write requests and only 219 merges happened. Three in a thousand. The knob is on and the mechanism is idle. The reason
is in the first lines of the function in block/blk-merge.c:

bool blk_attempt_plug_merge(struct request_queue *q, struct bio *bio,
                unsigned int nr_segs)
{
    struct blk_plug *plug = current->plug;
    struct request *rq;

    if (!plug || rq_list_empty(plug->mq_list))
        return false;
Enter fullscreen mode Exit fullscreen mode

Merging happens if there is already a request waiting in the list. In an application that submits
one at a time the list is empty: when each bio arrives its neighbour is either still on its way or
long gone to the device. So merging is not a knob, it is the reward for how you submit. What cut
requests 31-fold was not nomerges, it was the application handing over 32 requests at once.

nomerges=1 disabled nothing on my path

The second row of the table does not show the "half on" state the document describes: with 1 I
still got 63,510 merges, the same as with 0. The source explains why. nomerges=2 sets
QUEUE_FLAG_NOMERGES in the kernel, while nomerges=1 sets QUEUE_FLAG_NOXMERGES
(block/blk-sysfs.c). But only two call sites actually make a decision on that flag (a third merely reads it back for
sysfs), both inside block/elevator.c, and both sit immediately before a hash lookup.
The merge path that goes through the plug list — the one in the table above — never asks about that flag;
blk_mq_attempt_bio_merge only checks blk_queue_nomerges.

In practice that means nomerges=1 matters if the elevator's hash lookup is in play. On my
production server sda has none as the active scheduler — no elevator at all. Writing
nomerges=1 there does precisely nothing. A document written correctly in 2010 now tells half the
story, after merging's centre of gravity shifted to the plug list. I ran into this kind of
scheduler side effect before
on the ionice side;
take the scheduler away and the settings it carries quietly go out of service with it.

In the third repeat the nomerges=1 arm reported a 242 KiB average request and a sector delta
three times the others. Something else — the host's own writes — got into that measurement. I left
the row in the table because the merge count was still 63,844, but I am not using that 242 KiB
cell as a finding.

If the lever is the submission batch, how far does it go?

The obvious next question: if I keep growing the batch, does the request keep growing with it? The
device allows a 1,280 KiB request (max_sectors_kb) and accepts 254 segments. With nomerges=0
fixed, I varied only the submission batch. This is a separate series and each point is a single
run — the only claim here is that the plateau exists, not the decimals in the cells:

Submission batch Device requests Bios per request Average request
1 65,561 1.0 4.0 KiB
32 2,067 31.7 126.9 KiB
128 531 123.5 494.0 KiB
256 531 123.5 494.2 KiB

Going from 32 to 128 made the request four times bigger; going from 128 to 256 changed nothing.
The plateau is at 494 KiB, well below the 1,280 KiB the device allows — so the ceiling is not the
device's, it is somewhere in the submission path. The kernel flushes the plug list on two
conditions: when the list reaches BLK_MAX_REQUEST_COUNT requests (32, or 64 with multiple
queues), or when the last request grows past BLK_PLUG_FLUSH_SIZE (128 KiB). Which of the two
produces 494 I did not trace; all that goes on the record is that the plateau exists and that it
is not the device's.

So what did 31× fewer requests buy? I could not measure it

Here is where I have to be honest. On the bandwidth side I have no usable result. Running the same
setting three times, fio reported 462.9 — 800.0 — 879.7 MiB/s. Same arm. Same file size. A 90%
spread. Behind this virtual disk sit the host's APFS and its own cache; what I am measuring there
is a file, not a disk.

Inside a distribution like that, saying "merging made it this much faster" would be selling noise
as a finding. Across every arm, bandwidth scatters between 373.7 (merging on, one-at-a-time) and 879.7 (merging
on, batched) MiB/s with no consistent ordering — the gap between two repeats of one setting is
larger than the gap between two different settings. So
the only sentence that goes on the record is this: cutting the number of requests reaching the
device to a thirty-first produced no measurable change in speed in this setup.

That was not a disappointment; it meant the question had to be built again from scratch. The value of merging
shows up not in speed but in where the request gets counted. So I added something that counts.

With a ceiling in place, the answer got clear

cgroup v2's io.max file takes four limits per device: rbps, wbps, riops and wiops.
Docker sets this up with the --device-write-iops flag. I started the container with
/dev/vda:2000 — 2000 write operations per second — and measured the same workload in two arms
with time-based 15-second runs.

Arm Device requests Merged Total bios fio IOPS Written
nomerges=0 (merging on) 943 / 1,074 / 2,970 29,231 / 29,554 / 28,416 30,174 / 30,628 / 31,386 2,002 / 2,002 / 1,841 117 / 117 / 108 MiB
nomerges=2 (merging off) 33,992 / 30,140 / 30,481 0 / 0 / 0 33,992 / 30,140 / 30,481 2,003 / 2,002 / 2,002 117 / 117 / 117 MiB

In the merging-on arm 30,174 bios fit into 943 device requests — a thirty-two-fold compression;
in the merging-off arm the request count is whatever the bio count is. Even so, in two of the
three repeats both arms land in exactly the same place, pinned to the ceiling: 2,002 IOPS,
7.8 MiB/s, 117 MiB.

In the third repeat the merging-on arm stopped at 1,841 IOPS. Device requests in that run had
climbed to 2,970 (943 and 1,074 in the other two), so the host's own writes got into the
measurement. I am not presenting that as merging doing harm — but in none of the three repeats did
the merging-on arm get ahead of the merging-off one. There is no negotiating with the ceiling.

One more detail: the limit is 2000 and the measurement is 2,002. The reason sits in
tg_within_iops_limit in block/blk-throttle.c: the allowance is computed by rounding the
elapsed time up to the next throttle slice (roundup(jiffy_elapsed + 1, throtl_slice)) and adding
any carryover_ios. That is where the one-in-a-thousand overshoot comes from; the ceiling is hard,
but its edge is not smooth.

The arithmetic closes too: 2000 ops/s × 15 s = 30,000 bios, × 4 KiB = 117.2 MiB. In the
merging-on arm, device requests (943) plus merged bios (29,231) add up to 30,174. So the ceiling
counted the ~30,000 bios the application submitted, not the 943 requests the device saw.

Why it works that way is written on a single line in block/blk-throttle.c:

static void throtl_charge_bio(struct throtl_grp *tg, struct bio *bio)
{
        ...
        tg->io_disp[rw]++;
Enter fullscreen mode Exit fullscreen mode

The counter goes up by one per bio; it does not care whether that bio is 4 KiB or 126 KiB. And the
charge is taken before merging: in block/blk-core.c, submit_bio_noacct first calls
blk_throtl_bio(bio); if the bio hits the ceiling it is queued there and the path stops, and only
otherwise does it continue through submit_bio_noacct_nocheck. Merging is attempted much later,
inside blk_mq_submit_bio. Putting two parcels in the same box after the invoice is paid does not
refund the shipping.

Diagram

When does merging actually pay off?

Neither regime in this article showed a gain, but that does not mean merging is useless. The gain
accumulates where cost is charged per request: head movement on a spinning disk, a device queue
with shallow depth, network storage with a per-request round trip. I have none of those — no
spinning disk, no real NVMe hardware. So what this article pins down is not the limits of
merging's usefulness; it is which kind of limit is indifferent to it.

The same holds for byte-based ceilings (wbps), from the opposite direction: that counter counts
bytes (bytes_disp[rw] += bio_size — unlike the io_disp increment, that line sits inside a
condition so an already-charged bio is not counted twice), and since merging does not change the
byte count, it is irrelevant there too. In other words, none of io.max's four limits negotiates with merging. I
covered the division of labour between io.max, io.weight and io.latency
in a separate article before;
this measurement is where I saw how hard that "hard ceiling" really is.

What I changed on my own server

Nothing. And that is the result of the measurement: sda keeps nomerges=0, the scheduler is
none, and the write merge rate is 23.6%. The default is the right value; there is nothing to
change. In 11.5 hours it wrote 64.5 GiB at an average of 337 bios per second — nowhere near any
ceiling.

The lesson is forward-looking. If a service writing to an IOPS-limited volume (a cloud disk, an
io.max rule, shared storage) is slow, fiddling with nomerges or the scheduler is a lost
evening. The only real lever there is the size of the bios the application submits: larger blocks,
buffered writes, fewer fsync calls. Kernel merging cannot do that for you, because you have
already paid the bill before you submit.

Questions for your own setup

  1. cat /sys/block/<disk>/queue/scheduler — if the active scheduler is none, nomerges=1 is already meaningless, while nomerges=2 still bites.
  2. Divide field 9 by fields 8 plus 9 in /proc/diskstats (writes merged over total write bios): that is your merge rate. Mine is 3,282,202 / 13,918,839 = 23.6%.
  3. Divide field 10 by field 8, then by 2: that is the average request size reaching your device, in KiB. Mine is 6.36 KiB — so much for "large requests".
  4. cat /sys/fs/cgroup/<path>/io.max — if there is a wiops limit, the answer to your speed problem is not in merging but in the number of bios you submit.
  5. If you are going to measure, run the same setting at least three times. My bandwidth column is missing from this article for exactly that reason.

What I did not prove

I did not measure a spinning disk; that is merging's true homeland and I do not own that device. I
did not measure how request count affects latency on real NVMe hardware — on the virtual disk in this lab, bandwidth swung 90% between repeats, so that question stays open. I did not get into how
the io.latency and io.cost controllers view merging; because io.cost's cost model carries
per-byte and per-operation coefficients, the story there is probably different — but "probably"
does not count as a finding on this blog.

The one sentence I am left with is this: the block layer's merge counter tells you how hard the
kernel is working, not how fast your system is. And if you do not know where a limit is charged,
every setting you believe improves that limit is really just making the ledger look nicer.

Official Sources

Top comments (2)

Collapse
 
ndcodes profile image
Nnamdi Felix Ibe •

The parcel line is the whole article, and I'd push on one sentence in your conclusion, because your own framework says it should be split into two.

You group "a cloud disk, an io.max rule, shared storage" as places where fiddling with merging is a lost evening. For io.max you proved it. For EBS it's the opposite, and for the reason you identified.

cgroup charges per bio, upstream of merging. EBS counts downstream, in fixed units. AWS documents a maximum counted I/O size of 256 KiB on SSD volumes, so your 65,536 writes of 4 KiB are 65,536 counted operations one at a time, and roughly 2,074 once merged into 126 KiB requests.

HDD is starker still. AWS's FAQ is explicit that a non-sequential I/O on st1 is processed as a full megabyte however small it actually is. Merging isn't an optimisation there, it's the performance model.

So your thesis survives and gains a second worked example. The invoice is paid before merging under cgroup and after merging on EBS. Nothing about the kernel changed. Only where the meter sits.

Which makes "divide field 10 by field 8, then by 2" the most useful line in your checklist for anyone on AWS.

Collapse
 
merbayerp profile image
Mustafa ERBAY •

That’s a very good catch.

You’re right — I stretched the conclusion further than the experiment supports.

What I actually demonstrated is that merging cannot help an io.max IOPS ceiling because the accounting happens at the BIO level before the merge. Extending that directly to “cloud disks” as a category was too broad.

EBS is a great counterexample precisely because it reinforces the underlying point rather than contradicting it: the accounting boundary matters more than the merge itself.

Under cgroup io.max:

BIO → charge → merge → request

So once the BIO has been charged, merging requests later cannot recover those IOPS.

With EBS, if the service accounts I/O after the kernel has produced the requests — and according to AWS’s documented I/O accounting rules it does — then changing the request shape before that boundary can absolutely change what gets counted.

So the stronger version of my conclusion should probably be:

Before optimizing request merging, find out where the system enforcing your limit counts I/O.

If it counts before merging, merging cannot buy back IOPS.
If it counts after merging, request shape may be part of the performance model itself.

And yes, that makes the diskstats average-request-size check considerably more useful for EBS than I gave it credit for.

I’m going to correct that wording in the article. Thanks for catching the boundary I overgeneralized — this is exactly the kind of comment that makes a measurement article better. 🙂