DEV Community

Why Amazon FSx for NetApp ONTAP block performance differed by about 2x on the same specified values, isolated by measurement

Why Amazon FSx for NetApp ONTAP block performance differed by about 2x on the same specified values, isolated by measurement

While measuring, I hit a case where the templates and the specified values matched, yet the number
moved every time I measured. Rather than let it go, I wanted to isolate whether it was a
measurement mistake or something about the environment, so I eliminated the causes one at a time.

The same thing shows up when a post-migration performance test does not agree, and when you want to
settle on one representative value while sizing an initial build. It is a question that eats a lot of
time, so this article is the record of that isolation.

What this article does and does not cover: it covers investigating the cause of the performance
swing, and how to isolate the bottleneck using four CloudWatch metrics. It does not cover mount
procedures, how to confirm multipathing, or figures usable for sizing (details under "What this
article does not cover" below). All figures are from a Single-AZ configuration; Multi-AZ is
unverified.

dev.to series: FSx for ONTAP Block Protocols, Measured.

A 1 MiB sequential read on Amazon FSx for NetApp ONTAP (hereafter FSx for ONTAP), measured with the
same template, the same specified values and the same workload, varied by 2.64x across
environments.
Each one was stable within its own run. This is not measurement scatter.

This is the record of eliminating the causes one at a time. It was not the path, not how the
connection was made, not the queue count, not the distribution, and not the iopolicy (Linux's
multipath setting for distributing I/O across paths to an NVMe device). What was
acting on it was the layout on disk, which breaks after 300 seconds of 4 KiB random writes and
comes back when the data is rewritten sequentially.

Three times along the way I nearly settled on the wrong explanation. First: "if it exceeds the
specified value it must be served from memory." Second: "the layout must also be setting the
amplification factor for 4 KiB random reads." Third: "the amplification factor must be decided by
whether 4 KiB random writes ever ran on that device." All three were disproved once I aimed a
measurement that could decide them.
All three are in here. The third had five observations
lined up cleanly behind it and survived until a control experiment killed it.

What was measured

Assuming no prior knowledge of FSx for ONTAP, here's what varies between the deployments.

  • Deployment 1: AWS Cloud, single VPC, single AZ. One EC2 client -> (NVMe/TCP, 1 MiB sequential read) -> Amazon FSx for NetApp ONTAP (same template, same specified values)
  • Deployment 2: same AWS Cloud, single VPC, single AZ. One EC2 client -> (same connection, same workload) -> Amazon FSx for NetApp ONTAP (same template, same specified values)
  • The invisible difference between Deployment 1 and 2: layout on disk. It never appears in the template or the list of specified values

Two FSx for ONTAP file systems built from the same template and the same specified values<br>
(Deployment 1, Deployment 2), each measured from one EC2 client with the same connection and the<br>
same workload. Shows what can be held identical (path, connection setup, queue count,<br>
distribution, iopolicy) against what cannot (layout on disk)

Figure: **matching the CloudFormation template and the specified values doesn't make the layout
on disk match across a redeploy.
* Path, how the connection is made, queue count, distribution,
and iopolicy can all be held to the same value across deployments. Layout is the one condition
that changes every deployment and never appears in the list of specified values.*

Item Role in this article
Deployment 1 / Deployment 2 Separate FSx for ONTAP file systems, built from the same template and the same specified values
Conditions that can be matched Path, how the connection is made, queue count, distribution, iopolicy
Condition that can't be matched Layout on disk. Breaks after 300 seconds of 4 KiB random writes, comes back when rewritten sequentially

What this article does not cover. Mount procedure is out of scope
(repository record, Japanese).
Confirming whether multipathing is usable is also out of scope (separate article).
The figures in this article cannot be used for sizing either. Why they cannot is the subject of
the article. All figures are from a Single-AZ configuration; Multi-AZ is unverified.

One of three, but there is no reading order. The block-protocol measurements are split by
reader.

Article Reader
A Sessions and queues People about to mount
B Confirming whether multipathing is usable People who followed the procedure and got stuck
C What moves the numbers (this article) People who measure the numbers, and people who cite them

None of the three presupposes the other two. Read only the part you need.

The same property on the file-protocol side (NFS / SMB / S3 API) is covered in
S3 Burst Part 2
, with the primary record of those figures in
throughput, IOPS and concurrency (Japanese).
Those figures are also iorate=max values from this environment, so the way of measuring
described here applies to them unchanged.

The environment measured

Item Value
File system FSx for ONTAP second generation SINGLE_AZ_2, 6,144 MBps x 1 HA pair, SSD 4,096 GiB, provisioned 200,000 IOPS
Region ap-northeast-1, single AZ
ONTAP Not one version. It spans 9.18.1P5 (the early deployments), 9.18.1P6 (the per-path measurements) and 9.18.1 (the last three deployments). This is exactly the unresolved candidate left at the end of this article (the cross-article version mapping is in the ONTAP version matrix)
NVMe read cache Disabled (confirmed on both nodes)
LUN / namespace 600 GiB, on a 900 GiB volume
Before measuring All 600 GiB was written once. Unwritten blocks on a thin-provisioned volume return zeros, so reading without writing first measures nothing
Client c5n.9xlarge; 50 Gbps network is the guaranteed figure
Instrument VDBENCH 5.04.07 directly, iorate=max (unlimited), 512 threads, 60 s warm-up + 300 s measurement
Test data / efficiency settings Not recorded whether the VDBENCH payload was compressible, or whether volume inline efficiency (StorageEfficiencyEnabled) was enabled. This article's subject is on-disk layout, so whether efficiency settings affected results is an unconfirmed variable

The window is 300 seconds, so the figures in this article include burst. They cannot be cited as a
baseline.

Every figure in this article is an iorate=max saturation point, not an operating point.
No target IOPS was given; it is what the path would take. This distinction bears on the subject
of the article itself.
A saturation point answers "where does it jam", and does not answer "at
what response time can it be used".

The file-protocol side of the same repository measured that gap. With a target of 4,400 IOPS the
result was 4,406 MB/s at 8.86 ms; unlimited, it was 4,204 MB/s at 121.79 ms.
Unlimited is 5% lower in throughput and 14x worse in response time.
Recording one iorate=max point as "the ceiling" leaves behind a figure that is lower than what
is reachable and an order of magnitude worse in response time.

Response time was recorded as measured in only the two series under "not concurrency" below.
Elsewhere the values VDBENCH printed were not kept, and the environments are torn down. But with
512 threads fixed and no rate limit, response time is the identity 512 / IOPS (derived 216.6 ms
against measured 216.588 ms, confirmed to agree). The figures below can be read through that
identity
: 1,288.87 MB/s is about 397 ms, 3,407.60 MB/s about 150 ms, 546.02 MB/s about
937 ms. All three are derived, not measured.

The response-time-versus-load curve was not measured. The file-protocol side produced a curve
across target IOPS values and a cutoff (the IOPS at which a given latency is reached) with
auto_vdbench, but the same thing could not be
produced on the block side.
auto_vdbench creates test files on a mount path and cannot drive a
raw device or an NVMe namespace. Open work.

The same configuration varied widely

This is the part of the article I would most like read carefully.

As F-3, the same shape was measured three times within one connection.

Run 1 MiB sequential read
1st 3,401.49 MB/s
2nd 3,407.60 MB/s
3rd 3,407.59 MB/s

Within one connection it reproduces to 0.18%. The second and third differ by 0.01 MB/s.

But three environments built with the same procedure, the same template and the same negotiated
result (4 queues)
produced different values.

Environment 1 MiB sequential read Stability within the window
First 1,288.87 MB/s stable
Second 2,366.49 MB/s stable
Third 3,401-3,408 MB/s stable (0.18%)

The spread is 2.64x. Each was stable within its run, so this is not measurement scatter.

So you cannot size from a single figure of the form "NVMe/TCP throughput on this configuration is
N MB/s".
At the point where I had measured the second environment, I was about to write 2,366.49
as the representative value. I had not followed my own rule of repeating three times, so I noticed
one environment later than I should have.

I had put forward one candidate at the time: Asymmetric Namespace Access (the mechanism that
tells the host which of several paths to the same namespace is the shortest one, referred to as
multipathing throughout this article) is unusable, so
which of the two controllers the device under measurement sits behind can change per deployment.
Measurement said otherwise.

The path candidate is gone

I brought in a kernel with multipathing enabled (Rocky Linux 9.7) and measured within one deployment, on
one namespace, with the paths narrowed to one at a time.

Path 1 MiB sequential read
optimized only (the node holding the namespace) 1,278.11 MB/s
non-optimized only 1,279.07 MB/s

0.08% apart. Either controller gives the same figure, so the path cannot explain that spread.

How the connection is made was eliminated too. Connecting from optimized first, from non-optimized
first, and using connect-all as the AWS procedure has it, gave 1,276.60 / 1,277.38 / 1,277.28
MB/s. A spread of 0.06%. The queue count was the same in every configuration
(queue_count=5, 10 established sockets to :4420).

What was acting on it was when the data being read was last written

Within one deployment, one device, one path and one policy, rewriting the 600 GiB brings it back.

State 1 MiB sequential read
Immediately before the rewrite (control) 1,239.15 MB/s
Immediately after the rewrite 2,298.54 MB/s
Once more, right after that 2,296.12 MB/s

1.85x. And the second read does not fall back. This is not a property that is consumed by one
read.

This is the only place in this article where a ratio is written. The numerator and denominator
come from the same file system, the same device, the same path and the same policy: the reason a
ratio can be written here is the same reason none is written in the other sections.

CloudWatch answers what the cold side is running into.

State DiskIopsUtilization Disk throughput utilization Network utilization
Cold read avg 57-60 / max 95.0-95.4 30 / 42 8.7 / 11.6
Warm read avg 13.6-14.4 40 / 41 19.2 / 19.5

The ceiling on the cold side is SSD IOPS. It is not the 6,144 MBps of throughput capacity that
was specified (utilization 42% or below). It is not the network either (11% or below).
And at 2,298 MB/s on the warm side, none of the utilization figures FSx for ONTAP publishes is at
its ceiling.
What is acting on it there is not known.
(What is being read here is DiskIopsUtilization / disk throughput utilization / network
utilization from the
second-generation file system metrics.
The utilization metrics emit one point per HA pair and per aggregate.)

Not reads, and not elapsed time

Since it moves with "when it was written", I aimed separate measurements at what brings it down.
On top of one fill, reads, elapsed time and writes were applied in turn.

Immediately after 1 MiB sequential read
The fill 2,263.8 MB/s
A second read straight after 2,267.5
A third read straight after 2,268.4
30 minutes idle 2,258.5
300 seconds of 4 KiB random writes 1,272.8
Refilling 2,235.4
  • Reads do not bring it down. Three full passes over 600 GiB, a spread of 0.2%
  • Time does not bring it down. 30 minutes with nothing running, within 0.4%
  • 4 KiB random writes bring it down. 2,258.5 to 1,272.8, a fall of 1.77x

Do not generalize this to "writes bring it down". What was aimed was one point: 300 seconds of
4 KiB random writes. The result of separating shape from duration and aiming again is in
the section below, and the generalization
did not hold.

What falls is not cache but the layout on disk

I first explained this as "cache retention". That was wrong. Read CloudWatch interval by
interval and the mechanism comes out plainly.

State Client One read on the disk side Disk IOPS Against the specified 200,000 DiskRead / DataRead
Right after the fill (5 intervals) 2,219-2,263 MB/s 87.3-89.6 KiB 25,743-27,404 13% 97-105%
After 4 KiB random writes (6 intervals) 1,151-1,267 MB/s 11.7-11.8 KiB 177,527-189,629 89-95% 169-188%

DiskRead / DataRead is about 100%, so it is reading from disk right after the fill too.
It is not being served from memory. What changed is the size of one read on the disk side, by a
factor of 7.5. Disk IOPS rises 7.3x as a result, and saturates at 89-95% of the specified value.

So: 4 KiB random writes break the layout on disk, and the same 1 MiB read becomes many small
reads. Reads do not change the layout, so they have no effect, and neither does time. Rewriting
sequentially restores a contiguous layout. This is also why only 4 KiB random reads are
unaffected
: a random 4 KiB is one read on the disk side whatever the layout is.

DiskIopsUtilization alone does not get you here. It tells you whether something is
saturated; it does not tell you why it saturated (because each read is small).

And it goes back and forth within a single run. I think this is the strongest evidence.

What was running One write One disk read
Fill, then 3 reads 1,024 KiB 89.1 / 89.6 / 88.3 / 87.4 KiB
The read after 30 minutes idle — 87.9 KiB
4 KiB random writes, then read 4.0 KiB 58.7 to 11.8 to 11.6 KiB
Sequential refill, then read 1,024 KiB 88.2 / 87.7 KiB

30 minutes idle does not move it, a 4 KiB write brings it down, a sequential write brings it back.
This is neither a difference between deployments nor measurement scatter.

Note: this d_io depends on the workload too. In 1 MiB sequential read intervals it is 88 or
11.7 KiB; in 4 KiB random read intervals it is 4-5 KiB. When comparing, match the client-side
size of one operation.
Without that, the layout difference and the workload difference mix.

As a benchmarking practice, I think this is the most portable thing here.
If you publish a sequential read figure, record whether small random writes were interposed
between writing and reading.
Without them, the value does not move however many times you read or
however long you wait. With them, it changes by nearly 1.8x. As the next section shows, 300 seconds
of sequential writes does not move it.

What breaks it is not writes but small random writes

This section was measured later. The wording above reads as "writes break it", and that was
too broad. On the same device, I separated the shape of the write from its duration.

Step 1 MiB sequential read Against the reference
Reference (not broken) 2,364.78 MB/s —
4 KiB random writes for 30 seconds 2,343.63 MB/s -0.9% (not broken)
4 KiB random writes for 300 seconds 546.02 MB/s -77% (broken)
After a sequential refill 2,364.18 MB/s -0.03% (fully restored)
Sequential writes for 300 seconds 2,365.59 MB/s +0.03% (not broken)
  • Sequential writes do not break it. 300 seconds of writing, 0.03%. What breaks it is small random writes only, not "writes" in general
  • 30 seconds does not break it, 300 seconds does. The threshold is somewhere in between, and where is not known
  • The unbroken side does not depend on the specified IOPS. 2,364.78 at 100,000, 2,366.23 at 200,000
  • For the broken side, proportionality cannot be written. The 546.02 MB/s here (specified 100,000) and 1,272.8 MB/s on another deployment (specified 200,000) give a ratio of 2.33 — it moves in the direction of the specified IOPS but not by 2.00, and the two are not from the same deployment. No run varied the specified IOPS while in the broken state, so this is not a control experiment

In practice, record whether you "read after small random writes", not whether you "read after
writes".
A workload that writes sequentially and reads sequentially is not subject to this.

The conclusion that you cannot size from a single point is unchanged. What changed is the
reason.
Not "an unexplained 2.64x", but "even on the same configuration it moves 1.8x
depending on whether small random writes were interposed after the data being read was last
written".

Why the 3,401-3,408 MB/s environment came out higher still remains unexplained.

The warm-side ceiling sits on the client side

In the warm state, none of the
published utilization metrics
was at its ceiling. I used two EC2 clients to decide which side it stops on.

What was measured 1 MiB sequential read
Rocky 9.7 (concurrent) 2,040.69 MB/s
RHEL 9.7 (concurrent) 2,163.86 MB/s
Total 4,204.55 MB/s
One EC2 client alone (same state) 2,265.7 MB/s

The total is 1.86x one client. The ceiling sits on the client side. The storage side still had
headroom; it was stopping at what one client could produce.

What determines the per-client ceiling was not measured. It is about 16-17 Gbps per client, short
of the 50 Gbps of a c5n.9xlarge.

Before you write "I measured the storage ceiling", try adding a client.
If the total grows, what you were measuring was the client's ceiling.

The same shape appeared on a path that is not block. On 2026-09-19, on the NFS path (first
generation 2,048 MBps, SSD IOPS 40,000, two c5n.9xlarge reading 300 GiB that does not overlap),
1,195.27 alone gave a total of 1,901.48 MB/s, 1.59x. Alone, disk throughput utilization was
63% and not saturated; with two clients it reaches 102-103%.

There, the single-client side was followed up the next day, and it turned out to be two settings
rather than a resource.
Transfer size (rsize 64 KiB to 1 MiB, +21%) and concurrency (8 to 16
threads, +14%) took a single client to 1,646.17 MB/s, at which point the file system side pins at
102.6%
(measurement, Japanese).

It was followed further (2026-09-20), and "the client side" split into parts. On second
generation 6,144 MBps / 200,000 IOPS with 4 KiB random reads, varying only how the working set is
divided, five ways:

Division Against the reference Disk side / specified 200,000
1 EC2 client / 1 SVM / 1 volume (reference) — 60.7%
Threads doubled only +8.9% 67.6%
2 volumes (same SVM) +1.0% 66.8%
2 SVMs +26.3% 83.1%
2 EC2 clients +52.7% 96.6%

What acts on it is neither thread count nor volume count, but the number of IP addresses the client
talks to and the number of clients.
With two clients the disk side reaches 96.6% of the specified
value, and only there does the file system side become the limit
(measurement, Japanese).

The same separation was not done for the 16-17 Gbps per client in this section. There is no
block-side measurement varying transfer size, concurrency and the number of peer IP addresses.
Had I followed it that far instead of stopping at "the client side", the block side might also
have turned out to be settings and division.

"Measuring with one client does not determine the kind of ceiling" holds across protocols, so the
procedure in this section is portable.

At least it is not concurrency

The candidate "there must not be enough outstanding requests" fell. In the unbroken state, I
varied thread count alone by 4x.

Threads 1 MiB sequential read Response time (measured)
512 2,363.91 MB/s 216.588 ms
1,024 2,364.28 MB/s 433.283 ms
2,048 2,365.58 MB/s 860.258 ms

Throughput moves by 0.07% and only response time scales precisely. Applying 4x the threads does
not raise the per-client ceiling; only the time requests spend waiting goes up 4x.

This table is also where the hazard of iorate=max is most visible in this article.
"2,364 MB/s" is the same on all three rows, but the response time behind it spans 216 ms to
860 ms.
Write throughput alone and that difference disappears.

The file-protocol side was the same. NFS 4 KiB random reads at 512 / 1,024 / 2,048 threads gave
94,949.6 / 97,915.3 / 97,014.5 IOPS (a spread of 2.2%), with response times of
5.391 / 10.456 / 21.108 ms (measured). That ceiling is not concurrency either.
Note that 94,949.6 is from a deployment with half the SSD capacity (4,096 against 8,192 GiB), and is
0.7% from the original measurement of 95,630. Between those two points, capacity was not
acting on it.

Do not cite those two figures as a same-conditions reproduction. Re-measured on 2026-09-20 with
a 600 GiB working set, the reference was 81,313.9, 14% below the figures above.
The original measurement did not record its working set, so the conditions cannot be matched.

Match the time of day when comparing configurations

My first comparison measured four configurations in sequence, across time. And as the section
above shows, a sequential read figure moves 1.85x with elapsed time since the write. I thought I
was changing the policy; I was changing elapsed time as well.

The same shape is not confined to policies. I once wrote a 23% difference for multipathing on versus off,
and that too measured the two at different times. When comparing configuration A against
configuration B, match the elapsed time from write to read, and if you cannot match it, it is better
not to compare.

On choosing between protocols

There is no protocol difference between iSCSI and NVMe/TCP at one connection. That conclusion
is easy to miss if you compare figures measured on different file systems, so the comparison is kept
below to show why.

Comparison iSCSI NVMe/TCP
Different file systems side by side (misleading) 1,135.18 MB/s 591.64 MB/s
As the documented procedure has it 1,135.18 MB/s 1,288.87-3,407.60 MB/s (3 deployments)
4 KiB random reads under that procedure 44,593 IOPS 126,964 / 128,154 IOPS (2 deployments)

Read on its own, the top row looks like "the ordering inverts at one connection." But those two
figures were measured on different file systems, and side by side the difference reads as a protocol
difference that isn't there.

I re-measured on one file system. A 600 GiB LUN and a 600 GiB namespace were placed together on
a single 1,800 GiB volume, leaving lun= as the only line differing between the parameter files.

One connection, same file system 1 MiB sequential read
iSCSI, 1 session 1,135.19 MB/s
NVMe/TCP, 1 I/O queue 1,135.88 MB/s
Difference 0.06%

The ordering does not invert. There was no protocol difference at one connection.
591.64 belongs to the deployment tier: one-connection figures appear at only two positions,
4.73 Gbps and 9.08 Gbps, and within one deployment iSCSI and NVMe/TCP appear at the same
position. The connect-all side gave 2,366.32 against the published 2,366.49, agreeing to 0.007%
(measurement, Japanese).

The question left over is what 591.64 was. In the single mode as published, all four workloads
landed between 591.5 and 591.7, the shape of hitting the ceiling of a single path. The
re-measured run does not pin there. The cause has not been confirmed.

No ratio is written. The as the documented procedure has it row still has its numerator and
denominator from different file systems, and the NVMe/TCP side scatters across deployments.
Computed from one pair it comes to "2.08x", but recomputed over the measured spread it is
1.14-3.00x.
Pick one figure out of that spread and write "N times", and the picking is itself the
conclusion.

The iSCSI side cannot be treated as a fixed value either. For sequential reads, one session and
both portals agreed across two deployments (1,135.18 and 1,135.19, 0.01 MB/s apart), but two
points are not a distribution.
And 16 sessions is still one deployment only. There is no
basis for "iSCSI is stable". The evidence is simply not there.

One thing can be kept as a comparison. 4 KiB random reads agreed to within 1% on the same two
deployments where sequential reads scattered (126,964 and 128,154). It is the sequential read
that scatters, and nothing else.

This was later corroborated. On the same deployment where a rewrite moved sequential reads by
1.85x, 4 KiB random reads gave 131,506 (warm) against 131,550 / 131,749 / 131,736 (cold), a spread
of 0.2%. Whatever moves sequential reads by 1.85x does not act on 4 KiB random reads.

4 KiB random reads do, however, scatter across deployments for a different reason (three tiers
observed: 131k / 203k / 242k). That cause could not be confirmed. Two things were established.

One is that 4 KiB random reads scale precisely with the specified SSD IOPS.
Only the specified value was changed, within one deployment.

Provisioned SSD IOPS 4 KiB random read Ratio to the specified value
200,000 (before lowering) 131,413.5 IOPS 0.657
100,000 (after lowering) 66,016.9 IOPS 0.660
100,000 (created that way on another deployment) 65,070-65,947 IOPS 0.651-0.659
200,000 (after raising it on that deployment) 131,529.7 IOPS 0.658

The ratio going down is 1.990 and going up is 2.02. The measured value lands at about 66% of the
specified value, the same ratio at all four points.
The 1 MiB sequential read in the same
states was 2,265.7 / 2,366.23 MB/s and does not change: sequential reads do not depend on the
specified IOPS.

The reverse control was taken differently, in the following run. The first attempt lowered from
200,000, and the request to restore it was refused with "cannot start until at least 6 hours have
passed since the last storage capacity update".
A provisioned IOPS change is treated as a storage capacity update, with a six-hour cooldown.
Billing continues while you wait, so the next run was created at 100,000 and raised to 200,000
partway through
: in the increasing direction it takes one update and does not hit the cooldown.
An update stays in UPDATED_OPTIMIZING for about 18 minutes (about 17 minutes when lowering).
If you measure by lowering, plan to throw the environment away; if you need a round trip, design
it in the increasing direction.

The other is that what sets the spread between deployments is the disk read amplification factor.

The specified value was 200,000 on every deployment. What differed was how many disk-side reads one
client read becomes.
CloudWatch retains 5-minute values for 63 days, so the intervals whose
average read I/O is 4 KiB can be pulled out and checked without rebuilding the environment.

Deployment Client IOPS Disk IOPS Against the specified 200,000 Amplification (byte ratio)
High side 203,961 162,562 81% 96.5%
High side (another deployment) 206,582 172,158 86% 83.3%
Low side 131,617 195,905 98% 193.7%
Low side (another deployment) 130,080 195,941 98% 190.2%

The IOPS the client sees is the specified IOPS divided by that environment's amplification factor.
131,617 x 1.49 = 196,109; 203,961 x 0.797 = 162,557. Both check out.

  • At an amplification of about 1.5, the disk saturates at 98% of the specified value and the client stops at 131k
  • At an amplification of about 0.8, the disk has headroom at 81% and the client reaches 204k

The "about 66% of the specified value" from the proportionality above is the reciprocal of that
1.5.
The same relation was being seen two different ways.

Here I nearly settled on a wrong explanation once. I thought "if it exceeds the specified value
it must be served from memory", but the way to decide that is the ratio of DiskReadBytes to
DataReadBytes
(a few percent or less means it is served from cache). Applied for real it was
104-181% across all five deployments: none of them was being served from cache. The cause
was looking only at DiskIopsUtilization. It tells you whether the disk is saturated; it tells
you nothing about the relation to what the client sees.

If you measure, take four: DataReadBytes / DataReadOperations / DiskReadBytes /
DiskReadOperations. Taken in 5-minute intervals, the average I/O size lets you cut out workload
intervals after the fact. You do not have to rely on the timestamps in your own notes.

So what is the reason the amplification factor splits between 0.8 and 1.5? This is what remained to
the end.

Here another of my hypotheses was disproved. I thought "the layout above must also be setting the
amplification factor", but even right after a sequential refill had restored the layout to 88 KiB,
random read amplification stayed at 190-208%. The layout for sequential reads and the
amplification factor for random reads are separate variables.

There was one candidate. Line up the five deployments by "did 4 KiB random writes ever run on that
device"
and they split cleanly.

Deployment and interval 4 KiB random write history Amplification Client IOPS
Never ran (a run that measured reads only) none 96.5% 203,961
An interval before it had run none 83.3% 206,582
After it ran yes 170.1% 140,423
After it ran yes 193.7% 131,617
After it ran (a sequential refill does not restore it) yes 190-208% 97,769-130,080

All five followed that ordering. Even so, this is an ordering of observations, not a control
experiment.
I wrote it down as something that would be settled by applying the steps in turn on a
single device.

And that candidate was disproved too

I applied it. It was disproved. On a new deployment I began measuring right after a fill, with
4 KiB random writes never having run.

State 4 KiB random read Against the specified 100,000
Right after a fill on a new file system (no random write history) 65,947.6 IOPS 0.659
After 300 seconds of 4 KiB random writes 65,430.8 IOPS 0.654
After a sequential refill 65,070.0 IOPS 0.651

It is already 0.659 with no history. Where the "none" rows in the table above were 0.83-0.97,
the same "none" condition produced a value on the 1.5 side. The spread across the three states is
1.3%, and the amplification factor moves with neither layout nor write history.

Five observations lining up cleanly was a coincidence. That is the third disproved hypothesis.

An ordering of observations is not a substitute for a control experiment. All five followed it
without exception, the mechanism was plausible, and it was wrong. When you write "all n cases
follow" as a basis, check at the same time how many unvaried variables were lining up in the same
direction.
In these five, the patch version and the client were lined up the same way too.

What remains is those two. The difference between the 0.83-0.97 of F-6 / F-7 and the 1.5 side
later on comes down to the ONTAP patch version (9.18.1P6 against 9.18.1) and the client used
for the measurement. Neither was varied within a single run, so they are named as candidates and
nothing more.

One of those two cannot be varied by a reader either. CreateFileSystem's
FileSystemTypeVersion
is documented as being for FSx for Lustre, and FSx for ONTAP has no parameter for specifying a
version. The list of updatable properties in
Updating file systems
does not include the ONTAP version either. So the patch version can be neither chosen nor
pinned
, and a control experiment varying that variable cannot be assembled on the user side.
The client is the only one that can be varied. If you want to press on here, what you can do is
show that the amplification factor does not move when the client is changed, narrowing the
candidate to the patch version. It does not establish causation.

The policy, incidentally, had nothing to do with it. Switching queue-depth to numa to
round-robin to queue-depth during a 4 KiB random read produced no step, and against the 132,044.8
IOPS average of the run containing all three policies, queue-depth alone gave 131,413.5 IOPS —
a difference of 0.5%.

Five things I hit on the instrument and the teardown

It took nine failures to reach a measurement. All of them gates, none of them numbers.
Four of those happened before reaching a mount, so they are in a
separate article
(A: sessions and queues); what belongs here are the five on the instrument and the teardown.
The detail is in the
block measurement runbook
as a table of symptom, cause and how to check.

  1. vdbench resolves include= against the current directory, not against the path given to -f. The fill "completed successfully without writing anything". Had I gone on without noticing, I would have been measuring reads against unwritten blocks (which return zeros).

6. key=value on a vdbench command line is not an override of the workload definition.
It is an assignment to a placeholder inside the parameter file. Passing xfersize=1024k failed with
Unused parameter substitution.

  1. I reported an SSM wait limit as a failure. A 600 GiB write did not fit inside a 7.5-minute
    wait, I saw status: InProgress, judged it a failure and tore the environment down.
    A wait limit is not a failure.

  2. A sentinel string matched my own explanatory text. Against an implementation that "prints
    MISSING if the module is absent", I was deciding with grep -q 'MISSING', and it matched the
    explanatory line "A MISSING line means…" in the same output, so F-2 was skipped although NVMe/TCP
    was available. A sentinel that your own explanatory text satisfies is not a sentinel.

  3. The teardown trap died partway through the teardown. Under set -euo pipefail, reading
    describe-stacks after a delete always fails once the stack has finished disappearing, and that
    failure kills the trap itself. The delete ran but the retry did not, and a DELETE_FAILED stack was
    left behind. Put the delete call first in the trap, and set +e immediately after it.

Three more hit during the rewrite

The nine above are from the first measurement. The run that recounted things per path added three
more. All of them the same shape as the nine: looking successful while measuring nothing, or
not being able to tell that nothing was measured.

  1. A counter table returned empty. nvmf_lif had zero rows even after 600 GiB had been written,
    and reading zero rows as zero bytes gets you the right conclusion for the wrong reason. Which
    table to look at, and the detail of that trap, are in the separate article
    (B).

  2. A socket aggregation matched my own metric. To sum bytes per destination I treated any field
    ending in :4420 as a destination, and bytes_acked:4420 (a byte count) happened to end in
    4420, growing one spurious path named bytes_acked. The value was 0 so the sum was not broken, but
    it is the same mistake as the eighth.

12. "It passed locally so it will pass in CI" did not hold. This is the repository side rather
than measurement, but the structure is the same. A check scans tracked files, and a new file is
outside the scan while it is untracked. Green locally, red in CI. "Check after the last edit" is
not enough; it is also "after the last git add".

Summary

  • The same configuration varied by 2.64x across environments. Do not size from a single measured value
  • What is acting on it is the layout on disk. One disk-side read goes from 87-89 KiB to 11.7 KiB (a factor of 7.5), disk IOPS rises 7.3x and saturates at 89-95% of the specified value. DiskRead / DataRead is about 100%, so the faster side is reading from disk too. It was not being served from memory
  • What breaks it is small random writes only. Three reads in a row, a spread of 0.2%; 30 minutes idle, within 0.4%; 300 seconds of sequential writes, 0.03%. 4 KiB random writes cost 0.9% at 30 seconds and 77% at 300 seconds, and a sequential refill restores it. It goes back and forth within a single run. Where the threshold sits between 30 and 300 seconds is not known
  • The warm-side ceiling sits on the client side, and it is not concurrency. Two clients concurrently total 4,204.55 MB/s, 1.86x one client. Meanwhile 512 to 2,048 threads moves throughput by only 0.07% and only scales response time from 216.6 to 860.3 ms. NFS 4 KiB random reads are the same, a spread of 2.2% (5.391 to 21.108 ms). Before you write "I measured the storage ceiling", try adding a client. Adding threads will not tell you
  • 4 KiB random reads scale precisely with the specified SSD IOPS (about 66% of the specified value, the same ratio at four points). It holds in both the decreasing and the increasing direction. Sequential reads do not depend on the specified IOPS
  • What splits IOPS between 131k and 204k across deployments is the disk read amplification factor. Client IOPS = specified IOPS / amplification factor. DiskIopsUtilization alone will not tell you. Take DataReadBytes / DataReadOperations / DiskReadBytes / DiskReadOperations. What sets that amplification factor is not known. That it is neither layout nor write history is settled; the remaining candidates are the ONTAP patch version and the client. The patch version cannot be specified on FSx for ONTAP, so a control experiment varying that variable cannot be assembled on the user side
  • Observation order is not a substitute for a control experiment. I had a candidate that five cases followed without exception; I aimed a measurement at it and it was disproved. Unvaried variables were simply lined up in the same direction
  • When comparing configuration A against B, match the elapsed time from write to read. I compared without matching it, wrote "1.73x from the policy", and withdrew it later
  • Which protocol comes out ahead changes with the procedure and the I/O shape. Publish which procedure measured what, not which one is faster
  • Do not write throughput alone. The figures in this article are iorate=max saturation points, and the same 2,364 MB/s sits at a 216 ms point and at an 860 ms point. Without response time alongside it, that difference disappears

Notes on the AWS documentation

This article contains no remarks about the documentation, and there is nothing to submit.

What it contains are errors in my own way of measuring and measurements that the published
specification cannot account for. Neither is a claim that the documentation is wrong, so I am
setting them out separately.

Content Kind
2.64x across the same configuration / 1.85x from layout Measurement. The published specification does not promise this granularity
The cause of amplification 0.8 against 1.5 is unknown Left unresolved. The candidates are the ONTAP patch version and the client, and the former cannot be chosen on FSx for ONTAP, so no control experiment can be assembled
The per-client ceiling (block side) Not measured. On the NFS side it was followed through on 2026-09-20 (the table above), so the same way of varying things should apply to the block side
Comparing policies across time, writing 1.73x, and withdrawing it My error
nvmf_lif returning no rows ONTAP-side behaviour. A question for the vendor, not a matter of AWS documentation

The remarks about the documentation are in the other two: A and
B, each with its own submission status.

Submission status across the three block-protocol articles and S3 Burst Part 2 is collected in
AWS documentation notes and inquiries across the block-protocol articles.

Primary record of the figures

The primary record of the figures and their conditions is in
per-protocol measurement results (Japanese),
and the reproduction procedure is in the
block measurement runbook (Japanese).
Whatever is unconfirmed is written as unconfirmed, so please do not drop that when citing.

Pitfalls with the VDBENCH measurement tool, and items to always record, are collected in
Considerations when measuring FSx for ONTAP's block protocols.

Top comments (0)