DEV Community

Amazon FSx for NetApp ONTAP NVMe/TCP assumes RHEL 9.3: why multipathing fails on Amazon Linux 2023

Amazon FSx for NetApp ONTAP's NVMe/TCP procedure assumes RHEL 9.3: why multipathing is unavailable on Amazon Linux 2023, and how to check before you are billed

Followed the AWS procedure and run cat /sys/module/nvme_core/parameters/multipath only to find
the file isn't there?

Moving your target Linux distribution to Amazon Linux 2023 can stop a documented command from
working exactly as written. You are reading the procedure correctly, yet your environment alone
gives a different result — and tracking the cause down to a kernel build setting takes time.
This article is for anyone who got stuck at that point.

What this article does and does not cover: it covers how to confirm multipathing is unusable on
AL2023, and how to count whether it is actually using the second path. It does not cover
performance comparisons or a generalized conclusion for distributions other than AL2023 (details
under "What this article does not cover" below). All figures are from a Single-AZ configuration;
Multi-AZ is unverified.

dev.to series: FSx for ONTAP Block Protocols, Measured.

What was checked

Assuming no prior knowledge of FSx for ONTAP, here's what this article is confirming.

  • Location: a single AWS VPC, single AZ
  • Client: one EC2 instance (Amazon Linux 2023)
  • Target: one Amazon FSx for NetApp ONTAP file system (second-generation SINGLE_AZ_2). The HA pair's two controllers hold two paths to one namespace
    • Path 1: to the optimized controller (shortest path)
    • Path 2: to the non-optimized controller. Only when multipathing is available can the host tell that path 2 is a standby

One EC2 client connects to an FSx for ONTAP HA pair's two controllers, over Path 1 (optimized,<br>
the shortest route) and Path 2 (non-optimized, which looks like a normal route without ANA), both<br>
reaching the same one namespace

Figure: an FSx for ONTAP HA pair has two paths to one namespace, and **Asymmetric Namespace
Access is the mechanism that tells the host which one is the shortest path
* (referred to as
multipathing throughout this article). Without this mechanism in the kernel, those two paths
don't appear as "one shortest path plus one standby": they appear as two separate devices
pointing at the same data.*

Item Role in this article
VPC / AZ Single AZ, so both paths to the HA pair's controllers sit on the same network
EC2 client Checks for multipathing and the kernel configuration behind it
optimized path The shortest path as determined by multipathing. Real data is expected to flow here
non-optimized path The failover path. Without multipathing it appears as an ordinary path too

AWS's
Provisioning NVMe/TCP for Linux
includes a step that confirms multipath has come up: the fourth item under "To discover the target
NVMe nodes". It says that a returned Y means success (wording as of 2026-09-17; every
reference to that page below was checked on the same day).

cat /sys/module/nvme_core/parameters/multipath
# Y means success
Enter fullscreen mode Exit fullscreen mode

On the Amazon Linux 2023 AMIs I used, that file did not exist. The kernel does not have
Asymmetric Namespace Access (a mechanism that tells the host which of several paths to the same
namespace is the shortest one) in it. The range I checked is the three kernel series returned by
al2023-ami-kernel-default-x86_64, so I cannot write "AL2023 always behaves this way". Checking
takes one command, and a later section gives it.

The procedure is written for RHEL 9.3. The same page's "Before you begin" says
Create an EC2 instance running Red Hat Enterprise Linux (RHEL) 9.3. Immediately after that there
is a statement about other distributions. In summary: on AMIs other than RHEL 9.3 some of the
utilities may already be present or the install command may differ, but apart from installing
packages, the commands in this section are valid on other EC2 Linux AMIs too
(the original wording
is at the end of "Before you begin" on that page).

What this article reports is two counterexamples to that statement. Neither of them is
explained by package installation.

Command from the documentation Result on AL2023 A packaging problem?
cat /etc/nvme/hostnqn File does not exist No. nvme-cli is installed; whether it creates this file differs by distribution
cat /sys/module/nvme_core/parameters/multipath File does not exist No. This is a kernel build option (CONFIG_NVME_MULTIPATH) and a package cannot change it

You can check this before any billing starts. That is the most practical part of this article,
I think.

This is not a claim that "the documentation is wrong". It is a report that these two do not
fall inside
the stated range (apart from installing packages, valid on other AMIs too).
The state of my feedback to AWS is at the end of this article.

Then, after bringing in a kernel that does have multipathing, I found that the optimized /
non-optimized output of nvme list-subsys says "there are two paths", not "two paths are in
use".
Whether they are in use has to be counted separately, and there is a trap in choosing
which table to count.

What this article does not cover. It does not compare performance. The figures here exist to
show that two clients agreed under the same conditions, and they cannot be used for sizing. What
moves the numbers is in a separate article. All figures are from a
Single-AZ configuration; Multi-AZ is unverified.

This is one of three, but there is no reading order. The block-protocol measurements are split
by reader.

Article Reader
A Session and queue counts Anyone about to mount
B Checking whether multipathing is available (this article) Anyone who followed the procedure and got stuck
C What moves the numbers Anyone measuring or citing figures

None of them assumes the other two. Read only the part you need.

The environment I measured

Item Value
File system Amazon FSx for NetApp ONTAP, second generation SINGLE_AZ_2, 6,144 MBps × 1 HA pair, 4,096 GiB SSD, 200,000 provisioned IOPS
Region ap-northeast-1, single AZ
ONTAP 9.18.1 (the reads and writes in the agreement table). For the 12 sequential-write points I failed to record the version. FileSystemTypeVersion returned null and I tore the environment down before re-reading it from the cluster API. The kernel investigation also spans deployments on other versions. The cross-article version mapping is in the ONTAP version matrix
NVMe read cache Disabled (confirmed on both nodes)
LUN / namespace 600 GiB, on a 900 GiB volume
Before measuring Wrote the full 600 GiB once. Unwritten blocks on a thin-provisioned volume return zeros, so reading without writing first measures nothing
Client c5n.9xlarge; 50 Gbps network is a guaranteed figure
Instrument VDBENCH 5.04.07 directly, iorate=max (unlimited), 512 threads, 60 s warm-up + 300 s measurement
Test data / efficiency settings Not recorded whether the VDBENCH payload was compressible, or whether volume inline efficiency (StorageEfficiencyEnabled) was enabled. On the file-protocol side, efficiency settings measurably affected figures; whether the same applies to block is unconfirmed

The window is 300 seconds, so the figures in this article include burst. They cannot be cited as a
baseline.

The figures in this article are iorate=max saturation points, not operating points. No target
IOPS was given; they are what the system took when everything available was pushed at it. On the
file-protocol side of the same repository, giving a target of 4,400 IOPS produced 4,406 MB/s at
8.86 ms, while unlimited produced 4,204 MB/s at 121.79 ms. Unlimited was 5% lower in
throughput and 14 times higher in response time.

With 512 threads fixed and no rate limit, response time becomes the identity 512 ÷ IOPS
(derived 216.6 ms against measured 216.588 ms in a separate run). The figures below can be read
through that identity: 1,277.48 MB/s is about 401 ms, and 131,736.5 IOPS is about
3.9 ms. Both are derived values and were not recorded as measurements.

For this article's purpose, being a saturation point is not a disadvantage. The figures are
here to show that Rocky and RHEL agreed under the same conditions, and as long as both are
saturated the same way, the agreement test holds. Those values still cannot be used for
sizing.

Behaviour on this AMI

On this AMI the file named at the top does not exist. Reading the kernel configuration gives
# CONFIG_NVME_MULTIPATH is not set, and modinfo nvme_core has no multipath parameter either.
So multipathing is not in the kernel.

As a result the two controllers are not merged into one device; they appear as separate devices
pointing at the same namespace UUID
(/dev/nvme2n1 and /dev/nvme3n1). nvme list-subsys shows
no optimized / non-optimized either.

The AWS procedure is written for RHEL 9.3, where it is present.

Please do not generalise this to "multipathing is not available on AL2023". What I observed were the
AMIs that the al2023-ami-kernel-default-x86_64 SSM parameter returned on 2026-09-12 and
09-13, and I did not record the kernel versions. The template said "record the kernel version
at measurement time" in place of pinning the AMI, and the block phase did not implement that.
That is a gap in my records.

Checking in your own environment is more reliable, and it is one command.

grep -i NVME_MULTIPATH /boot/config-$(uname -r)

I later widened the observation to three kernel series. All three returned by
al2023-ami-kernel-default-x86_64 (6.1.186 / 6.12.103 / 6.18.48) gave
# CONFIG_NVME_MULTIPATH is not set. You can check this before creating a file system, by
extracting the config from the kernel package in the repository. It is known before billing
starts, so checking in that order is cheaper.

To measure multipathing I then brought in another distribution. The AWS procedure is written for RHEL 9.3,
so I put both Rocky Linux 9.7, a rebuild, and RHEL 9.7 itself on the same namespace and compared
them.

Workload RHEL 9.7 Rocky 9.7
1 MiB sequential read 1,277.48 MB/s 1,276.60–1,279.22 MB/s
4 KiB random read 131,736.5 IOPS 131,550.6 / 131,749.4 IOPS
4 KiB random write 195,439.8 IOPS 194,960.6 IOPS

Only the reads agreed in the same state (within 0.1%). Same minor version, same nvme-cli
(2.16-1.el9), same cold state. The kernels differ in z-stream (5.14.0-611.55.1 against 611.5.1).
For reads, there is now one piece of grounds for reading a rebuild's figure as the figure for the
product itself.

Both write workloads had the Rocky side measured warm, so the states are not aligned. Counting
the 0.3% difference on 4 KiB random write as "agreement" was my error, and the note in that same
table contradicted my own conclusion
(corrected 2026-09-20). For that shape I can say neither
that it holds nor that it does not.

Sequential write was re-measured later, in the same state, for that shape alone. It does not
agree (2026-09-17).
On a different file system (second generation, 1,536 MBps, 50,000 provisioned
IOPS), same namespace, same state, same parameters, six runs each.

Client Min Median Max Spread
Rocky 9.7 842.70 1,040.26 1,092.50 29.6%
RHEL 9.7 1,105.75 1,112.65 1,129.30 2.1%

Units MB/s (1 MiB sequential write, 8 threads, o_direct, 120-second window). These 12
points were taken under conditions that differ from the environment described at the top: a
different file system (second generation, 1,536 MBps, 50,000 provisioned IOPS), 8 threads rather
than 512, and 120 seconds rather than 300. The 512 ÷ IOPS reading given at the top cannot be
applied to these 12 points. RHEL's minimum is above Rocky's maximum: six against six, with no
overlap in range. RHEL's median is 7.0% above Rocky's, and the asymmetry in spread is larger than
the difference in medians.

So "a rebuild's figure can be read as the figure for the product itself" is limited to the
reads. When citing a sequential-write figure, say which of the two it was measured on.
4 KiB random write was not re-measured in the same state, so for that shape I can say neither
that it holds nor that it does not. I cannot state a cause. The only difference on record is
the kernel z-stream, and I did not vary it on its own. During the runs the file system side was at
72–81% disk throughput, 16–19% IOPS and 31–38% network, none of them saturated (so this is not a
case of the two meeting at a ceiling). For these 12 points I failed to record the ONTAP version.
FileSystemTypeVersion returned null and I tore the environment down before re-reading it from
the cluster API.

iopolicy on a kernel with multipathing, and the value the procedure states

This is the third counterexample. iopolicy (Linux's multipath setting that decides how I/O
is distributed across multiple paths to an NVMe device: round-robin cycles through them,
queue-depth sends to whichever has the shortest queue) is what step 5 of the same procedure
says to confirm as round-robin (and the example nvme list-subsys output at step 7 also shows
iopolicy=round-robin).

Value
Value the documentation says to confirm round-robin
Measured on Rocky Linux 9.7 / RHEL 9.7 queue-depth
What set it the udev rule 71-nvmf-netapp.rules shipped by nvme-cli (2.16-1.el9)

The kernel default is numa, neither round-robin nor queue-depth. It was queue-depth in
this environment because the NetApp-oriented udev rule bundled with the nvme-cli package set it.
Following the confirmation step returns a value other than the expected one.

I measured the effect on performance. It is negligible: switching policy mid-run produced no
step change, and the difference was 0.5%. The figures and the method are in
a separate article (C). So this is not a story about round-robin being
required for speed. The only part worth reporting is that the confirmation step returns a value
that differs from the documentation.

This can change with the nvme-cli version. What I observed here is one version, 2.16-1.el9,
and I did not confirm whether other versions ship the same udev rule.

How to check whether multipathing is using the second path

This section was added afterwards. It became necessary in order to do the analysis above, and it
is also the reason I once published a wrong conclusion.

nvme list-subsys shows the two controllers as optimized / non-optimized, and prints iopolicy
as well. That is a display of "there are two paths", not of "two paths are in use".

At first I varied iopolicy three ways, compared total throughput, saw that only queue-depth was
faster, and wrote that "only queue-depth uses the second path". That was wrong. I had not
counted bytes per path. A client-side total that does not increase is consistent with two different
facts: "the second path is not being used", and "it is being used but the ceiling is nearer than
the path".

Bytes per path can be read from two places

On the ONTAP side they can be read from a counter table. There is a trap here.

Table Result in this configuration
nvmf_lif Zero rows. Still empty after writing 600 GiB over NVMe/TCP
lif Reports 0 bytes for the LIF in question. It does not count NVMe-oF
nvmf_tcp_port Has the real data. read_data / write_data / total_ops per LIF

"The table exists but has no rows" and "this version does not have that table" add up to the same
appearance. Return 0 and move on, and you reach the correct conclusion, "the second path saw
0 bytes", for the wrong reason.

It can also be counted on the client side. NVMe/TCP controllers have different destination
addresses, so aggregating established sockets by destination separates the paths. That is a second
source, independent of the ONTAP side.

ss -tin state established | awk '/:4420/{...}'   # sum bytes_sent / bytes_received per destination
Enter fullscreen mode Exit fullscreen mode

Counted properly, the second path was not running

I ran four workloads under queue-depth and took before/after differences.

Path Read bytes Write bytes Total ops
optimized +about 1,012 GiB +about 730 GiB 641,908 → 120,008,866
non-optimized ±0 ±0 4 → 4

Neither reads nor writes went to the second path. The ONTAP side and the client side agreed, and
it was the same on Rocky and on RHEL (the second path's increment across all workloads was about
20 KB).

Further, switching iopolicy within a single run produces no step change. With a sequential read
running I changed queue-depth → numa → queue-depth, and the values sampled at one-second
intervals stayed inside the same 1,150–1,400 MB/s range both before and after each switch.

Summary

  • All three kernel series returned by al2023-ami-kernel-default-x86_64 (6.1.186 / 6.12.103 / 6.18.48) had CONFIG_NVME_MULTIPATH unset. Measuring multipathing requires a different distribution. Asking AWS why AL2023 is excluded from the procedure (reply 2026-09-21) surfaced a likely reason: nvme connect-all establishes both paths and auto-discovers the alternate LIF even when only one is given, but without CONFIG_NVME_MULTIPATH the kernel never merges them into one namespace device. The procedure's later steps assume that merge already happened, so proceeding unmerged leaves the other path writable against the same namespace with no multipath software mediating it: a data-corruption risk
  • You can check before creating a file system. Extract the config from the kernel package in the repository and you know before billing starts. On a running instance it is the one command grep -i NVME_MULTIPATH /boot/config-$(uname -r)
  • Rocky Linux 9.7's figures can be read as RHEL 9.7's figures, but only for the reads. On the same namespace in the same state they agreed within 0.1%. Both write workloads had the Rocky side measured warm, so the states are not aligned. Counting the 0.3% difference on 4 KiB random write as agreement was my error, and for that shape I can say neither that it holds nor that it does not (corrected 2026-09-20). Sequential write does not agree: re-measured later in the same state, six runs each, RHEL's median was 7.0% higher and RHEL's minimum was above Rocky's maximum (spread 29.6% for Rocky against 2.1% for RHEL). When citing, say which of the two it was measured on
  • optimized / non-optimized in nvme list-subsys is a display of path existence. For whether they are in use, count with the nvmf_tcp_port counters. nvmf_lif returns no rows, and lif does not count NVMe-oF
  • In this configuration, not one byte went to the second path even under queue-depth. Two places agreed (the ONTAP-side counters and the client-side sockets), and it was the same on Rocky and on RHEL

AWS documentation notes and support inquiries

This article touches three places in the documentation. All three were submitted on 2026-09-17.

The source for all of them is
Provisioning NVMe/TCP for Linux.

Remark State
Against "apart from installing packages, valid on other EC2 Linux AMIs too": cat /etc/nvme/hostnqn finds no file on AL2023 Submitted (2026-09-17). Reproduced in the same environment, confirming the wording doesn't hold as-is for the AL2023 environment checked. The two proposed fixes (scoping the wording, adding a generation step) have been shared as an improvement request. Adoption, content, and timing are unconfirmed
Against the same statement: cat /sys/module/nvme_core/parameters/multipath finds no file on AL2023 (kernel build option) Submitted (2026-09-17). The observation was confirmed and filed as documentation feedback. Fix timing undecided. A follow-up question ("why is AL2023 excluded from scope?") also got a technical explanation: proceeding with the paths unmerged leaves the other path writable against the same namespace with no multipath software mediating it, a data-corruption risk, so the exclusion has a technical rationale
Step 5 directs confirming that iopolicy is round-robin, but the measurement on Rocky 9.7 / RHEL 9.7 is queue-depth. nvme-cli's version boundary is 2.11 -> 2.12 (see below) Submitted (2026-09-17). On 2026-09-30 it was acknowledged that a note is needed stating whether iopolicy resolves to round-robin or queue-depth depends on the environment, and this was fed to the responsible team (with a note that NetApp's own published documentation carries the same point). The content and timing of any change are not disclosed in advance, so they are unconfirmed

The form I submitted was this. Verbatim quotation (which line of which step), reproduction steps,
expected and measured values, and the range I confirmed across three kernel series, written as
such. Not generalising to "multipathing is not available on AL2023" is the same in the feedback: the
observation is limited to the three series returned by al2023-ami-kernel-default-x86_64.
For the second one I also included the control under which the procedure does hold (on a kernel
with CONFIG_NVME_MULTIPATH=y, Y is returned). It is not a report that the procedure itself is
wrong.

The third is not submitted as a performance problem. The difference between policies is 0.5% in
this configuration, and the report is the single point that the confirmation step returns a value
differing from the documentation.

The third changed shape before submission, after I looked at the package side. Reading
71-nvmf-netapp.rules out of four builds of the public nvme-cli package (the SRPMs for 2.11 /
2.13 / 2.16, plus the installed 2.16-1.el9) showed that 2.11 has round-robin and 2.13 changed it
to queue-depth (confirmed 2026-09-17). At that point 2.12 had not been examined, so the boundary
could only be written as "somewhere between 2.11 and 2.13".

Submission status across the three block-protocol articles and S3 Burst Part 2 is collected in
AWS documentation notes and inquiries across the block-protocol articles.

A pointer to the upstream v2.11 and v2.12 sources followed, showing the boundary is v2.12.
Fetching and checking that commit (2026-09-19, 71-nvmf-netapp.rules.in) confirmed that
queue-depth was already in place as of v2.12. So the version boundary is settled at
2.11 -> 2.12
(the original "somewhere between 2.11 and 2.13" narrowed to 2.12, from both that
pointer and my own re-check). This is agreed to be a version drift rather than an error in the
documentation.

On 2026-09-30, it was acknowledged for this point that a note is needed stating that whether
iopolicy resolves to round-robin or queue-depth depends on the environment, and it was fed to
the responsible team.
NetApp's NVMe-oF procedure for RHEL 9.x
carries the same point (from RHEL 9.7 the default is queue-depth; on 9.6 round-robin is the
default with queue-depth selectable), consistent with this understanding. The content and timing
of any change are not disclosed in advance, so its reflection in the documentation is unconfirmed.

Primary record of the figures

The primary record of the figures and their conditions is in
the protocol measurement results,
and the reproduction steps are in
the block measurement runbook.
What is unconfirmed is written as unconfirmed there, so please do not drop that when citing.

How to check whether multipathing is available, counter-table traps, and iopolicy version differences are
collected in
Considerations when measuring FSx for ONTAP's block protocols.

Top comments (0)