Moving on-premises block storage to AWS? The "divide by 625 MBps" sizing approach comes out wrong — measured session and queue counts
What this article does and does not cover: it covers the measured failure of the "divide by
625 MBps" sizing approach, and four configuration failures hit before mounting. It does not
cover performance conclusions, a comparison against file protocols, or distributing block storage
to multiple sites (details under "What this article does not cover" below).
All figures are from a Single-AZ configuration; Multi-AZ is unverified.
When using FSx for ONTAP over iSCSI or NVMe/TCP, how many sessions should you open? AWS's
procedures surface one number: "Amazon EC2 single client maximum of 5 Gbps (~625 MBps)",
which appears on both the
iSCSI and
NVMe/TCP pages as
the lead-in to the procedure for adding sessions.
One thing first: AWS does not instruct you to divide. The source text says "if you want
throughput beyond this value, follow the EC2 network bandwidth page's procedure to add
sessions," and it's marked(Optional). No formula is given. "required bandwidth ÷ 625 =
session count" is a sizing approach I built after reading that figure. This article tests
my sizing approach, not a formula AWS published.
Measuring it, that sizing approach missed in both directions. 1.82x with one session, 0.20x
with sixteen. And going from one session to two moved sequential throughput by zero bytes.
So this figure can't be used as a divisor to determine session count. What's documented as a
single flow's ceiling in EC2 network bandwidth
turned out to be a separate thing from what a block protocol session measures per session.
NVMe/TCP has nothing equivalent to nr_sessions; the analogous quantity is queue count. That
doesn't honor a request either — asking for 36 produced 4.
What was measured
Assuming no prior knowledge of FSx for ONTAP, here's where this article's measurement happens.
Figure: what this article measures is the connection count between one EC2 client and one FSx for
ONTAP file system, not multiple clients pooled together.
The figure states the same thing as the table below. What this article measures is the
connection count between one EC2 client and one FSx for ONTAP file system, not multiple
clients pooled together.
| Item | Role in this article |
|---|---|
| VPC / AZ | Single AZ, no network detour, so what's measured is the protocol itself |
| EC2 client | Varies session/queue count and measures 1 MiB sequential reads |
| FSx for ONTAP | Provides the iSCSI LUN or the NVMe/TCP namespace |
| Connection count | For iSCSI, TCP sessions counted with ss; for NVMe/TCP, requested vs. effective queue count |
What this article does not cover. No performance conclusion is drawn; the numbers here only
serve as evidence that the division misses. The same configuration swings 2.64x across
environments, so don't use these figures for sizing. Why, and what actually drives the numbers,
is in a separate article.
There's no comparison against file protocols (NFS/SMB) either — there's a reason they don't belong
in the same table. And there's no discussion of distributing block storage to multiple sites — this
architecture's distribution layer (FlexCache) distributes volumes, not LUNs or namespaces.
This is one of a set of three, but there's no order to read them in. The block protocol
measurements are split by reader.
Article Reader A Session and queue counts (this article) Anyone about to mount B Checking whether ANA is available Anyone who followed the procedure and got stuck C What moves the numbers Anyone measuring or citing figures None assumes the other two. Read only what you need.
Measurement environment
| Item | Value |
|---|---|
| File system | FSx for ONTAP second-generation SINGLE_AZ_2, 6,144 MBps x1 HA pair, SSD 4,096 GiB, provisioned 200,000 IOPS |
| Region | ap-northeast-1, single AZ |
| ONTAP | 9.18.1P5 (the three divisor-table data points). The queue negotiation section used a separate deployment, and its version was not recorded — it wasn't read at measurement time and can't be recovered after the fact. I can't write "all figures in this article are the same version." (Multiple versions coexisting across this configuration is covered in the separate article (C). The cross-article version mapping is in the ONTAP version matrix.) |
| NVMe read cache | Disabled (confirmed on both nodes) |
| LUN / namespace | 600 GiB, on a 900 GiB volume |
| Before measuring | Wrote the full 600 GiB once. Thin-provisioned unwritten blocks return zero, so reading without writing first measures nothing |
| Client | c5n.9xlarge; 50 Gbps network is the guaranteed value |
| Measurement tool | VDBENCH 5.04.07 direct, iorate=max (unlimited), 512 threads, 60s warmup + 300s measurement |
| Test data / efficiency settings |
Not recorded whether the VDBENCH payload was compressible, or whether volume inline efficiency (StorageEfficiencyEnabled) was enabled. On the file-protocol side, efficiency settings measurably affected figures; whether the same applies to block is unconfirmed |
Because the window is 300 seconds, this article's figures include bursts. They can't be cited as
a baseline.
This article's figures are the saturation point of
iorate=max, not an operating point in
production. No target IOPS was given — this is however much could be pushed through. On the
file-protocol side of the same repository, giving a target of 4,400 IOPS produced 4,406 MB/s at
8.86 ms, while unlimited produced 4,204 MB/s at 121.79 ms. Unlimited is 5% lower
throughput and 14x the response time.With threads fixed at 512 and unlimited, response time is an identity of
512 / IOPS(confirmed
against a separate measurement: 216.6 ms derived versus 216.588 ms measured). The values below
can be re-read through that identity — 1,135.18 MB/s is about 451 ms, 1,970.43 MB/s is
about 260 ms. Both are derived values, not recorded measurements.
The divisor division doesn't hold
Over iSCSI, I varied the counted number of TCP connections and measured 1 MiB sequential
reads. These are actual counts from ss, not the requested value.
| TCP counted | Measured | Predicted from 625 MBps | Measured / predicted |
|---|---|---|---|
| 1 | 1,135.18 MB/s | 625 MB/s | 1.82x |
| 2 | 1,135.18 MB/s | 1,250 MB/s | 0.91x |
| 16 | 1,970.43 MB/s | 10,000 MB/s | 0.20x |
With one session it's faster than predicted; with sixteen it's a fifth of the prediction.
And going from one session to two moves sequential throughput by zero bytes — 1,135.18 appears
twice in a row. So the ceiling hit with one session is not a single flow's capacity.
1,135 MB/s is 9.08 Gbps. In the same environment, the file protocols landed near the documented
5 Gbps on a single connection — NFS 591.62 MB/s (4.73 Gbps), SMB 574.24 MB/s (4.59 Gbps).
A single iSCSI session carries roughly twice that. Why is unconfirmed. (This comparison is
against single-connection figures from other protocols in the same environment, not a direct A/B
test against iSCSI.)
Here's the split between what's documented and what's an unverified candidate. The single-flow
5 Gbps ceiling itself is documented. AWS defines a single flow as "a TCP or UDP flow identified
by a unique 5-tuple (source IP, destination IP, source port, destination port, protocol)" and
states that outside a cluster placement group, a single flow is limited to 5 Gbps
(Amazon EC2 instance network bandwidth).
The same page documents the workarounds too — up to 10 Gbps inside a cluster placement group,
up to 25 Gbps within the same AZ with ENA Express, or spreading traffic over multiple paths with
Multipath TCP (MPTCP)
(ENA Express).From here on, these are unverified candidates. Whether this measurement environment was
inside a cluster placement group was not recorded and is unconfirmed. 9.08 Gbps would fit inside
a 10 Gbps allowance, but whether this configuration was inside that allowance is unknown.
Whether iSCSI, NVMe/TCP, NFS, and SMB all count as "one" the same way against the 5-tuple
definition is also unconfirmed. ENA elastic network interfaces are documented as having up to 8
RSS queues
(Optimize network performance on EC2 Windows instances —
that page covers Windows; I haven't checked the equivalent for Linux), and if protocols
distribute across CPU cores differently, the apparent ceiling could differ as a result. This
article does not test these candidates. A separate measurement would be needed.
The requested queue count doesn't hold
NVMe/TCP has nothing equivalent to nr_sessions. It opens a TCP connection per queue, so the
quantity that corresponds to session count is queue count.
Requesting via --nr-io-queues produced this. NCQA (Number of Completion Queues Allocated) and
NSQA (Number of Submission Queues Allocated) are the actual completion and submission queue
counts the NVMe Get Feature command reports back to the host.
| Mode | Requested queue count | Effective NCQA / NSQA | Controller count |
|---|---|---|---|
| Single flow | 1 | 1 | 1 |
AWS's procedure as written (one connect-all) |
unspecified | 4 | 2 |
| Explicit queue count | 36 (vCPU count) | 4 (unchanged) | 2 |
Requesting 36 produced 4 regardless.
This isn't speculation — it's documented on the ONTAP side. vserver nvme subsystem show's
description of -default-io-queue-count states "The actual value used when a connection is
established may vary depending on the host and transport protocol used," and
vserver nvme show-host-priority
states that I/O queue count is determined per combination of node, transport, and host
priority. NetApp's KB has an example where setting 15 on a subsystem produced 2 on the host
side.
Don't use the example's value as your own environment's value. In the example above,
nvme-tcp/regularis 2, but in this environment it was 4.
As a byproduct, one thing left as "unconfirmed" in planning got resolved. queue_count includes
the admin queue and matches NCQA + 1 (1 -> 2, 4 -> 5). With one queue, there was also exactly
one established TCP connection carrying I/O.
Four failures hit before mounting
Nine failures were hit before reaching a measurement. All were gates, not numbers. Each one
looks like it succeeded while measuring nothing. Four of those happened before mounting, so
they're here; the rest concern the measurement tool and teardown and are in the
separate article (C:
what moves the numbers). Details, as a table of symptom / cause / how to check, are in the
block measurement runbook.
-
JunctionPathturned out to be required. I had assumed it wasn't needed since a LUN isn't read through a junction, but the API rejects that. Worse, the property reference's entry forJunctionPathsaysThis parameter is required.in the body while the same property'sRequired:column saysNo. That's a literal self-contradiction (still present as of 2026-09-17). Trust the API. It fails 25 minutes after file system creation starts.
The same page has an example of writing this without contradiction. StorageEfficiencyEnabled
uses Required: Conditional and explains the condition in the body (required when the volume is
RW). Conditional is a value this reference already uses, so JunctionPath could take the
same form.
I can't state exactly which condition makes it required. What I measured is one case,
OntapVolumeType: RW, where it was required. I did not test whether it's optional for
DP. Determining the condition belongs to whoever holds the API's implementation;
what I can report stops at "the body and theRequired:column disagree."
2. AL2023 does not create /etc/nvme/hostnqn. AWS's procedure assumes this file exists and
reads it with cat (RHEL creates it). Generating it with nvme gen-hostnqn was necessary.
3. iSCSI login is asynchronous. iscsiadm --login returns before iscsid actually connects. In
this environment it took about 100 seconds. Counting with ss right after returned 0, so
every alias and fill attempt right after also came up empty. "The command returned" is not
"the connection is up."
4. find_multipaths yes does not create a map for a single-path device. When intentionally
measuring a single flow, the path count is 1 by design, and with the default setting
/dev/mapper/<alias> never appears. Explicit registration via multipath -a <wwid> was
necessary. That wwid is also more reliably read from the device with scsi_id than assembled as
3600a0980+serial (assembling it is useful as a cross-check, not as the primary source).
Summary
- Sizing session count by dividing required bandwidth by 625 MBps missed in both directions. 1.82x with one session, 0.20x with sixteen.
- Going from one session to two doesn't move sequential throughput. The ceiling hit with one session is not a single flow's capacity
-
NVMe/TCP queue count is the result of negotiation. Requesting 36 produced 4. Read the
effective value with
nvme get-feature --feature-id 7.queue_countincludes the admin queue, so it equalsNCQA + 1 -
All four configuration failures looked like success while measuring nothing.
JunctionPathis required, AL2023 doesn't createhostnqn, iSCSI login is asynchronous at about 100 seconds, andfind_multipaths yesdoesn't map a single-path device. - Don't use this article's figures for sizing. The same configuration swings 2.64x across environments
- What drives that 2.64x swing is out of scope for this article. It's covered in a separate article
AWS documentation notes and support inquiries
This article touches two points in the documentation. One has been filed; the other is not
being filed.
| Finding | Source | Status |
|---|---|---|
JunctionPath's body and Required: column disagree. Traced to the API reference, which propagates into CloudFormation and the CDK |
same |
Filed (2026-09-17). AWS's reply (2026-09-22): explained that when OntapVolumeType is RW, JunctionPath is required; when DP, it cannot be specified at all — so "Required: No" has a rationale when read across both volume types combined. The proposed change to Required: Conditional was described as feedback to be shared with the internal team, with no commitment on adoption. The question of whether the API reference is upstream of CFN/CDK went unanswered ("internal AWS information, cannot disclose"); instead AWS said it would file feedback against the API reference, CFN, and CDK pages individually |
| Sizing by dividing 5 Gbps (~625 MBps) doesn't match measurement | iSCSI / NVMe/TCP procedures | Not filed. AWS never states a formula, so this isn't a documentation error |
Why the second one isn't filed. This is a record of my own sizing approach missing, not a
finding about the documentation. The figure itself is correctly documented as a single flow's
ceiling; I was the one who used it as a divisor for session count, and that's where it broke
down.
The first was filed with AWS Support on 2026-09-17. The submission included the observed
requirement (omitting it fails stack creation 25 minutes in) and the fact that the CDK inherits
it as an optional type, requesting a fix on the API reference side.
AWS's 2026-09-22 reply settles only the rationale behind "Required: No" (it holds when RW and
DP are read as one combined attribute). It doesn't change the underlying fact that JunctionPath is
required when creating an RW volume, and whether Required: Conditional will actually be adopted
remains unconfirmed.
Primary record of the figures
The primary record of the figures and conditions is in the
per-protocol measurement results
(Japanese); the reproduction steps are in the
block measurement runbook
(Japanese). Anything marked unconfirmed there is marked unconfirmed for a reason — don't drop
that qualifier when citing it.
Considerations before sizing session and queue counts, and pitfalls hit before mounting, are
collected in
Considerations when measuring FSx for ONTAP's block protocols.
Submission status across the three block-protocol articles and S3 Burst Part 2 is collected in
AWS documentation notes and inquiries across the block-protocol articles.

Top comments (0)