After 16 years operating platforms on AWS, I've learned that GPU capacity and training data rarely live in the same Region, and the question that matters isn't 'can you train far from the data?', it's 'what does each batch cost while the cache hasn't converged?'. The validation AWS published with SageMaker HyperPod and Cloud Native Qumulo answers with a number: a cluster in us-west-2 reading from us-east-2 at 60 ms RTT matched the co-located cluster: 115β116 versus 116β117 samples/s, after a warmup of 100 to 150 batches. This article takes apart the mechanism, the failure modes, and the conditions under which that result holds.
The real problem: GPU where it shows up, data where it stayed
Anyone who has hunted ml.p5.48xlarge capacity knows it doesn't wait for a committee decision. HyperPod is available in 18 Regions, but the H100 queue isn't uniform, capacity shows up where it shows up, and the pretraining dataset, tens or hundreds of terabytes of it, stayed in the Region where the data pipeline built it.
Until now, the two known exits were bad in different ways. Replicating the dataset: copying petabytes across Regions pays per-GB transfer, doubles storage cost and, the cost nobody budgets, creates two truths to keep in sync for the life of the project. Reading remotely with no cache: every NFS read crosses 60 ms of RTT, the data loader can't feed the GPU at pace, and you pay H100 hours to watch it wait on I/O.
The approach validated in the post is the third way: the dataset stays in us-east-2 on a Cloud Native Qumulo hub, and a spoke in us-west-2 exposes the same filesystem to the training cluster over NFS, with Cloud Data Fabric replicating metadata within seconds and NeuralCache, the predictive prefetch engine, pulling blocks before the job asks for them. No upfront copy, no code change: the PyTorchJob mounts a 50 Ti ReadWriteMany PV as if the data were local.
The batch's path: local hit, miss crosses the Region, prefetch runs ahead
Read flow of the cross-Region training validated in the post, compute in us-west-2, data in us-east-2, and NeuralCache turning 60 ms of RTT into sub-5 ms local reads.
π§ AWS, us-west-2 (spoke, treino)
- PyTorchJob FSDP 2Γ ml.p5.48xlarge (16 H100) (compute)
- Data loader 64 workers, prefetch_factor=4 (compute)
- CNQ spoke NeuralCache em NVMe local (storage)
π‘ Rede, inter-Region
- VPC peering 60 ms RTT Β· MTU 8500 (network)
π§ AWS, us-east-2 (hub, dados)
- CNQ hub (4Γ r5.8xlarge) C4 tokenizado, PV 50 Ti (storage)
- HyperPod hub baseline co-localizado (compute)
Flows
- loader -> gpu: batches 24 Γ 4,096 tokens
- loader -> spoke: NFSv3: 4 KB blocks
- spoke -> peer: miss (4β6% after warmup)
- peer -> hub: remote read TCP 2049
- hub -> spoke: predictive prefetch + metadata within seconds
- hubgpu -> hub: baseline: 116β117 samples/s
How the mechanism actually works
Cloud Data Fabric doesn't replicate the data, it replicates the filesystem metadata to the spoke within seconds, with strict consistency between the ends. Reads are served from the nearest valid source: a hit lands on the spoke's local NVMe, a miss goes to the hub over the peering. The transport protocol between portals uses pacing-based congestion control rather than loss-based: Qumulo states near-line-rate throughput on paths of up to 900 ms RTT, which makes the 60 ms between Ohio and Oregon a comfortable case.
NeuralCache is the piece that changes the equation: it observes the sequential 4 KB block reads the data loader issues and learns the access pattern: LLM training is one of the most predictable workloads there is, the loader sweeps the tokenized dataset in order with prefetch_factor=4 and 256 batches in flight. Pattern learned, the cache fetches blocks before the request and reaches a 94β96% hit rate, serving from local NVMe in under 5 ms.
Convergence has measured phases: during batches 0β100, every read pays the 60 ms and GPU utilization drops to 80β90%; between batches 100 and 150, latency collapses to 5β10 ms; from there the spoke runs identical to the hub, GPUs above 98%. The warmup isn't configurable, it's learning, and its price shows up once per access pattern.
The validation numbers
- 94β96%: NeuralCache hit rate after warmup. Reads served from local NVMe at sub-5 ms, only 4β6% cross the Region
- 60 ms: RTT between us-east-2 and us-west-2. Paid in full during batches 0β100; invisible after convergence
- 0,08%: cold-start impact on a 100,000-batch run. The post's extrapolation: 0.81% at 10,000 batches, 0.008% at 1 million
The benchmark under scrutiny: what was measured and what wasn't
The experimental design is honest about its scope: the same job, a 1.02-billion-parameter LLaMA v3 over C4 tokenized in 4,096-token windows, 48.4 million sequences, batch 24 per GPU, PyTorch 2.1 with FSDP, ran independently on both clusters, each with 2Γ ml.p5.48xlarge orchestrated by EKS via a Kubeflow PyTorchJob. Hub: 116β117 samples/s and 18.5 minutes for 999 batches. Spoke with a warm cache: 115β116 samples/s, the same 18.5 minutes, GPUs above 99% from the first batch. A technical tie, measured, not promised.
That said, what the benchmark does not cover matters as much as what it does. It's 2 nodes per cluster: 16 GPUs reading 1.0β1.3 GBps; nothing guarantees 32 nodes contending for the same 4Γ r5.8xlarge hub hold the curve, and CNQ scales performance independently of storage precisely because it will need to. The workload is pure sequential read, the ideal case for a predictive prefetcher. And the write path was left out: FSDP checkpoints are heavy write bursts, and the post doesn't measure what happens when they cross the fabric. The 0.08% impact extrapolation at 100,000 batches assumes a single warmup, true as long as the access pattern doesn't change mid-run.
What actually holds the result up: The throughput tie doesn't come from magic networking, it comes from the fact that training data loading is one of the most predictable I/O patterns in computing. The loader sweeps tokenized sequences in order, in 4 KB blocks, with 256 batches of lookahead declared in the code itself. Any decent predictive cache nails that pattern; NeuralCache's merit is nailing it across 60 ms of WAN without strangling the pipeline. The implication is the contrapositive: random-access workloads, fine-tuning with aggressive per-epoch shuffling, feature stores, sparse reads, inherit none of these numbers.
Failure modes and limits of the network path
The network has fine print: inter-Region peering runs at an MTU of 8,500 bytes, not the 9,001 of intra-Region, and it's not transitive: each hub-spoke pair needs its own connection, so three spokes become three peerings and three pairs of route tables with CIDRs that must not overlap. NFS crosses on TCP 2049, allowed by security group between the VPCs, and inter-Region peering traffic is encrypted in transit by AWS itself, no VPN to operate.
A hub outage is the structural failure mode: the spoke serves hits from local cache, but the 4β6% of misses depend on the hub, unavailability in us-east-2 degrades the remote run in minutes, not hours. The peering stays Active even when a regional event blocks traffic; detection is on you, not AWS.
Cost is dominated by the miss: inter-Region transfer is billed per GB in both directions of the peering, and the good math is exactly the hit rate, at 94β96%, you pay the dataset's crossing roughly once (warmup) plus the residual fraction, versus paying the full copy plus duplicated storage under classic replication. Qumulo states over 30% reduction in WAN traffic and egress from prefetch aggregating reads.
Minimum observability: cache hit rate, GPU utilization and samples/s tell the same story from different angles: GPUs dropping from 98% to 85% with no code change is the cache losing the pattern.
Anti-patterns
- Replicating petabytes 'just in case': duplicating the dataset across Regions before measuring the cache pays full transfer, doubles storage and creates two truths to sync for years, the cost that matters is not the copy, it's keeping the copies honest.
- Extending the numbers to random access: the 94β96% hit rate was measured over sequential 4 KB sweeps; aggressive per-epoch shuffling or sparse reads defeat the prefetcher and hand the 60 ms back on every miss, measure with your own loader before committing the schedule.
- Ignoring the write path: the benchmark measures reads; FSDP checkpoints are write bursts the post does not validate cross-Region, land checkpoints on storage local to the compute Region (or same-Region S3) until you've measured otherwise.
-
Treating the hub as an invisible dependency: without alarms on hit rate and GPU utilization, a hub degradation in
us-east-2shows up first as slow training inus-west-2, and the peering will readActivethe whole time.
Through the Well-Architected lens
-
security: Inter-Region peering encrypts traffic in transit and keeps everything on private IPs; the surface to govern is the security group allowing
TCP 2049only between the two VPCs' CIDRs, plus the awareness that NFSv3 withnolockcarries no authentication of its own. -
reliability: The hub is a single point for misses: a spoke survives partial degradation by serving hits, but not a hub outage. Treat hub and peering health as part of the training SLO, and remember the peering's
Activestate does not attest traffic is flowing.
Curator's note: I would adopt this architecture today in one specific scenario: GPU capacity secured in a Region that isn't the data's, a run long enough to dilute the warmup, and sequential reads dominating. Before committing a schedule, I'd run exactly what the post ran, the same job on both ends, 999 batches, comparing samples/s and GPU utilization, because vendor benchmarks are reproduced, not inherited. And I'd instrument the hit rate from day one: the hard-won lesson from years operating distributed storage is that a cache degrading silently becomes an incident named after something else, here, 'training got slow' three layers above where the problem lives.
Verdict
Use HyperPod with CNQ and Cloud Data Fabric when: the GPU capacity you secured is in a different Region from the dataset, the run has tens of thousands of batches to amortize the warmup, and access is sequential data-loader reads. Stay co-located (or on FSx for Lustre with same-Region data) when: the job is short, the access pattern is random, or a cross-Region write path is critical to the pipeline, none of that was validated. The central result is solid within its scope: 60 ms of RTT becomes a 0.08% warmup cost on a production run, and that changes the question from 'where is the data?' to 'where are the GPUs?'. But the validation is 2 nodes per cluster and measures reads only, treat the numbers as a floor to reproduce, not a ceiling under contract.
Rating: adopt-with-conditions
References
- Multi-Region training with Amazon SageMaker HyperPod and Qumulo (AWS ML Blog, 25/09/2026)
- Amazon SageMaker HyperPod: Developer Guide
- How VPC peering connections work (MTU inter-Region, limites e lifecycle)
- Qumulo NeuralCache, predictive caching engine
- Qumulo Cloud Data Fabric, hub/spoke e consistΓͺncia estrita
Originally published at fernando.moretes.com. By Fernando F. Azevedo: Senior Solutions Architect.
Top comments (0)