DEV Community

AWS Transform Now Supports Block Storage Migration to FSx for ONTAP — Benefits and Pitfalls from a Hands-On EC2 Test

Introduction

This is Part 2 of the series "The Data Foundation for AWS Modernization." Part 1 set out the lens: migration is not the goal but the entry point, and putting FSx for ONTAP at the data foundation lets you re-choose the compute. This part takes that entry point through AWS Transform, the core of modernization, and verifies hands-on that the data area lands on FSx for ONTAP as part of the migration on AWS. Migration from an on-premises VMware environment will come separately; the idea here is to start modernizing from the tightly coupled EC2 × EBS configuration by decoupling the data store.

Until now, when you migrated servers with AWS Transform (MGN), every disk was placed on Amazon EBS. If the source was not ONTAP and you wanted the data area on Amazon FSx for NetApp ONTAP (FSx for ONTAP below), you needed a "two-step migration" — migrate to EBS first, then move the data — or a third-party block storage migration tool.

With the update of 30 August 2026, FSx for ONTAP can now be chosen directly as the MGN target storage type.

That looked useful, so I tried it hands-on straight away.
The source in this test was not an on-premises VMware environment but an EC2 instance already on AWS (Amazon Linux 2023). Using that EC2 as the source, I ran the new path that migrates the data area (EBS) directly to an FSx for ONTAP iSCSI LUN.

Running it for real turned up several pitfalls that reading the documentation alone does not reveal.
Besides sorting out what this new feature can and cannot do, this article records the traps I actually stepped on during the test — errors hidden behind a successful job, and an unexpected capacity spike.

Here are the highlights up front, especially the traps to watch for.

  • Only data volumes can be targeted. Boot is always Amazon EBS.
  • What became GA is "the MGN target storage type". It is a different thing from presenting FSx for ONTAP as a VMware datastore (the Amazon EVS side).
  • In MGN's phases, "SNAPSHOT" was a volume Snapshot done as a metadata operation, and "LAUNCH" was the creation of a FlexClone (about 44 seconds for 8 GiB).
  • [Trap 1] Finalize is not just clean-up; it is the step where physical capacity temporarily peaks (about 2x).
  • [Trap 2] Even when the job status becomes COMPLETED, a failed final snapshot (fallback) or the creation of an unbootable target can be hiding behind it.

Test Assumptions and Scope

The test was run with the following environment and scope.
(Region: ap-northeast-1 / Source: Amazon EC2 (AL2023) / ONTAP: 9.18.1P3D1)

[What was run]
The whole flow: agent installation → replication → test launch → cutover → Finalize → teardown.

[Out of scope for this article (not verified)]

  • Migration with VMware as the source (no vCenter was prepared, so EC2 stood in as the source)
  • FSx for ONTAP as an Amazon EVS datastore (a different feature)
  • Durations at production-scale data volumes (the test used a tiny 16 GiB environment, dominated by fixed overhead)
  • Windows sources, and access over SMB / NFS
  • Post-migration optimization with lun move start

The Shape of the Configuration

A table-shaped diagram pairing three Amazon EC2 instances with their boot / root volume and their data volume. The rows are, from the top, the source (Amazon EC2 in this test), staging and after cutover. The columns are the boot / root volume and the data volume. The boot column is Amazon Elastic Block Store in all three rows: boot 8 GiB at the source, the boot staging at staging, and root 8 GiB after cutover. In the data column, only the source row is Amazon Elastic Block Store (data 4 GiB × 2); staging and after cutover are Amazon FSx for NetApp ONTAP (data staging, then the target FlexVol over iSCSI). The rows are joined downwards by block replication and cutover arrows
Figure 1: three EC2 instances and where each volume goes: the boot column is EBS in all three rows, and only the data column changes to FSx for ONTAP from staging on (dark theme)

The control path. AWS Transform reaches the ONTAP management endpoint (REST API / 443) inside the customer VPC through AWS PrivateLink and a Network Load Balancer. The management endpoint belongs to Amazon FSx for NetApp ONTAP, whose icon and service name sit directly above the box. The client certificate goes from AWS Secrets Manager, outside the VPC, to AWS Transform
Figure 2: the control path, from AWS Transform to the ONTAP management endpoint (dark theme)

The point is that the boot area and the data area go to different places. The boot column stays Amazon EBS in all three rows, and only the data column changes to FSx for ONTAP from the staging row on. The target instance boots from EBS and receives the data over iSCSI. The source in this test was EC2, but the lower two rows of the configuration are the same whether the source is a VMware datastore or on-premises block storage.

For the control path in Figure 2, the official blog says a PrivateLink connection is established automatically, but in practice a Network Load Balancer (NLB) and a VPC endpoint service are created inside your VPC. They incur hourly and LCU charges.

Layer Normal operation (replicating) After cutover
Source EC2 Running. The agent keeps sending blocks Can be stopped
Replication server EC2 MGN launches it automatically Terminated by Finalize
Amazon EBS Boot staging Becomes the target's root volume
FSx for ONTAP Staging FlexVol and LUN Target FlexVol (FlexClone)
Network Load Balancer Running Remains after Finalize (manual deletion needed)

Table 1: each layer after cutover: the one that needs manual deletion is the Network Load Balancer


The Most Important Distinction: What Became GA Is the MGN Target

Mix this up and you will evaluate the wrong product. It is easily confused with "FSx for ONTAP can be used as a VMware datastore".

Question Current support
Can FSx for ONTAP be used as the MGN target storage type? Yes (GA with this release)
Is FSx for ONTAP presented as a VMware datastore? No. The datastore is a separate feature on the Amazon EVS side
Can the boot disk also go on FSx for ONTAP? No. Boot is always Amazon EBS
Can EBS and FSx for ONTAP be mixed within one server? No. All data volumes use the same type
Can you migrate agentlessly? No. Agent-based replication only
Which protocol do clients connect with? iSCSI. The data is placed as LUNs inside a FlexVol

Table 2: questions and current support: the agentless path is not supported

The Known limitations in the MGN documentation state Agent-based replication only. If you were planning around vCenter's agentless path, the plan has to change. The path depends on whether you leave VMs for EC2, or keep the VMs and move only the storage out.

The replication template edit screen of AWS Transform MGN. The Amazon FSx for NetApp ONTAP configuration section has two inputs: Storage Virtual Machine (SVM) ID and FSx Storage Secret ARN. The note above them reads: Data disks will be migrated to FSx for ONTAP. Boot disk is always migrated to EBS as required by Amazon EC2.
Figure 3: the FSx for ONTAP settings in the replication template. Only the SVM ID and the Secret ARN are specified, and the note states that boot goes to EBS


Environments This Configuration Suits

Whether a migration fits depends on the shape of the workload, not on the product of the source storage.

  • The boot / root area and the data area are separate, and the data area is on block storage
  • You want capabilities such as Snapshot and thin provisioning on the data area after migration
  • You want to turn "migrate, then move the data" into a single step
  • You already run ONTAP and want the same mechanisms (Snapshot / FlexClone / SnapMirror) after migration

Most of these apply whatever the source storage is. Even when the source is ONTAP, however, permissions, quotas and Snapshot policies are not migrated and need to be set up again.

Environments Where Another Option Fits Better for Now

  • Boot and data sit on the same volume
  • You cannot install an agent (agentless is not supported)
  • The source is Windows and you want to move the data as SMB shares
  • File-level migration is enough and block-level replication is not needed (a file transfer such as AWS DataSync is sufficient)
  • The source is ONTAP and you want to take the volumes as they are (SnapMirror takes fewer steps)

A Mini Glossary

Term Meaning Where it appears in this article
SVM A logical server inside the file system Where MGN creates the FlexVol and LUN
FlexVol A volume that holds data. LUNs live inside it too MGN creates one for staging and one for the target
LUN A unit of block storage, exposed over iSCSI One source disk becomes one LUN
Snapshot A point-in-time image inside a volume MGN's SNAPSHOT phase creates it
FlexClone A writable clone made from a Snapshot The target FlexVol is one
Split Detaching a FlexClone from its parent and materializing it Finalize runs it

The Procedure and the Measured Time for Each Step

The entry point is the AWS Transform workspace. MGN runs the migration itself, but it is AWS Transform that handles everything from planning to landing zone and network as a single job.

# Step Measured
1 Initialize MGN Done from the console. The CLI failed (see below)
4 Install the agent Failed once (see below)
5 Initial full sync (16 GiB / 3 disks) 226 seconds
6 Test launch SNAPSHOT 43 seconds, CONVERSION about 4 minutes 30 seconds
7 Cutover (correct order) MGN job 649 seconds; 817 seconds from issuing it to confirming boot
9 Finalize Split under 60 seconds; about 33 minutes until clean-up completed

(n=1 measurements at 16 GiB / 3 disks. Fixed overhead dominates, so they cannot be used as a basis for RTO.)


Snapshot and FlexClone: What AWS Transform's Phase Names Refer To

I lined up the ONTAP side by timestamp to see which ONTAP operation each AWS Transform phase name corresponds to.

AWS Transform phase Corresponding ONTAP operation
SNAPSHOT Creating a volume Snapshot of the staging FlexVol
LAUNCH Creating a FlexClone from that Snapshot

The SNAPSHOT phase is not a plain data copy but a metadata operation (44 seconds for 8 GiB). The target volume is a clone with is_flexclone: true.

What matters most is that a FlexClone uses almost no physical capacity when it is created.
Measured just before Finalize, the target volume's logical size was 7.91 GiB while its physical consumption was only 35.5 MiB. Capacity planning has to look at physical values, not logical ones.


Finalize Is Not Clean-up; It Is Where Capacity Peaks

Running Finalize starts a split of the FlexClone (the operation that detaches it and makes it independent), which materializes the data.

Elapsed from T0 Observed change
Immediately Lifecycle CUTOVER, replication DISCONNECTED
About 3 minutes FlexClone split starts. Physical consumption rising from 35.5 MiB → 3.94 GiB
About 4 minutes Split completes. Physical consumption 8.57 GiB
About 13 minutes Staging FlexVol deleted, replication server EC2 terminated

There was a gap of about 9 minutes between the split completing and the staging volume being deleted.
During those 9 minutes, physical capacity holds twice the migrated data.

If you run Finalize while the aggregate's free space is below the migrated data size, it is likely to get stuck here. Keep in mind that Finalize is a capacity risk, not an availability risk.


Four Places I Stumbled

1. A Failed Final Snapshot Behind a Successful Job

To measure downtime, I shut down the source OS and then ran cutover. The job returned COMPLETED and the target LAUNCHED. It looked like a success, but the job log recorded this:

09:14:21  SNAPSHOT_START
09:19:22  SNAPSHOT_FAIL            timed out after 300 seconds
09:19:23  USING_PREVIOUS_SNAPSHOT

Enter fullscreen mode Exit fullscreen mode

The final sync did not complete, and the job fell back to the previous snapshot. The cause was my procedure: shutting down the OS also stopped the replication agent, and taking the crash-consistent snapshot timed out.

The correct procedure is to stop only the application's writes and run cutover with the OS and the agent still running. "Stop writes on the source" in the official documentation does not mean shutting down the OS.

In this test there were no recent writes, so the data matched, but doing this in a real environment loses the latest data written after the snapshot it fell back to. What makes this risky is that it does not show up in the job status (status or launchStatus) at all. By the API's design, the job returns COMPLETED as long as processing finishes, even after a fallback.

Lesson: After cutover, always check event in describe-job-log-items. A COMPLETED job is not evidence that the final sync succeeded.

2. Two Unbootable Targets: NVMe Device Names Re-enumerated

I ran cutover in the correct order twice, and both produced instances that would not boot (a UEFI shell reboot loop). The job log showed no abnormal events, and the status was COMPLETED.

Examining the boot EBS volume of an unbootable instance, the 8 GiB boot volume held the contents of a 4 GiB data disk. The boot and data disk assignment had been swapped.

The cause: when the source was stopped and restarted as part of the test, NVMe device names were re-enumerated and the boot disk moved from /dev/nvme0n1 to /dev/nvme1n1. MGN's staging-type assignment appears to be keyed on the device name, and it is not re-evaluated after the name moves.

The official best practices say not to restart the source before cutover in the first place. If a restart does happen, check the disk assignment with a command like the one below and redo the test launch.

# Check which device is treated as the boot disk, and its assignment
aws mgn get-replication-configuration --source-server-id "$SRC" \
  --query 'replicatedDisks[?isBootDisk==`true`].[deviceName,stagingDiskType]' --output table

Enter fullscreen mode Exit fullscreen mode

3. No Repair API, and Three Contradicting Errors

I called UpdateReplicationConfiguration to fix the disk assignment, but whatever parameters I passed, it returned contradicting errors, and the API could not fix it.

In the end I deleted the source server and reinstalled the agent, which turned out to be unnecessary: rerunning the installer is the documented procedure.
Rerunning the installer re-establishes the list of replicated disks and their mapping against the actual machine. To avoid losing launch template customizations, try rerunning the installer before deleting anything.

4. The Installer's Hint Did Not Match the Real Cause

When the agent installation failed, the installer displayed Are kernel linux headers installed correctly?. Reading the log closely, the real cause was lack of space in /tmp.

On Amazon Linux 2023, /tmp is tmpfs; on a t3.small it has only 955 MiB, which does not meet the installer's requirement (1 GB or more free).

sudo env TMPDIR=/var/tmp/mgn-build TEMP=/var/tmp/mgn-build TMP=/var/tmp/mgn-build \
  ./aws-replication-installer-init --region ap-northeast-1 --no-prompt

Enter fullscreen mode Exit fullscreen mode

Pointing it at /var/tmp like this works around it. The installer's hint is not necessarily the real cause, so check the error code in the log.


Teardown Checklist

Finalize does not finish the clean-up. The correct teardown order for leaving nothing behind is as follows (steps 1–7 depend on each other).

  1. Terminate the target and source EC2 instances
  2. Delete the MGN source server
  3. Reject the VPC endpoint connection, then delete the endpoint service
  4. Delete the NLB and its target group (important: charges continue until you delete them by hand)
  5. Delete the target FlexVol with fsx delete-volume
  6. Delete the SVM with fsx delete-storage-virtual-machine
  7. Delete the ONTAP security login

The staging EBS volumes are deleted automatically with a delay, but the target's root EBS volume has no DeleteOnTermination and must be deleted manually.


Initializing AWS Transform from the CLI: Create the IAM Roles First

If running aws mgn initialize-service from the API or CLI fails, the cause is that the IAM roles have not been created.
Initializing from the console creates the IAM roles for you, but on the CLI path you need to create eight roles (including AWSApplicationMigrationFsxProxyRole for FSx for ONTAP) and attach their policies first.


Closing

With this GA, a migration now finishes with the data area already on FSx for ONTAP as iSCSI LUNs, and the step of moving the data afterwards is gone.

Any server whose boot and data are separate fits the same shape, whatever the original storage.
On the other hand, watch for the configuration spanning two kinds of storage, the physical capacity spike during Finalize, and MGN-specific traps such as "a job status of COMPLETED does not mean the final sync succeeded".

I hope this test record helps anyone migrating servers with separate boot and data to EC2 while considering FSx for ONTAP as the home for the data area.


Resources


All test environments have been deleted. The figures are single measurements under a specific environment and conditions, and vary with data size and configuration. The configuration details are in the verification report above.

This article is Part 2 of the series "The Data Foundation for AWS Modernization." The Japanese original is on hatenablog.

Top comments (1)