After I shared my VMware to OpenShift Virtualization lab, a consultant asked me a question I keep thinking about: “Is this about moving a VM, or about moving application stacks? It’s a lab, so what’s the overall context?”
It’s a fair challenge. My lab proves the mechanics on one VM: networking, storage, the Migration Toolkit for Virtualization (MTV) mappings and the cutover. But nobody migrates “a VM”. You migrate a payroll system, an order platform, a reporting stack, each made of several VMs that talk to each other, to databases, to load balancers and to things nobody has written down.
In my years in VMware support, the escalations that hurt most were rarely “the tool failed”. They were “we moved half of an application and the other half stopped working”. So this post is about the part the lab doesn’t show: how to go from one migrated VM to a migration programme, organised by application, in waves, with a way back.
One VM vs. a programme
| One-VM lab | A real migration programme | |
|---|---|---|
| Unit of work | A VM | An application (all of its tiers) |
| Main risk | Wrong mapping, ports, storage class | Broken dependencies, missed change windows, no rollback |
| Downtime | Doesn’t matter | Agreed per application with its owner |
| Network | One test VLAN | Keep IPs or re-IP, firewall rules, load balancers, DNS |
| Success | The VM boots and pings | The business process works end to end, and monitoring, backup and CMDB know about it |
The tooling is the same. What changes is the planning around it.
Step 1: Inventory and dependencies
Start with facts, not spreadsheets people remember.
- Inventory: export every VM with its CPU, memory, disks, datastore, port groups, guest OS, snapshots and VMware Tools state. MTV’s own inventory, RVTools or a PowerCLI export all work.
- Dependencies: who talks to whom. Use what you already have: network flow data from your monitoring or firewall logs, application owner interviews, and the CMDB as a starting point (not as truth). On a few critical VMs,
ss -tunp(Linux) ornetstat -ano(Windows) over a business day is surprisingly revealing. - Output: a list of application groups. Each group is the set of VMs that must move together, plus its external dependencies (shared databases, AD, file shares, licence servers, SMTP relays).
A rule I use: if two VMs exchange traffic every few seconds, they belong in the same wave. Splitting them puts latency, and sometimes a firewall, between them during the migration.
Step 2: Decide what happens to each VM
Not every VM should go through MTV. Give each one a disposition:
| Disposition | When |
|---|---|
| Rehost with MTV | Supported guest OS, standard disks and NICs. Most VMs. |
| Rehost later | Needs fixing first: unsupported OS, BitLocker, Windows Fast Startup, old snapshots, CBT off |
| Special handling | Shared-disk clusters, RDMs, very large or very busy databases, appliances with vendor support rules |
| Replatform | Stateless apps you plan to containerise anyway; do it once, not twice |
| Retire | Nobody owns it, nobody uses it. Every migration finds these. |
| Stay on vSphere (for now) | Licence-bound or vendor-certified only on VMware |
The readiness checks per VM are the boring part that saves weekends: a supported guest OS, VMware Tools running, no leftover snapshots, CBT enabled for warm migration, BitLocker suspended, Windows Fast Startup off, and firmware (BIOS or UEFI) recorded so the target VM matches.
Step 3: Network: keep the IPs or re-IP?
This decision drives everything else, so make it early.
- Keep the IPs (recommended for the first waves). Present the same VLANs to OpenShift with NMState bridges and NetworkAttachmentDefinitions, so a migrated VM lands on the same subnet. MTV keeps the NIC MAC addresses, so DHCP reservations and MAC-bound licences keep working. Firewall rules, load-balancer pools and DNS don’t change.
- Re-IP. Only when you’re redesigning the network anyway. Every re-IP touches DNS, firewall rules, load balancers, application configs and sometimes certificates, so it multiplies the testing.
The most common first-wave surprise I’ve seen in labs and in support is routing, not MTV. A VLAN that “exists” on OpenShift but isn’t trunked to every node, or a subnet mask that doesn’t match what the router actually routes. Test the network path with a throwaway VM on every target VLAN before wave 1.
Step 4: Build the waves
The SOP I published uses this structure:
| Wave | What goes in | Why |
|---|---|---|
| 0 | Test VMs: one Linux and one Windows, one per storage protocol | Prove each protocol, network and the runbook |
| 1 | Low-risk, stateless VMs (web servers, jump hosts) | Build confidence and real timing data |
| 2..n | Application groups, all tiers of one application together | Keep dependencies intact |
| Last | Databases, shared-disk clusters, very large VMs | Most planning, warm migration, longest windows |
Three rules make waves predictable:
- Group by application, not by datastore. Datastores are how vSphere stores things; applications are how the business notices outages.
- Keep each MTV plan to roughly 10–20 VMs. When something fails, you want to know which VM and why, not search through 200.
- Size the window from measured numbers. After wave 0 and 1 you know your real throughput. Then:
window ≥ (total GB to copy ÷ measured GB per hour) + conversion + validation + rollback buffer
For example, if wave 1 showed 300 GB per hour and an application has 1.2 TB of disks, a cold copy alone is about 4 hours, before conversion, testing and a rollback buffer. That’s usually the moment the application owner chooses warm migration.
Step 5: Cold or warm, per tier
| Cold | Warm | |
|---|---|---|
| Source VM during copy | Powered off | Keeps running |
| Downtime | Full copy + conversion | Final delta + conversion + boot |
| Needs | Nothing extra | CBT on, VDDK strongly recommended |
| Good for | Small VMs, test VMs, tiers that can be off for hours | Production tiers with short change windows |
In one application you’ll often mix them. The web tier can go cold on a Saturday morning; the database goes warm, with precopy running for days and a cutover scheduled in the change window. With warm migration, the downtime you agree with the business is the final delta, not the full disk size.
Step 6: Runbook and rollback
Every wave gets the same runbook. A shortened version of the one in the SOP:
- T-5 days: change approved; application owner and rollback criteria agreed. Readiness checks on every VM.
- T-3 days: backups of the source VMs verified (restored, not just “job succeeded”).
- T-2 days: MTV plan created and Ready, all critical concerns cleared; warm precopy started.
- T-1 hour: cluster and storage health checks, capacity checked.
- T-0: application owner stops the service if needed; cutover.
- T+: watch the pipeline, validate each VM, run the application’s smoke tests, then go / no-go.
- T+1 day: update the CMDB, monitoring and backup jobs for the new VMs.
- T+14 days: hypercare ends; only now delete the source VMs.
Rollback is the reason this works. MTV copies the disks and leaves the source VM in vCenter. If the go/no-go fails, you power off the new VM on OpenShift, power the source VM back on, and the application is where it was (same IPs if you kept them). Write the rollback steps down per application, and decide in advance who says no-go and by when.
The things that get forgotten
These aren’t migration steps, but each one has caused a “successful” migration to fail a week later:
- Backup: the vSphere backup jobs don’t follow the VM. Set up backup for OpenShift Virtualization VMs (OADP or your vendor’s tool) and test a restore before the source VMs are deleted.
- Monitoring and alerting: agents and dashboards that keyed on vCenter objects.
- Licensing: check guest OS and application licensing on the new platform with your vendors before wave 1, not after.
- Operations: your team now runs VMs as Kubernetes objects. Live migration, node maintenance and upgrades work differently, so train people before the first production wave, not during it.
- Storage for production: the lab used NFS. Production needs a CSI driver certified for OpenShift Virtualization, with RWX volumes so VMs can live-migrate during node maintenance.
Where my lab stands, honestly
The lab validated the mechanics on a single Ubuntu VM with cold migration. The wave structure, runbook and rollback above come from the production part of my SOP and from what I saw in VMware support escalations. They’re not yet proven on a multi-tier application in my lab.
That’s the next build: a three-tier application (web, app, database) moved as one wave, with warm migration on the database, a timed cutover and a deliberate rollback test. I’ll publish the results, including what goes wrong.
Your checklist
- Inventory every VM and map dependencies into application groups.
- Give each VM a disposition; fix the “rehost later” ones early.
- Decide keep-IP vs re-IP; test every target VLAN with a throwaway VM.
- Wave 0 with one VM per OS and storage protocol; measure throughput.
- Build waves by application, 10–20 VMs per plan, windows sized from real numbers.
- Choose cold or warm per tier; enable CBT and VDDK for warm.
- Run the same runbook every wave, with a written rollback and a named go/no-go owner.
- Move backup, monitoring, CMDB and licensing with the VMs.
Read more
- The 10-part lab series, from Linux basics to a migrated VM: start with Part 1, or jump to Part 8: MTV and Part 10: Troubleshooting and going to production.
- The full SOP (including FC, iSCSI and NFS storage, wave planning and the runbook): github.com/Willey2003/openshift-vmware-migration-sop
If you’re planning a VMware exit, I’d like to hear how you’re grouping applications into waves, and what surprised you. Leave a comment here or message me on LinkedIn.
Top comments (0)