<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gaganpreet Singh</title>
    <description>The latest articles on DEV Community by Gaganpreet Singh (@willeysingh).</description>
    <link>https://dev.to/willeysingh</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F387686%2F33ae1247-e556-4db9-bbcb-64a136821265.jpg</url>
      <title>DEV Community: Gaganpreet Singh</title>
      <link>https://dev.to/willeysingh</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/willeysingh"/>
    <language>en</language>
    <item>
      <title>Would your Kubernetes cluster survive losing one ESXi host? I built a read-only tool to find out</title>
      <dc:creator>Gaganpreet Singh</dc:creator>
      <pubDate>Sun, 11 Oct 2026 01:09:34 +0000</pubDate>
      <link>https://dev.to/willeysingh/would-your-kubernetes-cluster-survive-losing-one-esxi-host-i-built-a-read-only-tool-to-find-out-10o9</link>
      <guid>https://dev.to/willeysingh/would-your-kubernetes-cluster-survive-losing-one-esxi-host-i-built-a-read-only-tool-to-find-out-10o9</guid>
      <description>&lt;p&gt;Kubernetes sees &lt;em&gt;nodes&lt;/em&gt;. vSphere sees &lt;em&gt;VMs on hosts&lt;/em&gt;. Neither tells you when DRS has put two of your three etcd VMs on the same ESXi host. From then on, one host failure takes your control plane down, and every dashboard still says green.&lt;/p&gt;

&lt;p&gt;I spent nearly seven years in VMware support, and I couldn't find a tool that joins those two views and simulates host failures against Kubernetes semantics. So I built &lt;strong&gt;kube-hostfail&lt;/strong&gt;, then wrapped it in a working observability and service-mesh stack so you can see it in action.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/Willey2003/k8s-vsphere-ha-cockpit" rel="noopener noreferrer"&gt;https://github.com/Willey2003/k8s-vsphere-ha-cockpit&lt;/a&gt; (Apache-2.0)&lt;/p&gt;

&lt;h2&gt;
  
  
  What it checks
&lt;/h2&gt;

&lt;p&gt;kube-hostfail is read-only. It joins node and etcd-member data from the Kubernetes API with VM and host data from vCenter (pyvmomi), then simulates every single-host failure, and optionally every two-host failure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;etcd quorum&lt;/strong&gt;: external etcd (from &lt;code&gt;--etcd-servers&lt;/code&gt; on kube-apiserver) or stacked etcd pods, matched to VMs by name, guest hostname or guest IP&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API availability&lt;/strong&gt;: control-plane nodes left, quorum, and which host holds the kube-vip VIP lease&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workloads&lt;/strong&gt;: Deployments and StatefulSets losing every ready replica; PodDisruptionBudget violations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity&lt;/strong&gt;: can the surviving workers absorb the evicted pods' requests?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DRS&lt;/strong&gt;: is each etcd and control-plane VM covered by an enabled VM-VM anti-affinity rule?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Verdicts are &lt;code&gt;SAFE&lt;/code&gt;, &lt;code&gt;DEGRADED&lt;/code&gt;, &lt;code&gt;APP OUTAGE&lt;/code&gt; and &lt;code&gt;CLUSTER DOWN&lt;/code&gt;. For missing DRS rules it &lt;strong&gt;prints&lt;/strong&gt; the &lt;code&gt;govc&lt;/code&gt; or PowerCLI command and never applies it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it in 30 seconds, no vSphere needed
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/Willey2003/k8s-vsphere-ha-cockpit.git
&lt;span class="nb"&gt;cd &lt;/span&gt;k8s-vsphere-ha-cockpit/kube-hostfail &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
kube-hostfail report &lt;span class="nt"&gt;--snapshot&lt;/span&gt; examples/lab-drifted.json &lt;span class="nt"&gt;--depth&lt;/span&gt; 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bundled snapshot is a lab whose etcd members have drifted onto one host:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[DOWN] esx002.corp.example   CLUSTER DOWN  etcd 1/3  apiservers 1
       - etcd loses quorum: 1/3 members left, 2 needed (members on esx002: etcd-1, etcd-2)
DRS ANTI-AFFINITY
  etcd   MISSING   fix (review first): govc cluster.rule.create -cluster 'LAB-CL01' -name k8s-lab-etcd-anti-affinity -enable -anti-affinity etcd-1 etcd-2 etcd-3
RESULT: NOT highly available - losing esx002.corp.example takes the cluster down.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;kube-hostfail check&lt;/code&gt; exits with code 2 when any single host is a single point of failure, so you can put it in a pipeline or a cron job.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cockpit around it
&lt;/h2&gt;

&lt;p&gt;One script per stage, all idempotent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;scripts/00-preflight.sh      &lt;span class="c"&gt;# read-only: what exists, what will be created, what could block you&lt;/span&gt;
scripts/10-observability.sh  &lt;span class="c"&gt;# kube-prometheus-stack tuned for kubeadm&lt;/span&gt;
scripts/20-mesh.sh           &lt;span class="c"&gt;# Istio, Jaeger, Kiali, Gateway&lt;/span&gt;
scripts/30-apps.sh           &lt;span class="c"&gt;# kube-hostfail, launcher, demo shop, routes&lt;/span&gt;
scripts/40-verify.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prometheus + Grafana&lt;/strong&gt; with no permanently red targets, an &lt;code&gt;emptyDir&lt;/code&gt; fallback when there's no StorageClass, opt-in scraping of &lt;em&gt;external&lt;/em&gt; etcd, and a kube-hostfail dashboard with alert rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Istio + Kiali + Jaeger&lt;/strong&gt; with a three-service demo shop (one slow, flaky &lt;code&gt;pricing&lt;/code&gt; version) and a load generator, so the graph and the traces are never empty.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One Gateway API Gateway, path-routed&lt;/strong&gt;: &lt;code&gt;/grafana&lt;/code&gt;, &lt;code&gt;/kiali&lt;/code&gt;, &lt;code&gt;/jaeger&lt;/code&gt;, &lt;code&gt;/hostfail&lt;/code&gt;, &lt;code&gt;/shop&lt;/code&gt;. A single address or a single SSH tunnel reaches everything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An app launcher&lt;/strong&gt; that builds its tiles from your &lt;code&gt;HTTPRoute&lt;/code&gt;s, with a live health dot per tile. Publish an app, a tile appears.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What a real cluster taught me
&lt;/h2&gt;

&lt;p&gt;Unit tests and a vCenter simulator passed, so I installed it on a real kubeadm 1.32 cluster (3 managers, 4 workers, CRI-O, Calico, MetalLB, external etcd) that already ran another Prometheus stack. The first run broke in three places:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Port clash.&lt;/strong&gt; The existing stack's node-exporter used host port 9100 on the same nodes, so four of my pods stayed &lt;code&gt;Pending&lt;/code&gt; ("didn't have free ports"). Mine now uses 9101.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grafana scrape.&lt;/strong&gt; Serving Grafana from &lt;code&gt;/grafana&lt;/code&gt; means &lt;code&gt;/metrics&lt;/code&gt; redirects to &lt;code&gt;root_url&lt;/code&gt;, so the target went DOWN. The ServiceMonitor now scrapes &lt;code&gt;/grafana/metrics&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uppercase GitHub owner.&lt;/strong&gt; My CI built the image tag from &lt;code&gt;github.repository_owner&lt;/code&gt;, and container image names must be lowercase.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;After fixing those, the Gateway came up with a MetalLB address, Kiali reported Prometheus, Grafana and Jaeger reachable, and Jaeger held traces for all three demo services. Your existing cluster isn't modified: everything lands in new namespaces, and sidecars are injected only into the demo namespace.&lt;/p&gt;

&lt;p&gt;I also left one honest finding in the README: two managers block the node-exporter port at the host firewall, so those two targets show DOWN. That's a host-level change I wasn't going to make on a cluster I don't own.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is and isn't verified
&lt;/h2&gt;

&lt;p&gt;Verified: the simulation logic (20 unit tests), an end-to-end test against govmomi's &lt;code&gt;vcsim&lt;/code&gt;, chart and manifest validation in CI, and a full install on a real cluster. &lt;strong&gt;Not yet verified:&lt;/strong&gt; live mode against a production vCenter. The simulator isn't byte-for-byte vCenter, so I'd welcome reports from real environments.&lt;/p&gt;

&lt;p&gt;Not modelled yet: datastore/PV placement, network partitions, DaemonSets and vSphere HA admission control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it and tell me what breaks
&lt;/h2&gt;

&lt;p&gt;If you run Kubernetes on vSphere, run &lt;code&gt;kube-hostfail report&lt;/code&gt; against a read-only vCenter account and see what it says about your etcd members. Issues and PRs are welcome.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/Willey2003/k8s-vsphere-ha-cockpit" rel="noopener noreferrer"&gt;https://github.com/Willey2003/k8s-vsphere-ha-cockpit&lt;/a&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>vmware</category>
      <category>devops</category>
      <category>observability</category>
    </item>
    <item>
      <title>Private AI SOC Triage Lab: Can a Small Local LLM Triage Alerts Safely?</title>
      <dc:creator>Gaganpreet Singh</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:00:00 +0000</pubDate>
      <link>https://dev.to/willeysingh/private-ai-soc-triage-lab-can-a-small-local-llm-triage-alerts-safely-16l7</link>
      <guid>https://dev.to/willeysingh/private-ai-soc-triage-lab-can-a-small-local-llm-triage-alerts-safely-16l7</guid>
      <description>&lt;p&gt;Can a small, local LLM take tier-1 alert triage off a SOC team without missing attacks, and without being talked out of an escalation by the attacker? I built a lab to measure it instead of guessing. The code is on GitHub: &lt;a href="https://github.com/Willey2003/soc-triage-lab" rel="noopener noreferrer"&gt;github.com/Willey2003/soc-triage-lab&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three designs, same alerts
&lt;/h2&gt;

&lt;p&gt;The same 120 labelled alerts (60 malicious, 60 benign look-alikes) go through three designs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Rules:&lt;/strong&gt; a classic SIEM threshold.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;LLM only:&lt;/strong&gt; the model decides everything.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Hybrid:&lt;/strong&gt; rules at both ends, the model only in the grey zone, with a prompt-injection guard and fail-safe defaults.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything runs on one VM with no GPU, using &lt;code&gt;qwen2.5:1.5b&lt;/code&gt; through Ollama, and no alert data leaves the box.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;feeder -&amp;gt; API -&amp;gt; enrich -&amp;gt; rule score -&amp;gt; hybrid gate -&amp;gt; (grey zone) -&amp;gt; Ollama qwen2.5:1.5b
                                |                            |
                           SQLite cases &amp;lt;--------------------+    Prometheus &amp;lt;- /metrics -&amp;gt; Grafana
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Rules&lt;/th&gt;
&lt;th&gt;LLM only&lt;/th&gt;
&lt;th&gt;Hybrid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;td&gt;48%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;81%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recall (escalate)&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Missed attacks&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;False positives&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;55&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;23&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Escalated to an analyst&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;108&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;83&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt-injection success&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model calls&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;43&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The hybrid design cut the analyst queue by 17% and false positives by 43%, and missed no attacks. Letting the model decide everything was the worst design: it missed 7 attacks, closed foreign logins from new devices, and one injected alert talked it into closing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The guardrails that mattered
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Rule floor.&lt;/strong&gt; Any alert scoring 70 or above is escalated before the model sees it. This handled 54 of 120 alerts, including an injection written to evade the pattern list.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Close veto.&lt;/strong&gt; The model may close an alert only when a trusted CMDB or allowlist fact explains it (VPN egress, signed internal tool, approved scanner). Without it, the hybrid closed 3 malicious foreign logins.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Injection guard.&lt;/strong&gt; Alert text is untrusted: command lines, user agents and DNS labels are attacker-controlled. The guard strips instruction-like text and the prompt marks the alert as evidence only.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Fail-safe on invalid output.&lt;/strong&gt; 12% of model answers still failed validation, and every one was escalated rather than guessed.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What the model is actually good for
&lt;/h2&gt;

&lt;p&gt;Reading business context the rule does not have. Its wins were VPN logins (10/10 correctly closed in hybrid mode) and corporate-network password typos. It still did not trust an allowlist for telemetry DNS or the automation-host flag for SCCM PowerShell. A 1.5B model often restates the rule instead of weighing the context.&lt;/p&gt;

&lt;p&gt;Latency was about 10 seconds per alert on CPU (p95 18 s). That is fine for a steady queue and not fine for a burst, which is one more reason to keep the model in the grey zone only.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;The lab runs rootless with podman (tested on RHEL 9 with SELinux enforcing), and there is a Helm chart for Kubernetes with a NetworkPolicy so only the API can reach the model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/Willey2003/soc-triage-lab &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;soc-triage-lab
scripts/lab.sh up        &lt;span class="c"&gt;# build image, start the pod, pull the model&lt;/span&gt;
scripts/lab.sh &lt;span class="nb"&gt;eval &lt;/span&gt;120  &lt;span class="c"&gt;# rules vs llm vs hybrid on 120 labelled alerts&lt;/span&gt;
scripts/lab.sh &lt;span class="nb"&gt;test&lt;/span&gt;      &lt;span class="c"&gt;# unit tests&lt;/span&gt;
scripts/lab.sh down
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;The alerts are synthetic with clean labels, and the indicators use RFC 5737 documentation IP ranges and invented hashes and domains. It is one model, one seed, temperature 0. Larger models are the obvious next run. Treat the model’s output as a suggestion that a human reviews, and keep high-risk alerts on deterministic rules.&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>llm</category>
      <category>devops</category>
    </item>
    <item>
      <title>From One VM to a Whole Application: Planning VMware to OpenShift Migration Waves</title>
      <dc:creator>Gaganpreet Singh</dc:creator>
      <pubDate>Fri, 09 Oct 2026 13:23:41 +0000</pubDate>
      <link>https://dev.to/willeysingh/from-one-vm-to-a-whole-application-planning-vmware-to-openshift-migration-waves-171l</link>
      <guid>https://dev.to/willeysingh/from-one-vm-to-a-whole-application-planning-vmware-to-openshift-migration-waves-171l</guid>
      <description>&lt;p&gt;After I shared my VMware to OpenShift Virtualization lab, a consultant asked me a question I keep thinking about: &lt;em&gt;“Is this about moving a VM, or about moving application stacks? It’s a lab, so what’s the overall context?”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;It’s a fair challenge. My lab proves the &lt;strong&gt;mechanics&lt;/strong&gt; on one VM: networking, storage, the Migration Toolkit for Virtualization (MTV) mappings and the cutover. But nobody migrates “a VM”. You migrate a payroll system, an order platform, a reporting stack, each made of several VMs that talk to each other, to databases, to load balancers and to things nobody has written down.&lt;/p&gt;

&lt;p&gt;In my years in VMware support, the escalations that hurt most were rarely “the tool failed”. They were “we moved half of an application and the other half stopped working”. So this post is about the part the lab doesn’t show: &lt;strong&gt;how to go from one migrated VM to a migration programme&lt;/strong&gt;, organised by application, in waves, with a way back.&lt;/p&gt;

&lt;h2&gt;
  
  
  One VM vs. a programme
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;One-VM lab&lt;/th&gt;
&lt;th&gt;A real migration programme&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Unit of work&lt;/td&gt;
&lt;td&gt;A VM&lt;/td&gt;
&lt;td&gt;An application (all of its tiers)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main risk&lt;/td&gt;
&lt;td&gt;Wrong mapping, ports, storage class&lt;/td&gt;
&lt;td&gt;Broken dependencies, missed change windows, no rollback&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Downtime&lt;/td&gt;
&lt;td&gt;Doesn’t matter&lt;/td&gt;
&lt;td&gt;Agreed per application with its owner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Network&lt;/td&gt;
&lt;td&gt;One test VLAN&lt;/td&gt;
&lt;td&gt;Keep IPs or re-IP, firewall rules, load balancers, DNS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Success&lt;/td&gt;
&lt;td&gt;The VM boots and pings&lt;/td&gt;
&lt;td&gt;The business process works end to end, and monitoring, backup and CMDB know about it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The tooling is the same. What changes is the planning around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Inventory and dependencies
&lt;/h2&gt;

&lt;p&gt;Start with facts, not spreadsheets people remember.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Inventory:&lt;/strong&gt; export every VM with its CPU, memory, disks, datastore, port groups, guest OS, snapshots and VMware Tools state. MTV’s own inventory, RVTools or a PowerCLI export all work.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Dependencies:&lt;/strong&gt; who talks to whom. Use what you already have: network flow data from your monitoring or firewall logs, application owner interviews, and the CMDB as a starting point (not as truth). On a few critical VMs, &lt;code&gt;ss -tunp&lt;/code&gt; (Linux) or &lt;code&gt;netstat -ano&lt;/code&gt; (Windows) over a business day is surprisingly revealing.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Output:&lt;/strong&gt; a list of &lt;strong&gt;application groups&lt;/strong&gt;. Each group is the set of VMs that must move together, plus its external dependencies (shared databases, AD, file shares, licence servers, SMTP relays).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A rule I use: if two VMs exchange traffic every few seconds, they belong in the same wave. Splitting them puts latency, and sometimes a firewall, between them during the migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Decide what happens to each VM
&lt;/h2&gt;

&lt;p&gt;Not every VM should go through MTV. Give each one a disposition:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Disposition&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rehost with MTV&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Supported guest OS, standard disks and NICs. Most VMs.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rehost later&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Needs fixing first: unsupported OS, BitLocker, Windows Fast Startup, old snapshots, CBT off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Special handling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Shared-disk clusters, RDMs, very large or very busy databases, appliances with vendor support rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Replatform&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stateless apps you plan to containerise anyway; do it once, not twice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Retire&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Nobody owns it, nobody uses it. Every migration finds these.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Stay on vSphere (for now)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Licence-bound or vendor-certified only on VMware&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The readiness checks per VM are the boring part that saves weekends: a supported guest OS, VMware Tools running, no leftover snapshots, CBT enabled for warm migration, BitLocker suspended, Windows Fast Startup off, and firmware (BIOS or UEFI) recorded so the target VM matches.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Network: keep the IPs or re-IP?
&lt;/h2&gt;

&lt;p&gt;This decision drives everything else, so make it early.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Keep the IPs (recommended for the first waves).&lt;/strong&gt; Present the same VLANs to OpenShift with NMState bridges and NetworkAttachmentDefinitions, so a migrated VM lands on the same subnet. MTV keeps the NIC MAC addresses, so DHCP reservations and MAC-bound licences keep working. Firewall rules, load-balancer pools and DNS don’t change.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Re-IP.&lt;/strong&gt; Only when you’re redesigning the network anyway. Every re-IP touches DNS, firewall rules, load balancers, application configs and sometimes certificates, so it multiplies the testing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The most common first-wave surprise I’ve seen in labs and in support is routing, not MTV. A VLAN that “exists” on OpenShift but isn’t trunked to every node, or a subnet mask that doesn’t match what the router actually routes. Test the network path with a throwaway VM on every target VLAN before wave 1.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Build the waves
&lt;/h2&gt;

&lt;p&gt;The SOP I published uses this structure:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wave&lt;/th&gt;
&lt;th&gt;What goes in&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Test VMs: one Linux and one Windows, one per storage protocol&lt;/td&gt;
&lt;td&gt;Prove each protocol, network and the runbook&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Low-risk, stateless VMs (web servers, jump hosts)&lt;/td&gt;
&lt;td&gt;Build confidence and real timing data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2..n&lt;/td&gt;
&lt;td&gt;Application groups, all tiers of one application together&lt;/td&gt;
&lt;td&gt;Keep dependencies intact&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Last&lt;/td&gt;
&lt;td&gt;Databases, shared-disk clusters, very large VMs&lt;/td&gt;
&lt;td&gt;Most planning, warm migration, longest windows&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three rules make waves predictable:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Group by application, not by datastore.&lt;/strong&gt; Datastores are how vSphere stores things; applications are how the business notices outages.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Keep each MTV plan to roughly 10–20 VMs.&lt;/strong&gt; When something fails, you want to know which VM and why, not search through 200.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Size the window from measured numbers.&lt;/strong&gt; After wave 0 and 1 you know your real throughput. Then: &lt;code&gt;window ≥ (total GB to copy ÷ measured GB per hour) + conversion + validation + rollback buffer&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For example, if wave 1 showed 300 GB per hour and an application has 1.2 TB of disks, a cold copy alone is about 4 hours, before conversion, testing and a rollback buffer. That’s usually the moment the application owner chooses warm migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Cold or warm, per tier
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Cold&lt;/th&gt;
&lt;th&gt;Warm&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Source VM during copy&lt;/td&gt;
&lt;td&gt;Powered off&lt;/td&gt;
&lt;td&gt;Keeps running&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Downtime&lt;/td&gt;
&lt;td&gt;Full copy + conversion&lt;/td&gt;
&lt;td&gt;Final delta + conversion + boot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Needs&lt;/td&gt;
&lt;td&gt;Nothing extra&lt;/td&gt;
&lt;td&gt;CBT on, VDDK strongly recommended&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Good for&lt;/td&gt;
&lt;td&gt;Small VMs, test VMs, tiers that can be off for hours&lt;/td&gt;
&lt;td&gt;Production tiers with short change windows&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In one application you’ll often mix them. The web tier can go cold on a Saturday morning; the database goes warm, with precopy running for days and a cutover scheduled in the change window. With warm migration, the downtime you agree with the business is the &lt;strong&gt;final delta&lt;/strong&gt;, not the full disk size.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Runbook and rollback
&lt;/h2&gt;

&lt;p&gt;Every wave gets the same runbook. A shortened version of the one in the SOP:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;T-5 days:&lt;/strong&gt; change approved; application owner and rollback criteria agreed. Readiness checks on every VM.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;T-3 days:&lt;/strong&gt; backups of the source VMs verified (restored, not just “job succeeded”).&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;T-2 days:&lt;/strong&gt; MTV plan created and Ready, all critical concerns cleared; warm precopy started.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;T-1 hour:&lt;/strong&gt; cluster and storage health checks, capacity checked.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;T-0:&lt;/strong&gt; application owner stops the service if needed; cutover.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;T+:&lt;/strong&gt; watch the pipeline, validate each VM, run the application’s smoke tests, then &lt;strong&gt;go / no-go&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;T+1 day:&lt;/strong&gt; update the CMDB, monitoring and backup jobs for the new VMs.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;T+14 days:&lt;/strong&gt; hypercare ends; only now delete the source VMs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Rollback&lt;/strong&gt; is the reason this works. MTV copies the disks and leaves the source VM in vCenter. If the go/no-go fails, you power off the new VM on OpenShift, power the source VM back on, and the application is where it was (same IPs if you kept them). Write the rollback steps down per application, and decide in advance &lt;strong&gt;who&lt;/strong&gt; says no-go and &lt;strong&gt;by when&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The things that get forgotten
&lt;/h2&gt;

&lt;p&gt;These aren’t migration steps, but each one has caused a “successful” migration to fail a week later:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Backup:&lt;/strong&gt; the vSphere backup jobs don’t follow the VM. Set up backup for OpenShift Virtualization VMs (OADP or your vendor’s tool) and test a restore before the source VMs are deleted.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Monitoring and alerting:&lt;/strong&gt; agents and dashboards that keyed on vCenter objects.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Licensing:&lt;/strong&gt; check guest OS and application licensing on the new platform with your vendors before wave 1, not after.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Operations:&lt;/strong&gt; your team now runs VMs as Kubernetes objects. Live migration, node maintenance and upgrades work differently, so train people before the first production wave, not during it.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Storage for production:&lt;/strong&gt; the lab used NFS. Production needs a CSI driver certified for OpenShift Virtualization, with RWX volumes so VMs can live-migrate during node maintenance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where my lab stands, honestly
&lt;/h2&gt;

&lt;p&gt;The lab validated the mechanics on a single Ubuntu VM with cold migration. The wave structure, runbook and rollback above come from the production part of my SOP and from what I saw in VMware support escalations. They’re not yet proven on a multi-tier application in my lab.&lt;/p&gt;

&lt;p&gt;That’s the next build: a three-tier application (web, app, database) moved as one wave, with warm migration on the database, a timed cutover and a deliberate rollback test. I’ll publish the results, including what goes wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your checklist
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt; Inventory every VM and map dependencies into application groups.&lt;/li&gt;
&lt;li&gt; Give each VM a disposition; fix the “rehost later” ones early.&lt;/li&gt;
&lt;li&gt; Decide keep-IP vs re-IP; test every target VLAN with a throwaway VM.&lt;/li&gt;
&lt;li&gt; Wave 0 with one VM per OS and storage protocol; measure throughput.&lt;/li&gt;
&lt;li&gt; Build waves by application, 10–20 VMs per plan, windows sized from real numbers.&lt;/li&gt;
&lt;li&gt; Choose cold or warm per tier; enable CBT and VDDK for warm.&lt;/li&gt;
&lt;li&gt; Run the same runbook every wave, with a written rollback and a named go/no-go owner.&lt;/li&gt;
&lt;li&gt; Move backup, monitoring, CMDB and licensing with the VMs.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Read more
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  The 10-part lab series, from Linux basics to a migrated VM: &lt;a href="https://willeysingh.wordpress.com/2026/10/09/vmware-to-openshift-migration-lab-part-1-linux-for-the-migration-engineer-rhel-10/" rel="noopener noreferrer"&gt;start with Part 1&lt;/a&gt;, or jump to &lt;a href="https://willeysingh.wordpress.com/2026/10/09/vmware-to-openshift-migration-lab-part-8-migration-toolkit-for-virtualization-mtv/" rel="noopener noreferrer"&gt;Part 8: MTV&lt;/a&gt; and &lt;a href="https://willeysingh.wordpress.com/2026/10/09/vmware-to-openshift-migration-lab-part-10-troubleshooting-and-going-to-production/" rel="noopener noreferrer"&gt;Part 10: Troubleshooting and going to production&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;  The full SOP (including FC, iSCSI and NFS storage, wave planning and the runbook): &lt;a href="https://github.com/Willey2003/openshift-vmware-migration-sop" rel="noopener noreferrer"&gt;github.com/Willey2003/openshift-vmware-migration-sop&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you’re planning a VMware exit, I’d like to hear how you’re grouping applications into waves, and what surprised you. Leave a comment here or message me on &lt;a href="https://www.linkedin.com/in/storageelearner" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>openshift</category>
      <category>vmware</category>
      <category>kubernetes</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
