Kubernetes sees nodes. vSphere sees VMs on hosts. Neither tells you when DRS has put two of your three etcd VMs on the same ESXi host. From then on, one host failure takes your control plane down, and every dashboard still says green.
I spent nearly seven years in VMware support, and I couldn't find a tool that joins those two views and simulates host failures against Kubernetes semantics. So I built kube-hostfail, then wrapped it in a working observability and service-mesh stack so you can see it in action.
Repo: https://github.com/Willey2003/k8s-vsphere-ha-cockpit (Apache-2.0)
What it checks
kube-hostfail is read-only. It joins node and etcd-member data from the Kubernetes API with VM and host data from vCenter (pyvmomi), then simulates every single-host failure, and optionally every two-host failure:
-
etcd quorum: external etcd (from
--etcd-serverson kube-apiserver) or stacked etcd pods, matched to VMs by name, guest hostname or guest IP - API availability: control-plane nodes left, quorum, and which host holds the kube-vip VIP lease
- Workloads: Deployments and StatefulSets losing every ready replica; PodDisruptionBudget violations
- Capacity: can the surviving workers absorb the evicted pods' requests?
- DRS: is each etcd and control-plane VM covered by an enabled VM-VM anti-affinity rule?
Verdicts are SAFE, DEGRADED, APP OUTAGE and CLUSTER DOWN. For missing DRS rules it prints the govc or PowerCLI command and never applies it.
Try it in 30 seconds, no vSphere needed
git clone https://github.com/Willey2003/k8s-vsphere-ha-cockpit.git
cd k8s-vsphere-ha-cockpit/kube-hostfail && pip install .
kube-hostfail report --snapshot examples/lab-drifted.json --depth 2
The bundled snapshot is a lab whose etcd members have drifted onto one host:
[DOWN] esx002.corp.example CLUSTER DOWN etcd 1/3 apiservers 1
- etcd loses quorum: 1/3 members left, 2 needed (members on esx002: etcd-1, etcd-2)
DRS ANTI-AFFINITY
etcd MISSING fix (review first): govc cluster.rule.create -cluster 'LAB-CL01' -name k8s-lab-etcd-anti-affinity -enable -anti-affinity etcd-1 etcd-2 etcd-3
RESULT: NOT highly available - losing esx002.corp.example takes the cluster down.
kube-hostfail check exits with code 2 when any single host is a single point of failure, so you can put it in a pipeline or a cron job.
The cockpit around it
One script per stage, all idempotent:
scripts/00-preflight.sh # read-only: what exists, what will be created, what could block you
scripts/10-observability.sh # kube-prometheus-stack tuned for kubeadm
scripts/20-mesh.sh # Istio, Jaeger, Kiali, Gateway
scripts/30-apps.sh # kube-hostfail, launcher, demo shop, routes
scripts/40-verify.sh
-
Prometheus + Grafana with no permanently red targets, an
emptyDirfallback when there's no StorageClass, opt-in scraping of external etcd, and a kube-hostfail dashboard with alert rules. -
Istio + Kiali + Jaeger with a three-service demo shop (one slow, flaky
pricingversion) and a load generator, so the graph and the traces are never empty. -
One Gateway API Gateway, path-routed:
/grafana,/kiali,/jaeger,/hostfail,/shop. A single address or a single SSH tunnel reaches everything. -
An app launcher that builds its tiles from your
HTTPRoutes, with a live health dot per tile. Publish an app, a tile appears.
What a real cluster taught me
Unit tests and a vCenter simulator passed, so I installed it on a real kubeadm 1.32 cluster (3 managers, 4 workers, CRI-O, Calico, MetalLB, external etcd) that already ran another Prometheus stack. The first run broke in three places:
-
Port clash. The existing stack's node-exporter used host port 9100 on the same nodes, so four of my pods stayed
Pending("didn't have free ports"). Mine now uses 9101. -
Grafana scrape. Serving Grafana from
/grafanameans/metricsredirects toroot_url, so the target went DOWN. The ServiceMonitor now scrapes/grafana/metrics. -
Uppercase GitHub owner. My CI built the image tag from
github.repository_owner, and container image names must be lowercase.
After fixing those, the Gateway came up with a MetalLB address, Kiali reported Prometheus, Grafana and Jaeger reachable, and Jaeger held traces for all three demo services. Your existing cluster isn't modified: everything lands in new namespaces, and sidecars are injected only into the demo namespace.
I also left one honest finding in the README: two managers block the node-exporter port at the host firewall, so those two targets show DOWN. That's a host-level change I wasn't going to make on a cluster I don't own.
What is and isn't verified
Verified: the simulation logic (20 unit tests), an end-to-end test against govmomi's vcsim, chart and manifest validation in CI, and a full install on a real cluster. Not yet verified: live mode against a production vCenter. The simulator isn't byte-for-byte vCenter, so I'd welcome reports from real environments.
Not modelled yet: datastore/PV placement, network partitions, DaemonSets and vSphere HA admission control.
Try it and tell me what breaks
If you run Kubernetes on vSphere, run kube-hostfail report against a read-only vCenter account and see what it says about your etcd members. Issues and PRs are welcome.
Top comments (0)