DEV Community

Gaganpreet Singh
Gaganpreet Singh

Posted on

Would your Kubernetes cluster survive losing one ESXi host? I built a read-only tool to find out

Kubernetes sees nodes. vSphere sees VMs on hosts. Neither tells you when DRS has put two of your three etcd VMs on the same ESXi host. From then on, one host failure takes your control plane down, and every dashboard still says green.

I spent nearly seven years in VMware support, and I couldn't find a tool that joins those two views and simulates host failures against Kubernetes semantics. So I built kube-hostfail, then wrapped it in a working observability and service-mesh stack so you can see it in action.

Repo: https://github.com/Willey2003/k8s-vsphere-ha-cockpit (Apache-2.0)

What it checks

kube-hostfail is read-only. It joins node and etcd-member data from the Kubernetes API with VM and host data from vCenter (pyvmomi), then simulates every single-host failure, and optionally every two-host failure:

  • etcd quorum: external etcd (from --etcd-servers on kube-apiserver) or stacked etcd pods, matched to VMs by name, guest hostname or guest IP
  • API availability: control-plane nodes left, quorum, and which host holds the kube-vip VIP lease
  • Workloads: Deployments and StatefulSets losing every ready replica; PodDisruptionBudget violations
  • Capacity: can the surviving workers absorb the evicted pods' requests?
  • DRS: is each etcd and control-plane VM covered by an enabled VM-VM anti-affinity rule?

Verdicts are SAFE, DEGRADED, APP OUTAGE and CLUSTER DOWN. For missing DRS rules it prints the govc or PowerCLI command and never applies it.

Try it in 30 seconds, no vSphere needed

git clone https://github.com/Willey2003/k8s-vsphere-ha-cockpit.git
cd k8s-vsphere-ha-cockpit/kube-hostfail && pip install .
kube-hostfail report --snapshot examples/lab-drifted.json --depth 2
Enter fullscreen mode Exit fullscreen mode

The bundled snapshot is a lab whose etcd members have drifted onto one host:

[DOWN] esx002.corp.example   CLUSTER DOWN  etcd 1/3  apiservers 1
       - etcd loses quorum: 1/3 members left, 2 needed (members on esx002: etcd-1, etcd-2)
DRS ANTI-AFFINITY
  etcd   MISSING   fix (review first): govc cluster.rule.create -cluster 'LAB-CL01' -name k8s-lab-etcd-anti-affinity -enable -anti-affinity etcd-1 etcd-2 etcd-3
RESULT: NOT highly available - losing esx002.corp.example takes the cluster down.
Enter fullscreen mode Exit fullscreen mode

kube-hostfail check exits with code 2 when any single host is a single point of failure, so you can put it in a pipeline or a cron job.

The cockpit around it

One script per stage, all idempotent:

scripts/00-preflight.sh      # read-only: what exists, what will be created, what could block you
scripts/10-observability.sh  # kube-prometheus-stack tuned for kubeadm
scripts/20-mesh.sh           # Istio, Jaeger, Kiali, Gateway
scripts/30-apps.sh           # kube-hostfail, launcher, demo shop, routes
scripts/40-verify.sh
Enter fullscreen mode Exit fullscreen mode
  • Prometheus + Grafana with no permanently red targets, an emptyDir fallback when there's no StorageClass, opt-in scraping of external etcd, and a kube-hostfail dashboard with alert rules.
  • Istio + Kiali + Jaeger with a three-service demo shop (one slow, flaky pricing version) and a load generator, so the graph and the traces are never empty.
  • One Gateway API Gateway, path-routed: /grafana, /kiali, /jaeger, /hostfail, /shop. A single address or a single SSH tunnel reaches everything.
  • An app launcher that builds its tiles from your HTTPRoutes, with a live health dot per tile. Publish an app, a tile appears.

What a real cluster taught me

Unit tests and a vCenter simulator passed, so I installed it on a real kubeadm 1.32 cluster (3 managers, 4 workers, CRI-O, Calico, MetalLB, external etcd) that already ran another Prometheus stack. The first run broke in three places:

  1. Port clash. The existing stack's node-exporter used host port 9100 on the same nodes, so four of my pods stayed Pending ("didn't have free ports"). Mine now uses 9101.
  2. Grafana scrape. Serving Grafana from /grafana means /metrics redirects to root_url, so the target went DOWN. The ServiceMonitor now scrapes /grafana/metrics.
  3. Uppercase GitHub owner. My CI built the image tag from github.repository_owner, and container image names must be lowercase.

After fixing those, the Gateway came up with a MetalLB address, Kiali reported Prometheus, Grafana and Jaeger reachable, and Jaeger held traces for all three demo services. Your existing cluster isn't modified: everything lands in new namespaces, and sidecars are injected only into the demo namespace.

I also left one honest finding in the README: two managers block the node-exporter port at the host firewall, so those two targets show DOWN. That's a host-level change I wasn't going to make on a cluster I don't own.

What is and isn't verified

Verified: the simulation logic (20 unit tests), an end-to-end test against govmomi's vcsim, chart and manifest validation in CI, and a full install on a real cluster. Not yet verified: live mode against a production vCenter. The simulator isn't byte-for-byte vCenter, so I'd welcome reports from real environments.

Not modelled yet: datastore/PV placement, network partitions, DaemonSets and vSphere HA admission control.

Try it and tell me what breaks

If you run Kubernetes on vSphere, run kube-hostfail report against a read-only vCenter account and see what it says about your etcd members. Issues and PRs are welcome.

Repo: https://github.com/Willey2003/k8s-vsphere-ha-cockpit

Top comments (0)