Every time my homelab Kubernetes cluster broke, I had to redo everything by hand: create VMs, set up HAProxy, init kubeadm, install Calico, copy join tokens around, join control-plane nodes, join workers...
After the third or fourth time, I stopped fighting it and wrote a script that does the whole thing with one command.
What it builds
One KVM host → 6 VMs → a real kubeadm HA cluster:
┌─────────────────────────────┐
Client ──────▶ │ k8s-bastion (HAProxy) │
│ :6443 → control-plane API │
│ :80/:443 → ingress │
└─────────────┬───────────────┘
│
┌──────────────┬───────┴──────┬──────────────┐
▼ ▼ ▼
k8s-control1 k8s-control2 k8s-control3 (stacked etcd)
▼ ▼
k8s-worker1 k8s-worker2 (ingress-nginx)
- 3 control-plane nodes with stacked etcd
- 2 workers
- HAProxy in front of the API server (used as
controlPlaneEndpoint) - Calico CNI
- ingress-nginx, spread across workers and wired to ports 80/443
Not k3s, not kind. The same kubeadm layout you'd see in a real on-prem setup, which is exactly why I wanted it: I wanted to see how every piece fits together.
How it works
The script is split into stages, and all just runs them in order:
./k8s_install-ha.sh prep # SSH key, host packages, download Rocky 9 cloud image
./k8s_install-ha.sh vms # create 6 VMs with cloud-init
./k8s_install-ha.sh ip # collect VM IPs, save a state file
./k8s_install-ha.sh bastion # HAProxy + bootstrap script on the bastion
./k8s_install-ha.sh install # containerd, kubeadm init, Calico, joins
./k8s_install-ha.sh addons # ingress-nginx + HAProxy 80/443
./k8s_install-ha.sh verify # check everything
A few things I learned along the way that made it much less painful:
Cloud images + cloud-init instead of OS installs. Each VM boots straight from the Rocky Linux 9 GenericCloud image. Kernel modules, sysctl, swap-off and base packages are all done by cloud-init on first boot, so there's no installer to click through.
Finding VM IPs is surprisingly annoying. Depending on timing, the guest agent isn't ready yet, or the lease isn't visible. The script tries the guest agent first, then libvirt, then DHCP leases by MAC address, and retries for a few minutes before giving up.
Split it into stages and make each one safe to re-run. Existing VMs, installed packages and already-joined nodes are skipped. When something fails halfway, I just fix it and re-run that stage instead of starting from zero. This alone saved me hours.
Prepare nodes in parallel. Installing containerd and kubeadm on 5 nodes one by one is slow, so that part runs in parallel from the bastion.
The result
After about 10–20 minutes for the cluster part:
NAME STATUS ROLES VERSION
k8s-control1.k8s.local Ready control-plane v1.33.x
k8s-control2.k8s.local Ready control-plane v1.33.x
k8s-control3.k8s.local Ready control-plane v1.33.x
k8s-worker1.k8s.local Ready <none> v1.33.x
k8s-worker2.k8s.local Ready <none> v1.33.x
Being honest about the limits
- The bastion is a single point of failure. The control plane is HA, but HAProxy in front of it is one node. No keepalived/VIP yet.
-
The default spec is chunky: 30 vCPU / 88 GB RAM. Every node's CPU/RAM/disk is a variable at the top of the script, so you can shrink it for a smaller host, and the
prepstep compares the plan against your host before creating anything. - It's for homelab / learning, not production. Firewalld is off, SELinux is permissive, and SSH host key checks are skipped to keep the lab simple.
If you want to skip the setup
I cleaned it up with a full README and put it on Gumroad, in case it's useful for anyone studying for CKA or building a homelab that looks like real infrastructure:

Top comments (0)