Our on-prem lab cluster was a typical kubeadm build on VMware: one control-plane node, no backups, local accounts on every host, and clocks drifting by up to six minutes. It worked until it didn’t. Over a few weeks I rebuilt it into a production-style platform that a whole team can use safely, on plain VMware VMs with no cloud load balancers.
Every step is scripted and published here: github.com/Willey2003/k8s-ha-ad-gateway-lab.
What was built
| Area | Outcome |
|---|---|
| High availability | 3 control-plane nodes behind a floating API VIP (kube-vip, ARP mode). The running cluster was migrated from a single-node endpoint to the VIP, with automatic rollback. Measured failover: about 1 second. |
| External etcd | 3-member TLS etcd cluster with weekly snapshots from the bastion. |
| Backups and DR | Velero (including volume data) to a self-hosted S3 store, plus etcd snapshots, on a weekly systemd timer. A scripted restore drill proves the backups actually restore. |
| Time and DNS | All hosts sync time from the domain controllers; A/PTR records and an *.apps wildcard. |
| Active Directory for Linux | Every node joined to AD (realmd/SSSD), AD-group-based SSH and sudo, Kerberos SSO between nodes. |
| Active Directory for kubectl | LDAPS on the domain controllers, Dex as the OIDC provider, and the API server’s structured authentication config (CEL claim mappings). RBAC comes from AD groups. |
| Self-service | Every AD user can read the cluster and deploy into a quota-limited playground namespace (Pod Security baseline, no NodePorts or LoadBalancers). |
| Ingress | Ingress-NGINX is retired, so: MetalLB (L2) plus HAProxy Unified Gateway with Gateway API. One shared Gateway, wildcard TLS, and users publish apps with a single HTTPRoute. |
Stack: Kubernetes 1.32 (kubeadm), CRI-O, Calico, kube-vip, etcd 3.5, Velero 1.18, Dex 2.45, MetalLB 0.16, HAProxy Unified Gateway 1.0, Gateway API 1.3, Active Directory (SSSD, Kerberos, LDAPS), chrony.
Access model
| Role | Who | SSH | Kubernetes |
|---|---|---|---|
| Lab user | Every enabled AD account | Bastion only, no sudo |
view cluster-wide (no Secrets) and edit in playground
|
| Lab administrator | AD group K8s-Admins
|
All hosts with sudo | cluster-admin |
| Emergency | Local accounts | All hosts |
admin.conf on the control plane |
For a user it is simple: SSH to the bastion with an AD account, and kubectl is already configured. The first kubectl command asks for the AD password once, and the token refreshes for up to seven days.
Run order
-
01-access: a least-privilege deployer ServiceAccount and kubeconfig. -
02-ha-control-plane: prepare the new managers, migrate to the VIP in three steps, test failover, roll back if needed. -
03-time-sync: chrony against the domain controllers. -
04-backups: backup disk, S3 gateway, Velero, weekly timer and the restore drill. -
05-dns: PowerShell for A and PTR records, with conflict detection and-WhatIf. -
06-active-directory: DNS through the DCs, realm join, Kerberos SSH, then Dex, the API-server auth config and RBAC for kubectl. -
07-gateway: MetalLB, HAProxy Unified Gateway, the shared Gateway and a demo app with route RBAC.
Lessons learned: the things that actually bit
- Check the router’s subnet, not just the hosts’. Every VM used a /24, but the gateway routed only a /27. Load-balancer IPs above .31 worked from inside the subnet and silently timed out from everywhere else.
- Validate API-server auth config before a rolling restart.
claims.groups.map(...)fails CEL type-checking; it needsdyn(claims.groups). A localkube-apiserverbinary reproduces the error in seconds, before two managers crash-loop. - A health check must prove the new container is healthy. Waiting for
/readyzright after editing a static pod manifest can hit the old process. Wait for a new container ID, then require sustained health. - HAProxy sizes its memory budget from the container memory limit. Set the limit too low and TLS workers refuse to start with the default
maxconn. - Helm post-install CRD jobs can race the controller. Restart the controller once after the first install.
- Dex expands
$VARin config values. A$in the bind password silently broke LDAP; setDEX_EXPAND_ENV=false. -
realm joinwith a short hostname registers only short SPNs, so Kerberos SSO fails with “Server not found in Kerberos database”. Fix/etc/hostsand registerHOST/<fqdn>. - Never make a repair tool depend on the thing it repairs. The Dex password rotation writes through the control plane’s local admin config, not through Dex-based kubectl.
- Prove the restore, not the backup. My first restore drill “passed” for the wrong reason: the proof string lived in the pod spec. Writing it only into the volume made the test honest.
Get it
All IPs, hostnames and the domain in the repository are example values, and no secrets are stored: keys and passwords are generated at run time. Read the scripts before you run them, and run them in a lab first: github.com/Willey2003/k8s-ha-ad-gateway-lab.
Top comments (0)