Most multi-cluster HA setups run active/active, which means paying for full workloads on both clusters around the clock. For our production workloads, I wanted automatic disaster recovery without the standby cluster burning money every day.
So I built an active/passive multi-cluster setup with Karmada 👇
🔹 𝗢𝗻𝗲 𝗰𝗼𝗻𝘁𝗿𝗼𝗹 𝗽𝗹𝗮𝗻𝗲, 𝘁𝘄𝗼 𝗞𝘂𝗯𝗲𝗿𝗻𝗲𝘁𝗲𝘀 𝗰𝗹𝘂𝘀𝘁𝗲𝗿𝘀 Karmada runs on a lightweight K3s host and manages both member clusters. Deployments, Services, Ingresses, ConfigMaps, and Secrets are propagated using a single PropagationPolicy.
🔹 𝗪𝗼𝗿𝗸𝗹𝗼𝗮𝗱𝘀 𝗿𝘂𝗻 𝗼𝗻 𝘁𝗵𝗲 𝗽𝗿𝗶𝗺𝗮𝗿𝘆 𝗰𝗹𝘂𝘀𝘁𝗲𝗿 Using ClusterAffinities, I defined an ordered failover chain. The secondary cluster remains on standby until the primary actually becomes unavailable.
🔹 𝗔𝘂𝘁𝗼𝗺𝗮𝘁𝗶𝗰 𝗳𝗮𝗶𝗹𝗼𝘃𝗲𝗿 When the primary cluster becomes NotReady, a ClusterTaintPolicy applies a NoExecute taint. Karmada then moves the workloads to the backup cluster automatically — without manual intervention.
🔹 𝗔𝘂𝘁𝗼𝗺𝗮𝘁𝗶𝗰 𝗳𝗮𝗶𝗹-𝗯𝗮𝗰𝗸 This was one of the interesting parts. Karmada doesnt automatically move workloads back to the primary cluster after recovery. Once workloads have failed over, they can remain on the secondary cluster and continue consuming resources even after the primary has recovered. To solve this, I built a small CronJob-based fail-back mechanism. It: → Waits for the primary cluster to remain healthy for 15 minutes → Clears the schedulers affinity pin → Triggers a reschedule → Moves workloads back to the primary cluster This allows the secondary cluster to return to its standby role automatically.
🔹 𝗖𝗹𝗼𝘂𝗱𝗳𝗹𝗮𝗿𝗲 𝗟𝗼𝗮𝗱 𝗕𝗮𝗹𝗮𝗻𝗰𝗲𝗿 𝗳𝗼𝗿 𝘁𝗿𝗮𝗳𝗳𝗶𝗰 𝗳𝗮𝗶𝗹𝗼𝘃𝗲𝗿 I initially used a self-hosted HAProxy setup for traffic management. Later, I moved this responsibility to Cloudflare Load Balancer. The configuration uses: → Primary and backup pools → HTTPS /healthz health monitoring → Automatic traffic failover → Edge-level traffic management This removes the need to maintain another HAProxy layer just for external traffic failover.
𝗪𝗵𝗮𝘁 𝗜 𝘁𝗼𝗼𝗸 𝗮𝘄𝗮𝘆
𝗛𝗶𝗴𝗵 𝗮𝘃𝗮𝗶𝗹𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗶𝘀𝗻’𝘁 𝗷𝘂𝘀𝘁 𝗮𝗻 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲 𝗱𝗲𝗰𝗶𝘀𝗶𝗼𝗻 — 𝗶𝘁’𝘀 𝗮𝗹𝘀𝗼 𝗮 𝗰𝗼𝘀𝘁 𝗱𝗲𝗰𝗶𝘀𝗶𝗼𝗻.
For workloads that don't require active/active capacity, an active/passive architecture with automated failover and fail-back can provide disaster recovery while avoiding the cost of continuously running full production capacity on both clusters.
The full setup is documented step by step (installation, cluster join, failover hardening, the fail-back script, and both load balancer options):
🔗 https://github.com/noushadhasan/Karmada-Kubeadm-Cluster-Management.git

Top comments (0)