DEV Community

Remus Kalathil
Remus Kalathil

Posted on

Running Argo CD across regions: four topologies, one reference setup

Argo CD on one cluster is easy. Add a second region and a new set of questions appears. Where does Argo CD itself run? What happens when that region goes down? How do you stop one bad commit from landing in every region at once?

This guide covers the topologies, the trade-offs, and a reference setup you can adapt, with Amazon EKS as the example.

First, what Argo CD actually needs

  • Git is the source of truth.
  • Kubernetes objects in the Argo CD cluster hold the rest: Application, ApplicationSet and AppProject objects, cluster Secrets and repo credentials.
  • Redis is only a cache, so losing it is not data loss.
  • It is not in the data path. If Argo CD is down, your workloads keep running. You lose the ability to deploy and to self-heal until it is back.

That last point shapes everything below. Multi-region design for Argo CD is about blast radius and recovery time, not user traffic.

Four topologies

Topology Good for Watch out for
1. Central hub One pane of glass, up to a few dozen clusters, small team Cross-region API latency; the hub's region is a single point of failure for changes; wide credentials
2. Hub per region Failure isolation, data-residency or compliance boundaries, 50+ clusters More Argo CD instances to run; fleet view needs extra tooling
3. One per cluster Strongest isolation, air-gapped or regulated clusters N control planes to upgrade; you need a bootstrap story
4. Agents (pull) Clusters that can't be reached inbound Newer model (the Argo CD agent project); check its maturity for your version

Choose with four questions:

  1. How many clusters do you have?
  2. Can your compliance rules allow one control plane to hold credentials for all regions?
  3. How long can you tolerate being unable to deploy?
  4. Do teams need to own their own Argo CD?

My default for multi-region production is a hub per region (or per region group), driven from one Git repo, because it keeps a regional failure regional. That is an opinion, not a rule. A central hub is fine below a few dozen clusters if the hub region has a standby.

Building blocks

1. Clusters as labelled Secrets

Labels carry the metadata that everything else selects on: environment, region, and a rollout wave.

apiVersion: v1
kind: Secret
metadata:
  name: cluster-prod-uswest-2-01
  namespace: argocd
  labels:
    argocd.argoproj.io/secret-type: cluster
    env: prod
    region: us-west-2
    wave: "1"
type: Opaque
stringData:
  name: prod-us-west-2-01
  server: https://EXAMPLE1234567890.gr7.us-west-2.eks.amazonaws.com
  config: |
    {
      "awsAuthConfig": {
        "clusterName": "prod-us-west-2-01",
        "roleARN": "arn:aws:iam::111122223333:role/argocd-deployer"
      },
      "tlsClientConfig": { "caData": "BASE64_CLUSTER_CA" }
    }
Enter fullscreen mode Exit fullscreen mode

2. EKS authentication with no long-lived tokens

  • Give the Argo CD controller an AWS identity (IRSA or EKS Pod Identity).
  • Allow it to sts:AssumeRole into the roleARN above (one role per account).
  • Map that role in the target cluster with an EKS access entry to a Kubernetes role. Scope that role down where you can, instead of defaulting to cluster-admin.

Tokens come from awsAuthConfig at call time, so there is nothing to rotate by hand.

3. One repo layout, three layers of values

Don't copy a directory per region. Layer the differences:

addons/                          # the chart or kustomize base
values/global.yaml               # shared by everything
values/env/prod.yaml             # prod-only
values/region/us-west-2.yaml     # region-only (endpoints, quotas, replicas)
Enter fullscreen mode Exit fullscreen mode

4. A project to bound the blast radius

apiVersion: argoproj.io/v1alpha1
kind: AppProject
metadata: { name: platform, namespace: argocd }
spec:
  sourceRepos: ["https://github.com/example/platform-gitops"]
  destinations:
    - { server: "*", namespace: platform }
    - { server: "*", namespace: monitoring }
  clusterResourceWhitelist:
    - { group: "", kind: Namespace }
  orphanedResources: { warn: true }
Enter fullscreen mode Exit fullscreen mode

5. One ApplicationSet, one Application per cluster

The cluster generator reads the labels from step 1.

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata: { name: platform-addons, namespace: argocd }
spec:
  goTemplate: true
  goTemplateOptions: ["missingkey=error"]
  generators:
    - clusters:
        selector: { matchLabels: { env: prod } }
  strategy:
    type: RollingSync
    rollingSync:
      steps:
        - matchExpressions: [{ key: wave, operator: In, values: ["1"] }]
        - matchExpressions: [{ key: wave, operator: In, values: ["2"] }]
          maxUpdate: 50%
        - matchExpressions: [{ key: wave, operator: In, values: ["3"] }]
  template:
    metadata:
      name: "addons-{{.name}}"
      labels:
        wave: "{{.metadata.labels.wave}}"
        region: "{{.metadata.labels.region}}"
    spec:
      project: platform
      source:
        repoURL: https://github.com/example/platform-gitops
        targetRevision: stable
        path: addons
        helm:
          valueFiles:
            - values/global.yaml
            - "values/env/{{.metadata.labels.env}}.yaml"
            - "values/region/{{.metadata.labels.region}}.yaml"
      destination: { server: "{{.server}}", namespace: platform }
      syncPolicy: { syncOptions: [CreateNamespace=true] }
Enter fullscreen mode Exit fullscreen mode

If you run a hub per region, each hub gets the same ApplicationSet with an extra region selector in the generator, so every hub only manages its own clusters.

Rolling out region by region

  • Waves, not region names. The wave label orders the rollout (a canary region first, then the rest). You can rename or reorder regions without touching the ApplicationSet.
  • Progressive Syncs (RollingSync) let the ApplicationSet controller sync wave 1, wait for health, then sync wave 2. It sits behind a feature flag on the ApplicationSet controller in current versions. I left auto-sync out of the example on purpose, so check your version's docs on how the two interact.
  • Don't let the fleet track main. Point targetRevision at a branch or tag that moves only after a successful earlier wave (stable above). One merge to main should never reach every region.

High availability and disaster recovery

  • HA install. The HA manifests give you several API server and repo-server replicas and a Redis HA setup.
  • Everything that matters is declarative. Applications, projects and cluster Secrets live in Git (secrets via an external secrets operator, not in plain Git). A lost Argo CD can be rebuilt from Git.
  • Active/passive for a central hub. Run a second Argo CD in another region with the controllers scaled to zero. Two active controllers will fight over the same clusters. To fail over, scale the standby's controller up.
  • Per-region hubs make this mostly moot. Losing a region's Argo CD only pauses changes for that region.

Scaling knobs

  • Controller sharding. The application controller is a StatefulSet. Set its replica count and let it spread clusters across shards. Newer versions offer a round-robin and a consistent-hashing algorithm, via controller.sharding.algorithm in argocd-cmd-params-cm.
  • Cross-region API latency. The controller watches every managed cluster, so distance and jitter show up as timeouts. This is a main reason to put the controller near its clusters.
  • Polling vs webhooks. The default Git poll is about three minutes (timeout.reconciliation). Use webhooks, and put your repo-server behind a cache that survives restarts.
  • Large clusters. Use resource exclusions and ignoreDifferences (for example HPA-managed replicas) to cut noise and controller memory.

Security

  • Use AppProjects as the permission boundary: allowed repos, destinations and cluster-scoped kinds.
  • Use SSO via OIDC with group claims mapped in argocd-rbac-cm. Debug it in layers: is the claim in the token, does the IdP's group filter match, and only then RBAC.
  • Keep RBAC config in Git, so hotfixes don't get silently reverted.
  • Keep secrets out of Git and out of Applications.

Failure modes cheat sheet

Symptom Likely cause
Apps flip to "Unknown" or "Unreachable" for one region API-server reachability or latency from the controller, expired exec-auth credentials
Hub outage Nothing breaks at runtime; you can't deploy until it's back or you fail over
Everything resyncs after a restart Expected after losing the controller cache; plan for the load
One bad commit hits all regions The fleet tracks main; add promotion and waves
"permission denied" for users after SSO Missing group claim or a filter mismatch, before RBAC

Checklist

  1. Pick a topology from the four questions.
  2. Register clusters as labelled Secrets, with cross-account roles and no static tokens.
  3. One repo, layered values, one ApplicationSet, waves for ordering.
  4. The fleet tracks a promoted branch, never main.
  5. HA install, a rebuild-from-Git runbook, and an active/passive standby if you run a central hub.
  6. Shard the controller, use webhooks, and tune exclusions.
  7. Projects and RBAC in Git, SSO tested in layers.

How are you running Argo CD across regions? Hub per region, or something else? Let me know in the comments.

Top comments (0)