<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Remus Kalathil</title>
    <description>The latest articles on DEV Community by Remus Kalathil (@remus_kalathil_235e438778).</description>
    <link>https://dev.to/remus_kalathil_235e438778</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3821453%2F83501349-2eab-4dc0-bcb3-ad5304284c28.png</url>
      <title>DEV Community: Remus Kalathil</title>
      <link>https://dev.to/remus_kalathil_235e438778</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/remus_kalathil_235e438778"/>
    <language>en</language>
    <item>
      <title>Running Argo CD across regions: four topologies, one reference setup</title>
      <dc:creator>Remus Kalathil</dc:creator>
      <pubDate>Wed, 07 Oct 2026 23:33:49 +0000</pubDate>
      <link>https://dev.to/remus_kalathil_235e438778/running-argo-cd-across-regions-four-topologies-one-reference-setup-23op</link>
      <guid>https://dev.to/remus_kalathil_235e438778/running-argo-cd-across-regions-four-topologies-one-reference-setup-23op</guid>
      <description>&lt;p&gt;Argo CD on one cluster is easy. Add a second region and a new set of questions appears. Where does Argo CD itself run? What happens when that region goes down? How do you stop one bad commit from landing in every region at once?&lt;/p&gt;

&lt;p&gt;This guide covers the topologies, the trade-offs, and a reference setup you can adapt, with Amazon EKS as the example.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyzlj3mob08ng9yfy2xv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyzlj3mob08ng9yfy2xv.png" alt=" " width="800" height="418"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  First, what Argo CD actually needs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Git is the source of truth.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kubernetes objects in the Argo CD cluster hold the rest:&lt;/strong&gt; Application, ApplicationSet and AppProject objects, cluster Secrets and repo credentials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redis is only a cache,&lt;/strong&gt; so losing it is not data loss.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is not in the data path.&lt;/strong&gt; If Argo CD is down, your workloads keep running. You lose the ability to deploy and to self-heal until it is back.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point shapes everything below. Multi-region design for Argo CD is about blast radius and recovery time, not user traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four topologies
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Topology&lt;/th&gt;
&lt;th&gt;Good for&lt;/th&gt;
&lt;th&gt;Watch out for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Central hub&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One pane of glass, up to a few dozen clusters, small team&lt;/td&gt;
&lt;td&gt;Cross-region API latency; the hub's region is a single point of failure for changes; wide credentials&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Hub per region&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Failure isolation, data-residency or compliance boundaries, 50+ clusters&lt;/td&gt;
&lt;td&gt;More Argo CD instances to run; fleet view needs extra tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. One per cluster&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Strongest isolation, air-gapped or regulated clusters&lt;/td&gt;
&lt;td&gt;N control planes to upgrade; you need a bootstrap story&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4. Agents (pull)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Clusters that can't be reached inbound&lt;/td&gt;
&lt;td&gt;Newer model (the Argo CD agent project); check its maturity for your version&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Choose with four questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How many clusters do you have?&lt;/li&gt;
&lt;li&gt;Can your compliance rules allow one control plane to hold credentials for all regions?&lt;/li&gt;
&lt;li&gt;How long can you tolerate being unable to deploy?&lt;/li&gt;
&lt;li&gt;Do teams need to own their own Argo CD?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;My default for multi-region production is a hub per region (or per region group), driven from one Git repo, because it keeps a regional failure regional. That is an opinion, not a rule. A central hub is fine below a few dozen clusters if the hub region has a standby.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building blocks
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Clusters as labelled Secrets
&lt;/h3&gt;

&lt;p&gt;Labels carry the metadata that everything else selects on: environment, region, and a rollout wave.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Secret&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cluster-prod-uswest-2-01&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argocd&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;argocd.argoproj.io/secret-type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cluster&lt;/span&gt;
    &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;prod&lt;/span&gt;
    &lt;span class="na"&gt;region&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;us-west-2&lt;/span&gt;
    &lt;span class="na"&gt;wave&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1"&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Opaque&lt;/span&gt;
&lt;span class="na"&gt;stringData&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;prod-us-west-2-01&lt;/span&gt;
  &lt;span class="na"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://EXAMPLE1234567890.gr7.us-west-2.eks.amazonaws.com&lt;/span&gt;
  &lt;span class="na"&gt;config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;{&lt;/span&gt;
      &lt;span class="s"&gt;"awsAuthConfig": {&lt;/span&gt;
        &lt;span class="s"&gt;"clusterName": "prod-us-west-2-01",&lt;/span&gt;
        &lt;span class="s"&gt;"roleARN": "arn:aws:iam::111122223333:role/argocd-deployer"&lt;/span&gt;
      &lt;span class="s"&gt;},&lt;/span&gt;
      &lt;span class="s"&gt;"tlsClientConfig": { "caData": "BASE64_CLUSTER_CA" }&lt;/span&gt;
    &lt;span class="s"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. EKS authentication with no long-lived tokens
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Give the Argo CD controller an AWS identity (IRSA or EKS Pod Identity).&lt;/li&gt;
&lt;li&gt;Allow it to &lt;code&gt;sts:AssumeRole&lt;/code&gt; into the &lt;code&gt;roleARN&lt;/code&gt; above (one role per account).&lt;/li&gt;
&lt;li&gt;Map that role in the target cluster with an EKS access entry to a Kubernetes role. Scope that role down where you can, instead of defaulting to cluster-admin.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tokens come from &lt;code&gt;awsAuthConfig&lt;/code&gt; at call time, so there is nothing to rotate by hand.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. One repo layout, three layers of values
&lt;/h3&gt;

&lt;p&gt;Don't copy a directory per region. Layer the differences:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;addons/                          # the chart or kustomize base
values/global.yaml               # shared by everything
values/env/prod.yaml             # prod-only
values/region/us-west-2.yaml     # region-only (endpoints, quotas, replicas)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. A project to bound the blast radius
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argoproj.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AppProject&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;platform&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;argocd&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;sourceRepos&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://github.com/example/platform-gitops"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;destinations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;platform&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;monitoring&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;clusterResourceWhitelist&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;group&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;Namespace&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;orphanedResources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;warn&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;true&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  5. One ApplicationSet, one Application per cluster
&lt;/h3&gt;

&lt;p&gt;The cluster generator reads the labels from step 1.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argoproj.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ApplicationSet&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;platform-addons&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;argocd&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;goTemplate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;goTemplateOptions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;missingkey=error"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;generators&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;clusters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;selector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;prod&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RollingSync&lt;/span&gt;
    &lt;span class="na"&gt;rollingSync&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[{&lt;/span&gt; &lt;span class="nv"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;wave&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;In&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt; &lt;span class="pi"&gt;}]&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[{&lt;/span&gt; &lt;span class="nv"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;wave&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;In&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt; &lt;span class="pi"&gt;}]&lt;/span&gt;
          &lt;span class="na"&gt;maxUpdate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;50%&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[{&lt;/span&gt; &lt;span class="nv"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;wave&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;In&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt; &lt;span class="pi"&gt;}]&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;addons-{{.name}}"&lt;/span&gt;
      &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;wave&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{.metadata.labels.wave}}"&lt;/span&gt;
        &lt;span class="na"&gt;region&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{.metadata.labels.region}}"&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;project&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;platform&lt;/span&gt;
      &lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;repoURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://github.com/example/platform-gitops&lt;/span&gt;
        &lt;span class="na"&gt;targetRevision&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;stable&lt;/span&gt;
        &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;addons&lt;/span&gt;
        &lt;span class="na"&gt;helm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;valueFiles&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;values/global.yaml&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;values/env/{{.metadata.labels.env}}.yaml"&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;values/region/{{.metadata.labels.region}}.yaml"&lt;/span&gt;
      &lt;span class="na"&gt;destination&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{.server}}"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;platform&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;syncPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;syncOptions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;CreateNamespace=true&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you run a hub per region, each hub gets the same ApplicationSet with an extra &lt;code&gt;region&lt;/code&gt; selector in the generator, so every hub only manages its own clusters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rolling out region by region
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Waves, not region names.&lt;/strong&gt; The &lt;code&gt;wave&lt;/code&gt; label orders the rollout (a canary region first, then the rest). You can rename or reorder regions without touching the ApplicationSet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Progressive Syncs (RollingSync)&lt;/strong&gt; let the ApplicationSet controller sync wave 1, wait for health, then sync wave 2. It sits behind a feature flag on the ApplicationSet controller in current versions. I left auto-sync out of the example on purpose, so check your version's docs on how the two interact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't let the fleet track &lt;code&gt;main&lt;/code&gt;.&lt;/strong&gt; Point &lt;code&gt;targetRevision&lt;/code&gt; at a branch or tag that moves only after a successful earlier wave (&lt;code&gt;stable&lt;/code&gt; above). One merge to &lt;code&gt;main&lt;/code&gt; should never reach every region.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  High availability and disaster recovery
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;HA install.&lt;/strong&gt; The HA manifests give you several API server and repo-server replicas and a Redis HA setup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything that matters is declarative.&lt;/strong&gt; Applications, projects and cluster Secrets live in Git (secrets via an external secrets operator, not in plain Git). A lost Argo CD can be rebuilt from Git.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Active/passive for a central hub.&lt;/strong&gt; Run a second Argo CD in another region with the controllers scaled to zero. Two active controllers will fight over the same clusters. To fail over, scale the standby's controller up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-region hubs make this mostly moot.&lt;/strong&gt; Losing a region's Argo CD only pauses changes for that region.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Scaling knobs
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Controller sharding.&lt;/strong&gt; The application controller is a StatefulSet. Set its replica count and let it spread clusters across shards. Newer versions offer a round-robin and a consistent-hashing algorithm, via &lt;code&gt;controller.sharding.algorithm&lt;/code&gt; in &lt;code&gt;argocd-cmd-params-cm&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-region API latency.&lt;/strong&gt; The controller watches every managed cluster, so distance and jitter show up as timeouts. This is a main reason to put the controller near its clusters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Polling vs webhooks.&lt;/strong&gt; The default Git poll is about three minutes (&lt;code&gt;timeout.reconciliation&lt;/code&gt;). Use webhooks, and put your repo-server behind a cache that survives restarts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Large clusters.&lt;/strong&gt; Use resource exclusions and &lt;code&gt;ignoreDifferences&lt;/code&gt; (for example HPA-managed replicas) to cut noise and controller memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Security
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Use AppProjects as the permission boundary: allowed repos, destinations and cluster-scoped kinds.&lt;/li&gt;
&lt;li&gt;Use SSO via OIDC with group claims mapped in &lt;code&gt;argocd-rbac-cm&lt;/code&gt;. Debug it in layers: is the claim in the token, does the IdP's group filter match, and only then RBAC.&lt;/li&gt;
&lt;li&gt;Keep RBAC config in Git, so hotfixes don't get silently reverted.&lt;/li&gt;
&lt;li&gt;Keep secrets out of Git and out of Applications.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Failure modes cheat sheet
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Likely cause&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Apps flip to "Unknown" or "Unreachable" for one region&lt;/td&gt;
&lt;td&gt;API-server reachability or latency from the controller, expired exec-auth credentials&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hub outage&lt;/td&gt;
&lt;td&gt;Nothing breaks at runtime; you can't deploy until it's back or you fail over&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Everything resyncs after a restart&lt;/td&gt;
&lt;td&gt;Expected after losing the controller cache; plan for the load&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One bad commit hits all regions&lt;/td&gt;
&lt;td&gt;The fleet tracks &lt;code&gt;main&lt;/code&gt;; add promotion and waves&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"permission denied" for users after SSO&lt;/td&gt;
&lt;td&gt;Missing group claim or a filter mismatch, before RBAC&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Checklist
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Pick a topology from the four questions.&lt;/li&gt;
&lt;li&gt;Register clusters as labelled Secrets, with cross-account roles and no static tokens.&lt;/li&gt;
&lt;li&gt;One repo, layered values, one ApplicationSet, waves for ordering.&lt;/li&gt;
&lt;li&gt;The fleet tracks a promoted branch, never &lt;code&gt;main&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;HA install, a rebuild-from-Git runbook, and an active/passive standby if you run a central hub.&lt;/li&gt;
&lt;li&gt;Shard the controller, use webhooks, and tune exclusions.&lt;/li&gt;
&lt;li&gt;Projects and RBAC in Git, SSO tested in layers.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;How are you running Argo CD across regions? Hub per region, or something else? Let me know in the comments.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>argocd</category>
      <category>gitops</category>
      <category>aws</category>
    </item>
    <item>
      <title>EKS Pod Identity vs IRSA: A Practical Migration Guide and the Gotchas Nobody Mentions</title>
      <dc:creator>Remus Kalathil</dc:creator>
      <pubDate>Mon, 28 Sep 2026 00:54:03 +0000</pubDate>
      <link>https://dev.to/remus_kalathil_235e438778/eks-pod-identity-vs-irsa-a-practical-migration-guide-and-the-gotchas-nobody-mentions-4ej9</link>
      <guid>https://dev.to/remus_kalathil_235e438778/eks-pod-identity-vs-irsa-a-practical-migration-guide-and-the-gotchas-nobody-mentions-4ej9</guid>
      <description>&lt;p&gt;If you're running workloads on EKS, your IAM strategy was probably built around IRSA. At re:Invent 2023, AWS introduced EKS Pod Identity as a simpler alternative to it. This post walks through the practical migration path, and, more importantly, five gotchas I hit that aren't covered in the official docs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F657vwebpso2ddg1x6z54.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F657vwebpso2ddg1x6z54.png" alt="IRSA vs EKS Pod Identity architecture" width="800" height="550"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick context (skip if you already know this)
&lt;/h2&gt;

&lt;p&gt;IRSA relies on OIDC federation. You create a per-cluster identity provider, write trust policies with sub/aud conditions, and annotate the service account with the target role ARN. At runtime, a mutating webhook injects a projected OIDC token into the pod, and the AWS SDK exchanges it for temporary credentials via &lt;code&gt;sts:AssumeRoleWithWebIdentity&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Pod Identity replaces this flow with an EKS Pod Identity Agent daemonset running on every node, plus a &lt;code&gt;PodIdentityAssociation&lt;/code&gt; resource that maps a (cluster, namespace, service account) triple to an IAM role. Trust is established through the &lt;code&gt;pods.eks.amazonaws.com&lt;/code&gt; service principal instead of per-cluster OIDC conditions. No OIDC provider to create, rotate, or track, and no per-role trust-policy sprawl.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why migrate (the real wins)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No per-cluster OIDC identity provider lifecycle to manage.&lt;/strong&gt; One less resource to create, rotate certificates for, and keep in sync across clusters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uniform trust policy shape across every role.&lt;/strong&gt; Every role trusts the same &lt;code&gt;pods.eks.amazonaws.com&lt;/code&gt; principal, so auditing and policy-as-code checks get simpler. You're not diffing sub/aud conditions role by role.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Easier cross-account and cross-cluster role reuse.&lt;/strong&gt; One association model instead of juggling OIDC provider ARNs per cluster.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Faster credential rotation.&lt;/strong&gt; The agent proactively fetches and rotates credentials, versus the webhook-injected token flow which ties rotation to token TTL.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Migration steps
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fygepd4z62py7hjqz78br.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fygepd4z62py7hjqz78br.png" alt="EKS Pod Identity migration steps" width="800" height="702"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Enable the EKS Pod Identity Agent add-on&lt;/strong&gt; on the cluster (via console, &lt;code&gt;eksctl&lt;/code&gt;, or Terraform's &lt;code&gt;aws_eks_addon&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Update role trust policies&lt;/strong&gt; to include the Pod Identity principal. Important: this can coexist with the existing OIDC trust condition during migration. You don't need a hard cutover; a role can trust both &lt;code&gt;pods.eks.amazonaws.com&lt;/code&gt; and the old OIDC provider simultaneously.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create a &lt;code&gt;PodIdentityAssociation&lt;/code&gt;&lt;/strong&gt; per (cluster, namespace, service account) triple, mapping it to the target IAM role ARN.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify empirically.&lt;/strong&gt; Don't trust that the association exists in the console. Check CloudTrail's &lt;code&gt;userIdentity&lt;/code&gt; field or application logs to confirm the pod is actually assuming the new role via the new path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drop the old &lt;code&gt;eks.amazonaws.com/role-arn&lt;/code&gt; annotation&lt;/strong&gt; only after verification passes for that workload.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Recommendation: run both trust paths in parallel for at least one deploy cycle per workload. It costs nothing and gives you an instant rollback if something's wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reference: trust policy and association
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Trust policy — before (IRSA only):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2012-10-17"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Statement"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Effect"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Principal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"Federated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:iam::111122223333:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sts:AssumeRoleWithWebIdentity"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Condition"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"StringEquals"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:sub"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"system:serviceaccount:my-namespace:my-service-account"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:aud"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sts.amazonaws.com"&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Trust policy — during migration (both trusts active):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2012-10-17"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Statement"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Effect"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Principal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"Federated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:iam::111122223333:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sts:AssumeRoleWithWebIdentity"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Condition"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"StringEquals"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:sub"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"system:serviceaccount:my-namespace:my-service-account"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:aud"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sts.amazonaws.com"&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Effect"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Principal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"Service"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pods.eks.amazonaws.com"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"sts:AssumeRole"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"sts:TagSession"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Trust policy — after (Pod Identity only, drop the OIDC statement once verified):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2012-10-17"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Statement"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Effect"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Principal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"Service"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pods.eks.amazonaws.com"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"sts:AssumeRole"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"sts:TagSession"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Creating the association — AWS CLI:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws eks create-pod-identity-association &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cluster-name&lt;/span&gt; my-cluster &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; my-namespace &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--service-account&lt;/span&gt; my-service-account &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--role-arn&lt;/span&gt; arn:aws:iam::111122223333:role/my-app-role
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Creating the association — Terraform:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_eks_pod_identity_association"&lt;/span&gt; &lt;span class="s2"&gt;"my_app"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;cluster_name&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"my-cluster"&lt;/span&gt;
  &lt;span class="nx"&gt;namespace&lt;/span&gt;       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"my-namespace"&lt;/span&gt;
  &lt;span class="nx"&gt;service_account&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"my-service-account"&lt;/span&gt;
  &lt;span class="nx"&gt;role_arn&lt;/span&gt;        &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_iam_role&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;my_app_role&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Verifying via CloudTrail (gotcha #4 check):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws cloudtrail lookup-events &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--lookup-attributes&lt;/span&gt; &lt;span class="nv"&gt;AttributeKey&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;EventName,AttributeValue&lt;span class="o"&gt;=&lt;/span&gt;AssumeRole &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"Events[?contains(CloudTrailEvent, 'pods.eks.amazonaws.com')]"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-results&lt;/span&gt; 20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The gotchas nobody mentions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Native subprocess libraries don't inherit Pod Identity credentials automatically
&lt;/h3&gt;

&lt;p&gt;If your application forks its own binary (a native library, a CLI tool it shells out to, anything that doesn't go through your language's AWS SDK credential chain), that subprocess resolves its own default credential chain. Without explicit environment propagation, it can fall through to the node's IMDS role instead of the Pod Identity credentials.&lt;/p&gt;

&lt;p&gt;This fails &lt;strong&gt;silently&lt;/strong&gt;. No error, no warning log. The subprocess just authenticates as the wrong principal. You only catch it by explicitly checking which identity made a given API call: &lt;code&gt;aws cloudtrail lookup-events&lt;/code&gt; filtered on &lt;code&gt;userIdentity.arn&lt;/code&gt;, or equivalent logging in your observability stack. If you have any workload that shells out to a native binary for AWS calls (some data-processing tools, some legacy CLIs), audit this explicitly before you assume the migration is clean.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. One (cluster, namespace, service account) triple = exactly one association
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;PodIdentityAssociation&lt;/code&gt; enforces a strict 1:1 mapping. Try to create a second association for the same triple without deleting the first, and the API returns &lt;code&gt;ResourceInUseException&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is a real trap for IaC-managed associations. A Terraform or Crossplane "destroy and recreate" apply (the kind that happens when you change an immutable field, or when a module gets refactored) will hit this if the delete and create aren't sequenced correctly. Plan for an explicit delete-then-create step in your pipeline, not a blind &lt;code&gt;replace&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. IaC provider version drift leaves orphaned associations
&lt;/h3&gt;

&lt;p&gt;If you migrate between provider versions, for example moving from an older Crossplane AWS provider to the newer Upbound-maintained one, associations created under the old provider version can be left behind, unmanaged, after the upgrade. They don't get cleaned up automatically, and they don't show up as drift in the new provider's state.&lt;/p&gt;

&lt;p&gt;Before any cleanup pass, check for a tag identifying which provider or provider version created the association (most providers tag their managed resources). Don't assume "not in current Terraform state" means "safe to delete." Check the tag first.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Startup-time credential propagation race
&lt;/h3&gt;

&lt;p&gt;A freshly created &lt;code&gt;PodIdentityAssociation&lt;/code&gt;'s credentials aren't always instantly available to a pod that starts immediately after. If your application reads AWS credentials very early in its boot sequence, before the Pod Identity Agent has had a moment to establish the mapping, you can see intermittent failures on first start that resolve on restart.&lt;/p&gt;

&lt;p&gt;This looks exactly like flakiness, and it's easy to misdiagnose as an unrelated bug (a race in your own init code, a transient network blip). If you see credential failures that only happen on cold start and never on restart, this is the first thing to check. Mitigate with a startup retry/backoff around your first AWS call rather than chasing a phantom bug in application code.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Compliance and security tooling built for IRSA won't recognize Pod Identity's trust shape
&lt;/h3&gt;

&lt;p&gt;Any scanner or policy-as-code check that looks for OIDC provider ARN conditions in trust policies (a common IRSA-era check: "does this role's trust policy scope to a specific OIDC provider and sub condition?") won't recognize the &lt;code&gt;pods.eks.amazonaws.com&lt;/code&gt; principal pattern at all.&lt;/p&gt;

&lt;p&gt;Until that tooling is updated, it will misreport actively-used, correctly-scoped roles as "unused" or "unscoped by OIDC condition." Depending on how your organization handles those findings, this can generate false-positive remediation tickets, or worse, feed an automated cleanup process that deletes roles still in active use. Audit your IAM scanning rules for OIDC-specific logic before rolling Pod Identity out broadly, not after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision checklist
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Migrate now if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're managing OIDC providers across many clusters and the operational overhead is real.&lt;/li&gt;
&lt;li&gt;Your compliance tooling can be updated in the same timeframe as the migration (or already supports Pod Identity).&lt;/li&gt;
&lt;li&gt;You can tolerate a parallel-run verification window per workload.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Wait if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your security scanning pipeline hasn't been updated for the new trust policy shape.&lt;/li&gt;
&lt;li&gt;You have subprocess-heavy workloads you haven't audited for credential inheritance.&lt;/li&gt;
&lt;li&gt;Your IaC is already mid-provider-version-migration. Don't stack two migrations.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The one-line takeaway
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"The association exists" is not the same as "the pod is using it."&lt;/strong&gt; Verify via logs and CloudTrail identity, not console state.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written as an AWS Community Builder. If you've hit other gotchas in your own migration, I'd like to hear about them in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>devops</category>
      <category>infrastructure</category>
      <category>kubernetes</category>
    </item>
    <item>
      <title>I Got Selected as an AWS Community Builder – Containers on My First Try. Here's My Honest Story.</title>
      <dc:creator>Remus Kalathil</dc:creator>
      <pubDate>Sun, 15 Mar 2026 07:29:15 +0000</pubDate>
      <link>https://dev.to/remus_kalathil_235e438778/i-got-selected-as-an-aws-community-builder-containers-on-my-first-try-heres-my-honest-story-58a5</link>
      <guid>https://dev.to/remus_kalathil_235e438778/i-got-selected-as-an-aws-community-builder-containers-on-my-first-try-heres-my-honest-story-58a5</guid>
      <description>&lt;p&gt;I had been selected for the &lt;strong&gt;AWS Community Builder program Containers category&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;First-time applicant. No conference speaking history. No massive LinkedIn following. Just consistent writing and 9+ years of real production experience with AWS and Kubernetes.&lt;/p&gt;

&lt;p&gt;This is my honest writeup of how it happened, what I submitted, and what I'd tell anyone thinking about applying.&lt;/p&gt;




&lt;h2&gt;
  
  
  How I Even Found Out This Program Existed
&lt;/h2&gt;

&lt;p&gt;I'll be upfront. I had zero idea the AWS Community Builder program existed until I joined the &lt;strong&gt;SA (Solutions Architect) bootcamp by Rajdeep Saha&lt;/strong&gt; last year.&lt;/p&gt;

&lt;p&gt;A few people in my cohort already had the CB badge on their LinkedIn profiles. I asked one of them about it and got the full picture what the program is, what you actually get out of it, and that applications open on a rolling basis, with a waitlist.&lt;/p&gt;

&lt;p&gt;That one conversation changed everything. I had spent 9+ years building cloud infrastructure with my head down. Running Kubernetes at scale, migrating workloads to AWS, designing multi-account architectures at Expedia Group. I had been doing it mostly quietly.&lt;/p&gt;

&lt;p&gt;The idea that there was a program specifically designed to recognize practitioners who share knowledge publicly was genuinely new to me.&lt;/p&gt;

&lt;p&gt;So I didn't just apply immediately I was intentional about it. I started &lt;strong&gt;posting regularly on LinkedIn&lt;/strong&gt;, sharing real production insights and lessons from my work. I began writing longer-form technical articles on &lt;strong&gt;Hashnode&lt;/strong&gt;. I built a habit of sharing knowledge publicly before I ever submitted an application.&lt;/p&gt;

&lt;p&gt;I joined the waitlist. When the application window opened I applied. First attempt. Got in.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Bootcamp That Started It All
&lt;/h2&gt;

&lt;p&gt;I want to give proper credit here because it matters.&lt;/p&gt;

&lt;p&gt;The SA bootcamp was run by &lt;strong&gt;Rajdeep Saha&lt;/strong&gt; at &lt;a href="https://www.cloudwithraj.com/" rel="noopener noreferrer"&gt;Cloud With Raj&lt;/a&gt;. If you don't know Raj, here's some context: he's a Stealth EdTech Startup Founder and former L7 Principal Solutions Architect at AWS, has presented at AWS re:Invent, KubeCon, and AWS Summits, co-authored the official Karpenter 1.0 announcement blog on AWS, and has trained over 350,000 students globally. LinkedIn awarded him "Top Systems Design Voice."&lt;/p&gt;

&lt;p&gt;The bootcamp attracts serious practitioners. So when multiple people in my cohort already had the AWS Community Builder badge, I paid attention.&lt;/p&gt;

&lt;p&gt;Raj's program didn't just teach cloud architecture. It created a community of people who push each other. That peer exposure is what made me aware of the CB program in the first place, and what motivated me to start sharing publicly.&lt;/p&gt;

&lt;p&gt;If you're an SA or cloud engineer looking to level up, check out &lt;a href="https://www.cloudwithraj.com/" rel="noopener noreferrer"&gt;cloudwithraj.com&lt;/a&gt;. It's the real deal.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the AWS Community Builder Program Actually Is
&lt;/h2&gt;

&lt;p&gt;For anyone reading this in the same position I was here's the quick version.&lt;/p&gt;

&lt;p&gt;The AWS Community Builder program recognizes AWS enthusiasts and &lt;strong&gt;emerging thought leaders&lt;/strong&gt; who are actively creating content and contributing to the technical community. It is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Not a certification&lt;/li&gt;
&lt;li&gt;Not a job&lt;/li&gt;
&lt;li&gt;A recognition and support program with real, tangible perks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What you get:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;$500–$1,000 in AWS credits for personal projects and experimentation&lt;/li&gt;
&lt;li&gt;Certification exam vouchers (Associate, Professional, Specialty)&lt;/li&gt;
&lt;li&gt;Early access to new AWS services and features&lt;/li&gt;
&lt;li&gt;A private Slack community with other builders and AWS service teams&lt;/li&gt;
&lt;li&gt;Support to amplify your content. AWS often reshares CB posts&lt;/li&gt;
&lt;li&gt;Invitations to re:Invent builder sessions and AWS Summits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are multiple specialty categories. I applied for &lt;strong&gt;Containers&lt;/strong&gt;, which covers EKS, ECS, Kubernetes, Docker, Karpenter, and everything in the container ecosystem.&lt;/p&gt;




&lt;h2&gt;
  
  
  My Background Going Into the Application
&lt;/h2&gt;

&lt;p&gt;I'm a Solutions Architect at &lt;strong&gt;Expedia Group&lt;/strong&gt; with 9+ years in cloud infrastructure. My day-to-day involves:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Running &lt;strong&gt;6XX+ Kubernetes clusters&lt;/strong&gt; across multiple AWS accounts&lt;/li&gt;
&lt;li&gt;Supporting &lt;strong&gt;1X,XXX+ microservices&lt;/strong&gt; on platform infrastructure I own&lt;/li&gt;
&lt;li&gt;Karpenter-based autoscaling and GitOps with ArgoCD&lt;/li&gt;
&lt;li&gt;GPU infrastructure for AI/ML inference workloads&lt;/li&gt;
&lt;li&gt;Large-scale cloud migrations (150+ servers, RTO ≤ 5 min, RPO ≤ 1 min)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I also hold a NVIDIA-Certified AI Infrastructure &amp;amp; Operations certification and am completing a PG Diploma in AI at Texas McCombs.&lt;/p&gt;

&lt;p&gt;I mention this not to flex but because it directly shaped what I highlighted in my application, and why that specificity mattered.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Highlighted in My Application
&lt;/h2&gt;

&lt;p&gt;The application asks you to describe your community contributions and provide three public links. I kept mine focused on two things.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. My writing LinkedIn and Hashnode
&lt;/h3&gt;

&lt;p&gt;I had been posting regularly on &lt;strong&gt;LinkedIn&lt;/strong&gt; about cloud architecture, Kubernetes, and platform engineering. Nothing viral. Just consistent, technical posts grounded in real production experience. I also had a &lt;strong&gt;Hashnode blog&lt;/strong&gt; with longer-form technical content.&lt;/p&gt;

&lt;p&gt;I didn't have hundreds of thousands of followers. What I had was a body of work showing I was genuinely trying to help other practitioners not just building a personal brand.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Real production context in everything I wrote
&lt;/h3&gt;

&lt;p&gt;When I wrote about Karpenter, I was writing from experience running it across hundreds of clusters. When I wrote about migration, I was describing a project where we achieved RTO ≤ 5 minutes on 150+ servers.&lt;/p&gt;

&lt;p&gt;That specificity matters. AWS is looking for people whose community contributions come from genuine depth not people who summarize AWS documentation.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Did NOT Have
&lt;/h2&gt;

&lt;p&gt;I want to be honest about this because a lot of CB posts make the bar sound impossibly high.&lt;/p&gt;

&lt;p&gt;I did not have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A massive LinkedIn following or viral posts&lt;/li&gt;
&lt;li&gt;Conference speaking credits&lt;/li&gt;
&lt;li&gt;Open source project maintainer status&lt;/li&gt;
&lt;li&gt;Prior AWS recognition of any kind&lt;/li&gt;
&lt;li&gt;A perfectly polished content portfolio&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I was a &lt;strong&gt;first-time applicant&lt;/strong&gt; with consistent writing and real production experience in the Containers space. That was enough.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Moment I Found Out
&lt;/h2&gt;

&lt;p&gt;I read the email twice to make sure I was reading it correctly.&lt;/p&gt;

&lt;p&gt;Then I updated my LinkedIn headline. Then I messaged the people in my bootcamp cohort who had encouraged me to apply which felt like the right way to close that loop.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'm Doing With It
&lt;/h2&gt;

&lt;p&gt;Being selected is the beginning, not the end. The CB program is only valuable if you actually show up.&lt;/p&gt;

&lt;p&gt;Here's what I'm committing to:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write consistently&lt;/strong&gt; at least one deep technical post per month on Dev.to (primary) and Hashnode (mirror). The program encourages quarterly content but I want to do more than the minimum.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cover the Containers category with production depth&lt;/strong&gt; Karpenter at scale, GitOps patterns with ArgoCD, GPU scheduling on EKS, multi-cluster architectures, ECS-to-EKS migration lessons. These are topics I live in daily.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engage with the community&lt;/strong&gt; the CB Slack is active and the builders in there are sharp. I want to learn as much as I contribute.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use the credits wisely&lt;/strong&gt; I have specific GPU infrastructure experiments I've wanted to run for a while. The AWS credits make that possible without the overhead of approvals.&lt;/p&gt;




&lt;h2&gt;
  
  
  Should You Apply?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Yes if you're actively sharing knowledge in any AWS category, apply.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A few things I'd tell my pre-application self:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't wait until you feel ready.&lt;/strong&gt; The program is for &lt;em&gt;emerging&lt;/em&gt; thought leaders, not established ones. You don't need 10,000 followers. You need genuine contributions and real knowledge. I almost talked myself out of applying because I didn't feel like I had enough. I was wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Be specific about your experience.&lt;/strong&gt; Generic answers don't stand out. If you've run something at scale, migrated something complex, or solved a real production problem say so. Specific always beats impressive-sounding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pick the right category.&lt;/strong&gt; Apply where you have genuine depth, not where you think the competition is lighter. I chose Containers because I work in it every single day. That authenticity comes through.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your three links are everything.&lt;/strong&gt; The application asks for three public links showcasing your contributions. These should demonstrate technical depth, community value, and consistency. Start building that body of work before you apply, not during.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Just start writing.&lt;/strong&gt; The single thing that helped my application most was having a public body of work. You don't need to write perfectly. You need to write consistently. Start now, even if applications are not open yet.&lt;/p&gt;




&lt;h2&gt;
  
  
  Quick Tips by Category
&lt;/h2&gt;

&lt;p&gt;If you're targeting the &lt;strong&gt;Containers&lt;/strong&gt; category like I did:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Write about EKS, ECS, Karpenter, ArgoCD, Helm&lt;/li&gt;
&lt;li&gt;Share real architecture decisions and trade-offs&lt;/li&gt;
&lt;li&gt;Document production lessons, not just setup tutorials&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're targeting &lt;strong&gt;AI / ML&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Share posts about Bedrock, SageMaker, or AWS AI services&lt;/li&gt;
&lt;li&gt;Show practical projects with AWS AI services&lt;/li&gt;
&lt;li&gt;Participate in AI-focused AWS events or hackathons&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For &lt;strong&gt;any category&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Contribute at least 3 to 6 months before applying. The team looks for momentum, not a last-minute burst&lt;/li&gt;
&lt;li&gt;Engage with others' content commenting, answering questions, participating in forums&lt;/li&gt;
&lt;li&gt;Show a consistent voice across your contributions&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;This post is the first of many. I'll be writing regularly about what I'm actually doing in production Kubernetes at scale, AWS platform engineering, AI/ML infrastructure, GitOps patterns, and the real lessons from 9+ years in cloud.&lt;/p&gt;

&lt;p&gt;If you're an SA, cloud engineer, platform engineer, or just someone building on AWS follow along.&lt;/p&gt;

&lt;p&gt;And if you're thinking about applying for the CB program, drop a comment. Happy to answer questions from where I'm standing.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Remus Kalathil is a Solutions Architect at Expedia Group and an AWS Community Builder – Containers. He writes about Kubernetes, AWS, and production platform engineering.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;a href="https://linkedin.com/in/remus-r-kalathil" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/rkalathil" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://hashnode.com/@rkalathil" rel="noopener noreferrer"&gt;Hashnode&lt;/a&gt; · &lt;a href="https://rkalathil.github.io" rel="noopener noreferrer"&gt;Website&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>kubernetes</category>
      <category>cloud</category>
      <category>career</category>
    </item>
  </channel>
</rss>
