<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alok Ranjan Daftuar</title>
    <description>The latest articles on DEV Community by Alok Ranjan Daftuar (@aloknecessary).</description>
    <link>https://dev.to/aloknecessary</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3791551%2F62fbfeb5-1fba-4e79-bc4b-780b7ce52748.jpg</url>
      <title>DEV Community: Alok Ranjan Daftuar</title>
      <link>https://dev.to/aloknecessary</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aloknecessary"/>
    <language>en</language>
    <item>
      <title>Progressive Delivery with Argo Rollouts: Canary, Blue-Green, and AnalysisTemplates That Actually Gate</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Mon, 05 Oct 2026 11:40:13 +0000</pubDate>
      <link>https://dev.to/aloknecessary/progressive-delivery-with-argo-rollouts-canary-blue-green-and-analysistemplates-that-actually-560c</link>
      <guid>https://dev.to/aloknecessary/progressive-delivery-with-argo-rollouts-canary-blue-green-and-analysistemplates-that-actually-560c</guid>
      <description>&lt;p&gt;A standard Kubernetes &lt;code&gt;Deployment&lt;/code&gt; rolling update has no concept of "check whether this is actually working before replacing more Pods." It replaces old Pods with new ones at the pace you configured, and if the new version is broken, it finds out the same way your users do — by serving it to more and more of them until someone notices.&lt;/p&gt;

&lt;p&gt;Argo Rollouts replaces the &lt;code&gt;Deployment&lt;/code&gt; resource with a &lt;code&gt;Rollout&lt;/code&gt; — API-compatible in its Pod template, but with an explicit strategy for how traffic shifts and, critically, automated checks that can halt or reverse that shift before it reaches everyone. This post covers the two strategies that matter in practice and the &lt;code&gt;AnalysisTemplate&lt;/code&gt; mechanism that makes the gating real rather than cosmetic.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why a Deployment Isn't Enough
&lt;/h2&gt;

&lt;p&gt;A &lt;code&gt;Deployment&lt;/code&gt;'s rolling update is governed by &lt;code&gt;maxSurge&lt;/code&gt; and &lt;code&gt;maxUnavailable&lt;/code&gt; — how many extra Pods can exist during the rollout, how many can be unavailable. Neither is a correctness check. A &lt;code&gt;Deployment&lt;/code&gt; will happily roll a broken version to 100% of Pods because nothing in its model asks "is this new version actually healthy" — only "are enough Pods reporting ready," which a broken-but-still-passing-liveness-probe version satisfies just fine.&lt;/p&gt;

&lt;p&gt;The two strategies have fundamentally different blast-radius profiles:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Canary                                    Blue-Green
  100% ─┐                                   100% ─┐ old (active)
        │                                         │
   10%  │ old          10% new                new │ 100% (preview,
        │              ▲                           │      no live traffic)
   50%  │ old      50% │ new                        │
        │              │ analysis gate               │ prePromotionAnalysis
    0%  │              │ new: 100%                    │ gate
        └──────────────┴──────────────►          cutover: 100% new,
        traffic shifts in steps,                  0% old, in one move
        gated between each step
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Canary spreads risk across several small, gated exposures. Blue-green concentrates verification into one pre-cutover check, then moves all traffic at once.&lt;/p&gt;




&lt;h2&gt;
  
  
  Canary Strategy
&lt;/h2&gt;

&lt;p&gt;Canary shifts traffic to the new version incrementally, in named steps, with a pause (manual or timed) between each. The key YAML shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;canary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;setWeight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pause&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;5m&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;analysis&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;templates&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;templateName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-success-rate&lt;/span&gt;
          &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service-name&lt;/span&gt;
              &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service-canary&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;setWeight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;50&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pause&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;5m&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;analysis&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;templates&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;templateName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-success-rate&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;setWeight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each &lt;code&gt;setWeight&lt;/code&gt; step shifts that percentage of traffic to the new ReplicaSet via whichever traffic-management integration is configured — an Ingress controller, a service mesh, or a load balancer controller. The &lt;code&gt;analysis&lt;/code&gt; steps between weight increases are where this stops being "a slower rolling update" and becomes an actual gate: the rollout does not proceed to the next &lt;code&gt;setWeight&lt;/code&gt; unless the referenced &lt;code&gt;AnalysisTemplate&lt;/code&gt; reports success.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The traffic-router gap worth knowing:&lt;/strong&gt; if the Ingress controller isn't one of Rollouts' supported traffic-router plugins, it silently falls back to a ReplicaSet-ratio approximation. A single canary Pod out of 10 total is not the same as a precise 10% of requests. Confirming which traffic-router plugin is actually active — not just assumed — is worth doing before trusting &lt;code&gt;setWeight&lt;/code&gt; values as precise.&lt;/p&gt;




&lt;h2&gt;
  
  
  Blue-Green Strategy
&lt;/h2&gt;

&lt;p&gt;Blue-green skips the incremental traffic shift and instead runs the full new version alongside the full old version, with a single cutover once the new version is verified:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;blueGreen&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;activeService&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service-active&lt;/span&gt;
    &lt;span class="na"&gt;previewService&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service-preview&lt;/span&gt;
    &lt;span class="na"&gt;autoPromotionEnabled&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
    &lt;span class="na"&gt;prePromotionAnalysis&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;templates&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;templateName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-smoke-test&lt;/span&gt;
      &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service-name&lt;/span&gt;
          &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service-preview&lt;/span&gt;
    &lt;span class="na"&gt;scaleDownDelaySeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;300&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The new version scales up fully behind &lt;code&gt;previewService&lt;/code&gt; — reachable for testing, not yet receiving production traffic. &lt;code&gt;prePromotionAnalysis&lt;/code&gt; runs against the preview version before any cutover is possible. With &lt;code&gt;autoPromotionEnabled: false&lt;/code&gt;, cutover requires an explicit &lt;code&gt;kubectl argo rollouts promote&lt;/code&gt; even after analysis passes — a deliberate human checkpoint on top of the automated one.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;scaleDownDelaySeconds&lt;/code&gt; keeps the old ReplicaSet running for 5 minutes after cutover so a rollback is a service-selector flip back, not a fresh Pod startup. Setting it too low defeats the rollback advantage — if Pods take 90 seconds to become traffic-ready and &lt;code&gt;scaleDownDelaySeconds&lt;/code&gt; is 30, the old ReplicaSet is already gone by the time a problem surfaces.&lt;/p&gt;




&lt;h2&gt;
  
  
  AnalysisTemplates — What Actually Gates the Rollout
&lt;/h2&gt;

&lt;p&gt;Both strategies are only as good as the analysis behind them. An &lt;code&gt;AnalysisTemplate&lt;/code&gt; defines a metric query and a success condition; Rollouts polls it on an interval and compares the result against thresholds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argoproj.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AnalysisTemplate&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-success-rate&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service-name&lt;/span&gt;
  &lt;span class="na"&gt;metrics&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;success-rate&lt;/span&gt;
      &lt;span class="na"&gt;interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1m&lt;/span&gt;
      &lt;span class="na"&gt;count&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;
      &lt;span class="na"&gt;successCondition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;result[0] &amp;gt;= &lt;/span&gt;&lt;span class="m"&gt;0.95&lt;/span&gt;
      &lt;span class="na"&gt;failureLimit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;
      &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;prometheus&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;address&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://prometheus.monitoring.svc:9090&lt;/span&gt;
          &lt;span class="na"&gt;query&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
            &lt;span class="s"&gt;sum(rate(http_requests_total{service="{{args.service-name}}", status!~"5.."}[1m]))&lt;/span&gt;
            &lt;span class="s"&gt;/&lt;/span&gt;
            &lt;span class="s"&gt;sum(rate(http_requests_total{service="{{args.service-name}}"}[1m]))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This queries Prometheus for the canary's success rate every minute, five times, and requires each sample to hold at or above 95%. Rollouts supports the same pattern against Datadog, New Relic, CloudWatch, and others — the shape of "query a metric, define a success condition, sample N times" is the transferable part.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The two failure modes worth planning for explicitly:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;code&gt;successCondition&lt;/code&gt; that's too loose (&lt;code&gt;result[0] &amp;gt;= 0.50&lt;/code&gt;, left over from testing) will pass a canary that's failing half its requests. The gate is present but not gating.&lt;/li&gt;
&lt;li&gt;A &lt;code&gt;count&lt;/code&gt; too low to be statistically meaningful on a low-traffic service — one or two failures in a handful of requests swing the ratio well outside a 95% threshold on pure noise, triggering rollbacks on releases that were actually fine.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both point to the same discipline: AnalysisTemplate thresholds need to be tuned against the actual traffic volume and baseline error rate of the specific service, not copied from another service's template.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where This Sits Against CI/CD Pipeline Stages
&lt;/h2&gt;

&lt;p&gt;Kargo's Stage-level verification (or a CI-to-GitOps handoff) answers "should this artifact be allowed into this environment at all." Argo Rollouts answers a different question, one layer down: "given that this artifact is now allowed into this environment, how carefully should it be exposed to that environment's traffic."&lt;/p&gt;

&lt;p&gt;The two compose rather than compete. A Kargo &lt;code&gt;Stage&lt;/code&gt; promoting into prod can just as easily be updating a &lt;code&gt;Rollout&lt;/code&gt; resource as a plain &lt;code&gt;Deployment&lt;/code&gt;. Nothing about progressive delivery requires picking a different CI-to-GitOps mechanism.&lt;/p&gt;

&lt;p&gt;Running both isn't redundant defense in depth for its own sake — each catches a class of failure the other structurally cannot. A service can pass every pre-promotion check and still fail an in-flight Rollouts analysis: a load-dependent memory leak, a connection pool exhausting under real concurrency, a downstream dependency behaving differently under production traffic patterns than under a synthetic staging test.&lt;/p&gt;




&lt;h2&gt;
  
  
  Decision Framework
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Recommended strategy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Stateless service, traffic-splitting infrastructure already in place&lt;/td&gt;
&lt;td&gt;Canary — gradual blast-radius limiting, most granular control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No traffic-splitting infrastructure available&lt;/td&gt;
&lt;td&gt;Blue-green, or canary via ReplicaSet ratio approximation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Release needs real-traffic-pattern validation before any production exposure&lt;/td&gt;
&lt;td&gt;Blue-green — preview service gives you that validation window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workload has strict version-coupling to a shared schema&lt;/td&gt;
&lt;td&gt;Neither — fix the deployment pattern first&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;The full post covers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Complete annotated YAML for both canary and blue-green &lt;code&gt;Rollout&lt;/code&gt; resources&lt;/li&gt;
&lt;li&gt;The exact traffic-router gap that makes &lt;code&gt;setWeight&lt;/code&gt; values imprecise without a supported plugin&lt;/li&gt;
&lt;li&gt;Why &lt;code&gt;scaleDownDelaySeconds&lt;/code&gt; sizing directly determines whether blue-green's rollback advantage is real or theoretical&lt;/li&gt;
&lt;li&gt;How AnalysisTemplate &lt;code&gt;count&lt;/code&gt; and &lt;code&gt;successCondition&lt;/code&gt; interact with low-traffic services to produce false rollbacks&lt;/li&gt;
&lt;li&gt;The precise relationship between Kargo/CI promotion gates and in-flight Rollouts analysis — why both are required and what each catches that the other cannot&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/argo-rollouts-progressive-delivery/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=argo-rollouts-progressive-delivery" rel="noopener noreferrer"&gt;Progressive Delivery with Argo Rollouts — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>progressivedelivery</category>
      <category>gitops</category>
    </item>
    <item>
      <title>The MCP Stateless Revolution</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Thu, 01 Oct 2026 10:58:38 +0000</pubDate>
      <link>https://dev.to/aloknecessary/the-mcp-stateless-revolution-57p9</link>
      <guid>https://dev.to/aloknecessary/the-mcp-stateless-revolution-57p9</guid>
      <description>&lt;p&gt;For two years, running an MCP server in production meant building infrastructure that had nothing to do with what the server was actually supposed to do. The protocol's session model forced sticky routing, shared session stores, and affinity cookies onto teams who otherwise had no reason to build stateful infrastructure at all. On July 28, 2026, the MCP specification removed the session entirely — and with it, the justification for that entire category of workaround.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the old transport forced on you
&lt;/h2&gt;

&lt;p&gt;An MCP &lt;code&gt;initialize&lt;/code&gt; handshake returned an &lt;code&gt;Mcp-Session-Id&lt;/code&gt; header, and every subsequent request in that conversation had to carry the same ID back to the same server instance. Keeping that session alive in production meant one of three things: sticky sessions at the load balancer, a shared session store like Redis so any instance could serve any request, or a client-side retry strategy that tolerated session loss and re-initialized when a pod restarted.&lt;/p&gt;

&lt;p&gt;Each solution adds a moving part that fails in a specific, familiar way during a rolling deployment. A pod gets terminated, its in-memory session state goes with it, and every client pinned to that pod either gets a connection reset mid-conversation or silently starts talking to a session that no longer exists.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Pre-2026-07-28: sticky routing required&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Service&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mcp-server&lt;/span&gt;
  &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;alb.ingress.kubernetes.io/target-group-attributes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
      &lt;span class="s"&gt;stickiness.enabled=true,&lt;/span&gt;
      &lt;span class="s"&gt;stickiness.type=lb_cookie,&lt;/span&gt;
      &lt;span class="s"&gt;stickiness.lb_cookie.duration_seconds=3600&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  What SEP-2567 actually removes
&lt;/h2&gt;

&lt;p&gt;The core change removes protocol-level sessions and the &lt;code&gt;Mcp-Session-Id&lt;/code&gt; header from the Streamable HTTP transport entirely — not deprecates, removes. Every request is now self-describing: protocol version, client identity, and capabilities travel in &lt;code&gt;_meta&lt;/code&gt; on that request rather than being negotiated once and remembered by a pinned server instance.&lt;/p&gt;

&lt;p&gt;This does not mean MCP servers can no longer carry state across calls. If a server needs to remember something between tool calls, it mints an explicit handle and returns it from the first call. The client passes that handle back as an ordinary argument — the same way any REST API has always handled state. What disappears is the requirement that the &lt;em&gt;same pod&lt;/em&gt; be the one to look it up.&lt;/p&gt;




&lt;h2&gt;
  
  
  Header-based routing: the other half of the change
&lt;/h2&gt;

&lt;p&gt;Statelessness solves scaling, but creates a new problem: how does a gateway make intelligent routing decisions without parsing every JSON-RPC body? SEP-2243 answers this: Streamable HTTP requests must now carry &lt;code&gt;Mcp-Method&lt;/code&gt; and &lt;code&gt;Mcp-Name&lt;/code&gt; headers. A load balancer can now route, throttle, and meter MCP traffic on headers — without ever being MCP-aware at the body level.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# AWS ALB post-2026-07-28 — no stickiness, header-based routing&lt;/span&gt;
&lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;alb.ingress.kubernetes.io/conditions.mcp-tasks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;[{"field":"http-header","httpHeaderConfig":{"httpHeaderName":"Mcp-Method","values":["tasks/*"]}}]&lt;/span&gt;
  &lt;span class="na"&gt;alb.ingress.kubernetes.io/actions.mcp-tasks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;{"type":"forward","forwardConfig":{"targetGroups":[{"serviceName":"mcp-tasks-pool","servicePort":8443}]}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Long-running task calls route to a differently-sized backend pool. Ordinary tool calls go elsewhere. No custom MCP-parsing logic in the gateway.&lt;/p&gt;




&lt;h2&gt;
  
  
  Per-request auth: the security implication
&lt;/h2&gt;

&lt;p&gt;Removing the session also removes the one place authentication used to be established once and implicitly trusted for the rest of the conversation. Every request now carries its own client identity in &lt;code&gt;_meta&lt;/code&gt;, which means authorization is evaluated per request. A revoked credential takes effect on the very next call instead of only once the session naturally expires. The infrastructure implication: your gateway now performs an identity check on every request rather than once per conversation — real added latency on every call, worth budgeting for explicitly when sizing the gateway layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  The migration window
&lt;/h2&gt;

&lt;p&gt;The specification's deprecation policy gives a twelve-month compatibility runway for features being phased out. What is removed outright — with no grace period — is the session header and &lt;code&gt;initialize&lt;/code&gt; handshake. The realistic migration path is running both transport paths side by side behind the same ingress until every client you support has moved. The sticky-routing configuration doesn't disappear the day you upgrade — it stays live serving the old path while new traffic routes to the stateless pool through header-based rules.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;This summary covers the core protocol changes and their infrastructure implications. The full article includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Complete before/after Kubernetes Service and Ingress YAML for both AWS ALB and Azure Application Gateway&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;ttlMs&lt;/code&gt;/&lt;code&gt;cacheScope&lt;/code&gt; caching signals (SEP-2549) and what they mean for CDN/edge caching&lt;/li&gt;
&lt;li&gt;W3C Trace Context propagation in &lt;code&gt;_meta&lt;/code&gt; (SEP-414) and distributed tracing implications&lt;/li&gt;
&lt;li&gt;How the removal of &lt;code&gt;Last-Event-ID&lt;/code&gt; stream resumability simplifies liveness probe design&lt;/li&gt;
&lt;li&gt;Detailed dual-path migration strategy to avoid mid-migration outages&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/the-mcp-stateless-revolution/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=the-mcp-stateless-revolution" rel="noopener noreferrer"&gt;The MCP Stateless Revolution — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>cloud</category>
      <category>kubernetes</category>
    </item>
    <item>
      <title>VPC Endpoints: Gateway vs Interface, and the AWS Traffic That Shouldn't Touch the Internet</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Mon, 28 Sep 2026 05:31:14 +0000</pubDate>
      <link>https://dev.to/aloknecessary/vpc-endpoints-gateway-vs-interface-and-the-aws-traffic-that-shouldnt-touch-the-internet-2cnj</link>
      <guid>https://dev.to/aloknecessary/vpc-endpoints-gateway-vs-interface-and-the-aws-traffic-that-shouldnt-touch-the-internet-2cnj</guid>
      <description>&lt;p&gt;Traffic from a private subnet to S3 routes through your NAT gateway by default — and gets billed per GB for the privilege — when a free alternative has existed for years. VPC endpoints give AWS services a private, in-VPC address path that never touches the internet. There are two fundamentally different mechanisms, and understanding how they differ changes how you decide which services actually need one.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why private subnet traffic hits NAT at all
&lt;/h2&gt;

&lt;p&gt;A private subnet instance calling &lt;code&gt;s3.ap-south-1.amazonaws.com&lt;/code&gt; does exactly what it would do calling any internet host: routes to the NAT gateway, gets source-translated, exits through the IGW, and comes back the same way. Both the instance and the S3 bucket are inside AWS — the traffic never needed to leave AWS's network. VPC endpoints fix this by giving the traffic a direct in-VPC path.&lt;/p&gt;




&lt;h2&gt;
  
  
  Gateway endpoints: a route table entry, nothing more
&lt;/h2&gt;

&lt;p&gt;Gateway endpoints exist for exactly two services: &lt;strong&gt;S3 and DynamoDB&lt;/strong&gt;. They are not network devices — they are route table targets. AWS injects a managed prefix list route automatically; traffic matching S3's IP ranges routes to the endpoint instead of NAT.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_vpc_endpoint"&lt;/span&gt; &lt;span class="s2"&gt;"s3"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt;            &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;service_name&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"com.amazonaws.ap-south-1.s3"&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_endpoint_type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"Gateway"&lt;/span&gt;
  &lt;span class="nx"&gt;route_table_ids&lt;/span&gt;   &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;aws_route_table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;private&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Cost: free.&lt;/strong&gt; No hourly charge, no per-GB processing charge. This makes the S3 Gateway endpoint close to a strict upgrade — there is essentially no reason not to add it to every VPC with private subnets that touch S3.&lt;/p&gt;

&lt;p&gt;The one real limitation: Gateway endpoints are associated with specific route tables, not the whole VPC. A new private subnet with its own route table falls back to NAT silently unless you explicitly associate the endpoint with that table too. Worth adding to any subnet-creation checklist.&lt;/p&gt;




&lt;h2&gt;
  
  
  Interface endpoints: a real ENI, backed by PrivateLink
&lt;/h2&gt;

&lt;p&gt;Every other AWS service that supports VPC endpoints uses an &lt;strong&gt;Interface endpoint&lt;/strong&gt; — a real Elastic Network Interface provisioned in a subnet you specify, with a private IP from that subnet's range, backed by AWS PrivateLink.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_vpc_endpoint"&lt;/span&gt; &lt;span class="s2"&gt;"secretsmanager"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt;             &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;service_name&lt;/span&gt;       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"com.amazonaws.ap-south-1.secretsmanager"&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_endpoint_type&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"Interface"&lt;/span&gt;
  &lt;span class="nx"&gt;subnet_ids&lt;/span&gt;         &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;aws_subnet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;private_a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;aws_subnet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;private_b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="nx"&gt;security_group_ids&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;aws_security_group&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;vpce_sm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="nx"&gt;private_dns_enabled&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because it's a real ENI, it has its own security group — Interface endpoint access is controlled by security group rules, not route table association. &lt;code&gt;private_dns_enabled = true&lt;/code&gt; makes the endpoint transparent to application code: AWS overrides the public DNS name to resolve to the endpoint's private IP inside the VPC, so no SDK or application configuration change is needed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost: the opposite of Gateway endpoints.&lt;/strong&gt; Interface endpoints bill hourly per AZ plus per-GB data processed. "Add Interface endpoints for everything" is not automatically correct — it requires actual usage data.&lt;/p&gt;




&lt;h2&gt;
  
  
  Which services actually justify an Interface endpoint
&lt;/h2&gt;

&lt;p&gt;Start with services every private instance calls constantly regardless of application logic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;STS&lt;/strong&gt; — IAM role assumption is called far more often than people expect under the hood&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CloudWatch Logs&lt;/strong&gt; — continuous log shipping from busy applications&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ECR&lt;/strong&gt; — image pulls from private subnets, especially during autoscaling cold starts or frequent CI/CD deployments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These three tend to justify their hourly cost quickly in any account with meaningful private-subnet compute. Lower-traffic services — Secrets Manager called once at startup, SNS for occasional notifications — are worth evaluating with real Cost Explorer data rather than provisioning preemptively.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ce get-cost-and-usage &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--time-period&lt;/span&gt; &lt;span class="nv"&gt;Start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2026-07-01,End&lt;span class="o"&gt;=&lt;/span&gt;2026-08-01 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--granularity&lt;/span&gt; MONTHLY &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter&lt;/span&gt; file://nat-gateway-filter.json &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--metrics&lt;/span&gt; &lt;span class="s2"&gt;"UsageQuantity"&lt;/span&gt; &lt;span class="s2"&gt;"BlendedCost"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Side-by-side comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Gateway Endpoint&lt;/th&gt;
&lt;th&gt;Interface Endpoint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mechanism&lt;/td&gt;
&lt;td&gt;Route table entry&lt;/td&gt;
&lt;td&gt;ENI backed by PrivateLink&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Services&lt;/td&gt;
&lt;td&gt;S3, DynamoDB only&lt;/td&gt;
&lt;td&gt;Most other AWS services&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;Hourly per AZ + per-GB processed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security control&lt;/td&gt;
&lt;td&gt;Existing SG/NACL on the instance&lt;/td&gt;
&lt;td&gt;Own security group on the endpoint ENI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DNS&lt;/td&gt;
&lt;td&gt;No change needed&lt;/td&gt;
&lt;td&gt;Private DNS override, if enabled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-subnet&lt;/td&gt;
&lt;td&gt;Associate route table per subnet&lt;/td&gt;
&lt;td&gt;Deploy ENI per AZ needed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Failure modes worth knowing before you're debugging under pressure
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Private DNS conflicts in Transit Gateway / peered VPC designs&lt;/strong&gt; — if two connected VPCs both have Interface endpoints for the same service with private DNS enabled, DNS resolution can become ambiguous. An instance in VPC-A may resolve the service name to VPC-B's endpoint ENI. Test this explicitly in hub-and-spoke designs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security group too narrow&lt;/strong&gt; — a security group that misses the calling subnet's CIDR produces a connection timeout that looks identical to a routing failure, but is actually a security group issue at the endpoint itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Endpoint deployed in wrong AZ&lt;/strong&gt; — the endpoint still functions cross-AZ, but incurs cross-AZ data transfer charges on top of the endpoint's per-GB cost. Deploy into every AZ that has calling instances.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom DHCP option set or Route 53 Resolver rule taking precedence&lt;/strong&gt; — if the VPC has custom DNS resolution configured, the endpoint's private hosted zone override can be superseded and traffic reverts to NAT despite correct endpoint configuration.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;This summary covers the core mechanisms and decision framework. The full article includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Complete Terraform for both Gateway and Interface endpoints with security group scoping&lt;/li&gt;
&lt;li&gt;The full AWS CLI route table inspection workflow&lt;/li&gt;
&lt;li&gt;Detailed cost modelling for Interface endpoint break-even analysis&lt;/li&gt;
&lt;li&gt;Security group and NACL implications in depth&lt;/li&gt;
&lt;li&gt;The specific action to take this week if you haven't added the S3 Gateway endpoint yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/aws-vpc-endpoints-gateway-vs-interface/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=aws-vpc-endpoints-gateway-vs-interface" rel="noopener noreferrer"&gt;VPC Endpoints: Gateway vs Interface — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>networking</category>
      <category>devops</category>
      <category>cloud</category>
    </item>
    <item>
      <title>Secrets Management in a GitOps World: Sealed Secrets vs. External Secrets Operator vs. Vault</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:03:22 +0000</pubDate>
      <link>https://dev.to/aloknecessary/secrets-management-in-a-gitops-world-sealed-secrets-vs-external-secrets-operator-vs-vault-262c</link>
      <guid>https://dev.to/aloknecessary/secrets-management-in-a-gitops-world-sealed-secrets-vs-external-secrets-operator-vs-vault-262c</guid>
      <description>&lt;p&gt;GitOps means everything a cluster needs lives in Git. Secrets very much don't belong in Git — not even in a private repo, because Git history is forever and "who had read access eighteen months ago" is a much harder question than "who has read access today."&lt;/p&gt;

&lt;p&gt;Three approaches resolve that contradiction. Each breaks the "it's all just committed to Git" assumption in a different place, and each solves a different subset of the actual problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "Solving" Secrets in GitOps Actually Means
&lt;/h2&gt;

&lt;p&gt;Before comparing tools, it's worth being precise about what's being solved — because the three approaches solve different parts of it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Never having plaintext in Git&lt;/strong&gt; — not even encrypted-and-committed, if the encryption key itself can leak&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rotating secrets without a Git commit&lt;/strong&gt; — if rotation requires a PR, it happens less often than it should&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scoping which cluster or namespace can see which secret&lt;/strong&gt; — the same multi-tenancy question AppProjects answer for Applications, but for secret values&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No single approach scores well on all three by default. Each is a different set of tradeoffs.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sealed Secrets
&lt;/h2&gt;

&lt;p&gt;Bitnami's Sealed Secrets takes the most literal path: encrypt the secret client-side with a public key, commit the resulting &lt;code&gt;SealedSecret&lt;/code&gt; resource, and let an in-cluster controller — the only holder of the matching private key — decrypt it at apply time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl create secret generic checkout-db-creds &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;client &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--from-literal&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;password&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'REPLACE_ME'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-o&lt;/span&gt; yaml | kubeseal &lt;span class="nt"&gt;--format&lt;/span&gt; yaml &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; checkout-db-creds-sealed.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# safe to commit — only the in-cluster controller can decrypt this&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bitnami.com/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;SealedSecret&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-db-creds&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;encryptedData&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;password&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AgBy3i4OJSWK+PiTySYZZA9rO43cGDEQ...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What it gets you:&lt;/strong&gt; nothing sensitive ever touches Git. No external dependency — self-contained controller, no Vault cluster or cloud secrets service required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it doesn't get you:&lt;/strong&gt; rotation still requires a commit. Changing a database password means re-sealing and re-committing — same PR friction as any other manifest change. At 40 services rotating every 90 days, that's 40 separate re-seal-and-commit operations per quarter, each needing someone with &lt;code&gt;kubeseal&lt;/code&gt; access to the right cluster's public key at the time rotation is due.&lt;/p&gt;

&lt;p&gt;It's also cluster-scoped by design: a &lt;code&gt;SealedSecret&lt;/code&gt; sealed for cluster A's public key is meaningless on cluster B. At multi-cluster scale, that means either re-sealing per cluster or restricting Sealed Secrets to genuinely cluster-local secrets.&lt;/p&gt;




&lt;h2&gt;
  
  
  External Secrets Operator
&lt;/h2&gt;

&lt;p&gt;ESO inverts the model entirely: the secret's actual value never enters Git at all. What's committed is a reference — "fetch this key from AWS Secrets Manager" — and a controller resolves that reference into a real &lt;code&gt;Secret&lt;/code&gt; by calling the secrets backend at sync time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;external-secrets.io/v1beta1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ExternalSecret&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-db-creds&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;refreshInterval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;15m&lt;/span&gt;
  &lt;span class="na"&gt;secretStoreRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;aws-secrets-manager&lt;/span&gt;
    &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;SecretStore&lt;/span&gt;
  &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-db-creds&lt;/span&gt;
  &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;secretKey&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;password&lt;/span&gt;
      &lt;span class="na"&gt;remoteRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service/prod/db-password&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This manifest is entirely safe to commit — there's no ciphertext, no encoded value. The sensitive part is the &lt;em&gt;value&lt;/em&gt; at &lt;code&gt;checkout-service/prod/db-password&lt;/code&gt;, which lives only in AWS Secrets Manager.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it gets you:&lt;/strong&gt; rotation with zero Git commits — update the value in Secrets Manager and every &lt;code&gt;ExternalSecret&lt;/code&gt; referencing it picks up the change on its next &lt;code&gt;refreshInterval&lt;/code&gt; cycle. Centralized secret storage that extends your existing cloud secrets service into Kubernetes rather than maintaining a separate story.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it doesn't get you:&lt;/strong&gt; the secrets backend must be reachable from every cluster running ESO — reintroducing the network-reachability planning from multi-cluster architecture, this time for every spoke's path to Secrets Manager or Key Vault.&lt;/p&gt;

&lt;p&gt;One scoping risk worth calling out: a &lt;code&gt;SecretStore&lt;/code&gt; backed by an overly broad IAM role — one that can read every key under &lt;code&gt;checkout-service/*&lt;/code&gt; rather than the specific key a given &lt;code&gt;ExternalSecret&lt;/code&gt; needs — means a misconfigured &lt;code&gt;ExternalSecret&lt;/code&gt; in one namespace can pull secrets it was never meant to see. Scope the IAM policy behind each &lt;code&gt;SecretStore&lt;/code&gt; to the narrowest key path prefix that namespace actually owns.&lt;/p&gt;




&lt;h2&gt;
  
  
  Vault Agent Injector
&lt;/h2&gt;

&lt;p&gt;Where ESO pulls a secret into a Kubernetes &lt;code&gt;Secret&lt;/code&gt; object, Vault's Kubernetes integrations more often avoid creating a &lt;code&gt;Secret&lt;/code&gt; object at all. The Vault Agent Injector mounts secrets directly into a Pod's filesystem as files, at Pod startup, authenticated via Kubernetes' own service account token exchanged against Vault's Kubernetes auth method.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;vault.hashicorp.com/agent-inject&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true"&lt;/span&gt;
        &lt;span class="na"&gt;vault.hashicorp.com/role&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service&lt;/span&gt;
        &lt;span class="na"&gt;vault.hashicorp.com/agent-inject-secret-db-creds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;secret/data/checkout-service/prod/db&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;serviceAccountName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service&lt;/span&gt;
      &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service&lt;/span&gt;
          &lt;span class="c1"&gt;# application reads from /vault/secrets/db-creds — never from an env var&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What it gets you:&lt;/strong&gt; the closest thing to "the secret was never a Kubernetes API object" — no &lt;code&gt;Secret&lt;/code&gt; resource exists to be over-broadly RBAC'd, listed via &lt;code&gt;kubectl get secrets&lt;/code&gt;, or accidentally dumped in a debug script.&lt;/p&gt;

&lt;p&gt;More importantly: Vault's database secrets engine can generate a unique, short-lived database username and password &lt;em&gt;per Pod&lt;/em&gt;, on demand, with a lease duration Vault itself enforces. A 1-hour lease means Vault revokes that specific credential at the database level after an hour, independent of whether anyone remembered to rotate anything. This changes rotation from "a scheduled human task" into "an enforced property of every credential issued" — a meaningfully different compliance posture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it doesn't get you:&lt;/strong&gt; simplicity. Running Vault well — unsealing, storage backend, HA, policy management — is a larger operational commitment than either of the other two approaches. If Vault isn't already part of your platform, adopting it purely to solve the GitOps secrets problem is usually the wrong-sized tool.&lt;/p&gt;




&lt;h2&gt;
  
  
  Decision Framework
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Recommended approach&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Small team, few clusters, low rotation cadence&lt;/td&gt;
&lt;td&gt;Sealed Secrets — simplest, no external dependency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Already using AWS Secrets Manager or Azure Key Vault&lt;/td&gt;
&lt;td&gt;External Secrets Operator — extends existing storage into Kubernetes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-cluster fleet, secrets needed consistently across clusters&lt;/td&gt;
&lt;td&gt;External Secrets Operator — reference model isn't cluster-scoped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Already running Vault for dynamic credentials or PKI&lt;/td&gt;
&lt;td&gt;Vault Agent Injector — extends a tool you're already operating&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Secrets where even a &lt;code&gt;Secret&lt;/code&gt; object existing in the cluster is too much exposure&lt;/td&gt;
&lt;td&gt;Vault injector — no &lt;code&gt;Secret&lt;/code&gt; object is ever created&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These aren't exclusive. Running ESO for the bulk of a fleet's secrets and reserving Vault's injector for the small set of credentials needing dynamic short-lived leases is a reasonable production posture — matching each secret's sensitivity to the approach that handles it well.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;The full article covers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The precise three-part definition of what "solving" secrets in GitOps actually means — and why no single tool covers all three&lt;/li&gt;
&lt;li&gt;The 90-day rotation cost at 40 services with Sealed Secrets, and why it surfaces at compliance audit time&lt;/li&gt;
&lt;li&gt;The cluster-scoping limitation of Sealed Secrets at multi-cluster scale&lt;/li&gt;
&lt;li&gt;Full &lt;code&gt;SecretStore&lt;/code&gt; + &lt;code&gt;ExternalSecret&lt;/code&gt; YAML with IRSA/workload-identity auth chain&lt;/li&gt;
&lt;li&gt;The IAM scoping risk with ESO and how it mirrors AppProject destinations restriction&lt;/li&gt;
&lt;li&gt;Vault's dynamic credential model explained — why it changes rotation from a task to an enforced property&lt;/li&gt;
&lt;li&gt;How all three approaches compose in a real fleet rather than requiring a single tool choice&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/gitops-secrets-management/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=gitops-secrets-management" rel="noopener noreferrer"&gt;Secrets Management in a GitOps World — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>security</category>
      <category>kubernetes</category>
      <category>gitops</category>
    </item>
    <item>
      <title>Migrating Ubuntu Servers from Azure to AWS: The linux-azure Kernel Trap</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Mon, 21 Sep 2026 07:34:51 +0000</pubDate>
      <link>https://dev.to/aloknecessary/migrating-ubuntu-servers-from-azure-to-aws-the-linux-azure-kernel-trap-5902</link>
      <guid>https://dev.to/aloknecessary/migrating-ubuntu-servers-from-azure-to-aws-the-linux-azure-kernel-trap-5902</guid>
      <description>&lt;p&gt;You replicate a Ubuntu 24.04 server from Azure to AWS using MGN. The replication completes cleanly. The test launch succeeds. EC2 status checks pass. And then you try to SSH in and get nothing — no response, no timeout, just silence.&lt;/p&gt;

&lt;p&gt;This is a specific failure mode that only affects Ubuntu servers that started life on Azure, and it's deceptive precisely because every layer of the migration tooling reports success. The problem is three layers down, in the kernel package itself.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Azure Ubuntu Images Are Different
&lt;/h2&gt;

&lt;p&gt;Ubuntu's official Azure images ship with &lt;code&gt;linux-azure&lt;/code&gt; — a kernel package built and tuned specifically for Hyper-V. It includes Hyper-V-specific drivers and, critically, does &lt;strong&gt;not&lt;/strong&gt; include AWS's &lt;code&gt;ena&lt;/code&gt; (Elastic Network Adapter) driver.&lt;/p&gt;

&lt;p&gt;MGN replicates the disk exactly as it is, including that kernel. When the Nitro hypervisor presents network hardware to the guest OS, the running kernel has no &lt;code&gt;ena&lt;/code&gt; module to handle it. The instance comes up with no functioning network interface — not because of security groups or routing, but because the OS literally cannot talk to the hardware AWS gave it.&lt;/p&gt;

&lt;p&gt;EC2 status checks pass because they verify the hypervisor can reach the instance at a basic level, not that the guest OS has a working network stack. "Instance healthy" and "instance reachable" are not the same claim.&lt;/p&gt;




&lt;h2&gt;
  
  
  Confirming the Cause via EC2 Serial Console
&lt;/h2&gt;

&lt;p&gt;Since SSH is unavailable, diagnosis happens through the EC2 Serial Console. Three commands confirm the root cause:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;
&lt;span class="c"&gt;# 6.8.0-1015-azure    &amp;lt;- the tell: "-azure" suffix on an AWS instance&lt;/span&gt;

dpkg &lt;span class="nt"&gt;-l&lt;/span&gt; | &lt;span class="nb"&gt;grep &lt;/span&gt;linux-image
&lt;span class="c"&gt;# linux-image-6.8.0-1015-azure   installed&lt;/span&gt;
&lt;span class="c"&gt;# (no linux-image-generic or linux-image-aws present)&lt;/span&gt;

lsmod | &lt;span class="nb"&gt;grep &lt;/span&gt;ena
&lt;span class="c"&gt;# (no output — module not loaded)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;-azure&lt;/code&gt; suffix on &lt;code&gt;uname -r&lt;/code&gt; while running as an EC2 instance is close to a definitive signal on its own. The absent &lt;code&gt;ena&lt;/code&gt; entry in &lt;code&gt;lsmod&lt;/code&gt; confirms the mechanism.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Fix: Kernel Swap Before Cutover
&lt;/h2&gt;

&lt;p&gt;Install &lt;code&gt;linux-generic&lt;/code&gt; — Ubuntu's broad-coverage kernel that includes &lt;code&gt;ena&lt;/code&gt; support — alongside the existing kernel, update GRUB, and only then remove the Azure-specific one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; linux-generic linux-headers-generic

&lt;span class="c"&gt;# Confirm GRUB will boot the generic kernel first&lt;/span&gt;
&lt;span class="nb"&gt;grep &lt;/span&gt;GRUB_DEFAULT /etc/default/grub
&lt;span class="nb"&gt;sudo &lt;/span&gt;update-grub
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The removal step comes &lt;strong&gt;after&lt;/strong&gt; confirming the generic kernel boots successfully:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Only after a successful boot on linux-generic:&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt remove &lt;span class="nt"&gt;--purge&lt;/span&gt; linux-azure linux-image-&lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="nt"&gt;-azure&lt;/span&gt; linux-headers-&lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="nt"&gt;-azure&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;update-grub
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Removing the old kernel before confirming the new one boots is how a routine kernel swap turns into a boot-loop with no fallback entry in the GRUB menu.&lt;/p&gt;

&lt;p&gt;The cleanest sequencing is doing this on the &lt;strong&gt;source Azure VM before MGN replication begins&lt;/strong&gt; — so the replicated disk already has the correct kernel. If the source can't be touched, do it on a test-launched instance and validate before cutover.&lt;/p&gt;




&lt;h2&gt;
  
  
  Baking It Into a Pre-Flight Check
&lt;/h2&gt;

&lt;p&gt;For a migration wave with multiple servers, this is worth scripting into a pre-migration audit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/bin/bash&lt;/span&gt;
&lt;span class="nv"&gt;KERNEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$KERNEL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt;azure&lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"WARNING: &lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;hostname&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt; is running &lt;/span&gt;&lt;span class="nv"&gt;$KERNEL&lt;/span&gt;&lt;span class="s2"&gt; — Azure-specific kernel detected."&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"This server needs linux-generic installed and GRUB updated before MGN cutover."&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;else
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"OK: &lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;hostname&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt; running &lt;/span&gt;&lt;span class="nv"&gt;$KERNEL&lt;/span&gt;&lt;span class="s2"&gt; — no Azure-specific kernel detected."&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running this across every source server during wave planning — before agent installation — turns a per-incident debugging exercise into a one-line pre-flight check.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Second Layer: Network Interface Naming
&lt;/h2&gt;

&lt;p&gt;Even after &lt;code&gt;ena&lt;/code&gt; loads correctly, the network interface name the OS expects can be wrong. Azure and AWS use different predictable naming schemes, and the &lt;code&gt;netplan&lt;/code&gt; config on the source server was written against Azure's naming. After the kernel swap, the &lt;code&gt;ena&lt;/code&gt;-backed interface may come up under a different name than the config expects.&lt;/p&gt;

&lt;p&gt;The robust fix is matching by driver rather than hardcoded name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;network&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
  &lt;span class="na"&gt;ethernets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;ena-primary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;driver&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ena&lt;/span&gt;
      &lt;span class="na"&gt;dhcp4&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This survives future re-platforming without another manual edit.&lt;/p&gt;




&lt;h2&gt;
  
  
  Other Azure Dependencies to Audit in the Same Pass
&lt;/h2&gt;

&lt;p&gt;While fixing the kernel, check for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;walinuxagent&lt;/code&gt;&lt;/strong&gt; — the Azure Linux Agent retries calls to &lt;code&gt;169.254.169.254&lt;/code&gt;, the same link-local address AWS's instance metadata service uses. This collision can interfere with &lt;code&gt;cloud-init&lt;/code&gt; and AWS tooling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Azure-specific &lt;code&gt;udev&lt;/code&gt; rules&lt;/strong&gt; — stale rules referencing Azure device paths, usually harmless but worth clearing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;cloud-init&lt;/code&gt; datasource&lt;/strong&gt; — if pinned to &lt;code&gt;Azure&lt;/code&gt; in &lt;code&gt;/etc/cloud/cloud.cfg.d/&lt;/code&gt;, it will hang or fail on AWS trying to find a datasource that doesn't exist.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Don't Stay on linux-generic — Move to linux-aws
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;linux-generic&lt;/code&gt; is the right kernel to get unblocked during migration. It's not the right long-term kernel for a production AWS instance. Ubuntu publishes &lt;code&gt;linux-aws&lt;/code&gt;, tuned specifically for Nitro, with platform-specific patches and its own update cadence separate from the generic track.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; linux-aws linux-headers-aws
&lt;span class="nb"&gt;sudo &lt;/span&gt;update-grub
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same discipline: confirm it boots and networks correctly before removing &lt;code&gt;linux-generic&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;This summary covers the core diagnosis and fix. The full article includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why EC2 status checks pass even when the network is completely broken&lt;/li&gt;
&lt;li&gt;The exact GRUB default entry approach vs. &lt;code&gt;GRUB_DEFAULT=0&lt;/code&gt; trade-offs&lt;/li&gt;
&lt;li&gt;The boot-loop scenario that results from removing the old kernel too early&lt;/li&gt;
&lt;li&gt;Validating through MGN test launch vs. serial console — why they're not equivalent&lt;/li&gt;
&lt;li&gt;The full audit checklist for Azure-sourced servers entering a migration wave&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/aws-migration-ubuntu-kernel-fix/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=aws-migration-ubuntu-kernel-fix" rel="noopener noreferrer"&gt;Migrating Ubuntu 24.04 from Azure to AWS: The linux-azure Kernel Trap — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>devops</category>
      <category>migration</category>
      <category>ubuntu</category>
    </item>
    <item>
      <title>Why Agent Infrastructure Is Its Own Discipline</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Fri, 18 Sep 2026 05:39:20 +0000</pubDate>
      <link>https://dev.to/aloknecessary/why-agent-infrastructure-is-its-own-discipline-4778</link>
      <guid>https://dev.to/aloknecessary/why-agent-infrastructure-is-its-own-discipline-4778</guid>
      <description>&lt;p&gt;Ask a platform team how they're going to run their first production agent and you'll get a confident answer within ten seconds: containerize it, put it behind an ingress, wire up a Horizontal Pod Autoscaler, done. It's the same playbook that's shipped every stateless service for the last decade, and there's no obvious reason an agent should be different. It accepts a request. It returns a response. It's just a container.&lt;/p&gt;

&lt;p&gt;That answer is wrong — and it's wrong in a way that doesn't show up in a demo. It shows up three weeks into production, when a pod gets killed mid-reasoning because a liveness probe decided a 40-second tool call was a hang, or when the autoscaler adds five replicas because CPU spiked on a single request calling six tools in sequence, or when two "sessions" turn out to share a K8s Service without anyone having reasoned about what that means for isolation.&lt;/p&gt;




&lt;h2&gt;
  
  
  The shape of an agent request vs. a microservice request
&lt;/h2&gt;

&lt;p&gt;A typical microservice request has a shape platform engineers have spent fifteen years optimizing around: bounded latency, a single unit of compute per request, statelessness between requests, and a clear success/failure signal at the HTTP layer. An agent request breaks all four at once.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Standard microservice request:
  client -&amp;gt; service -&amp;gt; [DB/cache lookup] -&amp;gt; response
  Duration: milliseconds to low seconds
  Compute: roughly constant per request
  State: none carried between requests
  Outcome signal: HTTP status code

Agent request:
  client -&amp;gt; agent -&amp;gt; [reason] -&amp;gt; tool call 1 -&amp;gt; [reason] -&amp;gt; tool call 2
        -&amp;gt; [reason] -&amp;gt; tool call N -&amp;gt; [reason] -&amp;gt; response
  Duration: seconds to minutes, highly variable
  Compute: proportional to reasoning depth and tool fan-out
  State: conversation/task context carried across the entire chain
  Outcome signal: HTTP 200 with a semantically wrong answer is common
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is the one platform teams underestimate most. An agent that loops on a tool, calls the wrong one, or returns a plausible-but-incorrect result will still return a healthy status code. The infrastructure layer has no way to distinguish a correct run from a confidently wrong one.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where the standard playbook breaks
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Microservice assumption&lt;/th&gt;
&lt;th&gt;Agent reality&lt;/th&gt;
&lt;th&gt;Practical implication&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Liveness = "process is alive and responsive"&lt;/td&gt;
&lt;td&gt;A live agent can be legitimately unresponsive for 30–90 seconds mid-tool-call&lt;/td&gt;
&lt;td&gt;Naive liveness probes kill healthy pods mid-reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU/memory tracks load&lt;/td&gt;
&lt;td&gt;Load tracks reasoning depth and tool fan-out, not CPU&lt;/td&gt;
&lt;td&gt;HPA on CPU/memory over- or under-scales unpredictably&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requests are independent&lt;/td&gt;
&lt;td&gt;A single task often spans multiple round trips carrying shared context&lt;/td&gt;
&lt;td&gt;Session/state handling needs an explicit design, not an assumption&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One replica serves any request&lt;/td&gt;
&lt;td&gt;Tool credentials and context may be tenant-specific&lt;/td&gt;
&lt;td&gt;Routing and isolation boundaries need to be architectural, not incidental&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timeout = failure&lt;/td&gt;
&lt;td&gt;Timeout at 30s might just mean the agent is still reasoning&lt;/td&gt;
&lt;td&gt;Timeout budgets have to be set per tool-call chain, not per request&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  The liveness probe problem, concretely
&lt;/h2&gt;

&lt;p&gt;With &lt;code&gt;periodSeconds: 10&lt;/code&gt; and &lt;code&gt;failureThreshold: 3&lt;/code&gt;, a standard probe kills a pod if &lt;code&gt;/health&lt;/code&gt; doesn't respond within roughly 30 seconds — which is an entirely normal duration for a reasoning step that includes a tool call to a slow downstream API.&lt;/p&gt;

&lt;p&gt;The fix isn't a bigger &lt;code&gt;failureThreshold&lt;/code&gt;. It's separating "the process is alive" from "the process is making forward progress":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;livenessProbe&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;httpGet&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/health&lt;/span&gt;          &lt;span class="c1"&gt;# answers: is the process itself alive?&lt;/span&gt;
    &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8080&lt;/span&gt;
  &lt;span class="na"&gt;periodSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;15&lt;/span&gt;
  &lt;span class="na"&gt;timeoutSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;
  &lt;span class="na"&gt;failureThreshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;      &lt;span class="c1"&gt;# ~45s of true unresponsiveness before restart&lt;/span&gt;

&lt;span class="na"&gt;readinessProbe&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;httpGet&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/ready&lt;/span&gt;            &lt;span class="c1"&gt;# answers: can this pod accept new work right now?&lt;/span&gt;
    &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8080&lt;/span&gt;
  &lt;span class="na"&gt;periodSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;
  &lt;span class="na"&gt;failureThreshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;/health&lt;/code&gt; should do nothing more than confirm the process's event loop is running — never block on the status of an in-progress tool call. &lt;code&gt;/ready&lt;/code&gt; reflects capacity: a pod mid-reasoning can report itself not-ready for new work without being treated as dead. For a request/response service these two questions have the same answer. For an agent, they routinely don't.&lt;/p&gt;




&lt;h2&gt;
  
  
  What changed in the protocol layer
&lt;/h2&gt;

&lt;p&gt;A large part of the awkwardness in running MCP-based agents on Kubernetes came from the original protocol design, which required persistent, pinned sessions between a client and a specific server instance — the opposite of what horizontally scaled infrastructure wants. That constraint forced teams into sticky routing and shared session stores just to keep a conversation coherent across requests.&lt;/p&gt;

&lt;p&gt;The July 2026 MCP specification revision removed the protocol-level session entirely. Any request can now land on any server instance, and applications that need to carry state across calls do it the way HTTP APIs always have — by minting an explicit handle passed back as an ordinary argument, rather than relying on the transport to remember. That single change removes an entire category of infrastructure workaround (sticky routing, pinned sessions, shared session stores) that used to be treated as unavoidable.&lt;/p&gt;

&lt;p&gt;The second shift is A2A (Agent-to-Agent protocol), which reached v1.0 in early 2026. Where MCP governs how an agent talks to tools and data sources, A2A governs how agents talk to each other — and that distinction matters for infrastructure design. A multi-agent system where agents call each other over A2A has different routing, identity, and isolation requirements than one where a single orchestrator calls tools over MCP. Both patterns are in production today.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the cost of getting this wrong compounds quietly
&lt;/h2&gt;

&lt;p&gt;A liveness probe that kills healthy pods doesn't fail loudly — it shows up as an elevated error rate attributed to "model flakiness" or "the tool API being unreliable," and teams spend weeks tuning retry logic against a problem that's actually a probe misconfiguration one layer down. An autoscaler tuned on the wrong signal doesn't fail either — it just runs 30% more replicas than the workload needs, indefinitely, because nobody has a reason to suspect the scaling metric itself.&lt;/p&gt;

&lt;p&gt;Getting the infrastructure layer right doesn't guarantee correct agent behavior, but getting it wrong guarantees you can't tell the difference between an agent that's actually failing and one that's simply being run on infrastructure that wasn't built for it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;The summary covers the core argument. The full article goes deeper on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The full breakdown of each assumption failure — autoscaling signal, session and state design, isolation boundaries, and timeout budget — with the specific production failure mode each one produces&lt;/li&gt;
&lt;li&gt;Why EKS and AKS don't solve the shape problem by default, and how the implementations diverge (IRSA vs. workload identity federation, ALB Ingress vs. Application Gateway) even when the underlying design goal is identical&lt;/li&gt;
&lt;li&gt;The complete naive vs. corrected deployment YAML side-by-side, with the exact reasoning behind each change&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/why-agent-infrastructure-is-its-own-discipline/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=why-agent-infrastructure-is-its-own-discipline" rel="noopener noreferrer"&gt;Why Agent Infrastructure Is Its Own Discipline — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>kubernetes</category>
      <category>devops</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>VPC Peering vs Transit Gateway: Choosing the Right AWS Inter-VPC Connectivity Model</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Tue, 15 Sep 2026 11:35:40 +0000</pubDate>
      <link>https://dev.to/aloknecessary/vpc-peering-vs-transit-gateway-choosing-the-right-aws-inter-vpc-connectivity-model-58i2</link>
      <guid>https://dev.to/aloknecessary/vpc-peering-vs-transit-gateway-choosing-the-right-aws-inter-vpc-connectivity-model-58i2</guid>
      <description>&lt;p&gt;The default assumption most teams start with is that VPC peering is the "simple" option and Transit Gateway is the "enterprise" option you graduate into once you're big enough. That framing is wrong often enough to cause real architectural pain. The actual decision isn't about company size — it's about topology shape.&lt;/p&gt;

&lt;p&gt;Before comparing features, answer one question: does your connectivity requirement look like a &lt;strong&gt;mesh&lt;/strong&gt; or a &lt;strong&gt;hub-and-spoke&lt;/strong&gt;? That single test resolves most of the debate.&lt;/p&gt;




&lt;h2&gt;
  
  
  The math that kills peering at scale
&lt;/h2&gt;

&lt;p&gt;A peering connection is non-transitive — A cannot reach C through B even if both connections exist. For N VPCs in a full mesh, the connections required are &lt;code&gt;N * (N - 1) / 2&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;VPCs&lt;/th&gt;
&lt;th&gt;Peering connections&lt;/th&gt;
&lt;th&gt;TGW attachments&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;190&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 10 VPCs you're managing 45 separate connections, each with its own route table entries on both sides. Transit Gateway needs exactly N attachments regardless of how many VPCs you have.&lt;/p&gt;




&lt;h2&gt;
  
  
  Setting up VPC Peering in Terraform
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_vpc_peering_connection"&lt;/span&gt; &lt;span class="s2"&gt;"app_to_shared"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;peer_vpc_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;shared_services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;auto_accept&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_route"&lt;/span&gt; &lt;span class="s2"&gt;"app_to_shared"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;route_table_id&lt;/span&gt;            &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_route_table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_private&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;destination_cidr_block&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;shared_services&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cidr_block&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_peering_connection_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc_peering_connection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_to_shared&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_route"&lt;/span&gt; &lt;span class="s2"&gt;"shared_to_app"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;route_table_id&lt;/span&gt;            &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_route_table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;shared_private&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;destination_cidr_block&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cidr_block&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_peering_connection_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc_peering_connection&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_to_shared&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both directions need an explicit route — peering doesn't propagate routes automatically. Every new VPC added to a peered mesh means touching route tables on every existing member.&lt;/p&gt;




&lt;h2&gt;
  
  
  Setting up Transit Gateway with segmented routing
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_ec2_transit_gateway"&lt;/span&gt; &lt;span class="s2"&gt;"main"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;description&lt;/span&gt;                     &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"tgw-central-networking"&lt;/span&gt;
  &lt;span class="nx"&gt;default_route_table_association&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"disable"&lt;/span&gt;
  &lt;span class="nx"&gt;default_route_table_propagation&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"disable"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_ec2_transit_gateway_vpc_attachment"&lt;/span&gt; &lt;span class="s2"&gt;"app"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;subnet_ids&lt;/span&gt;         &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;aws_subnet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_private_a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;aws_subnet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_private_b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="nx"&gt;transit_gateway_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_ec2_transit_gateway&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt;             &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;disable&lt;/code&gt; pair on both defaults is deliberate — it enables route table segmentation instead of every attachment automatically seeing every other attachment.&lt;/p&gt;




&lt;h2&gt;
  
  
  Route table segmentation: the feature peering can't replicate
&lt;/h2&gt;

&lt;p&gt;A single Transit Gateway can have multiple route tables. Each attachment associates with one, giving you actual network segmentation without touching security groups or NACLs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TGW Route Table: "production"
├── app-vpc attachment    (associated)
├── shared-vpc attachment (associated, propagated)
└── sandbox-vpc attachment (NOT associated — isolated by design)

TGW Route Table: "sandbox"
├── sandbox-vpc attachment (associated)
└── shared-vpc attachment  (associated, propagated)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the pattern most multi-account AWS Organizations setups converge on: one Transit Gateway per region, a small number of route tables representing trust boundaries, and VPCs associated into the appropriate table.&lt;/p&gt;




&lt;h2&gt;
  
  
  Cross-account via AWS RAM
&lt;/h2&gt;

&lt;p&gt;Transit Gateway solves the multi-account case cleanly through Resource Access Manager — one account owns the TGW, shares it via RAM, and spoke accounts create their own attachments:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_ram_resource_share"&lt;/span&gt; &lt;span class="s2"&gt;"tgw_share"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;                      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"tgw-network-share"&lt;/span&gt;
  &lt;span class="nx"&gt;allow_external_principals&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_ram_resource_association"&lt;/span&gt; &lt;span class="s2"&gt;"tgw"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;resource_arn&lt;/span&gt;       &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_ec2_transit_gateway&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
  &lt;span class="nx"&gt;resource_share_arn&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_ram_resource_share&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tgw_share&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_ram_principal_association"&lt;/span&gt; &lt;span class="s2"&gt;"spoke_account"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;principal&lt;/span&gt;          &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"222233334444"&lt;/span&gt;
  &lt;span class="nx"&gt;resource_share_arn&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_ram_resource_share&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tgw_share&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arn&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Peering works cross-account too, but the N(N-1)/2 problem gets worse because you're also managing acceptance workflows and cross-account IAM permissions for every connection.&lt;/p&gt;




&lt;h2&gt;
  
  
  DNS and security group gotchas
&lt;/h2&gt;

&lt;p&gt;Routing connectivity doesn't automatically mean DNS works across it. For peering, DNS resolution support must be enabled explicitly on both sides — without it, private hosted zone records resolve to public IPs or fail entirely, which looks identical to a routing problem.&lt;/p&gt;

&lt;p&gt;Security group referencing across peering connections is also limited — most cross-account or cross-region setups fall back to CIDR-based rules rather than SG-to-SG references. Transit Gateway doesn't change this constraint.&lt;/p&gt;




&lt;h2&gt;
  
  
  Troubleshooting in priority order
&lt;/h2&gt;

&lt;p&gt;When "VPC-A can't reach VPC-B" comes in:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Route table on both sides&lt;/strong&gt; — both directions need explicit routes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connection/attachment state&lt;/strong&gt; — &lt;code&gt;pending-acceptance&lt;/code&gt; on peering, non-&lt;code&gt;available&lt;/code&gt; TGW attachment&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Security group and NACL on both ends&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DNS resolution settings&lt;/strong&gt; — easy to mistake for a routing failure&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TGW route table association&lt;/strong&gt; — an attachment not associated with the expected route table silently fails&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;The full post covers everything above in depth, plus:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The DNS resolution strategy for hub-and-spoke setups using Route 53 Resolver endpoints — why it needs to be planned alongside routing, not retrofitted&lt;/li&gt;
&lt;li&gt;The full side-by-side comparison table across all dimensions (routing model, cost, cross-region, segmentation, best fit)&lt;/li&gt;
&lt;li&gt;The migration path from peering to Transit Gateway — running both in parallel, using most-specific-match routing for a controlled cutover without an all-or-nothing switch&lt;/li&gt;
&lt;li&gt;When peering is genuinely the right answer and why you shouldn't add TGW complexity to a two-VPC problem&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/aws-vpc-peering-vs-transit-gateway/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=aws-vpc-peering-vs-transit-gateway" rel="noopener noreferrer"&gt;VPC Peering vs Transit Gateway — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>networking</category>
      <category>cloud</category>
      <category>terraform</category>
    </item>
    <item>
      <title>CI to GitOps Handoff: Argo Image Updater vs. CI-Writes-the-Commit vs. Kargo</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:23:01 +0000</pubDate>
      <link>https://dev.to/aloknecessary/ci-to-gitops-handoff-argo-image-updater-vs-ci-writes-the-commit-vs-kargo-468</link>
      <guid>https://dev.to/aloknecessary/ci-to-gitops-handoff-argo-image-updater-vs-ci-writes-the-commit-vs-kargo-468</guid>
      <description>&lt;p&gt;ArgoCD only reconciles what's in Git. A CI pipeline builds an image, pushes it to a registry — and then what? Something has to turn "a new image exists" into "a commit exists in the GitOps repo referencing it." That handoff is where a surprising number of GitOps implementations quietly go wrong: either it's done manually, or it's automated in a way that reintroduces the exact imperative-deploy risk GitOps was supposed to remove.&lt;/p&gt;

&lt;p&gt;Three approaches cover almost every real setup. The choice isn't just operational preference — it interacts directly with your repo structure and how much promotion-stage enforcement you actually need.&lt;/p&gt;




&lt;h2&gt;
  
  
  The structural difference in one picture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Approach 1 — Argo Image Updater
  CI: build → push image ──────────────► registry
                                              │ (polled every 1-2 min)
                                              ▼
                                    Image Updater controller
                                              │ writes tag
                                              ▼
                                     GitOps repo commit ──► ArgoCD syncs

Approach 2 — CI writes the commit
  CI: build → push image ──► registry
       │
       └── same workflow, next job: bump tag, commit, push ──► GitOps repo ──► ArgoCD syncs

Approach 3 — Kargo
  CI: build → push image ──► registry
                                 │
                                 ▼
                            Kargo Warehouse (watches registry)
                                 │
                                 ▼
                          Stage: staging ──(verify)──► Stage: prod
                                 │                          │
                                 ▼                          ▼
                          GitOps repo commit          GitOps repo commit
                          (staging path)               (prod path, only if
                                                         staging verified)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Approaches 1 and 2 go straight from "new image" to "commit" with no concept of one environment gating another. Kargo inserts an explicit, enforced step between them. Everything below is detail on why you'd pick one over another.&lt;/p&gt;




&lt;h2&gt;
  
  
  Argo Image Updater
&lt;/h2&gt;

&lt;p&gt;Image Updater runs alongside ArgoCD, polls registries for new tags matching a pattern, and writes the updated tag back into the GitOps repo. CI's job stays exactly "build, test, push image" — no GitOps-repo awareness required.&lt;/p&gt;

&lt;p&gt;Configuration lives as annotations on the Application itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;argocd-image-updater.argoproj.io/image-list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout=111122223333.dkr.ecr.us-east-1.amazonaws.com/checkout-service&lt;/span&gt;
  &lt;span class="na"&gt;argocd-image-updater.argoproj.io/checkout.update-strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;semver&lt;/span&gt;
  &lt;span class="na"&gt;argocd-image-updater.argoproj.io/checkout.allow-tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;regexp:^v[0-9]+\.[0-9]+\.[0-9]+$'&lt;/span&gt;
  &lt;span class="na"&gt;argocd-image-updater.argoproj.io/write-back-method&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;git&lt;/span&gt;
  &lt;span class="na"&gt;argocd-image-updater.argoproj.io/git-branch&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;main&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;write-back-method: git&lt;/code&gt; setting is the important one. The alternative (&lt;code&gt;argocd&lt;/code&gt;) writes to the Application object only, meaning the running state and the Git repo can silently diverge — undermining the entire premise of GitOps as the source of truth. If you use Image Updater, &lt;code&gt;git&lt;/code&gt; write-back is the only mode consistent with a real GitOps model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trade-off:&lt;/strong&gt; polling-based (1-2 minute delay), and no native concept of promotion stages. Getting a build to flow dev → staging → prod with any gate means bolting that logic on separately — which is exactly the gap Kargo exists to close.&lt;/p&gt;




&lt;h2&gt;
  
  
  CI writes the manifest commit
&lt;/h2&gt;

&lt;p&gt;No separate controller. The CI job itself pushes the commit after a successful build. The auth chain here deserves the same scrutiny as cloud credentials — a long-lived PAT with write access to the GitOps repo is the wrong answer for the same reasons a long-lived AWS access key is wrong.&lt;/p&gt;

&lt;p&gt;The right pattern: a GitHub App scoped to exactly the one GitOps repo, with an installation token minted per workflow run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;update-gitops-repo&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;needs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;build&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Generate GitHub App token for GitOps repo&lt;/span&gt;
        &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;app-token&lt;/span&gt;
        &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/create-github-app-token@v1&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;app-id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ vars.GITOPS_BOT_APP_ID }}&lt;/span&gt;
          &lt;span class="na"&gt;private-key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.GITOPS_BOT_PRIVATE_KEY }}&lt;/span&gt;
          &lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;example-org&lt;/span&gt;
          &lt;span class="na"&gt;repositories&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-team-gitops&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v6&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;repository&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;example-org/checkout-team-gitops&lt;/span&gt;
          &lt;span class="na"&gt;token&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ steps.app-token.outputs.token }}&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Bump image tag&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;yq -i '.image.tag = "${{ github.sha }}"' apps/checkout-service/prod/values.yaml&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Commit and push&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;git config user.name "gitops-bot"&lt;/span&gt;
          &lt;span class="s"&gt;git config user.email "gitops-bot@example.org"&lt;/span&gt;
          &lt;span class="s"&gt;git commit -am "checkout-service: bump to ${{ github.sha }}"&lt;/span&gt;
          &lt;span class="s"&gt;git push&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The token is scoped to one installation (one repo) and expires roughly an hour after issuance — no long-lived credential sitting in a secrets store.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trade-off:&lt;/strong&gt; no delay, no extra controller. But CI now has push access to the thing that controls production. A GitHub App token scoped org-wide instead of to one repo turns a compromised CI secret from a one-service incident into an org-wide one. The scoping isn't optional hardening — it's the difference between those two outcomes.&lt;/p&gt;




&lt;h2&gt;
  
  
  Kargo
&lt;/h2&gt;

&lt;p&gt;Image Updater and CI-writes-the-commit both update one Application at a time with no native concept of "this build must pass staging before prod." Kargo models that as a first-class API.&lt;/p&gt;

&lt;p&gt;Two core resources: a &lt;code&gt;Warehouse&lt;/code&gt; (watches for new artifacts, like Image Updater) and a &lt;code&gt;Stage&lt;/code&gt; (one step in a promotion pipeline). The key field:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kargo.akuity.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Stage&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;prod&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-team&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;requestedFreight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;origin&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Warehouse&lt;/span&gt;
        &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-service&lt;/span&gt;
      &lt;span class="na"&gt;sources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;stages&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;staging&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;   &lt;span class="c1"&gt;# prod can only promote freight that passed through staging&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;sources: stages: [staging]&lt;/code&gt; means prod is structurally incapable of promoting an image that hasn't gone through staging — enforced by Kargo, not by convention. Verification gates (smoke tests, analysis runs) attach to each Stage, and Kargo tracks exactly which artifact made it through which Stages.&lt;/p&gt;

&lt;p&gt;The distinction from bolting a manual approval onto Image Updater: the gate is a property of the Stage graph itself, checked by Kargo before it constructs a promotion for the next Stage. Not a step someone has to remember to run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trade-off:&lt;/strong&gt; another controller with its own resource model on top of ArgoCD's. Earns its cost specifically when multi-stage promotion with gated, auditable progression is a real requirement — not as a default for every service.&lt;/p&gt;




&lt;h2&gt;
  
  
  Decision framework
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Recommended approach&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single environment, or environments promote independently&lt;/td&gt;
&lt;td&gt;Argo Image Updater — simplest, no CI changes needed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Simple promotion logic already visible in CI (merge to main = staging, tag = prod)&lt;/td&gt;
&lt;td&gt;CI writes the commit — keeps logic in one place, no extra controller&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard requirement that prod only receives artifacts that passed staging with automated verification&lt;/td&gt;
&lt;td&gt;Kargo — this is the exact problem it's built to solve&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regulated workloads needing an auditable promotion trail as a first-class object&lt;/td&gt;
&lt;td&gt;Kargo — Stage/Freight model gives you that, not something reconstructed from CI logs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These aren't mutually exclusive across a platform — a team running mostly independent services on Image Updater might still put its most compliance-sensitive service through Kargo specifically.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;The summary covers the shape of each approach and the key trade-offs. The full article goes deeper on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The complete Image Updater Application manifest — &lt;code&gt;image-list&lt;/code&gt;, &lt;code&gt;update-strategy&lt;/code&gt;, &lt;code&gt;allow-tags&lt;/code&gt;, and &lt;code&gt;write-back-method&lt;/code&gt; in context, not just the annotation snippet&lt;/li&gt;
&lt;li&gt;The full two-job GitHub Actions workflow (build + update-gitops-repo), including the step that verifies the App token &lt;em&gt;fails&lt;/em&gt; against repos it shouldn't reach — confirming blast radius, not just that it works&lt;/li&gt;
&lt;li&gt;The complete Kargo &lt;code&gt;Warehouse&lt;/code&gt; + &lt;code&gt;Stage&lt;/code&gt; YAML including the &lt;code&gt;verification&lt;/code&gt; block that wires an &lt;code&gt;AnalysisTemplate&lt;/code&gt; smoke test to the staging Stage — the part that makes "staging verified" mean something concrete&lt;/li&gt;
&lt;li&gt;Why a Matrix generator with a too-broad cluster label selector connects directly to the auth chain here: AppProject destinations as the backstop when the ApplicationSet generates more than intended&lt;/li&gt;
&lt;li&gt;Article 4 preview: secrets management (Sealed Secrets, External Secrets Operator, Vault plugin) — what each actually protects against and what it doesn't&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/ci-to-gitops-handoff/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=ci-to-gitops-handoff" rel="noopener noreferrer"&gt;CI to GitOps Handoff: Argo Image Updater vs. CI-Writes-the-Commit vs. Kargo — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>githubactions</category>
      <category>argocd</category>
      <category>devops</category>
    </item>
    <item>
      <title>Always Encrypt in Transit: The Gap Between TLS Everywhere and Actual Transport Security</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Tue, 08 Sep 2026 07:26:21 +0000</pubDate>
      <link>https://dev.to/aloknecessary/always-encrypt-in-transit-the-gap-between-tls-everywhere-and-actual-transport-security-2mib</link>
      <guid>https://dev.to/aloknecessary/always-encrypt-in-transit-the-gap-between-tls-everywhere-and-actual-transport-security-2mib</guid>
      <description>&lt;p&gt;"We have TLS everywhere." It appears in every architecture review and every compliance questionnaire. The gap is not in the intention — it is in the implementation. Most systems that claim TLS everywhere have TLS at the edge and plaintext everywhere else: the ingress controller to the pod is HTTP, pod-to-pod traffic is unencrypted, and the application-to-database connection string never had &lt;code&gt;Encrypt=True&lt;/code&gt; or &lt;code&gt;sslmode=require&lt;/code&gt; set. None of these gaps appear in architecture diagrams.&lt;/p&gt;

&lt;p&gt;This post maps where TLS is actually absent, why the gaps exist, and the implementation path that closes them — without defaulting to "add a service mesh" as the answer to every transport security question.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Four Layers Where TLS Is Actually Absent
&lt;/h2&gt;

&lt;p&gt;Before solutions, the precise inventory of where plaintext exists in systems claiming TLS everywhere:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Scenario: "Secure" architecture
  ✅ HTTPS on the load balancer
  ✅ Private subnets, no public IPs
  ✅ Security groups restricting access
  ❌ Ingress Controller → Pod: plaintext HTTP
  ❌ Pod → Pod: plaintext HTTP/gRPC
  ❌ Application → RDS: plaintext TCP
  ❌ Cluster infrastructure certs: expire annually, no automated rotation

What "TLS everywhere" actually means in this architecture:
  TLS on the edge. Plaintext the rest of the way.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Layer 1 — Ingress-to-pod:&lt;/strong&gt; nginx-ingress and Traefik terminate TLS at the edge and forward plain HTTP to backend pods by default. The gap is invisible in architecture diagrams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2 — Pod-to-pod:&lt;/strong&gt; Kubernetes does not encrypt data plane traffic. A compromised pod on a shared node can observe plaintext traffic from other pods on the same node — the VPC boundary is not the relevant perimeter here, the pod boundary is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 — Application-to-database:&lt;/strong&gt; Enabling TLS on RDS makes the database &lt;em&gt;capable&lt;/em&gt; of TLS connections. It does not enforce them. A connection string without &lt;code&gt;sslmode=require&lt;/code&gt; or &lt;code&gt;Encrypt=True&lt;/code&gt; connects over plaintext regardless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 4 — Cluster infrastructure certs:&lt;/strong&gt; kubeadm cluster CA certificates expire after one year by default with no automated renewal. A missed rotation takes down the entire cluster — not an application outage, a cluster-level failure where &lt;code&gt;kubectl&lt;/code&gt; stops working entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  Termination Architecture: Where the Decision Gets Made
&lt;/h2&gt;

&lt;p&gt;Where TLS terminates determines which traffic is encrypted. Three patterns:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Edge termination only&lt;/strong&gt; — TLS terminates at the ingress controller, all internal traffic is plaintext. Legitimate for single-tenant clusters with no regulated data. Not legitimate to call "TLS everywhere."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Re-encryption&lt;/strong&gt; — the ingress controller terminates external TLS and re-encrypts before forwarding to the pod. Closes the ingress-to-pod gap without a service mesh:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# NGINX Ingress: re-encrypt to backend over TLS&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;networking.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Ingress&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;secure-ingress&lt;/span&gt;
  &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;nginx.ingress.kubernetes.io/backend-protocol&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HTTPS"&lt;/span&gt;
    &lt;span class="na"&gt;nginx.ingress.kubernetes.io/ssl-redirect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true"&lt;/span&gt;
    &lt;span class="na"&gt;cert-manager.io/cluster-issuer&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;letsencrypt-prod"&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;ingressClassName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
  &lt;span class="na"&gt;tls&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;hosts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;api.yourdomain.com&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;secretName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api-tls-secret&lt;/span&gt;
  &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;host&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api.yourdomain.com&lt;/span&gt;
    &lt;span class="na"&gt;http&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/&lt;/span&gt;
        &lt;span class="na"&gt;pathType&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Prefix&lt;/span&gt;
        &lt;span class="na"&gt;backend&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;service&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api-service&lt;/span&gt;
            &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;number&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8443&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;End-to-end mTLS via service mesh&lt;/strong&gt; — the mesh intercepts all pod-to-pod traffic and wraps it in mTLS. Strongest posture, highest operational overhead. Justified in regulated multi-tenant clusters; disproportionate for smaller service estates.&lt;/p&gt;




&lt;h2&gt;
  
  
  cert-manager: The Implementation That Actually Automates It
&lt;/h2&gt;

&lt;p&gt;cert-manager handles issuance, renewal, and storage in Kubernetes Secrets with zero manual steps in the rotation path. One ClusterIssuer and one Certificate resource per service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# ClusterIssuer: Let's Encrypt production&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cert-manager.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterIssuer&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;letsencrypt-prod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;acme&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://acme-v02.api.letsencrypt.org/directory&lt;/span&gt;
    &lt;span class="na"&gt;email&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;platform@yourdomain.com&lt;/span&gt;
    &lt;span class="na"&gt;privateKeySecretRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;letsencrypt-prod-account-key&lt;/span&gt;
    &lt;span class="na"&gt;solvers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;http01&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;ingress&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;ingressClassName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="c1"&gt;# Certificate: 90-day validity, auto-renewed at 60 days remaining&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cert-manager.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Certificate&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api-certificate&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;secretName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api-tls-secret&lt;/span&gt;
  &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;2160h&lt;/span&gt;
  &lt;span class="na"&gt;renewBefore&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;720h&lt;/span&gt;
  &lt;span class="na"&gt;dnsNames&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;api.yourdomain.com&lt;/span&gt;
  &lt;span class="na"&gt;issuerRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;letsencrypt-prod&lt;/span&gt;
    &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterIssuer&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For internal service-to-service certificates, use a self-signed CA issuer with 24-hour validity — cert-manager makes short-lived internal certs operationally trivial, and a 24-hour certificate has a dramatically smaller blast radius than a one-year certificate if the key is compromised:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Internal service certificate — 24h validity, rotated automatically&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;cert-manager.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Certificate&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;order-service-cert&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;secretName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;order-service-tls&lt;/span&gt;
  &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;24h&lt;/span&gt;
  &lt;span class="na"&gt;renewBefore&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;8h&lt;/span&gt;
  &lt;span class="na"&gt;dnsNames&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;order-service.production.svc.cluster.local&lt;/span&gt;
  &lt;span class="na"&gt;issuerRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;internal-ca-issuer&lt;/span&gt;
    &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterIssuer&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Certificate Lifecycle: Where Security Controls Become Outages
&lt;/h2&gt;

&lt;p&gt;The failure mode is consistent: a certificate is issued, configured, and forgotten. The alert that was supposed to fire before expiry never got configured, or fired into a channel nobody watches. The certificate expires. The service goes down. The postmortem recommends "better monitoring" — until the same thing happens eighteen months later with a different certificate.&lt;/p&gt;

&lt;p&gt;Alert on time-to-expiry, not on expiry. A Prometheus alert that fires at 15 days remaining gives the team time to investigate before the outage. An alert that fires when the certificate has already expired is a notification of an ongoing incident:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;groups&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;certificate-expiry&lt;/span&gt;
  &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;alert&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;CertificateExpiringIn15Days&lt;/span&gt;
    &lt;span class="na"&gt;expr&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
      &lt;span class="s"&gt;certmanager_certificate_expiration_timestamp_seconds&lt;/span&gt;
        &lt;span class="s"&gt;- time() &amp;lt; (15 * 24 * 3600)&lt;/span&gt;
    &lt;span class="na"&gt;for&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1h&lt;/span&gt;
    &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;warning&lt;/span&gt;
    &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Certificate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;expiring&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;soon"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;alert&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;CertificateExpired&lt;/span&gt;
    &lt;span class="na"&gt;expr&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
      &lt;span class="s"&gt;certmanager_certificate_expiration_timestamp_seconds - time() &amp;lt; 0&lt;/span&gt;
    &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;critical&lt;/span&gt;
    &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Certificate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;has&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;EXPIRED"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both alerts are required. Most teams have only the expired alert — which is a notification of an ongoing incident, not a prevention.&lt;/p&gt;




&lt;h2&gt;
  
  
  mTLS Without a Service Mesh
&lt;/h2&gt;

&lt;p&gt;For fewer than twenty services, cert-manager internal CA plus application-level TLS provides mutual authentication without sidecar injection or a mesh control plane. The application handles the TLS handshake directly using certificates cert-manager issues and rotates automatically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="c1"&gt;// .NET: present a client certificate and verify the server against the internal CA&lt;/span&gt;
&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;clientCert&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;X509Certificate2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;CreateFromPemFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;clientCertPath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;clientKeyPath&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;caCert&lt;/span&gt;     &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;X509Certificate2&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;caCertPath&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;handler&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nf"&gt;HttpClientHandler&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="n"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ClientCertificates&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;clientCert&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="n"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ServerCertificateCustomValidationCallback&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cert&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;!.&lt;/span&gt;&lt;span class="n"&gt;ChainPolicy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TrustMode&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="n"&gt;X509ChainTrustMode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CustomRootTrust&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="n"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ChainPolicy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CustomTrustStore&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;caCert&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Build&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cert&lt;/span&gt;&lt;span class="p"&gt;!);&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The certificates are mounted from Kubernetes Secrets that cert-manager manages. When cert-manager rotates the Secret, the volume mount is updated in place — no application restart required.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;This summary covers the core gaps, termination architecture patterns, cert-manager setup, and the mTLS-without-mesh approach. The full article includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The complete database enforcement gap — &lt;code&gt;rds.force_ssl&lt;/code&gt;, connection string &lt;code&gt;Encrypt=True&lt;/code&gt;, and why enabling TLS on the DB is not the same as enforcing it&lt;/li&gt;
&lt;li&gt;The VPC isolation misconception in full — why private subnets don't substitute for encryption&lt;/li&gt;
&lt;li&gt;The "Kubernetes encrypts cluster traffic" misconception — what it actually encrypts vs. what it doesn't&lt;/li&gt;
&lt;li&gt;Full decision framework mapping every traffic path to the right termination architecture&lt;/li&gt;
&lt;li&gt;When regulated environments (PCI-DSS, HIPAA) make full end-to-end TLS non-negotiable&lt;/li&gt;
&lt;li&gt;Eight key takeaways covering every layer of the in-transit security posture&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/always-encrypt-in-transit/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=always-encrypt-in-transit" rel="noopener noreferrer"&gt;Always Encrypt in Transit: The Gap Between TLS Everywhere and Actual Transport Security — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>kubernetes</category>
      <category>cloudnative</category>
      <category>devops</category>
    </item>
    <item>
      <title>AWS MGN Architecture: How Continuous Replication Actually Works</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Thu, 03 Sep 2026 07:29:32 +0000</pubDate>
      <link>https://dev.to/aloknecessary/aws-mgn-architecture-how-continuous-replication-actually-works-47f4</link>
      <guid>https://dev.to/aloknecessary/aws-mgn-architecture-how-continuous-replication-actually-works-47f4</guid>
      <description>&lt;p&gt;The first time I ran an AWS MGN migration, I trusted the console's green "Healthy" status without understanding what it actually meant. It took a stalled replication and a confusing lag metric mid-cutover to make me go back and learn the mechanics underneath the dashboard. This post is the breakdown I wish I'd had before that first wave.&lt;/p&gt;

&lt;p&gt;MGN does one job: continuous, block-level replication of a source server's disks to a staging area in your target AWS account, so that when you cut over, you launch a fully synced, bootable EC2 instance instead of doing a one-time data copy and hoping nothing changed since. The distinction that matters is &lt;strong&gt;block-level, not file-level&lt;/strong&gt; — MGN doesn't care about your filesystem or application state, which is what allows it to keep a target in near-real-time sync with a live, running source server throughout the migration.&lt;/p&gt;




&lt;h2&gt;
  
  
  The four components, in order
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The replication agent&lt;/strong&gt; — installed directly on the source server, hooks into the OS's disk I/O path at the kernel level. On Linux this means a kernel-level block device reader; on Windows, a filter driver. This low-level hook is why a kernel version mismatch on the source isn't a minor footnote — it's a direct threat to replication itself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;python3 ./aws-replication-installer-init.py &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--region&lt;/span&gt; ap-south-1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--aws-access-key-id&lt;/span&gt; &amp;lt;access-key-id&amp;gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--aws-secret-access-key&lt;/span&gt; &amp;lt;secret-access-key&amp;gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--no-prompt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;--no-prompt&lt;/code&gt; flag matters for scripted wave installations — without it, the installer pauses for disk selection confirmation and silently stalls automation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The staging area subnet&lt;/strong&gt; — a designated subnet in the target AWS account where MGN provisions its replication infrastructure. Needs outbound connectivity and enough IOPS headroom on EBS to keep up with the source's write rate. Undersizing this subnet or placing it behind restrictive route tables is a common cause of replication that starts but never stabilizes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The replication server and EBS volumes&lt;/strong&gt; — for each source server, MGN provisions a temporary EC2 instance in the staging subnet whose only job is to receive the block-level stream and write it to EBS volumes that mirror the source's disk layout. These EBS volumes are the actual replica — they're what gets snapshotted at cutover.&lt;/p&gt;

&lt;p&gt;The cost model people get wrong: you're paying for these replication servers and EBS volumes for the &lt;strong&gt;entire duration&lt;/strong&gt; of the migration project, not just at cutover. A wave that sits in "replicating, not yet cut over" for six weeks accumulates staging infrastructure cost for all six weeks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Continuous sync and the lag metric&lt;/strong&gt; — once initial replication completes, MGN switches to continuous incremental sync: every write on the source is captured and streamed as it happens. The metric to watch is &lt;strong&gt;replication lag&lt;/strong&gt; — how far behind the target EBS volumes are from the live source. Lag that climbs rather than staying flat is a leading indicator of a cutover that will either take much longer than expected or launch a target instance further behind the source than your rollback tolerance allows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws mgn describe-source-servers &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--region&lt;/span&gt; ap-south-1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s1"&gt;'items[*].{Server:sourceServerID, LagDuration:dataReplicationInfo.lagDuration, State:dataReplicationInfo.dataReplicationState}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; table
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use this as a scriptable pre-cutover gate rather than eyeballing a color indicator in the console.&lt;/p&gt;




&lt;h2&gt;
  
  
  What cutover actually triggers
&lt;/h2&gt;

&lt;p&gt;Cutover is not a data operation — the data is already synced. What it actually does:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Takes a final snapshot of the target EBS volumes at the moment of cutover&lt;/li&gt;
&lt;li&gt;Launches a new EC2 instance from that snapshot using the launch template you've configured&lt;/li&gt;
&lt;li&gt;Runs the MGN post-launch conversion process — driver injection, network configuration, and OS-level changes needed to make a disk image that ran on a different hypervisor boot correctly under AWS's Nitro hypervisor&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 3 is where kernel driver failures happen. The conversion process needs to inject or activate the correct network driver (&lt;code&gt;ena&lt;/code&gt;) for the instance to have connectivity after launch. If the source server's kernel doesn't have the module available or the boot configuration doesn't reference it correctly, the instance can come up without network connectivity — or fail to boot entirely.&lt;/p&gt;




&lt;h2&gt;
  
  
  Launch settings most people under-invest in
&lt;/h2&gt;

&lt;p&gt;Every source server has its own launch template in MGN. Worth setting deliberately before cutover, not left at defaults:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Target instance type&lt;/strong&gt; — set explicitly for anything performance-sensitive rather than trusting MGN's right-size recommendation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subnet and security groups&lt;/strong&gt; — should be decided by your VPC CIDR plan, not chosen ad hoc at cutover time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IAM instance profile&lt;/strong&gt; — doesn't carry over from the source server; must be set in the launch template explicitly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test launch&lt;/strong&gt; — MGN supports launching a test instance from current replication state without ending replication. Always run at least one test launch before the real cutover. Skipping this is the single most common way teams discover a launch-time problem during the actual cutover window instead of during a safe rehearsal&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Replication settings that matter under load
&lt;/h2&gt;

&lt;p&gt;Three settings worth understanding rather than leaving on defaults for source servers with high write throughput:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bandwidth throttling&lt;/strong&gt; — MGN doesn't cap replication bandwidth by default, which can compete with production traffic on a busy source server. Set a maximum throughput per source server for anything actively serving production load during the replication window&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Staging volume type&lt;/strong&gt; — the default EBS volume type for staging can itself become the bottleneck for high-IOPS sources, showing up as climbing lag even when network bandwidth is fine. Rule this in or out separately from the network-bandwidth cause when diagnosing a server that won't stabilize&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Point-in-time (PIT) snapshots&lt;/strong&gt; — MGN can retain a rolling window of snapshots of the staging volumes, giving you a recovery point earlier than "right now" if you need to launch from a known-good state. Configure this under &lt;strong&gt;Replication settings → Point-in-time snapshots&lt;/strong&gt; in the MGN console before replication starts — snapshots only accumulate from the point the setting is enabled&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why continuous replication beats snapshot-and-copy
&lt;/h2&gt;

&lt;p&gt;The snapshot approach has an unavoidable trade-off: the longer the gap between snapshot and cutover, the more source-side changes are missing from the target. Closing that gap means either accepting data loss or taking the source offline for the copy window.&lt;/p&gt;

&lt;p&gt;MGN's continuous model removes that trade-off entirely. The target stays within seconds of the source right up until cutover, and the source never has to go offline for the migration itself — only for the brief cutover window when traffic is actually redirected. For anything with a real availability requirement, this is the entire reason to use MGN over a manual export/import process.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;The summary covers the core mechanics. The full article goes deeper on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The troubleshooting priority order for "these are definitely connected but nothing works" tickets — route tables, attachment state, security layers, DNS — and why checking in that sequence resolves the large majority of cases without reasoning about the full topology from scratch&lt;/li&gt;
&lt;li&gt;The specific kernel driver failure mode that surfaces when migrating Azure-sourced Ubuntu images to AWS, why it manifests the way it does, and the fix&lt;/li&gt;
&lt;li&gt;The full cost model for staging infrastructure across a multi-week wave, with the specific line items to flag before the wave kicks off&lt;/li&gt;
&lt;li&gt;CIDR planning considerations for the staging subnet relative to the rest of the target VPC&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/aws-mgn-architecture-continuous-replication/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=aws-mgn-architecture-continuous-replication" rel="noopener noreferrer"&gt;AWS MGN Architecture: How Continuous Replication Actually Works — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>cloud</category>
      <category>aws</category>
      <category>migration</category>
    </item>
    <item>
      <title>Enforcing Modular Monolith Boundaries in .NET: NDepend, Parallel Pipelines, and the Architecture That Holds</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Mon, 31 Aug 2026 06:03:32 +0000</pubDate>
      <link>https://dev.to/aloknecessary/enforcing-modular-monolith-boundaries-in-net-ndepend-parallel-pipelines-and-the-architecture-37e1</link>
      <guid>https://dev.to/aloknecessary/enforcing-modular-monolith-boundaries-in-net-ndepend-parallel-pipelines-and-the-architecture-37e1</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;A modular monolith without enforcement is not an architecture — it is a monolith with good intentions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Most teams skip the modular monolith and jump straight to microservices. The ones that do attempt a modular monolith rely on convention — "don't cross module boundaries" — which fails the moment deadlines hit.&lt;/p&gt;

&lt;p&gt;The difference between a well-structured modular monolith and a mess is whether boundaries are maintained by &lt;strong&gt;tooling&lt;/strong&gt; or by convention.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Solution Structure
&lt;/h2&gt;

&lt;p&gt;Each module is a pair of .NET projects:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;src/Modules/
  Orders/
    YourApp.Orders/              ← internal: domain, application, infrastructure
    YourApp.Orders.Contracts/    ← public: DTOs, interfaces, events
  Payments/
    YourApp.Payments/
    YourApp.Payments.Contracts/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The rule&lt;/strong&gt;: modules may only reference each other's &lt;code&gt;*.Contracts&lt;/code&gt; projects. The compiler enforces this physically — no project reference means no type access.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four Layers of Enforcement
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Compiler&lt;/strong&gt; — project references prevent cross-module type access&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NetArchTest&lt;/strong&gt; — architecture tests fail the build on namespace-level violations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;NDepend CQLinq&lt;/strong&gt; — catches dependency cycles and coupling the compiler can't see&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality Gates&lt;/strong&gt; — block PRs that introduce &lt;em&gt;new&lt;/em&gt; boundary violations&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Module-Scoped Data
&lt;/h2&gt;

&lt;p&gt;Each module owns a dedicated &lt;code&gt;DbContext&lt;/code&gt; with a schema prefix (&lt;code&gt;orders.*&lt;/code&gt;, &lt;code&gt;payments.*&lt;/code&gt;). No module queries another module's tables.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross-Module Communication
&lt;/h2&gt;

&lt;p&gt;Modules communicate via MediatR in-process events. Orders publishes &lt;code&gt;OrderPlaced&lt;/code&gt;; Payments subscribes — without Orders knowing Payments exists.&lt;/p&gt;

&lt;p&gt;This is also the extraction seam: when you eventually extract a module into a service, MediatR becomes a message broker. The event contract stays the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  Parallel CI
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;matrix&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;module&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Orders&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;Payments&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;Inventory&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;fail-fast&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each module's tests run in parallel. CI time scales with the slowest module, not the total count.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Extraction Path
&lt;/h2&gt;

&lt;p&gt;When a module genuinely needs independence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add outbox table → publish to real broker&lt;/li&gt;
&lt;li&gt;Replace MediatR handlers with broker consumers&lt;/li&gt;
&lt;li&gt;Deploy module as separate service&lt;/li&gt;
&lt;li&gt;Publish &lt;code&gt;*.Contracts&lt;/code&gt; as NuGet package&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The boundary was already clean. Extraction is a deployment change, not a redesign.&lt;/p&gt;




&lt;p&gt;The full post covers NDepend CQLinq rule examples, Quality Gate configuration, GitHub Actions pipeline YAML, test isolation patterns, and a production checklist.&lt;/p&gt;

&lt;p&gt;👉 &lt;a href="https://aloknecessary.in/blogs/modular-monolith-dotnet-ndepend/?utm_source=devto&amp;amp;utm_medium=crosspost&amp;amp;utm_campaign=modular-monolith-dotnet" rel="noopener noreferrer"&gt;Read the complete implementation guide&lt;/a&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>devops</category>
      <category>dotnet</category>
      <category>csharp</category>
    </item>
    <item>
      <title>AWS VPC Networking Fundamentals: VPCs, Subnets, CIDR, Route Tables, IGW, and NAT Gateways</title>
      <dc:creator>Alok Ranjan Daftuar</dc:creator>
      <pubDate>Fri, 28 Aug 2026 03:49:00 +0000</pubDate>
      <link>https://dev.to/aloknecessary/aws-vpc-networking-fundamentals-vpcs-subnets-cidr-route-tables-igw-and-nat-gateways-19h1</link>
      <guid>https://dev.to/aloknecessary/aws-vpc-networking-fundamentals-vpcs-subnets-cidr-route-tables-igw-and-nat-gateways-19h1</guid>
      <description>&lt;p&gt;If you've provisioned a VPC from a Terraform module without fully internalising what each piece is doing, that's fine — right up until something breaks. An instance that should be reachable isn't. A private instance can't pull a package update. And you're left checking five different resources with no clear mental model of how they connect.&lt;/p&gt;

&lt;p&gt;This post builds that mental model from the ground up. Not just definitions — the &lt;em&gt;why&lt;/em&gt; behind each piece, so troubleshooting becomes deduction instead of guesswork.&lt;/p&gt;




&lt;h2&gt;
  
  
  CIDR math you actually need
&lt;/h2&gt;

&lt;p&gt;A CIDR block is &lt;code&gt;IP address / prefix length&lt;/code&gt;. The prefix length fixes the network portion; the remaining bits are your host space.&lt;/p&gt;

&lt;p&gt;Formula: &lt;code&gt;2^(32 - prefix) = total addresses&lt;/code&gt;. AWS reserves 5 per subnet (network address, VPC router, DNS, reserved, broadcast).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CIDR&lt;/th&gt;
&lt;th&gt;Total addresses&lt;/th&gt;
&lt;th&gt;Usable&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;/16&lt;/td&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;td&gt;65,531&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;/20&lt;/td&gt;
&lt;td&gt;4,096&lt;/td&gt;
&lt;td&gt;4,091&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;/24&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;251&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;/28&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;To reverse-engineer a prefix from a required host count: round up to the next power of two, subtract the exponent from 32. Need 300 hosts? Next power of two is 512 (2⁹), so prefix = 32 - 9 = &lt;code&gt;/23&lt;/code&gt;. Run this before sizing any subnet that will host an autoscaling group or EKS node group.&lt;/p&gt;

&lt;p&gt;Start with &lt;code&gt;/16&lt;/code&gt; for the VPC itself. VPC CIDR is difficult to resize after the fact — once you have subnets, peering connections, or Transit Gateway attachments built against it, renumbering becomes a migration project. &lt;code&gt;/16&lt;/code&gt; costs nothing up front and avoids that corner.&lt;/p&gt;




&lt;h2&gt;
  
  
  Subnet allocation: carving up the VPC
&lt;/h2&gt;

&lt;p&gt;A practical three-AZ production layout from &lt;code&gt;10.0.0.0/16&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;AZ-a&lt;/th&gt;
&lt;th&gt;AZ-b&lt;/th&gt;
&lt;th&gt;AZ-c&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Typical use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Public&lt;/td&gt;
&lt;td&gt;10.0.0.0/24&lt;/td&gt;
&lt;td&gt;10.0.1.0/24&lt;/td&gt;
&lt;td&gt;10.0.2.0/24&lt;/td&gt;
&lt;td&gt;/24&lt;/td&gt;
&lt;td&gt;ALB, NAT gateway, bastion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Private/app&lt;/td&gt;
&lt;td&gt;10.0.16.0/20&lt;/td&gt;
&lt;td&gt;10.0.32.0/20&lt;/td&gt;
&lt;td&gt;10.0.48.0/20&lt;/td&gt;
&lt;td&gt;/20&lt;/td&gt;
&lt;td&gt;EKS nodes, ECS, EC2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data&lt;/td&gt;
&lt;td&gt;10.0.64.0/24&lt;/td&gt;
&lt;td&gt;10.0.65.0/24&lt;/td&gt;
&lt;td&gt;10.0.66.0/24&lt;/td&gt;
&lt;td&gt;/24&lt;/td&gt;
&lt;td&gt;RDS, ElastiCache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reserved&lt;/td&gt;
&lt;td&gt;10.0.128.0/17&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;/17&lt;/td&gt;
&lt;td&gt;Future tiers, Transit Gateway, VPN&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The jump from &lt;code&gt;/24&lt;/code&gt; in the public tier to &lt;code&gt;/20&lt;/code&gt; in the app tier is intentional. ALBs and NAT gateways consume very few IPs; the app tier is where consumption scales with autoscaling groups, rolling deployments, and pod density.&lt;/p&gt;

&lt;p&gt;For EKS specifically: with the VPC CNI, every pod can consume an ENI-backed IP. IP exhaustion is one of the most common EKS production incidents. &lt;code&gt;/20&lt;/code&gt; per AZ for worker subnets is the standard starting point.&lt;/p&gt;

&lt;p&gt;The deliberate gaps between tiers (0–2, then 16–48, then 64–66) leave room to insert new tiers later without renumbering anything already deployed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Route tables: the actual decision maker
&lt;/h2&gt;

&lt;p&gt;A subnet is "public" or "private" because of its route table — not any inherent property of the subnet itself. The table is a list of &lt;code&gt;destination → target&lt;/code&gt; rules evaluated by &lt;strong&gt;most specific match&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Every route table gets an implicit &lt;code&gt;local&lt;/code&gt; route for the full VPC CIDR — this can't be removed, and it's what lets every subnet reach every other subnet inside the VPC by default.&lt;/p&gt;

&lt;p&gt;A public subnet route table in Terraform:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_route_table"&lt;/span&gt; &lt;span class="s2"&gt;"public"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;vpc_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_vpc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;tags&lt;/span&gt;   &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"rtb-public"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;Tier&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"public"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_route"&lt;/span&gt; &lt;span class="s2"&gt;"public_internet"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;route_table_id&lt;/span&gt;         &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_route_table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;public&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;destination_cidr_block&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"0.0.0.0/0"&lt;/span&gt;
  &lt;span class="nx"&gt;gateway_id&lt;/span&gt;             &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_internet_gateway&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;main&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"aws_route_table_association"&lt;/span&gt; &lt;span class="s2"&gt;"public_a"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;subnet_id&lt;/span&gt;      &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_subnet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;public_az_a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="nx"&gt;route_table_id&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;aws_route_table&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;public&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Creating the route table does nothing on its own — the association step is what binds it to a subnet and makes routing take effect.&lt;/p&gt;




&lt;h2&gt;
  
  
  Internet Gateway and the three conditions for inbound access
&lt;/h2&gt;

&lt;p&gt;An IGW is horizontally scaled, redundant, and AZ-agnostic — one per VPC, no capacity to configure. It does two things: 1:1 NAT between public and private IPs (the public IP mapping lives at the IGW, not on the instance — which is why &lt;code&gt;ip addr&lt;/code&gt; on an EC2 instance never shows its public IP), and serves as a route table target.&lt;/p&gt;

&lt;p&gt;All three of these must be true simultaneously for inbound internet access to work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The instance has a public or Elastic IP on its ENI.&lt;/li&gt;
&lt;li&gt;The subnet's route table has &lt;code&gt;0.0.0.0/0 → IGW&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Both the security group and the NACL allow the inbound traffic on that port.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Any one missing produces the same symptom: a silent timeout with no obvious pointer to the actual cause. This is where most "why can't I reach my instance" tickets originate.&lt;/p&gt;




&lt;h2&gt;
  
  
  NAT Gateway: outbound only
&lt;/h2&gt;

&lt;p&gt;A NAT gateway lives in a specific subnet in a specific AZ, performs source NAT for private instances, and has real hourly and per-GB cost. The packet walk for a private instance at &lt;code&gt;10.0.2.15&lt;/code&gt; requesting a public registry:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Private subnet route table: &lt;code&gt;0.0.0.0/0 → nat-0abc...&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;NAT gateway rewrites source to its own Elastic IP + ephemeral port&lt;/li&gt;
&lt;li&gt;NAT gateway's public subnet route table: &lt;code&gt;0.0.0.0/0 → IGW&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;IGW performs its own separate 1:1 NAT translation&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two distinct NAT translations — easy to collapse into one mental step, but they're separate resources doing separate jobs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it lives in the public subnet:&lt;/strong&gt; the NAT gateway needs its own route to the IGW, so it must sit in a subnet whose route table already points to the IGW. The private subnet's route table then points &lt;code&gt;0.0.0.0/0&lt;/code&gt; at the NAT gateway. Two different route tables, two different subnets, one resource bridging them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;HA pattern:&lt;/strong&gt; one NAT gateway per AZ, each AZ's private subnet routing to the NAT gateway in its own AZ. One NAT gateway for the whole VPC is cheaper but creates a single point of failure — if that AZ has an outage, every private subnet in every other AZ loses outbound internet access.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost trap worth auditing:&lt;/strong&gt; traffic to S3 and DynamoDB from private subnets doesn't need to go through NAT at all if you use VPC Gateway Endpoints. Routing S3 traffic through NAT is billed per GB with no benefit over a free Gateway Endpoint.&lt;/p&gt;




&lt;h2&gt;
  
  
  Read the Full Article
&lt;/h2&gt;

&lt;p&gt;The summary covers the core mental model. The full article goes deeper on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The reverse-engineering formula for subnet sizing applied to autoscaling groups and EKS node groups, with the specific IP exhaustion failure mode explained&lt;/li&gt;
&lt;li&gt;The full route table example with VPC peering and S3 Gateway Endpoint entries, and why most-specific-match matters for overlapping routes&lt;/li&gt;
&lt;li&gt;IGW statelessness and why NACLs require explicit ephemeral port rules (&lt;code&gt;1024–65535&lt;/code&gt;) that security groups handle automatically&lt;/li&gt;
&lt;li&gt;NAT Gateway connection tracking limits: 55,000 concurrent connections per unique destination, &lt;code&gt;PortAllocationErrors&lt;/code&gt; in CloudWatch as the signal, and when to reconsider architecture vs. adding more NAT gateways&lt;/li&gt;
&lt;li&gt;The full security group vs. NACL comparison — stateful vs. stateless evaluation, allow-only vs. allow-and-deny, and why the default NACL and default security group behave differently out of the box&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;👉 &lt;a href="https://aloknecessary.in/blogs/aws-vpc-networking-fundamentals/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=blog_syndication&amp;amp;utm_content=aws-vpc-networking-fundamentals" rel="noopener noreferrer"&gt;AWS VPC Networking Fundamentals — Full Article&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>networking</category>
      <category>vpc</category>
      <category>terraform</category>
      <category>aws</category>
    </item>
  </channel>
</rss>
