<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sathpal Singh</title>
    <description>The latest articles on DEV Community by Sathpal Singh (@sathpal_singh).</description>
    <link>https://dev.to/sathpal_singh</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1595943%2F19b522c7-82ed-4bf5-952e-72baa23f1857.png</url>
      <title>DEV Community: Sathpal Singh</title>
      <link>https://dev.to/sathpal_singh</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sathpal_singh"/>
    <language>en</language>
    <item>
      <title>Running GitHub Actions Runner Controller on AKS Automatic: the blog post, executed end to end</title>
      <dc:creator>Sathpal Singh</dc:creator>
      <pubDate>Mon, 21 Sep 2026 17:37:04 +0000</pubDate>
      <link>https://dev.to/sathpal_singh/running-github-actions-runner-controller-on-aks-automatic-the-blog-post-executed-end-to-end-24bp</link>
      <guid>https://dev.to/sathpal_singh/running-github-actions-runner-controller-on-aks-automatic-the-blog-post-executed-end-to-end-24bp</guid>
      <description>&lt;p&gt;On 17 September 2026 the AKS engineering blog published &lt;a href="https://blog.aks.azure.com/2026/09/17/github-actions-runner-controller-aks-automatic" rel="noopener noreferrer"&gt;Running GitHub Actions Runner Controller on AKS Automatic&lt;/a&gt; by Steve Griffith. It is a clean guide: create an AKS Automatic cluster, install Actions Runner Controller (ARC), register a runner scale set against a GitHub repository, run a workflow, watch it scale.&lt;/p&gt;

&lt;p&gt;I followed it from a subscription that had never run AKS Automatic, and this post is what happened, command by command, with the real output. Most of it went exactly as written. Three things did not, and those are the useful parts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/sathpal/arc-aks-automatic-demo" rel="noopener noreferrer"&gt;github.com/sathpal/arc-aks-automatic-demo&lt;/a&gt;. One &lt;code&gt;make&lt;/code&gt; target per step, a screenshot of each, a troubleshooting section for everything that went wrong, and the validation workflow. Fork it, set your GitHub user in &lt;code&gt;.env&lt;/code&gt;, and the runners register against your fork.&lt;/p&gt;

&lt;h2&gt;
  
  
  What AKS Automatic changes
&lt;/h2&gt;

&lt;p&gt;A standard AKS cluster gives you a Kubernetes API and lets you decide the rest. AKS Automatic decides for you: node auto-provisioning (Karpenter) instead of hand-sized node pools, Azure RBAC for Kubernetes with local accounts disabled, managed Prometheus and Grafana, and Deployment Safeguards, which are Gatekeeper policies that block or warn on common anti-patterns such as pods without resource requests or images tagged &lt;code&gt;latest&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;ARC is a good test of that. It has a controller, a listener per runner set, and ephemeral runner pods that appear when a job is queued and vanish when it finishes. The runner pods are exactly the kind of bursty workload node auto-provisioning is for, and the charts are exactly the kind of upstream Helm charts that Safeguards will have opinions about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 0: two things the blog assumes you already have
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A current Azure CLI.&lt;/strong&gt; The guide's first command is &lt;code&gt;az extension add --name aks-preview --upgrade&lt;/code&gt;. On my machine that pulled an aks-preview build that needs a newer CLI than the 2.72 I had, and every &lt;code&gt;az aks&lt;/code&gt; command then died with a Python traceback. The &lt;code&gt;--sku automatic&lt;/code&gt; flag is GA and lives in the core CLI, so the fix was not the extension but the CLI itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az extension remove &lt;span class="nt"&gt;-n&lt;/span&gt; aks-preview
brew upgrade azure-cli          &lt;span class="c"&gt;# 2.72.0 -&amp;gt; 2.90.0&lt;/span&gt;
az aks create &lt;span class="nt"&gt;--help&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-A2&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;--sku&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Sixteen vCPUs of quota in one VM family that AKS Automatic likes.&lt;/strong&gt; My first create attempt failed after a minute with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AKS Automatic could not find a suitable VM size. The subscription may not have the
required quota of '16' vCPUs, may have restrictions, or location 'australiaeast' may
not support three availability zones for the following VM sizes: 'standard_d4lds_v5,
standard_d4ads_v5, standard_d4ds_v5, standard_d4d_v5, standard_d4d_v4, standard_ds3_v2,
standard_ds12_v2, standard_d4alds_v6, standard_d4lds_v6, standard_d4alds_v5'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every D-family in every region I checked was at the default limit of 10 cores. The fix was a quota request through the CLI, which was approved instantly for one family:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az extension add &lt;span class="nt"&gt;-n&lt;/span&gt; quota
az quota update &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-name&lt;/span&gt; standardDaldv6Family &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--scope&lt;/span&gt; &lt;span class="s2"&gt;"/subscriptions/&amp;lt;sub-id&amp;gt;/providers/Microsoft.Compute/locations/australiaeast"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--limit-object&lt;/span&gt; &lt;span class="nv"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;48 &lt;span class="nt"&gt;--resource-type&lt;/span&gt; dedicated
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same request for the DSv5 family came back "ContactSupport", so pick a family that is cheap and new; the v6 AMD sizes were the ones that went through for me.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: create the cluster
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;LOCATION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;australiaeast &lt;span class="nv"&gt;RG&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;rg-arc-auto-lab &lt;span class="nv"&gt;CLUSTER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;arc-auto-lab
az group create &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="nv"&gt;$RG&lt;/span&gt; &lt;span class="nt"&gt;--location&lt;/span&gt; &lt;span class="nv"&gt;$LOCATION&lt;/span&gt;
az aks create &lt;span class="nt"&gt;--resource-group&lt;/span&gt; &lt;span class="nv"&gt;$RG&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="nv"&gt;$CLUSTER&lt;/span&gt; &lt;span class="nt"&gt;--location&lt;/span&gt; &lt;span class="nv"&gt;$LOCATION&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--sku&lt;/span&gt; automatic &lt;span class="nt"&gt;--no-ssh-key&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This took 33 minutes, not the ten I had budgeted from experience with standard AKS. Automatic creates a three-node system pool spread across zones, a hosted pool for its own add-ons, Karpenter, managed monitoring, and the policy stack, then waits for all of it to be healthy.&lt;/p&gt;

&lt;p&gt;Local accounts are disabled, so you give yourself a Kubernetes RBAC role through Azure and use kubelogin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;AKS_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;az aks show &lt;span class="nt"&gt;-g&lt;/span&gt; &lt;span class="nv"&gt;$RG&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nv"&gt;$CLUSTER&lt;/span&gt; &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; tsv&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;ME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;az ad signed-in-user show &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; tsv&lt;span class="si"&gt;)&lt;/span&gt;
az role assignment create &lt;span class="nt"&gt;--assignee&lt;/span&gt; &lt;span class="nv"&gt;$ME&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--role&lt;/span&gt; &lt;span class="s2"&gt;"Azure Kubernetes Service RBAC Cluster Admin"&lt;/span&gt; &lt;span class="nt"&gt;--scope&lt;/span&gt; &lt;span class="nv"&gt;$AKS_ID&lt;/span&gt;
az aks get-credentials &lt;span class="nt"&gt;-g&lt;/span&gt; &lt;span class="nv"&gt;$RG&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nv"&gt;$CLUSTER&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt; &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;--overwrite-existing&lt;/span&gt;
kubectl get nodes &lt;span class="nt"&gt;-o&lt;/span&gt; wide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffdwzsppttqcs8spr2461.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffdwzsppttqcs8spr2461.png" alt="Cluster profile: Automatic SKU, node provisioning Auto, Azure RBAC on, local accounts disabled, six nodes" width="800" height="246"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Six nodes on day one: three &lt;code&gt;nodepool1&lt;/code&gt; system nodes and three &lt;code&gt;hostedpool&lt;/code&gt; nodes that AKS runs for its managed components. All on Azure Linux 3.0.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: install the controller
&lt;/h2&gt;

&lt;p&gt;The controller chart has no resource requests by default, and Deployment Safeguards will refuse a pod without them. The blog's values file fixes that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# arc-controller-values.yaml&lt;/span&gt;
&lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;100m&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;128Mi&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
  &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;   &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;500m&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;512Mi&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm upgrade &lt;span class="nt"&gt;--install&lt;/span&gt; arc &lt;span class="se"&gt;\&lt;/span&gt;
  oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set-controller &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; arc-systems &lt;span class="nt"&gt;--create-namespace&lt;/span&gt; &lt;span class="nt"&gt;--wait&lt;/span&gt; &lt;span class="nt"&gt;--timeout&lt;/span&gt; 10m &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-f&lt;/span&gt; arc-controller-values.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmkc75o6mqz5trogtyja4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmkc75o6mqz5trogtyja4.png" alt="helm install of the controller succeeds; controller pod Running" width="800" height="242"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It installs, but Safeguards has more to say than the blog mentions. The install printed warnings that the controller container has no liveness or readiness probe, and that anti-affinity and topology spread constraints were added to the deployment by mutation. Those are warn-level policies, so the install goes through. The ones in deny mode, such as resource requests, are the ones that would have stopped it:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyt7evf2ncbgt34gwal0f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyt7evf2ncbgt34gwal0f.png" alt="Deployment Safeguards warnings from the helm install, and the Gatekeeper constraints with their enforcement actions" width="800" height="418"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you want the probe warning gone, the controller chart accepts &lt;code&gt;livenessProbe&lt;/code&gt; and &lt;code&gt;readinessProbe&lt;/code&gt; values. I left it as the blog has it so the output matches.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: token and runner scale set
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl create namespace arc-runners &lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;client &lt;span class="nt"&gt;-o&lt;/span&gt; yaml | kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; -
&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$GITHUB_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | kubectl create secret generic github-pat &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; arc-runners &lt;span class="nt"&gt;--from-file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;github_token&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/dev/stdin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A classic PAT with &lt;code&gt;repo&lt;/code&gt; scope is enough for a repository-level runner set. The blog says, and I agree, that anything beyond a lab should use a GitHub App so the permission is narrow and rotation is not a person's job.&lt;/p&gt;

&lt;p&gt;The runner set values need three things to be explicit for Safeguards: resource requests on the listener, resource requests on the runner, and a pinned runner image tag. The default chart uses &lt;code&gt;ghcr.io/actions/actions-runner:latest&lt;/code&gt;, and Safeguards rejects floating tags.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# arc-runner-set-values.yaml&lt;/span&gt;
&lt;span class="na"&gt;githubConfigUrl&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://github.com/&amp;lt;owner&amp;gt;/&amp;lt;repo&amp;gt;&lt;/span&gt;
&lt;span class="na"&gt;githubConfigSecret&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;github-pat&lt;/span&gt;
&lt;span class="na"&gt;minRunners&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="na"&gt;maxRunners&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;
&lt;span class="na"&gt;listenerTemplate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;listener&lt;/span&gt;
        &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;100m&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;128Mi&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
          &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;   &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;500m&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;512Mi&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;runner&lt;/span&gt;
        &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io/actions/actions-runner:2.337.0&lt;/span&gt;
        &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/home/runner/run.sh"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;4Gi&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
          &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;   &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;4Gi&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The blog pins 2.336.0. By the time I ran this, 2.337.0 was current, and GitHub refuses runners more than a few versions behind, so check &lt;a href="https://github.com/actions/runner/releases" rel="noopener noreferrer"&gt;the releases page&lt;/a&gt; before you copy a tag.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm upgrade &lt;span class="nt"&gt;--install&lt;/span&gt; arc-auto-runners &lt;span class="se"&gt;\&lt;/span&gt;
  oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; arc-runners &lt;span class="nt"&gt;--wait&lt;/span&gt; &lt;span class="nt"&gt;--timeout&lt;/span&gt; 10m &lt;span class="nt"&gt;-f&lt;/span&gt; arc-runner-set-values.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffaikbjhzdsx1hldc80y0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffaikbjhzdsx1hldc80y0.png" alt="Runner scale set installed: min 0, max 3, listener pod Running next to the controller" width="800" height="294"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A listener pod appears in &lt;code&gt;arc-systems&lt;/code&gt; and long-polls GitHub's broker for jobs. With &lt;code&gt;minRunners: 0&lt;/code&gt; nothing else runs until a job arrives, and the repository's runner list stays empty. That is normal and briefly alarming.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1y6arowonj7d5k73mcqf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1y6arowonj7d5k73mcqf.png" alt="Listener pod Running; log shows the ephemeral runner set scaled and the listener waiting on the broker" width="799" height="263"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: run a workflow and watch it scale
&lt;/h2&gt;

&lt;p&gt;The validation workflow from the blog, unchanged, committed to &lt;code&gt;.github/workflows/arc-automatic-validation.yml&lt;/code&gt; with &lt;code&gt;runs-on: arc-auto-runners&lt;/code&gt;, which is the Helm release name and therefore the runner label.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh workflow run arc-automatic-validation.yml &lt;span class="nt"&gt;--repo&lt;/span&gt; &amp;lt;owner&amp;gt;/&amp;lt;repo&amp;gt; &lt;span class="nt"&gt;--ref&lt;/span&gt; main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I polled the run status, the runner pods and the node count every ten seconds:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd5lvjpll3ngsiuky7kwt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd5lvjpll3ngsiuky7kwt.png" alt="Timeline: job queued, runner pod Pending, node count 7 to 8, ContainerCreating, Running, completed success, back to 0 runners" width="800" height="208"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What happened in those 190 seconds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;t+19s&lt;/strong&gt; the listener saw the queued job and created one ephemeral runner pod. It went Pending: the system nodes are tainted for system workloads and the one schedulable node did not have 4Gi free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;t+53s&lt;/strong&gt; node auto-provisioning launched a new node for it. Karpenter nominated the pod onto the node claim before the VM existed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;t+121s&lt;/strong&gt; the node was ready and the pod pulled the runner image, which is about a gigabyte and took 52 seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;t+173s&lt;/strong&gt; the job ran. It took 8 seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;t+190s&lt;/strong&gt; the run reported success, the pod was gone, and the runner count was back to zero.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx2talovu55wlj934fzsc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx2talovu55wlj934fzsc.png" alt="kubectl get nodeclaims shows the D4as_v6 Karpenter provisioned, then a consolidation candidate event offering to replace it for a cheaper size with the savings in dollars" width="799" height="358"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The node claim output is the part worth staring at. Karpenter picked a &lt;code&gt;Standard_D4as_v6&lt;/code&gt; on demand for the runner, and within minutes of the job finishing it flagged that node as a consolidation candidate, launched a cheaper &lt;code&gt;D4als_v6&lt;/code&gt; replacement and moved the controller and listener onto it, quoting the saving. Nobody configured any of that.&lt;/p&gt;

&lt;p&gt;On the GitHub side it looks like any other run:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F41tw98tq2oyjb31ytqzc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F41tw98tq2oyjb31ytqzc.png" alt="GitHub Actions run page: ARC AKS Automatic validation, Success, validate job 8s, total 2m 48s" width="800" height="514"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;Current runner version: '2.337.0'
Runner name: 'arc-auto-runners-j4kzc-runner-x8p8j'
ARC runner reached workflow execution
&lt;/span&gt;&lt;span class="gp"&gt;Linux arc-auto-runners-j4kzc-runner-x8p8j 6.6.150.1-1.azl3 #&lt;/span&gt;1 SMP x86_64 GNU/Linux
&lt;span class="go"&gt;Docker version 29.7.2, build a7dcaa6
Validation complete
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9064rxb20ohe4gfeuth1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9064rxb20ohe4gfeuth1.png" alt="Job log from gh run view: runner version, runner name, uname, df, docker version" width="800" height="720"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The total of 2 minutes 48 seconds is almost entirely cold start: a new VM and a cold image pull. A second job arriving while the node is still there would start in seconds. If that latency matters, set &lt;code&gt;minRunners: 1&lt;/code&gt;, and accept that you now pay for one idle D4.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: clean up
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm uninstall arc-auto-runners &lt;span class="nt"&gt;-n&lt;/span&gt; arc-runners
helm uninstall arc &lt;span class="nt"&gt;-n&lt;/span&gt; arc-systems
kubectl delete namespace arc-runners arc-systems
az group delete &lt;span class="nt"&gt;--name&lt;/span&gt; rg-arc-auto-lab &lt;span class="nt"&gt;--yes&lt;/span&gt; &lt;span class="nt"&gt;--no-wait&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The resource group delete also removes the node resource group AKS created.&lt;/p&gt;

&lt;p&gt;One more thing I learned by doing it wrong. I moved the runner set from one repository to another with &lt;code&gt;helm upgrade&lt;/code&gt; and a new &lt;code&gt;githubConfigUrl&lt;/code&gt;. The listener crashed on every start with "No runner scale set found with identifier 1": the set had kept the ID it registered under the first repository, and the second one had no such ID. Uninstall and reinstall is the only clean path, and if a stale &lt;code&gt;AutoscalingListener&lt;/code&gt; object survives that, delete it and the controller recreates it in seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I landed
&lt;/h2&gt;

&lt;p&gt;If someone on my team asked whether to run our Actions runners this way, I would say yes, with two caveats that are about Azure rather than ARC.&lt;/p&gt;

&lt;p&gt;The ARC side is boring in the best sense. Steve's values files are exactly what the cluster demands, the charts install first time, and the runner behaves like a hosted runner from the workflow's point of view. I did not change a line of the workflow. What I spent my time on was getting a cluster at all: the quota round trip and the CLI upgrade took longer than everything after them combined, and neither is mentioned in the post because a Microsoft engineer's subscription does not have those problems. Yours might. Run &lt;code&gt;az vm list-usage&lt;/code&gt; for your region before you type &lt;code&gt;az aks create&lt;/code&gt;, and check the CLI version before you add any extension.&lt;/p&gt;

&lt;p&gt;The thing that changed my mind about AKS Automatic was not the cluster creation, which is slow, or Safeguards, which mostly nag. It was watching Karpenter after the job finished. It had bought a D4as_v6 for the runner, noticed twenty minutes later that the node was underused, priced a cheaper D4als_v6, launched it, moved the pods, and deleted the original. It printed the saving in the event log. I have written that logic by hand for other clusters and never got it this tidy.&lt;/p&gt;

&lt;p&gt;The caveat on the ARC side is the token. I used my own PAT because it was a lab and I wanted to get to the interesting part. Do not do that for a team. A GitHub App gives the runner set its own identity with the permissions it needs and nothing else, and rotation stops being someone's memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh repo fork sathpal/arc-aks-automatic-demo &lt;span class="nt"&gt;--clone&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;arc-aks-automatic-demo
make tools &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env         &lt;span class="c"&gt;# set GITHUB_OWNER to your user&lt;/span&gt;
make quota &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make cluster &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make access &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make arc
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GITHUB_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;gh auth token&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make runners &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make &lt;span class="nb"&gt;test
&lt;/span&gt;make cleanup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>kubernetes</category>
      <category>azure</category>
      <category>github</category>
      <category>devops</category>
    </item>
    <item>
      <title>Same app, two base images: 178 CVEs vs 5. A hands-on Chainguard on AKS walkthrough</title>
      <dc:creator>Sathpal Singh</dc:creator>
      <pubDate>Tue, 15 Sep 2026 18:24:22 +0000</pubDate>
      <link>https://dev.to/sathpal_singh/same-app-two-base-images-178-cves-vs-5-a-hands-on-chainguard-on-aks-walkthrough-3fc0</link>
      <guid>https://dev.to/sathpal_singh/same-app-two-base-images-178-cves-vs-5-a-hands-on-chainguard-on-aks-walkthrough-3fc0</guid>
      <description>&lt;p&gt;I watched the &lt;a href="https://www.youtube.com/watch?v=-dMyVMPeUug&amp;amp;t=320s" rel="noopener noreferrer"&gt;&lt;em&gt;Cloud Native Partner Showcase&lt;/em&gt; episode&lt;/a&gt; where Microsoft's David Giard talks to Hannah Hawken and Manfred Moser from Chainguard about secure-by-default container images on Azure Kubernetes Service. Good conversation, but I wanted numbers I produced myself. So I built the smallest possible demo that proves or disproves the pitch, and this post is that demo, step by step, with the actual output.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/-dMyVMPeUug" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;The destination is Azure Kubernetes Service. The first steps build the evidence on your machine so you can see the difference before it reaches a cluster; the later steps put both images on AKS behind public load balancers and a Kyverno admission policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/sathpal/chainguard-aks-demo" rel="noopener noreferrer"&gt;github.com/sathpal/chainguard-aks-demo&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What we are testing
&lt;/h2&gt;

&lt;p&gt;The claim from the episode: Chainguard images are minimal, rebuilt continuously from source, signed, and ship with an SBOM, so most CVEs never reach your cluster in the first place.&lt;/p&gt;

&lt;p&gt;The test: one 40-line FastAPI app, built two ways.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;upstream&lt;/th&gt;
&lt;th&gt;chainguard&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;base image&lt;/td&gt;
&lt;td&gt;&lt;code&gt;python:3.13-slim&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;cgr.dev/chainguard/python&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;user&lt;/td&gt;
&lt;td&gt;root&lt;/td&gt;
&lt;td&gt;nonroot (65532)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;shell / package manager&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Then scan both, inspect both, deploy both, and see what the difference actually is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;brew &lt;span class="nb"&gt;install &lt;/span&gt;trivy grype syft cosign crane      &lt;span class="c"&gt;# scanners, SBOM, signing&lt;/span&gt;
&lt;span class="c"&gt;# also: docker, kubectl, helm, azure-cli&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 1: the app
&lt;/h2&gt;

&lt;p&gt;The app deliberately reports what it is running on: OS, uid, whether a shell exists, how many OS packages are installed. Same code in both images.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;facts&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_flavor&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;IMAGE_FLAVOR&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;os&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;os_release&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;                     &lt;span class="c1"&gt;# /etc/os-release PRETTY_NAME
&lt;/span&gt;        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;uid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getuid&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;shell_present&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/bin/sh&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/bin/bash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;package_manager_present&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/usr/bin/apt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/sbin/apk&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/usr/bin/pip&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;os_packages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;package_count&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;         &lt;span class="c1"&gt;# dpkg or apk database
&lt;/span&gt;    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 2: two Dockerfiles
&lt;/h2&gt;

&lt;p&gt;The upstream one is what most tutorials show you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; python:3.13-slim&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; app/requirements.txt .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-dir&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; app/main.py .&lt;/span&gt;
&lt;span class="k"&gt;ENV&lt;/span&gt;&lt;span class="s"&gt; IMAGE_FLAVOR=upstream&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python", "main.py"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Chainguard one is multi-stage. The &lt;code&gt;-dev&lt;/code&gt; tag has pip and a shell so you can build; the plain tag has neither, so you only copy the result in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;cgr.dev/chainguard/python:latest-dev&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;AS&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;builder&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /home/nonroot&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv venv
&lt;span class="k"&gt;ENV&lt;/span&gt;&lt;span class="s"&gt; PATH="/home/nonroot/venv/bin:$PATH"&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; app/requirements.txt .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-dir&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt

&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; cgr.dev/chainguard/python:latest&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --from=builder --chown=nonroot:nonroot /home/nonroot/venv /app/venv&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; --chown=nonroot:nonroot app/main.py .&lt;/span&gt;
&lt;span class="k"&gt;ENV&lt;/span&gt;&lt;span class="s"&gt; PATH="/app/venv/bin:$PATH" IMAGE_FLAVOR=chainguard&lt;/span&gt;
&lt;span class="k"&gt;USER&lt;/span&gt;&lt;span class="s"&gt; 65532&lt;/span&gt;
&lt;span class="k"&gt;ENTRYPOINT&lt;/span&gt;&lt;span class="s"&gt; ["python", "main.py"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things bite people here. There is no shell in the runtime image, so &lt;code&gt;CMD python main.py&lt;/code&gt; (string form) fails; use the exec form. &lt;code&gt;pip install&lt;/code&gt; must happen in the builder stage, never in the runtime stage. And the builder stage works inside &lt;code&gt;/home/nonroot&lt;/code&gt;, a directory that already exists and belongs to the nonroot user. The &lt;code&gt;-dev&lt;/code&gt; image runs as nonroot, and while BuildKit creates a new &lt;code&gt;WORKDIR&lt;/code&gt; owned by the current user, the legacy builder that ACR Tasks still uses creates it as root, so &lt;code&gt;WORKDIR /app&lt;/code&gt; followed by &lt;code&gt;python -m venv venv&lt;/code&gt; fails with permission denied the moment you build in the cloud instead of on your laptop. I found that one the hard way later in this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: build
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker build &lt;span class="nt"&gt;-f&lt;/span&gt; docker/Dockerfile.upstream   &lt;span class="nt"&gt;-t&lt;/span&gt; demo-app:upstream &lt;span class="nb"&gt;.&lt;/span&gt;
docker build &lt;span class="nt"&gt;-f&lt;/span&gt; docker/Dockerfile.chainguard &lt;span class="nt"&gt;-t&lt;/span&gt; demo-app:chainguard &lt;span class="nb"&gt;.&lt;/span&gt;
docker images demo-app
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7euxw70bxt9w0dc3y6a0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7euxw70bxt9w0dc3y6a0.png" alt="docker images: 248MB vs 140MB" width="800" height="124"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: scan
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;grype demo-app:upstream   &lt;span class="nt"&gt;-o&lt;/span&gt; table
grype demo-app:chainguard &lt;span class="nt"&gt;-o&lt;/span&gt; table
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3mps1h4k44mwghvxv1fk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3mps1h4k44mwghvxv1fk.png" alt="grype output for both images" width="800" height="349"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A small script summarises the two JSON reports into one table:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbt1ok4p0brbn1e697xcs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbt1ok4p0brbn1e697xcs.png" alt="comparison table: 178 CVEs vs 5" width="800" height="131"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Same app. 178 known CVEs versus 5. Zero critical on the Chainguard side, and the five that remain are Wolfi packages with a fix already queued (zlib) or with no upstream fix yet (python). Rebuild tomorrow and the count moves. That is the "continuously rebuilt" part of the pitch working as described.&lt;/p&gt;

&lt;p&gt;An honest note: my first scan showed 12 on the Chainguard side, not 5. Seven of those were in &lt;code&gt;starlette&lt;/code&gt;, because I had pinned an old FastAPI in &lt;code&gt;requirements.txt&lt;/code&gt;. The base image cannot fix your dependency file. Application-level CVEs stay your job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: poke around inside
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; demo-app:upstream sh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"id; which apt-get; ls /bin | wc -l"&lt;/span&gt;
docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; &lt;span class="nt"&gt;--entrypoint&lt;/span&gt; sh demo-app:chainguard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyva0ec2qx1tyfnxexxgh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyva0ec2qx1tyfnxexxgh.png" alt="upstream: root with apt-get and 259 binaries; chainguard: no sh at all" width="799" height="218"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The upstream container runs as root with apt-get and 259 binaries available to anyone who gets code execution. The Chainguard container cannot even start a shell. For debugging you use &lt;code&gt;kubectl debug&lt;/code&gt; with an ephemeral container, which is a workflow change worth planning for (see cons below).&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: verify the base image is really from Chainguard
&lt;/h2&gt;

&lt;p&gt;Chainguard signs every image with Sigstore keyless signing from their GitHub release workflow. You can check that without any keys:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;cosign verify cgr.dev/chainguard/python:latest &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--certificate-oidc-issuer&lt;/span&gt; https://token.actions.githubusercontent.com &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--certificate-identity&lt;/span&gt; https://github.com/chainguard-images/images/.github/workflows/release.yaml@refs/heads/main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fitnxp6iypnud96ive7w9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fitnxp6iypnud96ive7w9.png" alt="cosign verify output showing issuer and identity" width="799" height="276"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The signature is in the public Rekor transparency log. This is the difference between "we pulled an image called python" and "we pulled the image Chainguard's release pipeline built".&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 7: SBOMs, yours and theirs
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;syft demo-app:upstream   &lt;span class="nt"&gt;-o&lt;/span&gt; spdx-json &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; out/upstream.sbom.spdx.json
syft demo-app:chainguard &lt;span class="nt"&gt;-o&lt;/span&gt; spdx-json &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; out/chainguard.sbom.spdx.json

&lt;span class="c"&gt;# the SBOM Chainguard already attested for the base image&lt;/span&gt;
cosign verify-attestation cgr.dev/chainguard/python:latest &lt;span class="nt"&gt;--type&lt;/span&gt; https://spdx.dev/Document &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--certificate-oidc-issuer&lt;/span&gt; https://token.actions.githubusercontent.com &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--certificate-identity&lt;/span&gt; https://github.com/chainguard-images/images/.github/workflows/release.yaml@refs/heads/main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffyjjupksvo72hartzi8x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffyjjupksvo72hartzi8x.png" alt="syft SBOM package counts, 109 vs 47" width="798" height="145"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;109 packages to account for versus 47. When someone from compliance asks "are we affected by CVE-X", the smaller list answers faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 8: run both and look
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; 8081:8080 demo-app:upstream
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; 8082:8080 demo-app:chainguard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqu388ybuzi4cgfdbiv9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqu388ybuzi4cgfdbiv9.png" alt="upstream app page: Debian, uid 0, shell present, 87 packages" width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxjf0aujuuj9tijc3uz3j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxjf0aujuuj9tijc3uz3j.png" alt="chainguard app page: Wolfi, uid 65532, no shell, 26 packages" width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And the one-page report the scan script generates, useful for a slide:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpzi0xu0o3725c3skjl7s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpzi0xu0o3725c3skjl7s.png" alt="HTML report comparing both images" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 9: Azure Container Registry and AKS
&lt;/h2&gt;

&lt;p&gt;This part costs money. Two small nodes for an hour is about a dollar, but delete the resource group when done.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az login &lt;span class="nt"&gt;--tenant&lt;/span&gt; &amp;lt;tenant-id&amp;gt;
az group create &lt;span class="nt"&gt;-n&lt;/span&gt; rg-chainguard-demo &lt;span class="nt"&gt;-l&lt;/span&gt; centralindia
az acr create &lt;span class="nt"&gt;-n&lt;/span&gt; sathpalcgdemo &lt;span class="nt"&gt;-g&lt;/span&gt; rg-chainguard-demo &lt;span class="nt"&gt;--sku&lt;/span&gt; Basic
az aks create &lt;span class="nt"&gt;-n&lt;/span&gt; aks-chainguard-demo &lt;span class="nt"&gt;-g&lt;/span&gt; rg-chainguard-demo &lt;span class="nt"&gt;--node-count&lt;/span&gt; 2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--node-vm-size&lt;/span&gt; Standard_D2s_v4 &lt;span class="nt"&gt;--attach-acr&lt;/span&gt; sathpalcgdemo &lt;span class="nt"&gt;--generate-ssh-keys&lt;/span&gt;
az aks get-credentials &lt;span class="nt"&gt;-n&lt;/span&gt; aks-chainguard-demo &lt;span class="nt"&gt;-g&lt;/span&gt; rg-chainguard-demo

az acr login &lt;span class="nt"&gt;-n&lt;/span&gt; sathpalcgdemo
docker buildx build &lt;span class="nt"&gt;--platform&lt;/span&gt; linux/amd64 &lt;span class="nt"&gt;-f&lt;/span&gt; docker/Dockerfile.chainguard &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-t&lt;/span&gt; sathpalcgdemo.azurecr.io/demo-app:chainguard &lt;span class="nt"&gt;--push&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Attaching ACR to AKS means the nodes pull with managed identity, no image pull secrets. Build for &lt;code&gt;linux/amd64&lt;/code&gt; if you are on Apple Silicon, since the default AKS node pool is x86.&lt;/p&gt;

&lt;p&gt;Or skip local Docker entirely and let the registry build it. My laptop's disk filled up halfway through this post and Docker Desktop stopped working, so the Chainguard image in the cluster was actually built by ACR Tasks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az acr build &lt;span class="nt"&gt;-r&lt;/span&gt; sathpalcgdemo &lt;span class="nt"&gt;-f&lt;/span&gt; docker/Dockerfile.chainguard &lt;span class="nt"&gt;-t&lt;/span&gt; demo-app:chainguard &lt;span class="nt"&gt;--platform&lt;/span&gt; linux/amd64 &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is also where the &lt;code&gt;WORKDIR /home/nonroot&lt;/code&gt; gotcha from Step 2 bit me. ACR Tasks uses the legacy builder, and my original &lt;code&gt;WORKDIR /app&lt;/code&gt; failed with &lt;code&gt;Permission denied: '/app/venv'&lt;/code&gt; on the first cloud build.&lt;/p&gt;

&lt;p&gt;Two &lt;code&gt;Standard_D2s_v4&lt;/code&gt; nodes came up in about six minutes. If your subscription refuses the VM size, &lt;code&gt;az vm list-usage -l &amp;lt;region&amp;gt; -o table&lt;/code&gt; shows which families you have quota for; mine had zero for B-series in Central India.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl get nodes
&lt;span class="go"&gt;NAME                                STATUS   ROLES    AGE     VERSION
&lt;/span&gt;&lt;span class="gp"&gt;aks-nodepool1-13038109-vmss000000   Ready    &amp;lt;none&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;7m17s   v1.35.7
&lt;span class="gp"&gt;aks-nodepool1-13038109-vmss000001   Ready    &amp;lt;none&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;7m16s   v1.35.7
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F04wwif9icttmch4hrj08.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F04wwif9icttmch4hrj08.png" alt="az aks list, kubectl get nodes, pods and services with public IPs" width="799" height="252"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Chainguard deployment can turn on every hardening knob because the image cooperates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;securityContext&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;runAsNonRoot&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;runAsUser&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;65532&lt;/span&gt;
  &lt;span class="na"&gt;allowPrivilegeEscalation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;readOnlyRootFilesystem&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;capabilities&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;drop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ALL"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Try the same block on the upstream image and the pod fails &lt;code&gt;runAsNonRoot&lt;/code&gt; immediately, because the image runs as uid 0.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;runAsUser: 65532&lt;/code&gt; line is not decoration. The Chainguard base image declares its user numerically, but my first Dockerfile overrode it with &lt;code&gt;USER nonroot&lt;/code&gt;, a name, and the kubelet refuses to start a &lt;code&gt;runAsNonRoot&lt;/code&gt; container when it cannot prove the user is non-root from a name alone. The Dockerfile above now says &lt;code&gt;USER 65532&lt;/code&gt;, and the manifest pins the same uid so the check holds even if someone changes the image later. My first rollout sat in &lt;code&gt;CreateContainerConfigError&lt;/code&gt; with exactly that message:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: container has runAsNonRoot and image has non-numeric user (nonroot),
cannot verify user is non-root
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Give it the numeric uid and it starts. Both deployments, each behind its own LoadBalancer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl &lt;span class="nt"&gt;-n&lt;/span&gt; chainguard-demo get pods,svc
&lt;span class="go"&gt;NAME                              READY   STATUS    RESTARTS   AGE
app-chainguard-6f776d7959-nprs9   1/1     Running   0          89s
app-upstream-574978df4f-4pf9g     1/1     Running   0          5m6s
NAME             TYPE           CLUSTER-IP     EXTERNAL-IP      PORT(S)        AGE
app-chainguard   LoadBalancer   10.0.67.88     98.70.244.97     80:30308/TCP   5m5s
app-upstream     LoadBalancer   10.0.126.245   20.204.187.119   80:30890/TCP   5m6s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the app reporting on itself from inside the cluster, same code, two answers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://20.204.187.119/api
&lt;span class="go"&gt;{
    "image_flavor": "upstream",
    "hostname": "app-upstream-574978df4f-4pf9g",
    "os": "Debian GNU/Linux 13 (trixie)",
    "python": "3.13.15",
    "uid": 0,
    "running_as_root": true,
    "shell_present": true,
    "package_manager_present": true,
    "os_packages": 87
}

&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://98.70.244.97/api
&lt;span class="go"&gt;{
    "image_flavor": "chainguard",
    "hostname": "app-chainguard-6f776d7959-nprs9",
    "os": "Wolfi",
    "python": "3.14.7+",
    "uid": 65532,
    "running_as_root": false,
    "shell_present": false,
    "package_manager_present": false,
    "os_packages": 26
}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhrzda6x16hkxxi62lust.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhrzda6x16hkxxi62lust.png" alt="curl /api on both public IPs" width="760" height="624"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The same page from Step 8, now served by AKS through a public load balancer:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsyv9kzhv6572xu03wede.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsyv9kzhv6572xu03wede.png" alt="upstream app page served from AKS" width="800" height="524"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5jrbw1q43tn72g7fldth.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5jrbw1q43tn72g7fldth.png" alt="chainguard app page served from AKS" width="800" height="524"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And the on-call moment. &lt;code&gt;kubectl exec&lt;/code&gt; into the upstream pod gives you a root shell with &lt;code&gt;apt-get&lt;/code&gt;. The same command on the Chainguard pod has nothing to run:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblhck5xcxge965k7oxzx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblhck5xcxge965k7oxzx.png" alt="kubectl exec: no sh in the Chainguard pod, root shell in the upstream pod" width="800" height="140"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 10: enforce it with Kyverno
&lt;/h2&gt;

&lt;p&gt;Scanning tells you. Admission control stops you. Two Kyverno policies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# only trusted registries in this namespace&lt;/span&gt;
&lt;span class="na"&gt;validate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Images&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;must&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;come&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;from&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;cgr.dev&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*.azurecr.io"&lt;/span&gt;
  &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cgr.dev/*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;|&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*.azurecr.io/*"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# every cgr.dev image must carry Chainguard's keyless signature&lt;/span&gt;
&lt;span class="na"&gt;verifyImages&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;imageReferences&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cgr.dev/chainguard/*"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;attestors&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;entries&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;keyless&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;subject&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://github.com/chainguard-images/images/.github/workflows/release.yaml@refs/heads/main"&lt;/span&gt;
              &lt;span class="na"&gt;issuer&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://token.actions.githubusercontent.com"&lt;/span&gt;
              &lt;span class="na"&gt;rekor&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;https&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;//rekor.sigstore.dev&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm upgrade &lt;span class="nt"&gt;--install&lt;/span&gt; kyverno kyverno/kyverno &lt;span class="nt"&gt;-n&lt;/span&gt; kyverno &lt;span class="nt"&gt;--create-namespace&lt;/span&gt;
kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; k8s/policies/
kubectl &lt;span class="nt"&gt;-n&lt;/span&gt; chainguard-demo run bad &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;docker.io/library/nginx:latest   &lt;span class="c"&gt;# rejected&lt;/span&gt;
kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; k8s/pod-chainguard-nginx.yaml                              &lt;span class="c"&gt;# admitted&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rejection, verbatim from the API server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl &lt;span class="nt"&gt;-n&lt;/span&gt; chainguard-demo run bad &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;docker.io/library/nginx:latest
&lt;span class="go"&gt;Error from server: admission webhook "validate.kyverno.svc-fail" denied the request: 
resource Pod/chainguard-demo/bad was blocked due to the following policies 
restrict-image-registries:
  allowed-registries: 'validation error: Images must come from cgr.dev or *.azurecr.io (trusted, signed sources). rule allowed-registries failed at path /spec/containers/0/image/'
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The signed image goes straight through, and Kyverno checked the Rekor entry for &lt;code&gt;cgr.dev/chainguard/nginx&lt;/code&gt; on the way in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; k8s/pod-chainguard-nginx.yaml
&lt;span class="go"&gt;pod/nginx-signed created
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;kubectl &lt;span class="nt"&gt;-n&lt;/span&gt; chainguard-demo get pod nginx-signed
&lt;span class="go"&gt;NAME           READY   STATUS    RESTARTS   AGE
nginx-signed   1/1     Running   0          8s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff35m7hkovy0b4id9o8tp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff35m7hkovy0b4id9o8tp.png" alt="Kyverno rejects Docker Hub nginx, admits signed cgr.dev nginx, both policies Ready" width="800" height="213"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One more real-world note: the namespace also carries the Pod Security Standards labels (&lt;code&gt;enforce: baseline&lt;/code&gt;, &lt;code&gt;warn: restricted&lt;/code&gt;). Every &lt;code&gt;kubectl apply&lt;/code&gt; for the upstream deployment prints a warning that it would violate &lt;code&gt;restricted&lt;/code&gt;. The Chainguard deployment is silent. That warning line is the cheapest security audit you will ever run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 11: clean up
&lt;/h2&gt;

&lt;p&gt;Everything above lives in one resource group, including the node resource group AKS creates for itself. One command stops the meter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az group delete &lt;span class="nt"&gt;-n&lt;/span&gt; rg-chainguard-demo &lt;span class="nt"&gt;--yes&lt;/span&gt; &lt;span class="nt"&gt;--no-wait&lt;/span&gt;     &lt;span class="c"&gt;# or: make aks-down&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The public IPs in this post are gone by the time you read it. The repo is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 12: the "agent skill" part
&lt;/h2&gt;

&lt;p&gt;The episode ends on AI-assisted workflows. I turned the migration steps above into a Claude Code skill (&lt;code&gt;.claude/skills/chainguard-migrate/SKILL.md&lt;/code&gt;) plus a base-image mapping table. In the repo you say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;use the chainguard-migrate skill on docker/Dockerfile.upstream&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and the agent baselines the image with grype, rewrites the Dockerfile to the multi-stage Chainguard pattern, rebuilds, rescans, verifies the base image signature and prints the before/after table. The skill's last rule is the important one: &lt;em&gt;never claim zero CVEs without a fresh scan in the transcript.&lt;/em&gt; Agents that write security claims need receipts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pros
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The CVE delta is real, not marketing.&lt;/strong&gt; 178 to 5 on the same app, with the remaining 5 traceable to two packages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Smaller attack surface by construction.&lt;/strong&gt; No shell, no package manager, nonroot by default. A large class of post-exploitation tricks simply has nothing to run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Signatures and SBOMs are already there.&lt;/strong&gt; You verify with public tooling, no vendor account needed for the free images.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kubernetes hardening becomes possible.&lt;/strong&gt; &lt;code&gt;readOnlyRootFilesystem&lt;/code&gt;, &lt;code&gt;drop: ALL&lt;/code&gt;, PSS restricted profile. These fail on most upstream images.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Less scanner noise.&lt;/strong&gt; Fewer findings means the ones left get read. Security teams stop being the department of 178 tickets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fits Azure without glue.&lt;/strong&gt; Push to ACR, attach to AKS, listed on Azure Marketplace for procurement.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Cons
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Debugging changes.&lt;/strong&gt; No shell means &lt;code&gt;kubectl exec&lt;/code&gt; is gone. You need &lt;code&gt;kubectl debug&lt;/code&gt; with ephemeral containers, and your on-call runbooks need updating.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free tier is &lt;code&gt;latest&lt;/code&gt; only.&lt;/strong&gt; Version-pinned tags (&lt;code&gt;python:3.12&lt;/code&gt;) and the full catalog need a paid Chainguard account. For reproducible enterprise builds that is a budget conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wolfi is not Debian.&lt;/strong&gt; Different package names, &lt;code&gt;apk&lt;/code&gt; not &lt;code&gt;apt&lt;/code&gt;, paths differ. Every &lt;code&gt;RUN apt-get install&lt;/code&gt; in your Dockerfiles needs rethinking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-stage is mandatory.&lt;/strong&gt; Single-stage Dockerfiles with build tools in the runtime image will not work. Good discipline, but it is work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your dependencies are still yours.&lt;/strong&gt; My first scan had 12 findings and 7 were my outdated FastAPI pin. The base image does not fix &lt;code&gt;requirements.txt&lt;/code&gt;, &lt;code&gt;package.json&lt;/code&gt; or &lt;code&gt;pom.xml&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vendor concentration.&lt;/strong&gt; You are trusting one company's rebuild pipeline. The signatures let you verify that trust, but it is a new dependency in your supply chain.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Who benefits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Platform teams on AKS&lt;/strong&gt; who own the golden-image catalog and are tired of re-patching Debian bases every week.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regulated teams&lt;/strong&gt; (finance, health, public sector) who need SBOMs, provenance and a credible answer to "how do you know what is in production".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small teams with no security engineer.&lt;/strong&gt; Swapping the base image is the highest-leverage security change per hour of effort I know of.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anyone adopting Pod Security Standards restricted.&lt;/strong&gt; You cannot pass it with root images; this is the shortcut.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Teams introducing AI coding agents.&lt;/strong&gt; A skill like the one above turns "harden this container" into a repeatable, auditable action instead of a vibe.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Who should wait: teams whose apps genuinely need a shell or system packages at runtime, and teams that need pinned versions but have no budget for the paid catalog yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/sathpal/chainguard-aks-demo &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;chainguard-aks-demo
make tools &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make demo                                  &lt;span class="c"&gt;# build, scan, verify, sbom, report&lt;/span&gt;
&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env                                     &lt;span class="c"&gt;# set your registry name and region&lt;/span&gt;
make aks-up &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make acr-push &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make deploy &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make policy   &lt;span class="c"&gt;# the AKS part, then make aks-down&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your numbers differ from mine, that is expected. The whole point is that both images change every day, and only one of them changes in your favour.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>azure</category>
      <category>security</category>
      <category>devops</category>
    </item>
    <item>
      <title>Picking the Latest Kubernetes Release on AKS Without Shooting Yourself in the Foot</title>
      <dc:creator>Sathpal Singh</dc:creator>
      <pubDate>Fri, 24 Jul 2026 07:41:22 +0000</pubDate>
      <link>https://dev.to/sathpal_singh/picking-the-latest-kubernetes-release-on-aks-without-shooting-yourself-in-the-foot-1oij</link>
      <guid>https://dev.to/sathpal_singh/picking-the-latest-kubernetes-release-on-aks-without-shooting-yourself-in-the-foot-1oij</guid>
      <description>&lt;h1&gt;
  
  
  Picking the Latest Kubernetes Release on AKS Without Shooting Yourself in the Foot
&lt;/h1&gt;

&lt;p&gt;Every few months a platform engineer opens a Slack message that reads something like: &lt;em&gt;"Hey, our AKS cluster is on a version that's being retired — what do we do?"&lt;/em&gt; The answer is never "just click upgrade," but it also shouldn't be a three-week fire drill. This post walks through how AKS versions work, how to find what's actually available in your region, and how to set a defensible upgrade posture — without repeating the node pool mechanics we've already covered in earlier posts in this series.&lt;/p&gt;

&lt;h2&gt;
  
  
  How AKS Manages Version Support Windows
&lt;/h2&gt;

&lt;p&gt;AKS doesn't support every Kubernetes minor version forever. Microsoft maintains a sliding support window — commonly described as the three most recent minor versions (N, N-1, N-2) — and retires older releases on a rolling basis. &lt;sup id="fnref1"&gt;1&lt;/sup&gt; NEEDS_VALIDATION: confirm the exact N/N-1/N-2 window against current Microsoft docs, as the supported count has varied historically.&lt;/p&gt;

&lt;p&gt;What that means operationally: if 1.30 is the current stable release, you can expect 1.29 and 1.28 to remain supported, while 1.27 approaches end-of-life. Clusters still running a retired version don't immediately explode, but they stop receiving security patches and you lose Microsoft support SLA coverage — two things you do not want to explain to your CISO.&lt;/p&gt;

&lt;p&gt;Patch versions (the third digit) within a supported minor are more granular. AKS regularly ships node OS patches and Kubernetes patch releases; these are where most of the day-to-day CVE fixes land.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding What's Available in Your Region
&lt;/h2&gt;

&lt;p&gt;Version availability is &lt;strong&gt;per-region&lt;/strong&gt;. A version that GA'd last week in East US may still be in preview in Southeast Asia. Always query your actual deployment region before planning an upgrade:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az aks get-versions &lt;span class="nt"&gt;--location&lt;/span&gt; eastus &lt;span class="nt"&gt;--output&lt;/span&gt; table
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NEEDS_VALIDATION: confirm &lt;code&gt;--output table&lt;/code&gt; renders usefully for this subcommand in current CLI versions; &lt;code&gt;--output json&lt;/code&gt; is always safe if the table format shifts.&lt;/p&gt;

&lt;p&gt;The output lists minor versions, their patch variants, and whether each is in preview or GA. Treat anything marked preview as off-limits for production unless you have a specific reason and the risk tolerance to match.&lt;/p&gt;

&lt;p&gt;If you want the current default version AKS would pick if you didn't specify one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az aks get-versions &lt;span class="nt"&gt;--location&lt;/span&gt; eastus &lt;span class="nt"&gt;--output&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="s1"&gt;'[.values[] | select(.isDefault == true)]'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NEEDS_VALIDATION: the &lt;code&gt;isDefault&lt;/code&gt; field name — verify against current CLI JSON schema.&lt;/p&gt;

&lt;p&gt;Key insight: &lt;strong&gt;the AKS default is not necessarily the latest GA patch&lt;/strong&gt;. &lt;sup id="fnref1"&gt;1&lt;/sup&gt; Microsoft selects a "recommended" version that has had some soak time. If you want the latest, you have to ask for it explicitly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pinning a Version at Cluster Create Time
&lt;/h2&gt;

&lt;p&gt;To be explicit about your version at cluster creation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az aks create &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; my-rg &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; my-cluster &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--location&lt;/span&gt; eastus &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kubernetes-version&lt;/span&gt; 1.30.2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--node-count&lt;/span&gt; 3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--generate-ssh-keys&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NEEDS_VALIDATION: substitute the actual latest GA patch for your region; &lt;code&gt;1.30.2&lt;/code&gt; is illustrative only — do not copy-paste without checking &lt;code&gt;az aks get-versions&lt;/code&gt; first.&lt;/p&gt;

&lt;p&gt;Omitting &lt;code&gt;--kubernetes-version&lt;/code&gt; hands the decision to Microsoft's recommended default. That's fine for dev clusters. For production, pin it, document why you chose that version, and put the upgrade decision in your change process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Auto-Upgrade Channels: Declaring Upgrade Intent
&lt;/h2&gt;

&lt;p&gt;Manually upgrading clusters is fine until you have fifteen of them. Auto-upgrade channels let you express &lt;em&gt;intent&lt;/em&gt; rather than managing individual upgrade events. AKS supports the following channels &lt;sup id="fnref1"&gt;1&lt;/sup&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;none&lt;/strong&gt; — no automatic upgrades; you control everything manually.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;patch&lt;/strong&gt; — automatically upgrades to the latest patch within the current minor version.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;stable&lt;/strong&gt; — targets the latest patch on the N-1 minor version.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;rapid&lt;/strong&gt; — targets the latest patch on the latest supported minor version (N).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;node-image&lt;/strong&gt; — upgrades only the node OS image, not the Kubernetes version itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;NEEDS_VALIDATION: confirm exact current semantics for &lt;code&gt;stable&lt;/code&gt; and &lt;code&gt;rapid&lt;/code&gt; against Microsoft docs — the mapping of "N" vs "N-1" has been adjusted in past releases.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;node-image&lt;/code&gt; channel is worth calling out explicitly: &lt;strong&gt;node image upgrades and Kubernetes version upgrades are decoupled&lt;/strong&gt;. &lt;sup id="fnref1"&gt;1&lt;/sup&gt; You can roll fresh OS images (security patches, kernel updates) without touching your Kubernetes minor or patch version. This is the channel most production clusters should be running continuously — there's little reason to let node images go stale.&lt;/p&gt;

&lt;p&gt;Setting the channel on an existing cluster:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az aks update &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; my-rg &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; my-cluster &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--auto-upgrade-channel&lt;/span&gt; patch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Gating Upgrades with Planned Maintenance Windows
&lt;/h2&gt;

&lt;p&gt;Auto-upgrade without a maintenance window means upgrades can fire at any time AKS decides conditions are met. That's a bad day when it happens at 2 PM on a Tuesday during peak traffic. Planned Maintenance lets you constrain when auto-upgrades are allowed to execute. &lt;sup id="fnref1"&gt;1&lt;/sup&gt; NEEDS_VALIDATION: confirm current Planned Maintenance API capabilities and whether it gates both K8s and node-image upgrades or only one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;maintenance.azure.com/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;MaintenanceConfiguration&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;aksmaintenance&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;timeInWeek&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;day&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Saturday&lt;/span&gt;
      &lt;span class="na"&gt;hourSlots&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;
  &lt;span class="na"&gt;notAllowedTime&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;start&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2024-12-20T00:00:00Z"&lt;/span&gt;
      &lt;span class="na"&gt;end&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2025-01-03T00:00:00Z"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NEEDS_VALIDATION: the above YAML structure is illustrative — validate the exact CRD schema against current AKS documentation before applying.&lt;/p&gt;

&lt;p&gt;Alternatively via CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az aks maintenanceconfiguration add &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; my-rg &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cluster-name&lt;/span&gt; my-cluster &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; aksManagedAutoUpgradeSchedule &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--weekday&lt;/span&gt; Saturday &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--start-hour&lt;/span&gt; 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NEEDS_VALIDATION: confirm subcommand name and flags for the current CLI version.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Concrete Upgrade Decision Process
&lt;/h2&gt;

&lt;p&gt;Here's the numbered checklist I'd run through before touching a production cluster's Kubernetes version:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Query available versions&lt;/strong&gt; in your target region with &lt;code&gt;az aks get-versions&lt;/code&gt;. Confirm the version you want is GA, not preview.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check your current cluster version&lt;/strong&gt; with &lt;code&gt;az aks show --resource-group my-rg --name my-cluster --query kubernetesVersion&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review the Kubernetes changelog&lt;/strong&gt; for the target minor version. Focus on API deprecations — if your workloads use APIs removed in the target version, fix that first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run &lt;code&gt;kubectl deprecations&lt;/code&gt;&lt;/strong&gt; (via the &lt;code&gt;kubent&lt;/code&gt; tool or similar) against your cluster to catch any in-use deprecated APIs before the upgrade window. NEEDS_VALIDATION: confirm &lt;code&gt;kubent&lt;/code&gt; compatibility with your target version.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upgrade control plane first&lt;/strong&gt;, validate, then upgrade node pools — this is standard AKS sequencing. If you need the node pool mechanics, see our earlier posts in this series.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set your auto-upgrade channel&lt;/strong&gt; to &lt;code&gt;patch&lt;/code&gt; at minimum so you don't fall behind on patch releases between planned minor upgrades.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Configure a Planned Maintenance window&lt;/strong&gt; so automated patch upgrades don't surprise you in business hours.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Version Skew Problem
&lt;/h2&gt;

&lt;p&gt;One thing that catches teams off guard: AKS allows your control plane and node pools to differ by at most one minor version, with the control plane always ahead. &lt;sup id="fnref2"&gt;2&lt;/sup&gt; If you're running control plane 1.30 and node pools at 1.28, you have a problem — the node pools need to come up through 1.29 before they can reach 1.30. You cannot skip minor versions on node pools.&lt;/p&gt;

&lt;p&gt;This is why letting clusters drift is expensive. A cluster that hasn't been upgraded in a year may require sequential minor version upgrades, each with its own validation cycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting It Together
&lt;/h2&gt;

&lt;p&gt;Version management on AKS is less about clicking the upgrade button and more about having a policy that makes the decision for you at the right time. Set the &lt;code&gt;patch&lt;/code&gt; auto-upgrade channel, add a maintenance window, keep node images continuously current via &lt;code&gt;node-image&lt;/code&gt; channel, and plan minor version upgrades quarterly. That's a posture you can defend, automate, and sleep through.&lt;/p&gt;

&lt;p&gt;The clusters that cause incidents are the ones where version management was treated as a one-time task rather than a continuous process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
    A([Platform Engineer]) --&amp;gt; B[az aks get-versions\n--location region]
    B --&amp;gt; C{Version Selection\nStrategy}
    C --&amp;gt;|Pin explicit version| D[az aks create\n--kubernetes-version X.Y.Z]
    C --&amp;gt;|Declare upgrade intent| E[Set Auto-Upgrade Channel]
    E --&amp;gt; F{Channel Choice}
    F --&amp;gt;|Conservative| G[patch\nStay on current minor]
    F --&amp;gt;|Balanced| H[stable\nN-1 minor version]
    F --&amp;gt;|Aggressive| I[rapid\nLatest supported minor]
    D --&amp;gt; J[AKS Cluster]
    G --&amp;gt; J
    H --&amp;gt; J
    I --&amp;gt; J
    J --&amp;gt; K[Planned Maintenance\nWindow Gates Upgrades]
    K --&amp;gt; L{Upgrade Scope}
    L --&amp;gt;|K8s version| M[Control Plane\n+ Node Pools]
    L --&amp;gt;|OS patches only| N[Node Image Upgrade\nDecoupled from K8s]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/aks/what-is-aks" rel="noopener noreferrer"&gt;What is Azure Kubernetes Service (AKS)?&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/aks/concepts-clusters-workloads" rel="noopener noreferrer"&gt;Azure Kubernetes Service (AKS) Core Concepts&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>kubernetes</category>
      <category>aks</category>
      <category>azure</category>
      <category>platformengineering</category>
    </item>
    <item>
      <title>AKS and Near-Bare-Metal Workloads: What Platform Teams Can Responsibly Plan For</title>
      <dc:creator>Sathpal Singh</dc:creator>
      <pubDate>Thu, 23 Jul 2026 16:55:32 +0000</pubDate>
      <link>https://dev.to/sathpal_singh/aks-and-near-bare-metal-workloads-what-platform-teams-can-responsibly-plan-for-5927</link>
      <guid>https://dev.to/sathpal_singh/aks-and-near-bare-metal-workloads-what-platform-teams-can-responsibly-plan-for-5927</guid>
      <description>&lt;h1&gt;
  
  
  AKS and Near-Bare-Metal Workloads: What Platform Teams Can Responsibly Plan For
&lt;/h1&gt;

&lt;p&gt;Every few months a team shows up with a workload that has opinions about hardware. GPU inference, high-frequency packet processing, latency-sensitive financial systems — something that makes a shared-tenant VM feel like a traffic jam. The conversation usually ends up at "can we get bare-metal nodes on AKS?"&lt;/p&gt;

&lt;p&gt;The honest answer right now: &lt;em&gt;partially, with caveats, and you should not trust anyone who gives you a confident full answer without pointing at current documentation.&lt;/em&gt; This post is about what you can design around today, what requires validation before you commit to it, and where the gaps are that will bite you if you skip due diligence.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A note on sources:&lt;/strong&gt; The AKS core concepts and product overview documentation available at time of writing covers cluster and workload fundamentals&lt;sup id="fnref1"&gt;1&lt;/sup&gt;&lt;sup id="fnref2"&gt;2&lt;/sup&gt; but does not address bare-metal or near-bare-metal node pool specifics. Claims in that category are marked &lt;code&gt;NEEDS_VALIDATION&lt;/code&gt; throughout. Everything else is grounded in general AKS behavior that is broadly consistent across the platform.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Why "Bare-Metal on AKS" Is Complicated Terminology
&lt;/h2&gt;

&lt;p&gt;AKS runs on Azure VMs. Azure VMs run on Azure's physical fleet. The question is how much of the hypervisor overhead and hardware sharing model you can escape.&lt;/p&gt;

&lt;p&gt;Azure offers VM SKU families that are designed to minimize virtualization overhead — isolated VM sizes that occupy an entire physical host, giving you a single-tenant hardware environment. Whether specific SKUs in this category are available as AKS node pool targets depends on regional availability and AKS's supported VM size list. &lt;code&gt;NEEDS_VALIDATION&lt;/code&gt; — check the current AKS supported VM sizes documentation for your target region before designing around a specific SKU.&lt;/p&gt;

&lt;p&gt;What you get with isolated or near-bare-metal SKUs, if available:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Predictable NUMA topology&lt;/strong&gt; — no neighbor noise on the physical host&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full access to hardware features&lt;/strong&gt; — relevant for SR-IOV, DPDK, and accelerated networking scenarios (&lt;code&gt;NEEDS_VALIDATION&lt;/code&gt; for specific feature compatibility with AKS CNI configurations)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single-tenant physical isolation&lt;/strong&gt; — meaningful for compliance frameworks that require it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What you do NOT automatically get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A different version of Kubernetes&lt;/li&gt;
&lt;li&gt;Different node lifecycle behavior&lt;/li&gt;
&lt;li&gt;Escape from AKS's node auto-repair or OS upgrade mechanics&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point is where platform teams usually get surprised.&lt;/p&gt;




&lt;h2&gt;
  
  
  Node Lifecycle Is Still AKS's Domain
&lt;/h2&gt;

&lt;p&gt;AKS manages node health and upgrades at the node pool level&lt;sup id="fnref1"&gt;1&lt;/sup&gt;. This does not change because you picked a specialty VM SKU. If auto-repair decides a node is unhealthy and reprovisioned it, the replacement VM must come from available capacity of the same SKU. For isolated VM families, that capacity is often constrained — regionally and numerically.&lt;/p&gt;

&lt;p&gt;The practical consequence: &lt;strong&gt;your cluster auto-repair and upgrade assumptions need re-evaluation for specialty node pools.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For standard node pools on commodity SKUs, AKS node upgrade behavior is well-documented and we've covered it in depth in our prior upgrade series (#2, #42, #50). The short version is that AKS handles node image upgrades via cordon-drain-replace, and you control the surge buffer and max unavailable settings. That mechanical behavior is the same regardless of SKU. What changes is the risk profile:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Capacity risk:&lt;/strong&gt; If your isolated SKU has low regional availability, a surge node during upgrade may not provision. The upgrade stalls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drain timing risk:&lt;/strong&gt; Pods on high-performance hardware often have longer graceful termination windows. Make sure your &lt;code&gt;terminationGracePeriodSeconds&lt;/code&gt; reflects that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OS upgrade compatibility:&lt;/strong&gt; Node OS image updates on specialty hardware SKUs may lag behind standard SKU rollouts. &lt;code&gt;NEEDS_VALIDATION&lt;/code&gt; — check AKS release notes for your target SKU family.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Node Pool Isolation Strategy
&lt;/h2&gt;

&lt;p&gt;If you're mixing standard workloads and hardware-sensitive workloads in the same cluster (which is a reasonable cost efficiency decision), node pool isolation is non-negotiable. This is standard AKS practice&lt;sup id="fnref1"&gt;1&lt;/sup&gt; and applies regardless of whether the specialty pool is "bare-metal adjacent" or not.&lt;/p&gt;

&lt;p&gt;The pattern:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Separate node pools by workload class&lt;/strong&gt; — one pool for standard platform services, one for hardware-intensive workloads&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Taint specialty node pools&lt;/strong&gt; — prevent accidental scheduling of standard workloads onto expensive hardware&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use node selectors and tolerations explicitly&lt;/strong&gt; — don't rely on the absence of a node being "wrong enough" to avoid it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set resource requests and limits that reflect actual hardware&lt;/strong&gt; — on high-performance SKUs with full NUMA access, misconfigured limits can be worse than no limits&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here's a minimal node pool configuration pattern for a tainted specialty pool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Create a specialty node pool with a taint to prevent accidental scheduling&lt;/span&gt;
&lt;span class="c"&gt;# Replace &amp;lt;SKU_NAME&amp;gt; with your validated target SKU&lt;/span&gt;
az aks nodepool add &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; myRG &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cluster-name&lt;/span&gt; myCluster &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; specialtypool &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--node-vm-size&lt;/span&gt; &amp;lt;SKU_NAME&amp;gt; &lt;span class="se"&gt;\ &lt;/span&gt;&lt;span class="c"&gt;# NEEDS_VALIDATION: confirm SKU is AKS-supported&lt;/span&gt;
  &lt;span class="nt"&gt;--node-count&lt;/span&gt; 2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--node-taints&lt;/span&gt; workload-class&lt;span class="o"&gt;=&lt;/span&gt;hardware-intensive:NoSchedule &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--labels&lt;/span&gt; workload-class&lt;span class="o"&gt;=&lt;/span&gt;hardware-intensive &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--zones&lt;/span&gt; 1 2 3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the corresponding pod spec side:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;hardware-sensitive-workload&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;tolerations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;workload-class"&lt;/span&gt;
      &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Equal"&lt;/span&gt;
      &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hardware-intensive"&lt;/span&gt;
      &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NoSchedule"&lt;/span&gt;
  &lt;span class="na"&gt;nodeSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;workload-class&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;hardware-intensive&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;app&lt;/span&gt;
      &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;your-registry/your-image&lt;/span&gt; &lt;span class="c1"&gt;# pin your digest, don't use latest&lt;/span&gt;
      &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8"&lt;/span&gt;
          &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;32Gi"&lt;/span&gt;
        &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8"&lt;/span&gt;
          &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;32Gi"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Concrete Steps Before Committing to a Specialty Node Pool
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Validate SKU availability in your target region.&lt;/strong&gt; Run &lt;code&gt;az vm list-skus --location &amp;lt;region&amp;gt; --size &amp;lt;sku-prefix&amp;gt; --output table&lt;/code&gt; and cross-reference against the AKS supported VM sizes list. Do not assume a SKU that works for standalone VMs is available for AKS node pools.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Check current AKS release notes for the SKU family.&lt;/strong&gt; Look specifically for known issues with node auto-repair, OS image compatibility, or accelerated networking feature support. &lt;code&gt;NEEDS_VALIDATION&lt;/code&gt; — this is release-cadence dependent.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Model your capacity headroom.&lt;/strong&gt; For upgrade surge and auto-repair replacement, you need available quota &lt;em&gt;and&lt;/em&gt; available physical capacity. For isolated SKUs, physical capacity is the harder constraint. File a support ticket to confirm capacity commitments if the workload is business-critical.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Test node pool upgrade behavior in a non-production cluster first.&lt;/strong&gt; Don't discover that your specialty SKU has zero surge capacity during a production Kubernetes version upgrade. The upgrade series we've covered previously (#50 is the most current approved version) gives you the tooling — apply the same runbook to a test specialty pool before you need it for real.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Define your eviction and rescheduling strategy.&lt;/strong&gt; If a node is auto-repaired and the replacement takes longer than expected (capacity contention), what happens to your workload? PodDisruptionBudgets should reflect the actual tolerance of the workload, not the default.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Align with your golden path guardrails.&lt;/strong&gt; If your platform has golden path templates (covered in #25), specialty node pools should either be an explicit supported variant or explicitly out-of-scope. The worst outcome is an undocumented workaround that every team discovers independently and implements differently.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What Not To Do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Don't build production dependency on a feature that's in preview.&lt;/strong&gt; AKS preview features can change or be dropped. If a specific bare-metal adjacent capability is currently in preview, treat it as R&amp;amp;D, not production infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't assume SKU-level isolation means network-level isolation.&lt;/strong&gt; Physical host isolation and network security are different axes. Your CNI configuration and network policies still matter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't skip the PodDisruptionBudget.&lt;/strong&gt; On hardware-intensive workloads, the instinct is often "this thing is stateful and complex, let's not let Kubernetes touch it." That instinct leads to nodes that can't be drained and upgrades that time out.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The Responsible Summary
&lt;/h2&gt;

&lt;p&gt;AKS can host workloads on VM SKUs designed to reduce virtualization overhead and provide single-tenant hardware access. The exact SKUs available, their compatibility with specific AKS networking and acceleration features, and their behavior under node lifecycle operations all require validation against current documentation before you commit. &lt;code&gt;NEEDS_VALIDATION&lt;/code&gt; on the specifics is not a cop-out — it's the difference between a platform that works and a platform that works until it doesn't.&lt;/p&gt;

&lt;p&gt;What you &lt;em&gt;can&lt;/em&gt; design now: isolation strategy, taint/toleration patterns, upgrade runbooks, and capacity planning processes. Get those right and the hardware specifics slot in cleanly once validated.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
    A[Platform Team Evaluation] --&amp;gt; B{Workload Requirements}
    B --&amp;gt; C[Standard Latency\nGeneral Purpose VMs]
    B --&amp;gt; D[Low Latency / Hardware Intensive\nSpecialty VM SKUs]
    D --&amp;gt; E[Validate SKU Availability\nby Region]
    E --&amp;gt; F[Dedicated Node Pool\nIsolation Strategy]
    F --&amp;gt; G[Review Auto-Repair &amp;amp;\nOS Upgrade Behavior]
    G --&amp;gt; H[Node Pool Taints &amp;amp; Tolerations\nfor Workload Scheduling]
    C --&amp;gt; H
    H --&amp;gt; I[Mixed Cluster:\nStandard + Specialty Pools]
    I --&amp;gt; J[Ongoing: Monitor AKS\nRelease Notes &amp;amp; SKU Docs]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/aks/concepts-clusters-workloads" rel="noopener noreferrer"&gt;Azure Kubernetes Service (AKS) Core Concepts&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/aks/what-is-aks" rel="noopener noreferrer"&gt;What is Azure Kubernetes Service (AKS)?&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>kubernetes</category>
      <category>azurekubernetesservice</category>
      <category>nodepools</category>
      <category>platformengineering</category>
    </item>
    <item>
      <title>Karpenter on AKS in 2026: What Actually Works</title>
      <dc:creator>Sathpal Singh</dc:creator>
      <pubDate>Sun, 31 May 2026 05:21:06 +0000</pubDate>
      <link>https://dev.to/sathpal_singh/karpenter-on-aks-in-2026-what-actually-works-1meh</link>
      <guid>https://dev.to/sathpal_singh/karpenter-on-aks-in-2026-what-actually-works-1meh</guid>
      <description>&lt;h1&gt;
  
  
  Karpenter on AKS in 2026: What Actually Works
&lt;/h1&gt;

&lt;p&gt;Karpenter on AKS has gone from "interesting experiment" to "something you can actually run in production" with some caveats that will save you a weekend of pain if you read them now. This post is a field report, not a sales pitch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Short Version
&lt;/h2&gt;

&lt;p&gt;If you're running homogeneous, predictable workloads and you're happy with cluster-autoscaler (CAS), stay there. CAS is boring, it works, and Azure supports it fully. If you're running GPU workloads, spot-heavy batch pipelines, or you need bin-packing that doesn't require you to pre-define a node pool for every VM SKU you might want, Karpenter is now worth the operational overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Karpenter Actually Does on AKS
&lt;/h2&gt;

&lt;p&gt;Karpenter watches for unschedulable pods and provisions nodes directly via the Azure provider, no VMSS node pools required for every SKU combination. It provisions, consolidates, and terminates nodes based on pod requirements and your defined &lt;code&gt;NodePool&lt;/code&gt; and &lt;code&gt;AKSNodeClass&lt;/code&gt; resources.&lt;/p&gt;

&lt;p&gt;The AKS provider for Karpenter (&lt;code&gt;karpenter-provider-azure&lt;/code&gt;) is a separate project from the AWS provider. Same core Karpenter engine, different provider implementation. This matters because feature parity with AWS Karpenter is not guaranteed and the cadence of releases differs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites and Installation
&lt;/h2&gt;

&lt;p&gt;You need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An AKS cluster with &lt;code&gt;--network-plugin azure&lt;/code&gt; or &lt;code&gt;--network-plugin overlay&lt;/code&gt; (Azure CNI in either mode works; kubenet is not supported)&lt;/li&gt;
&lt;li&gt;A managed identity with the right RBAC, the provider needs to create and delete VMs and manage NICs, disks, and NSGs&lt;/li&gt;
&lt;li&gt;Workload identity enabled on the cluster&lt;/li&gt;
&lt;li&gt;Karpenter installed via Helm&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The official installation path uses Helm with values pulled from your cluster. Here's a stripped-down install sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Set environment variables: replace with your actual values&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CLUSTER_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"my-aks-cluster"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;RESOURCE_GROUP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"my-rg"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;LOCATION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"eastus2"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;KARPENTER_NAMESPACE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"kube-system"&lt;/span&gt;

&lt;span class="c"&gt;# Get cluster details needed for Karpenter config&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SUBSCRIPTION_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;az account show &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; tsv&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;NODE_RESOURCE_GROUP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;az aks show &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CLUSTER_NAME&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resource-group&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;RESOURCE_GROUP&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; nodeResourceGroup &lt;span class="nt"&gt;-o&lt;/span&gt; tsv&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# Install via Helm&lt;/span&gt;
&lt;span class="c"&gt;# NEEDS_VALIDATION: confirm chart version and repo URL against&lt;/span&gt;
&lt;span class="c"&gt;# https://github.com/Azure/karpenter-provider-azure at time of deployment&lt;/span&gt;
helm upgrade &lt;span class="nt"&gt;--install&lt;/span&gt; karpenter oci://mcr.microsoft.com/aks/karpenter/karpenter &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;KARPENTER_NAMESPACE&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--create-namespace&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; &lt;span class="s2"&gt;"settings.clusterName=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CLUSTER_NAME&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; &lt;span class="s2"&gt;"settings.location=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;LOCATION&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; &lt;span class="s2"&gt;"settings.subscriptionID=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SUBSCRIPTION_ID&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; &lt;span class="s2"&gt;"settings.resourceGroup=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;NODE_RESOURCE_GROUP&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--wait&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Defining Your First NodePool
&lt;/h2&gt;

&lt;p&gt;This is where Karpenter's model diverges most from node pools. Instead of pre-creating a pool for every SKU you might want, you define constraints and let Karpenter pick:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.sh/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;NodePool&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;general&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;workload-type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;general&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;nodeClassRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.azure.com/v1alpha2&lt;/span&gt;
        &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AKSNodeClass&lt;/span&gt;
        &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
      &lt;span class="na"&gt;requirements&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.sh/capacity-type&lt;/span&gt;
          &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
          &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;on-demand"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spot"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubernetes.io/arch&lt;/span&gt;
          &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
          &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amd64"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.azure.com/sku-family&lt;/span&gt;
          &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
          &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;D"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;E"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.azure.com/sku-version&lt;/span&gt;
          &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Gt&lt;/span&gt;
          &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;200"&lt;/span&gt;
    &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;800Gi&lt;/span&gt;
  &lt;span class="na"&gt;disruption&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;consolidationPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;WhenEmptyOrUnderutilized&lt;/span&gt;
    &lt;span class="na"&gt;consolidateAfter&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;30s&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.azure.com/v1alpha2&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AKSNodeClass&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;imageFamily&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AzureLinux&lt;/span&gt;
  &lt;span class="na"&gt;osDiskSizeGB&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;128&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note on API versions&lt;/strong&gt;: The &lt;code&gt;karpenter.azure.com&lt;/code&gt; API group is versioned separately from upstream Karpenter. &lt;code&gt;v1alpha2&lt;/code&gt; was current as of early 2026 but &lt;strong&gt;NEEDS_VALIDATION&lt;/strong&gt; check the CRD definitions in the installed chart before you copy this into a GitOps repo.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What Works Well
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Spot consolidation&lt;/strong&gt; is the headline win. When spot VMs get preempted or you have underutilized nodes, Karpenter's consolidation loop handles bin-packing without you writing any automation. With CAS you're responsible for node pool min/max sizing and the consolidation is coarse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-SKU scheduling&lt;/strong&gt; is the other real win. A pod requesting 8 vCPU and 64 GiB RAM will cause Karpenter to search the allowed SKU families for a node that fits, rather than failing because your single pre-configured node pool is exhausted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GPU node provisioning&lt;/strong&gt; works, including time-slicing scenarios, as long as your &lt;code&gt;AKSNodeClass&lt;/code&gt; uses an image family that ships the NVIDIA drivers. AzureLinux with GPU extensions does this. You still need to manage the device plugin separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Still Rough
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Node provisioning latency&lt;/strong&gt; is higher than AWS because Azure VM creation is slower than EC2. Plan for 3–5 minutes from pod pending to node ready on cold starts. This isn't a Karpenter problem per se, but it affects how you design your buffer capacity and PodDisruptionBudgets.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Windows node pools&lt;/strong&gt; are not supported by the Azure Karpenter provider. If you have Windows workloads, keep a static node pool managed by CAS or manual scaling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Custom VNet/subnet selection&lt;/strong&gt; requires care. The &lt;code&gt;AKSNodeClass&lt;/code&gt; lets you specify subnet IDs, but if you're using private clusters with complex network topologies, test thoroughly before rolling to production. Subnet exhaustion errors surface late and are annoying to debug.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Observability is immature&lt;/strong&gt;. Karpenter emits metrics to Prometheus and logs to stdout, but the AKS provider's specific actions (VM creation, NIC attachment) aren't surfaced as well as you'd want. You'll be reading &lt;code&gt;kubectl logs&lt;/code&gt; more than you'd like.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concrete Steps to a Safe Rollout
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Start with a non-production cluster.&lt;/strong&gt; Run Karpenter alongside CAS, not instead of it. CAS can manage your system node pool; Karpenter handles a &lt;code&gt;workload&lt;/code&gt; node pool namespace.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Define &lt;code&gt;limits&lt;/code&gt; on your NodePool.&lt;/strong&gt; Without a CPU/memory ceiling, a scheduling bug or runaway HPA can provision hundreds of nodes before you notice. Set limits conservatively and raise them deliberately.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Set &lt;code&gt;consolidateAfter&lt;/code&gt; to something sane for your workload.&lt;/strong&gt; 30 seconds is aggressive for stateful apps. Use 5–10 minutes for anything with slow startup or persistent volumes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Test spot preemption handling.&lt;/strong&gt; Deploy a test workload on spot nodes and manually deallocate a VM. Verify that Karpenter reprovisioned within your acceptable window and that your pod disruption budgets held.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Add Karpenter's node labels to your monitoring dashboards.&lt;/strong&gt; Specifically track &lt;code&gt;karpenter.sh/capacity-type&lt;/code&gt;, &lt;code&gt;karpenter.azure.com/sku-name&lt;/code&gt;, and &lt;code&gt;karpenter.sh/nodepool&lt;/code&gt; as label dimensions so you can see cost and performance breakdown by node type.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pin your Helm chart version in GitOps.&lt;/strong&gt; The provider is still in active development. Uncontrolled upgrades have broken NodePool CRD schemas between minor versions. Treat upgrades as a planned event.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  CAS vs. Karpenter: The Honest Comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;CAS&lt;/th&gt;
&lt;th&gt;Karpenter&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Operational maturity on AKS&lt;/td&gt;
&lt;td&gt;Production-grade&lt;/td&gt;
&lt;td&gt;Production-capable with caveats&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-SKU bin-packing&lt;/td&gt;
&lt;td&gt;Requires pre-defined pools&lt;/td&gt;
&lt;td&gt;Native&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spot handling&lt;/td&gt;
&lt;td&gt;Decent&lt;/td&gt;
&lt;td&gt;Better&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows nodes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debug tooling&lt;/td&gt;
&lt;td&gt;Mature&lt;/td&gt;
&lt;td&gt;Developing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure support&lt;/td&gt;
&lt;td&gt;First-party&lt;/td&gt;
&lt;td&gt;Community + Microsoft OSS&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;Karpenter on AKS in 2026 is the right choice if you have heterogeneous compute requirements and engineering capacity to own the operational model. It is not yet the "set it and forget it" experience that CAS is for straightforward clusters.&lt;/p&gt;

&lt;p&gt;The Azure team has been shipping at a reasonable pace and the GitHub issues backlog is actually getting shorter, which is a good sign. The API is stabilizing. The path from alpha to beta to stable is visible.&lt;/p&gt;

&lt;p&gt;Just don't copy that YAML into production without validating the API versions first. I warned you.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>aks</category>
      <category>karpenter</category>
      <category>autoscaling</category>
    </item>
  </channel>
</rss>
