<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Billy Walker</title>
    <description>The latest articles on DEV Community by Billy Walker (@bwlkr).</description>
    <link>https://dev.to/bwlkr</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3987873%2F8b891f89-a0f7-4af8-b25d-f746595125dc.png</url>
      <title>DEV Community: Billy Walker</title>
      <link>https://dev.to/bwlkr</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bwlkr"/>
    <language>en</language>
    <item>
      <title>GitOps with Helm for small teams: what</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:34:34 +0000</pubDate>
      <link>https://dev.to/coresolutions/gitops-with-helm-for-small-teams-what-59cc</link>
      <guid>https://dev.to/coresolutions/gitops-with-helm-for-small-teams-what-59cc</guid>
      <description>&lt;p&gt;Most small teams arrive at GitOps from the same place. Every service might already have a Helm chart, but production is whatever someone last ran &lt;code&gt;helm upgrade&lt;/code&gt; with from their laptop, plus the &lt;code&gt;--set replicaCount=4&lt;/code&gt; somebody added during last month's incident, plus a values file that only exists in one engineer's home directory. Then they go looking for help and find a wall of machinery: app-of-apps layouts, promotion pipelines, multi-cluster generators, progressive delivery. If you are a team of three trying to stop deploys happening from somebody's laptop, that is not a realistic starting point.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For small teams, the part of GitOps that pays off immediately is the reconciliation loop. Most of the machinery built around it can wait.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Strip GitOps back to first principles and it is genuinely useful. The &lt;a href="https://opengitops.dev/" rel="noopener noreferrer"&gt;OpenGitOps principles&lt;/a&gt; give you four: desired state is declarative, versioned and immutable, pulled automatically, and continuously reconciled. Those get you reviewed change, a real audit trail, drift detection and rollback by &lt;code&gt;git revert&lt;/code&gt;, and none of them needs an elaborate delivery platform.&lt;/p&gt;

&lt;p&gt;If you already use Helm, most of the pieces are in place. What you need on top is small: charts and values files in Git, one reconciler to apply them, a values layout that keeps environments honest, and a secrets plan that is not reckless. This post takes each in turn, then covers the machinery you can leave alone and the sharp edges you will meet once the loop is running.&lt;/p&gt;

&lt;h2&gt;
  
  
  The useful bit of GitOps is boring in the best way
&lt;/h2&gt;

&lt;p&gt;Once a tool such as Argo CD or Flux is watching a repository and rendering your charts from it, a few things improve straight away:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Change becomes reviewable.&lt;/strong&gt; Every deploy is a commit to a chart or a values file, and usually a pull request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rollback gets simpler.&lt;/strong&gt; In the happy path it is a &lt;code&gt;git revert&lt;/code&gt; rather than a reconstruction job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drift becomes visible.&lt;/strong&gt; When someone changes a live object out of band, the reconciler notices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cluster state becomes reproducible.&lt;/strong&gt; You stop relying on memory and shell history to explain which values produced what is running.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That rollback point changes one Helm habit. Under GitOps, &lt;code&gt;helm rollback&lt;/code&gt; stops being the tool for the job. On Argo CD there is no Helm release to roll back, and on Flux a manual rollback leaves the release out of step with what Git declares, so the controller upgrades it straight back. Revert the commit instead and let the loop do the deploy.&lt;/p&gt;

&lt;p&gt;All of this is a meaningful operational upgrade at any size, and it lines up with DORA's research. Its &lt;a href="https://dora.dev/capabilities/continuous-delivery/" rel="noopener noreferrer"&gt;continuous delivery capability&lt;/a&gt; lists version control for all production artefacts, and deployment automation, among the practices that drive continuous delivery, which in turn improves delivery performance. I would be careful not to overclaim here. DORA does not measure GitOps as a separate intervention with a tidy uplift figure. What it does back is the discipline GitOps happens to enforce well.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the loop
&lt;/h2&gt;

&lt;p&gt;If I were introducing GitOps to a team that already ships Helm charts, I would start with a setup that feels almost underwhelming:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;charts and their values in Git, with environment differences as values files on one branch&lt;/li&gt;
&lt;li&gt;one reconciler running in the cluster&lt;/li&gt;
&lt;li&gt;pull requests as the only normal path into production&lt;/li&gt;
&lt;li&gt;chart versions pinned, so the same commit always renders the same manifests&lt;/li&gt;
&lt;li&gt;automated sync switched on, with pruning and self-heal treated as things the team earns&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point deserves a closer look, because the two main tools handle it differently. Self-healing sounds great until an urgent out-of-band fix gets reverted mid-incident because Git still says otherwise. The controller is doing what you asked, not what you meant five minutes into an outage.&lt;/p&gt;

&lt;p&gt;Argo CD leaves both sharp edges off by default. An &lt;code&gt;Application&lt;/code&gt; with automated sync and nothing else set behaves like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;syncPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;automated&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;prune&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
    &lt;span class="na"&gt;selfHeal&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With those settings, Argo CD syncs when a new commit lands. It does not delete resources that disappear from Git, and it does not revert a live change on its own: the drift shows up as &lt;code&gt;OutOfSync&lt;/code&gt; for a human to deal with. Turn on &lt;code&gt;selfHeal&lt;/code&gt; and &lt;code&gt;prune&lt;/code&gt; once the team trusts the loop and, more to the point, once everyone has stopped fixing production by hand.&lt;/p&gt;

&lt;p&gt;Flux splits the same decision across its two controllers, and the split catches Helm users out. A plain &lt;code&gt;Kustomization&lt;/code&gt; re-applies what Git says on every &lt;code&gt;interval&lt;/code&gt; and &lt;a href="https://fluxcd.io/flux/components/kustomize/kustomizations/" rel="noopener noreferrer"&gt;corrects any drift it finds&lt;/a&gt;, with no switch to turn that off. A &lt;code&gt;HelmRelease&lt;/code&gt; is different: &lt;a href="https://fluxcd.io/flux/components/helm/helmreleases/#drift-detection" rel="noopener noreferrer"&gt;drift detection is opt-in&lt;/a&gt;, so by default a &lt;code&gt;kubectl edit&lt;/code&gt; against a Helm-managed Deployment stays put until the next chart or values change triggers an upgrade. You choose the behaviour explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;driftDetection&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;mode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;warn&lt;/span&gt; &lt;span class="c1"&gt;# report drift as an event; change to `enabled` to correct it&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;warn&lt;/code&gt; is the Flux equivalent of Argo CD with self-heal off, and a sensible place to start. The section on running the loop comes back to what to do when you need the reconciler to back off.&lt;/p&gt;

&lt;h2&gt;
  
  
  Argo CD vs Flux: pick by how your team works
&lt;/h2&gt;

&lt;p&gt;People love turning this into theology. It doesn't need to be.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.cncf.io/projects/flux/" rel="noopener noreferrer"&gt;Flux&lt;/a&gt; and &lt;a href="https://www.cncf.io/projects/argo/" rel="noopener noreferrer"&gt;Argo, the project that includes Argo CD&lt;/a&gt;, reached CNCF Graduated status within a week of each other in late 2022. Either can be the right answer, and neither is a toy. The first difference is how people expect to interact with deployments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Argo CD&lt;/strong&gt; suits teams that want an app-centric control plane with a strong UI. You declare &lt;code&gt;Application&lt;/code&gt; resources and get a visual model of sync status, health and drift, so people can see what is going on without living in the CLI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Flux&lt;/strong&gt; suits teams that prefer a composable set of controllers and an experience shaped around Git, the CLI and the Kubernetes API. It feels closer to assembling the controllers you need than to logging into a deployment console.&lt;/p&gt;

&lt;p&gt;The second difference is specific to Helm, and it is worth knowing before you pick. Argo CD &lt;a href="https://argo-cd.readthedocs.io/en/stable/faq/" rel="noopener noreferrer"&gt;uses Helm only as a template engine&lt;/a&gt;: it runs &lt;code&gt;helm template&lt;/code&gt; and applies the rendered manifests itself, so &lt;code&gt;helm list&lt;/code&gt; shows nothing and there is no Helm release history in the cluster. Flux's helm-controller performs real Helm installs and upgrades, so releases, history and &lt;code&gt;helm list&lt;/code&gt; all behave as they did before GitOps, and failed upgrades can be remediated with Helm's own rollback. Neither approach is wrong. If your team leans on Helm tooling and release history, Flux will feel familiar; if you would rather Helm stayed a rendering step and the Argo CD UI became the source of truth for what is deployed, Argo CD's model is simpler to reason about.&lt;/p&gt;

&lt;p&gt;So my heuristic is mostly a social one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;pick &lt;strong&gt;Argo CD&lt;/strong&gt; if people routinely want to look at deploy state and reason about it in one place&lt;/li&gt;
&lt;li&gt;pick &lt;strong&gt;Flux&lt;/strong&gt; if your team is happy in Git and the terminal, wants real Helm releases, and would rather compose small controllers than run a dashboard&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The secrets section below adds one more technical tie-breaker. If you want to go further with Argo, we have a hands-on post on &lt;a href="https://dev.to/blog/gitops-for-kubernetes-with-argo-cd"&gt;GitOps for Kubernetes with Argo CD&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  One repo per service, chart included
&lt;/h2&gt;

&lt;p&gt;Repository debates get oddly ideological. Keep it plain: each service is its own repo, and its Helm chart lives next to the code it deploys. The chart then versions with the application, so a change that needs a new environment variable and the code that reads it lands in one pull request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payments-api/
├── src/
├── Dockerfile
└── chart/
    ├── Chart.yaml
    ├── values.yaml          # base: what every environment shares
    ├── values/
    │   ├── staging.yaml     # only what differs in staging
    │   └── production.yaml  # only what differs in production
    └── templates/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The layout mirrors how Helm already merges values: a base that every environment shares, and a thin file per environment that only says what differs. In Argo CD, each environment's &lt;code&gt;Application&lt;/code&gt; points at the service repo's &lt;code&gt;chart/&lt;/code&gt; directory and names its environment file. The chart's own &lt;code&gt;values.yaml&lt;/code&gt; is always the base, and anything in &lt;code&gt;valueFiles&lt;/code&gt; is &lt;a href="https://argo-cd.readthedocs.io/en/stable/user-guide/helm/" rel="noopener noreferrer"&gt;layered on top of it&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;repoURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://github.com/example/payments-api.git&lt;/span&gt;
  &lt;span class="na"&gt;targetRevision&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1.4.2&lt;/span&gt;
  &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;chart&lt;/span&gt;
  &lt;span class="na"&gt;helm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;valueFiles&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;values/production.yaml&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Flux does the same with a &lt;code&gt;HelmRelease&lt;/code&gt; whose chart comes from the service's &lt;code&gt;GitRepository&lt;/code&gt;. Its &lt;code&gt;valuesFiles&lt;/code&gt; &lt;a href="https://fluxcd.io/flux/components/source/helmcharts/" rel="noopener noreferrer"&gt;replace the chart's default values&lt;/a&gt; and merge in order, with paths relative to the repo root, so list the base first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;chart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;chart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./chart&lt;/span&gt;
      &lt;span class="na"&gt;sourceRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;GitRepository&lt;/span&gt;
        &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payments-api&lt;/span&gt;
      &lt;span class="na"&gt;valuesFiles&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./chart/values.yaml&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./chart/values/production.yaml&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those &lt;code&gt;Application&lt;/code&gt; and &lt;code&gt;HelmRelease&lt;/code&gt; definitions are the only thing that isn't per service: a small deployments repo holding one per service and environment, which is also where the reconciler reads what runs where.&lt;/p&gt;

&lt;p&gt;Argo CD's &lt;a href="https://argo-cd.readthedocs.io/en/stable/user-guide/best_practices/" rel="noopener noreferrer"&gt;best-practices guide&lt;/a&gt; recommends keeping config in a separate repo from application source, and its reasons are worth handling rather than ignoring. A values-only change shouldn't trigger a full CI build, so path-filter the pipeline. CI shouldn't commit image tags back into the repo it builds from, or you get a build loop. And a code commit shouldn't reach production on its own, which is what the pinned &lt;code&gt;targetRevision&lt;/code&gt; above is for: staging tracks &lt;code&gt;main&lt;/code&gt;, production pins a release, and promotion is a reviewed pull request that bumps the pin.&lt;/p&gt;

&lt;p&gt;A few opinions, stated plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use values files for environments, not branches.&lt;/strong&gt; Branch-per-environment sounds tidy until promotion turns into cherry-picking and drift archaeology.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep environment files thin.&lt;/strong&gt; If &lt;code&gt;production.yaml&lt;/code&gt; is most of a copy of &lt;code&gt;values.yaml&lt;/code&gt;, the base isn't doing its job, and the next shared change will land in one file and not the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pin what production runs.&lt;/strong&gt; Pin the service's &lt;code&gt;targetRevision&lt;/code&gt; to a tag or commit, and pin upstream chart versions exactly rather than to a range such as &lt;code&gt;6.5.*&lt;/code&gt;. Let a bot such as &lt;a href="https://dev.to/blog/renovate-kubernetes-helm-terraform"&gt;Renovate&lt;/a&gt; raise the bumps as pull requests. A range means the same commit can render different manifests next week.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One Helm behaviour bites almost everyone who adopts this layout. Helm merges maps deeply, but it &lt;a href="https://github.com/helm/helm/blob/main/pkg/chart/common/util/coalesce.go" rel="noopener noreferrer"&gt;replaces lists outright&lt;/a&gt;: "scalar values and arrays are replaced, maps are merged". So a &lt;code&gt;production.yaml&lt;/code&gt; that adds one environment variable to an &lt;code&gt;env:&lt;/code&gt; list doesn't add it; it replaces every entry the base defined, and the base's variables silently vanish from production. Keep list-valued settings in one file, or have the chart accept a map and render the list from it.&lt;/p&gt;

&lt;p&gt;Shared infrastructure follows the same rule: cert-manager, ingress and External Secrets are each their own chart, deployed before the services that use them. Your charts will create custom resources, such as a cert-manager &lt;code&gt;Certificate&lt;/code&gt; or an &lt;code&gt;ExternalSecret&lt;/code&gt;, and those only work once the CRDs exist and the controller behind them is actually running. It is also a Helm limitation: Helm &lt;a href="https://helm.sh/docs/chart_best_practices/custom_resource_definitions/" rel="noopener noreferrer"&gt;never upgrades or deletes CRDs&lt;/a&gt; from a chart's &lt;code&gt;crds/&lt;/code&gt; directory, and its own docs suggest a separate chart for them. Flux handles the ordering with &lt;a href="https://fluxcd.io/flux/components/helm/helmreleases/" rel="noopener noreferrer"&gt;&lt;code&gt;dependsOn&lt;/code&gt;&lt;/a&gt;, so a service's release waits until the infrastructure it needs reports ready. Argo CD uses &lt;a href="https://argo-cd.readthedocs.io/en/stable/user-guide/sync-waves/" rel="noopener noreferrer"&gt;sync waves&lt;/a&gt;, set with the &lt;code&gt;argocd.argoproj.io/sync-wave&lt;/code&gt; annotation, which order resources within a single Application. Ordering whole Applications needs app-of-apps plus a custom health check for the &lt;code&gt;Application&lt;/code&gt; kind, and that is one of the few good reasons to adopt app-of-apps early.&lt;/p&gt;

&lt;h2&gt;
  
  
  Secrets in GitOps: SOPS, Sealed Secrets or External Secrets
&lt;/h2&gt;

&lt;p&gt;Every GitOps conversation eventually arrives at the part the tidy diagrams leave out, and with Helm the first place a secret leaks is a values file. "Just put them in Git" is not a serious answer. Git is durable, replicated, and very good at remembering things you wish it had forgotten.&lt;/p&gt;

&lt;p&gt;The usual options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SOPS&lt;/strong&gt; encrypts the values in the files you commit, so diffs stay reviewable. &lt;a href="https://fluxcd.io/flux/guides/mozilla-sops/" rel="noopener noreferrer"&gt;Flux decrypts SOPS natively&lt;/a&gt; through &lt;code&gt;spec.decryption&lt;/code&gt; on a Kustomization, and a &lt;code&gt;HelmRelease&lt;/code&gt; can read the decrypted Secret as values through &lt;code&gt;valuesFrom&lt;/code&gt;. Argo CD has no built-in support, so you add SOPS to its repo server with a plugin. That moves decryption into manifest generation, which Argo CD's own &lt;a href="https://argo-cd.readthedocs.io/en/stable/operator-manual/secret-management/" rel="noopener noreferrer"&gt;secret-management guide&lt;/a&gt; strongly cautions against, because generated manifests sit in plaintext in its Redis cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sealed Secrets&lt;/strong&gt; gives you a &lt;code&gt;SealedSecret&lt;/code&gt; that is safe to commit and that only the controller in the target cluster can decrypt. It is easy to explain and a fine early move, but it is cluster-bound. By default the controller renews its sealing key every 30 days and keeps the old ones, so disaster recovery means backing up every secret labelled &lt;code&gt;sealedsecrets.bitnami.com/sealed-secrets-key&lt;/code&gt;, and refreshing that backup after each renewal. Lose those keys along with the cluster and every &lt;code&gt;SealedSecret&lt;/code&gt; in Git is ciphertext nobody can open.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;External Secrets Operator&lt;/strong&gt; keeps references in Git and the values in a real secrets manager such as Vault, AWS Secrets Manager or Azure Key Vault. Your chart templates an &lt;code&gt;ExternalSecret&lt;/code&gt; instead of a &lt;code&gt;Secret&lt;/code&gt;, and the values file only ever holds the key's path. It is often the cleanest long-term model, and our post on &lt;a href="https://dev.to/blog/secrets-management-scale-external-secrets-operator"&gt;syncing secrets into Kubernetes with External Secrets&lt;/a&gt; walks through it. The trade-off is a hard runtime dependency on that external system, and on the operator itself: in 2025 the project paused releases for several weeks over maintainer burnout, until more maintainers joined. Anything on your runtime path deserves a look at who maintains it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Argo CD's guide recommends populating secrets on the destination cluster, which is what the last two do. Put that together and my bias looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;if you already have a cloud secrets manager, go straight to &lt;strong&gt;External Secrets Operator&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;if you don't and you run Flux, &lt;strong&gt;SOPS&lt;/strong&gt; is the cleanest Git-native option&lt;/li&gt;
&lt;li&gt;if you don't and you run Argo CD, start with &lt;strong&gt;Sealed Secrets&lt;/strong&gt; and treat the key backup as part of the install&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Picking the "wrong" tool is recoverable. The mistake that lingers is deferring the decision, migrating everything else, and leaving secrets as a loose end to tidy up later. Sort them out alongside the reconciler, because they are part of the same design.&lt;/p&gt;

&lt;h2&gt;
  
  
  The machinery to skip until it hurts
&lt;/h2&gt;

&lt;p&gt;This is where small teams get talked into complexity they have not earned yet. Plenty of the canonical GitOps add-ons are real solutions, just to later problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;app-of-apps sprawl&lt;/strong&gt; solves large estates with many related deployments&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ApplicationSet&lt;/code&gt; generators&lt;/strong&gt; solve repetition across fleets of charts and clusters&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;promotion pipelines&lt;/strong&gt; solve controlled movement across several environments&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;progressive delivery&lt;/strong&gt;, with a tool such as &lt;a href="https://dev.to/blog/argo-rollouts-kubernetes-canary-blue-green"&gt;Argo Rollouts&lt;/a&gt;, solves high-stakes releases where traffic shaping is worth another controller&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;multi-cluster GitOps&lt;/strong&gt; solves hard isolation, geography, tenancy or blast-radius constraints&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If those are not your current pressures, adopting them early is overhead you pay every week for a problem you don't have yet.&lt;/p&gt;

&lt;p&gt;It is the same argument as our post on &lt;a href="https://dev.to/blog/stop-copying-big-tech-platform-architecture"&gt;right-sizing platform architecture&lt;/a&gt;: do not borrow scale machinery before the scale arrives. GitOps has the trap every fashionable platform idea has. The sensible core gets wrapped in an architecture identity, and teams start buying components to prove they are doing it properly. You are doing GitOps properly if Git is the source of truth and a reconciler is continuously working to make the cluster match it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running the loop once it's live
&lt;/h2&gt;

&lt;p&gt;The machinery can wait. These can't, because every team meets them within a few weeks of switching the reconciler on, however small the estate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pause the loop before you fix production by hand.&lt;/strong&gt; During an incident the reconciler is on Git's side, not yours. In Flux, &lt;code&gt;flux suspend helmrelease &amp;lt;name&amp;gt;&lt;/code&gt; (or &lt;code&gt;flux suspend kustomization &amp;lt;name&amp;gt;&lt;/code&gt; for plain manifests) stops new revisions and drift correction, and the matching &lt;code&gt;flux resume&lt;/code&gt; turns them back on. In Argo CD, &lt;code&gt;argocd app set &amp;lt;app&amp;gt; --sync-policy none&lt;/code&gt; takes one Application off automated sync. There are two catches on the Argo side: an Application generated by an &lt;code&gt;ApplicationSet&lt;/code&gt; ignores changes to its own sync policy, and a parent app with self-heal on can put a child's policy straight back. Whatever you change by hand, commit the same change to the values file before you resume, or the loop will quietly undo your fix. Put those commands in the incident runbook now, while things are calm.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let the autoscaler own &lt;code&gt;replicas&lt;/code&gt;.&lt;/strong&gt; If a HorizontalPodAutoscaler manages a Deployment and your chart also renders &lt;code&gt;spec.replicas&lt;/code&gt;, the two fight, and every sync resets the count the HPA chose. The chart that &lt;code&gt;helm create&lt;/code&gt; scaffolds already handles this: its Deployment only renders &lt;code&gt;replicas&lt;/code&gt; when &lt;code&gt;autoscaling.enabled&lt;/code&gt; is false, so the fix is usually one value rather than a template change. Argo CD's best-practices guide and the &lt;a href="https://fluxcd.io/flux/faq/" rel="noopener noreferrer"&gt;Flux FAQ&lt;/a&gt; give the same advice: leave &lt;code&gt;replicas&lt;/code&gt; out of what you apply. If a third-party chart insists on rendering it, Argo CD's &lt;a href="https://argo-cd.readthedocs.io/en/stable/user-guide/diffing/" rel="noopener noreferrer"&gt;&lt;code&gt;ignoreDifferences&lt;/code&gt;&lt;/a&gt; on &lt;code&gt;/spec/replicas&lt;/code&gt; hides it from the diff, but the sync still applies it unless you also set the &lt;code&gt;RespectIgnoreDifferences=true&lt;/code&gt; sync option. On Flux, a &lt;code&gt;HelmRelease&lt;/code&gt; takes &lt;code&gt;driftDetection.ignore&lt;/code&gt; with &lt;code&gt;paths: ["/spec/replicas"]&lt;/code&gt;, which its docs suggest for exactly this case.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Preview what a pull request will actually render.&lt;/strong&gt; With Helm, reading a values diff is not the same as knowing what the manifests will look like, because one value can fan out across a dozen templates. The cheapest check needs no cluster at all: run &lt;code&gt;helm template&lt;/code&gt; with the base and environment files on your branch and on &lt;code&gt;main&lt;/code&gt;, and diff the two outputs. For the live comparison, &lt;code&gt;argocd app diff &amp;lt;app&amp;gt; --revision &amp;lt;branch&amp;gt;&lt;/code&gt; renders the chart and compares it with what is running, with one gap: it leaves Secrets out. &lt;code&gt;flux diff kustomization&lt;/code&gt; is less useful here, because for a &lt;code&gt;HelmRelease&lt;/code&gt; it shows the change to the &lt;code&gt;HelmRelease&lt;/code&gt; object rather than to the rendered manifests. Running a live diff in CI means giving CI read access to the cluster, which is a real trade for a pull-based setup. Running it locally before you open the pull request costs nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make the loop react faster than it polls.&lt;/strong&gt; By default Argo CD checks Git every two minutes plus up to a minute of jitter, and Flux fetches each &lt;code&gt;GitRepository&lt;/code&gt; on whatever &lt;code&gt;interval&lt;/code&gt; you give it. A merged pull request that appears to do nothing for a few minutes is usually waiting for the next poll. Both tools accept webhooks: Argo CD's API server takes &lt;a href="https://argo-cd.readthedocs.io/en/stable/operator-manual/webhook/" rel="noopener noreferrer"&gt;Git webhook events&lt;/a&gt; directly, and Flux uses a &lt;a href="https://fluxcd.io/flux/guides/webhook-receivers/" rel="noopener noreferrer"&gt;&lt;code&gt;Receiver&lt;/code&gt;&lt;/a&gt; from its notification controller. Keep polling as the fallback, so a missed webhook costs you minutes rather than a deploy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to start
&lt;/h2&gt;

&lt;p&gt;Pick one low-risk chart, ideally a stateless internal service, and move it to the base-plus-environment layout: &lt;code&gt;values.yaml&lt;/code&gt; for what is shared, &lt;code&gt;values/production.yaml&lt;/code&gt; for the handful of lines that differ. Put it under the reconciler with automated sync on and pruning off, and on Flux set &lt;code&gt;driftDetection.mode: warn&lt;/code&gt;. Make a pull request the only way it changes for a fortnight, and keep a note of every drift the reconciler reports, whether that is Argo CD marking the app &lt;code&gt;OutOfSync&lt;/code&gt; or Flux emitting a drift event on the &lt;code&gt;HelmRelease&lt;/code&gt;. Every one of those is a hand-edit or a &lt;code&gt;--set&lt;/code&gt; someone still reaches for, and each becomes either a runbook entry, a value that belongs in Git, or a field handed over to another controller. When a week goes by with nothing to explain, turn on self-heal and pruning, move the next chart across, and leave the rest of the GitOps catalogue on the shelf until a real problem asks for it.&lt;/p&gt;

&lt;p&gt;For a small team, most of the effort in this post goes into standing the loop up rather than using it. That is the part we built &lt;a href="https://kupe.cloud" rel="noopener noreferrer"&gt;Kupe Cloud&lt;/a&gt; to take away. It is our managed Kubernetes platform, and both halves of the setup are ready as soon as you spin up your first cluster. Argo CD is already running, with a project set up for your tenant, so once you connect a service repo laid out as above, its chart deploys as it is. Secrets work the same way from day one: you create a secret once in the console, it is stored encrypted in the platform's vault, and Kupe syncs it as a Kubernetes &lt;code&gt;Secret&lt;/code&gt; into the clusters and namespaces you choose. That is the External Secrets model from earlier, with nothing extra to install or run. Everything else in this post still applies; you just start with the loop already running.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>gitops</category>
      <category>helm</category>
      <category>argocd</category>
    </item>
    <item>
      <title>Falco on Kubernetes: install, tune, and route alerts</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Tue, 22 Sep 2026 18:28:51 +0000</pubDate>
      <link>https://dev.to/coresolutions/falco-on-kubernetes-install-tune-and-route-alerts-1la1</link>
      <guid>https://dev.to/coresolutions/falco-on-kubernetes-install-tune-and-route-alerts-1la1</guid>
      <description>&lt;p&gt;You scan images in CI, sign them before deploy, and gate manifests at admission. Every one of those controls answers the same question: should this workload be allowed to start? None of them can tell you what it did at 02:00 once it was running. Someone opening a shell in a production pod, a binary that was never in the image suddenly executing, a process reading &lt;code&gt;/etc/shadow&lt;/code&gt; for no good reason. Those are runtime events, and runtime is where Falco lives.&lt;/p&gt;

&lt;p&gt;The catch is that Falco tutorials age faster than most. The project has dropped its legacy eBPF probe, its gVisor engine and its gRPC output in recent releases, deprecated the &lt;code&gt;append: true&lt;/code&gt; rule syntax in favour of &lt;code&gt;override:&lt;/code&gt;, and moved rule distribution onto OCI artifacts pulled by a sidecar. A lot of the snippets still ranking on the first page will not load on a current install. So this walkthrough sticks to what the current chart and docs actually do: install, trigger one stock alert, write one custom rule, and then spend most of the time on tuning, which is the part that decides whether Falco is still switched on in six months.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where runtime detection fits
&lt;/h2&gt;

&lt;p&gt;A Kubernetes security stack has layers, and each one answers a different question:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;image scanning: what went into the artefact&lt;/li&gt;
&lt;li&gt;signing and provenance: whether you trust where it came from&lt;/li&gt;
&lt;li&gt;admission policy: whether the cluster should accept it&lt;/li&gt;
&lt;li&gt;runtime detection: what the workload is doing &lt;em&gt;now&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last layer is the one teams skip until something odd is already running in production. Falco fills it by watching syscalls from the kernel, matching them against a rules engine in userspace, and emitting an alert when a rule fires. If you want the eBPF side of that story in more depth, the &lt;a href="https://dev.to/blog/ebpf-powered-networking-on-kubernetes-with-cilium"&gt;Cilium networking post&lt;/a&gt; covers how the same machinery is used for the network path.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Falco collects events, and where it differs from Tetragon
&lt;/h2&gt;

&lt;p&gt;Search for "Falco vs eBPF" and you will find a lot of confused threads. Falco is not an alternative to eBPF. It uses eBPF as one way of collecting kernel events, and there are two supported drivers today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;modern_ebpf&lt;/code&gt;, a CO-RE probe embedded in the binary, which is what Falco itself defaults to&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;kmod&lt;/code&gt;, the kernel module, still supported and still the right answer on kernels the modern probe cannot run on&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Helm chart adds a third value, &lt;code&gt;auto&lt;/code&gt;, which is its default. It tries the modern probe first and falls back to the kernel module. The older &lt;code&gt;ebpf&lt;/code&gt; and &lt;code&gt;gvisor&lt;/code&gt; driver kinds are gone: the chart will not even render if you pass either, which is a kinder failure than a pod that starts and silently sees nothing.&lt;/p&gt;

&lt;p&gt;Tetragon, from the Cilium project, makes a different trade-off. It leans into in-kernel enforcement: a tracing policy can override a kernel function's return value or send &lt;code&gt;SIGKILL&lt;/code&gt; to the offending process before it gets any further. Falco keeps the decision in userspace, with a mature rules engine, a maintained default ruleset, and a lot of flexibility around outputs and tuning. If your first question is "what happened inside that container?", Falco is the more natural place to start. If it is "stop that from ever completing", look at Tetragon.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install Falco with the Helm chart
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm repo add falcosecurity https://falcosecurity.github.io/charts
helm repo update
helm &lt;span class="nb"&gt;install &lt;/span&gt;falco falcosecurity/falco &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--create-namespace&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; falco &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; falcosidekick.enabled&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; falcosidekick.webui.enabled&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--set&lt;/span&gt; collectors.kubernetes.enabled&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those three flags are worth understanding rather than copying:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;falcosidekick.enabled=true&lt;/code&gt; deploys Falcosidekick and switches Falco to JSON output over HTTP, so alerts can be forwarded to systems people already watch&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;falcosidekick.webui.enabled=true&lt;/code&gt; adds the web UI and a Redis instance behind it, which is handy while you are testing and worth turning off once alerts flow somewhere durable&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;collectors.kubernetes.enabled=true&lt;/code&gt; deploys the metadata collector and the &lt;code&gt;k8smeta&lt;/code&gt; plugin, which adds the owning Deployment, ReplicaSet and Service to each alert. Without it you still get the pod name, namespace and labels, because Falco reads those from the container runtime directly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The chart also deploys two containers you did not ask for by name: a &lt;code&gt;falcoctl&lt;/code&gt; init container that pulls the default rules as an OCI artifact, and a &lt;code&gt;falcoctl&lt;/code&gt; sidecar that checks for a newer artifact once a week and swaps it in. That has consequences for tuning, which we will come back to.&lt;/p&gt;

&lt;p&gt;If you need to be explicit about the driver, use one of the two current values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--set&lt;/span&gt; driver.kind&lt;span class="o"&gt;=&lt;/span&gt;modern_ebpf   &lt;span class="c"&gt;# or driver.kind=kmod&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then watch the Falco container start. The pod has several containers now, so name the one you want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl &lt;span class="nt"&gt;-n&lt;/span&gt; falco logs ds/falco &lt;span class="nt"&gt;-c&lt;/span&gt; falco
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You are looking for a clean driver load and a &lt;code&gt;Loading rules from file&lt;/code&gt; line for each rules file. If the driver is wrong for the kernel underneath, this is where you find out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trigger one stock alert first
&lt;/h2&gt;

&lt;p&gt;Before writing anything custom, prove the default path works. Start a throwaway pod:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl run runtime-test &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;busybox &lt;span class="nt"&gt;--restart&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Never &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nb"&gt;sleep &lt;/span&gt;3600
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then exec into it with a terminal attached:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; runtime-test &lt;span class="nt"&gt;--&lt;/span&gt; sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That fires the stock &lt;code&gt;Terminal shell in container&lt;/code&gt; rule. The &lt;code&gt;-t&lt;/code&gt; is doing real work here: the rule checks that the shell has a TTY, so &lt;code&gt;kubectl exec runtime-test -- sh -c 'id'&lt;/code&gt; without one will not trigger it. Check the Falco logs or the Falcosidekick UI and you should see the alert with &lt;code&gt;container_id&lt;/code&gt; and &lt;code&gt;container_name&lt;/code&gt; appended to the end of the output line. Those fields are not in the rule's own output string; the container plugin suggests them and Falco appends suggested fields automatically.&lt;/p&gt;

&lt;p&gt;This is also the moment to reset expectations about volume. The stock &lt;code&gt;falco_rules.yaml&lt;/code&gt; is a small, conservative file: 25 stable rules at the time of writing, every one enabled, none of them experimental. The noisy reputation comes from the incubating and sandbox rulesets, which the chart does not load by default, and from teams tuning carelessly once they do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write one custom rule
&lt;/h2&gt;

&lt;p&gt;The simplest useful rule is one you can explain in a sentence: tell me if a miner-like binary starts inside a container. With the chart, custom rules go in the &lt;code&gt;customRules&lt;/code&gt; value, and each key becomes a file under &lt;code&gt;/etc/falco/rules.d/&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;customRules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;miners.yaml&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|-&lt;/span&gt;
    &lt;span class="s"&gt;- rule: Crypto miner process in container&lt;/span&gt;
      &lt;span class="s"&gt;desc: Detect common miner binaries starting inside a container&lt;/span&gt;
      &lt;span class="s"&gt;condition: &amp;gt;&lt;/span&gt;
        &lt;span class="s"&gt;spawned_process and container and&lt;/span&gt;
        &lt;span class="s"&gt;proc.name = anyof (xmrig, minerd, cryptominer)&lt;/span&gt;
      &lt;span class="s"&gt;output: &amp;gt;&lt;/span&gt;
        &lt;span class="s"&gt;Crypto miner process in container | user=%user.name&lt;/span&gt;
        &lt;span class="s"&gt;process=%proc.name command=%proc.cmdline&lt;/span&gt;
      &lt;span class="s"&gt;priority: WARNING&lt;/span&gt;
      &lt;span class="s"&gt;tags: [container, malware, crypto-miner]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things about that condition. &lt;code&gt;spawned_process&lt;/code&gt; and &lt;code&gt;container&lt;/code&gt; are macros from the stock file, so you get exec events inside containers for free. And &lt;code&gt;anyof&lt;/code&gt; is a comparator modifier from the current rule language: one field compared against a short list, without a chain of &lt;code&gt;or&lt;/code&gt; clauses. It sits after the operator, with a space either side and the values in parentheses.&lt;/p&gt;

&lt;p&gt;One trap that is not in the tutorials: &lt;code&gt;proc.name&lt;/code&gt; is the kernel's process name, and the kernel truncates it to 15 characters. The stock rules file depends on this, which is why you will find entries like &lt;code&gt;mysql_install_d&lt;/code&gt; in its lists. If your binary has a long name, match on &lt;code&gt;proc.exepath&lt;/code&gt; or &lt;code&gt;container.image.repository&lt;/code&gt; instead.&lt;/p&gt;

&lt;p&gt;Load order matters. The default &lt;code&gt;rules_files&lt;/code&gt; list reads &lt;code&gt;falco_rules.yaml&lt;/code&gt;, then &lt;code&gt;falco_rules.local.yaml&lt;/code&gt;, then everything in &lt;code&gt;rules.d&lt;/code&gt;, and a file that overrides a stock rule has to load after it. Give the file a &lt;code&gt;.yaml&lt;/code&gt; or &lt;code&gt;.yml&lt;/code&gt; extension too, because anything else in that directory is ignored.&lt;/p&gt;

&lt;p&gt;Before you apply, validate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;falco &lt;span class="nt"&gt;-V&lt;/span&gt; miners.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That parses and compiles the rules file and exits, and the easiest place to run it is the same Falco image tag you deploy, so the engine version matches. It is worth the thirty seconds, because a rules error is fatal. Falco stops at the first file that fails to load, prints the reason, and exits non-zero, which on Kubernetes means the whole DaemonSet goes into &lt;code&gt;CrashLoopBackOff&lt;/code&gt; over one bad line in one file. Warnings are logged and tolerated. Errors are not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tuning is where most Falco deployments succeed or fail
&lt;/h2&gt;

&lt;p&gt;Install is the easy part. Tuning decides whether people trust the alerts next month, and the stock rules file is built for it in a specific way that most tutorials, including the official exceptions examples, obscure.&lt;/p&gt;

&lt;p&gt;The fact the docs bury is that &lt;strong&gt;none of the stable stock rules ships an &lt;code&gt;exceptions:&lt;/code&gt; block.&lt;/strong&gt; The docs' canonical example appends values to an exception on &lt;code&gt;Write below binary dir&lt;/code&gt;, and that rule lives in the sandbox ruleset, not the one you just installed. Copy it against a stable rule and the loader rejects it, because there is no existing exception to inherit fields from.&lt;/p&gt;

&lt;p&gt;What the stable rules give you instead is a set of hooks with names starting &lt;code&gt;user_&lt;/code&gt;, each defined as &lt;code&gt;(never_true)&lt;/code&gt;, plus empty image lists. Take &lt;code&gt;Contact K8S API Server From Container&lt;/code&gt;, which fires the first time an operator or controller you run outside &lt;code&gt;kube-system&lt;/code&gt; talks to the API server. Its condition ends with &lt;code&gt;and not user_known_contact_k8s_api_server_activities&lt;/code&gt;, and that macro is yours to replace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;macro&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;user_known_contact_k8s_api_server_activities&lt;/span&gt;
  &lt;span class="na"&gt;condition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="s"&gt;(k8s.ns.name = argocd and&lt;/span&gt;
     &lt;span class="s"&gt;container.image.repository = quay.io/argoproj/argocd)&lt;/span&gt;
  &lt;span class="na"&gt;override&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;condition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;replace&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same shape works for the other stable rules. &lt;code&gt;Terminal shell in container&lt;/code&gt; has &lt;code&gt;user_expected_terminal_shell_in_container_conditions&lt;/code&gt;. &lt;code&gt;Read sensitive file untrusted&lt;/code&gt; has a &lt;code&gt;read_sensitive_file_images&lt;/code&gt; list you can append to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;read_sensitive_file_images&lt;/span&gt;
  &lt;span class="na"&gt;items&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;quay.io/argoproj/argocd&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;override&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;items&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;append&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you would rather use exceptions, you can. Appending a &lt;em&gt;new&lt;/em&gt; named exception to a stable rule is supported as long as you supply &lt;code&gt;fields&lt;/code&gt; alongside &lt;code&gt;values&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;rule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Contact K8S API Server From Container&lt;/span&gt;
  &lt;span class="na"&gt;exceptions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argocd_controllers&lt;/span&gt;
      &lt;span class="na"&gt;fields&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;container.image.repository&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;k8s.ns.name&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;quay.io/argoproj/argocd&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;argocd&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;override&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;exceptions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;append&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Falco adds &lt;code&gt;and not (container.image.repository = ... and k8s.ns.name = ...)&lt;/code&gt; to the condition for you. Either approach beats disabling the rule, because the rule is still right for every other workload.&lt;/p&gt;

&lt;p&gt;When a rule really does not fit your threat model, turn it off explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;rule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Drop and execute new binary in container&lt;/span&gt;
  &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;override&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;replace&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;override:&lt;/code&gt; block is the part older posts get wrong. It states, per key, whether you are appending to or replacing the stock definition. &lt;code&gt;append: true&lt;/code&gt; still parses for now but is deprecated, so translate any snippet built around it before you paste.&lt;/p&gt;

&lt;p&gt;One more thing the file already did for you: the API-server rule's &lt;code&gt;k8s_containers&lt;/code&gt; macro exempts everything in &lt;code&gt;kube-system&lt;/code&gt;. If it is firing, the workload is somewhere else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route alerts somewhere people already look
&lt;/h2&gt;

&lt;p&gt;An alert that only exists in one DaemonSet's log stream changes nothing. Falcosidekick takes Falco's JSON output and fans it out to Slack, Loki, Elasticsearch, Splunk, Datadog, Alertmanager and a long list beyond that. The chart values map directly to its outputs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;falcosidekick&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;slack&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;webhookurl&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://hooks.slack.com/services/XXXX/YYYY/ZZZZ&lt;/span&gt;
      &lt;span class="na"&gt;minimumpriority&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;warning&lt;/span&gt;
    &lt;span class="na"&gt;loki&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;hostport&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://loki-gateway.monitoring:80&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no global minimum priority in Falcosidekick. Each output has its own &lt;code&gt;minimumpriority&lt;/code&gt;, and its severity ladder spells the level &lt;code&gt;informational&lt;/code&gt; rather than Falco's &lt;code&gt;info&lt;/code&gt;, which catches people out when a filter silently matches nothing.&lt;/p&gt;

&lt;p&gt;The pattern that holds up on platform teams is a split by urgency: Slack or Teams for the handful of rules that need a human now, a log backend such as Loki for searchable history, and the rest filtered out at the output. If Grafana is already your front door, the &lt;a href="https://dev.to/blog/enhance-your-kubernetes-monitoring-with-grafana"&gt;Grafana monitoring post&lt;/a&gt; covers getting Grafana itself stood up and the first dashboards in place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Operating it after week one
&lt;/h2&gt;

&lt;p&gt;The first three sections get Falco running. These are the three things that go wrong once it has been running for a while.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your rules are floating.&lt;/strong&gt; The &lt;code&gt;falcoctl&lt;/code&gt; sidecar follows a major-version tag of the rules artifact and installs whatever is newest every 168 hours, into an &lt;code&gt;emptyDir&lt;/code&gt; mounted over &lt;code&gt;/etc/falco&lt;/code&gt;. Two consequences: an edited &lt;code&gt;falco_rules.yaml&lt;/code&gt; inside the pod is overwritten on the next restart or the next follow, and a rule you tuned by name can change underneath you on a Tuesday. Keep every override in &lt;code&gt;customRules&lt;/code&gt;, and pin &lt;code&gt;falcoctl.config.artifact.install.refs&lt;/code&gt; and &lt;code&gt;follow.refs&lt;/code&gt; to an exact artifact tag or an &lt;code&gt;@sha256:&lt;/code&gt; digest once you are past the exploring stage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You are dropping events and nobody knows.&lt;/strong&gt; Falco emits an internal alert named &lt;code&gt;Falco internal: syscall event drop&lt;/code&gt; when the kernel ring buffer overflows, but it is a &lt;code&gt;debug&lt;/code&gt;-priority alert rate-limited to one message every 30 seconds. Raise Falco's minimum &lt;code&gt;priority&lt;/code&gt; to cut noise, or set &lt;code&gt;minimumpriority: warning&lt;/code&gt; on your Slack output, and it vanishes. The durable answer is metrics:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--set&lt;/span&gt; metrics.enabled&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--set&lt;/span&gt; serviceMonitor.create&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That turns on Falco's Prometheus endpoint and a ServiceMonitor for it. The two series to graph against each other are &lt;code&gt;falcosecurity_scap_n_drops_total&lt;/code&gt; and &lt;code&gt;falcosecurity_scap_n_evts_total&lt;/code&gt;. If drops climb with load, the documented remedies are a bigger ring buffer via &lt;code&gt;engine.modern_ebpf.buf_size_preset&lt;/code&gt;, fewer CPUs per buffer via &lt;code&gt;engine.modern_ebpf.cpus_for_each_buffer&lt;/code&gt;, and &lt;code&gt;base_syscalls&lt;/code&gt; to narrow what the driver captures. Which of those helps depends on your syscall mix, which is also why there is no universal "Falco overhead" figure worth quoting: it tracks what your workloads do, not a percentage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You want to act, not just alert.&lt;/strong&gt; Falco itself only detects. Falco Talon is the incubating response engine in the same ecosystem: it receives events from Falcosidekick's &lt;code&gt;talon&lt;/code&gt; output and applies rules such as terminating the pod or labelling it for quarantine. Its Kubernetes actions need &lt;code&gt;k8s.pod.name&lt;/code&gt; and &lt;code&gt;k8s.ns.name&lt;/code&gt; among the event's output fields, which the stock rule outputs do not carry, so expect to append those to any rule you want it to act on. It is receiving commits but has had long gaps between tagged releases, so treat it as something to evaluate rather than a default, and keep the enforcement question in mind when you weigh it against Tetragon.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure you will hit first
&lt;/h2&gt;

&lt;p&gt;Almost every first Falco tuning session ends one of two ways. Either the DaemonSet is crash-looping, in which case the Falco container's log has the offending file and line near the top and &lt;code&gt;falco -V&lt;/code&gt; would have told you first. Or the pods are healthy and the override seems to do nothing. For the second, check three things before you touch the rule. Confirm the file landed in &lt;code&gt;/etc/falco/rules.d/&lt;/code&gt; with a &lt;code&gt;.yaml&lt;/code&gt; extension and appears in the &lt;code&gt;Loading rules from file&lt;/code&gt; lines. Check the rule name you are overriding still exists with that exact spelling in the ruleset the &lt;code&gt;falcoctl&lt;/code&gt; sidecar just installed. Then check whether the workload is already exempt, the way &lt;code&gt;kube-system&lt;/code&gt; is for the API-server rule. Usually the rule was fine and the plumbing was not.&lt;/p&gt;

</description>
      <category>security</category>
      <category>kubernetes</category>
      <category>falco</category>
      <category>runtimesecurity</category>
    </item>
    <item>
      <title>CIS Benchmark for Kubernetes with kube-bench and Kubescape</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Thu, 17 Sep 2026 17:51:08 +0000</pubDate>
      <link>https://dev.to/coresolutions/cis-benchmark-for-kubernetes-with-kube-bench-and-kubescape-38je</link>
      <guid>https://dev.to/coresolutions/cis-benchmark-for-kubernetes-with-kube-bench-and-kubescape-38je</guid>
      <description>&lt;p&gt;Most teams get told to “harden Kubernetes to CIS” long before anyone explains what happens after the first scan. You run the tool, it prints a wall of &lt;code&gt;FAIL&lt;/code&gt; and &lt;code&gt;WARN&lt;/code&gt;, half the findings turn out to be somebody else’s problem on a managed control plane, and the rest sit in a ticket queue until the next audit comes round. There is a narrower version that works. Use &lt;strong&gt;kube-bench&lt;/strong&gt; for the host and control-plane checks only it can see, use &lt;strong&gt;Kubescape&lt;/strong&gt; for the posture checks you want to keep running from the API side, and treat the &lt;a href="https://www.cisecurity.org/benchmark/kubernetes" rel="noopener noreferrer"&gt;CIS Kubernetes Benchmark&lt;/a&gt; as a triage tool rather than a moral score.&lt;/p&gt;

&lt;p&gt;One thing to know before either tool runs: the scanners trail the benchmark. CIS revises the benchmark as new Kubernetes releases land, and each tool picks the revision up some time later, as a new profile that somebody has to write and merge. On a recently upgraded cluster there is a fair chance your scanner is running the closest older profile and hasn’t mentioned it. The result is still useful, but that caveat belongs in the first line of anything you hand to an auditor, and there is a quick way to check, which we’ll get to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two scanners, two views of the cluster
&lt;/h2&gt;

&lt;p&gt;The short version:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;kube-bench&lt;/strong&gt; runs on a node. It reads config files, file permissions and the flags on running processes, then compares them with the CIS checks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kubescape&lt;/strong&gt; talks to the API server. It evaluates the resources in your cluster, or the manifests in your repo, against frameworks such as NSA-CISA, MITRE ATT&amp;amp;CK and CIS.&lt;/li&gt;
&lt;li&gt;They overlap on the policy-style checks, and that is about it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is why “kube-bench or Kubescape?” is the wrong question. Run only kube-bench and you get a point-in-time audit with no idea what changed last Tuesday. Run only the Kubescape CLI and you miss everything that depends on a file mode or a kubelet flag on a real node.&lt;/p&gt;

&lt;p&gt;There is also one deliciously awkward detail to get out early. The standard kube-bench Job sets &lt;code&gt;hostPID: true&lt;/code&gt; and mounts host paths such as &lt;code&gt;/etc/kubernetes&lt;/code&gt; and &lt;code&gt;/var/lib/kubelet&lt;/code&gt;, so the compliance tool needs an exemption from the restricted pod policy it is about to tell you to enforce. A host-level audit has to see the host, so plan for the exemption now rather than discovering it when admission rejects the pod.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the CIS Kubernetes Benchmark is, and what it is not
&lt;/h2&gt;

&lt;p&gt;The benchmark is a hardening checklist for cluster configuration. Its sections cover the control plane components, etcd, control plane configuration, worker nodes, and policies: RBAC, pod security, network policies and secrets. Every recommendation comes with an audit procedure and a remediation, which is what makes it automatable.&lt;/p&gt;

&lt;p&gt;What it leaves out is where people over-read the score. A good CIS result tells you the configuration is in better shape than it was. It does &lt;em&gt;not&lt;/em&gt; tell you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;whether workloads behave safely at runtime&lt;/li&gt;
&lt;li&gt;whether your images are full of known CVEs&lt;/li&gt;
&lt;li&gt;whether your admission policies are broad enough for your actual risk model&lt;/li&gt;
&lt;li&gt;whether the control plane you rent from a provider was set up the way you would have done it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So treat the benchmark the way you would treat a good linter: very useful, occasionally annoying, and dangerous only when somebody mistakes it for a complete security strategy. If a score is going in front of leadership, put that caveat on the same slide. “100% CIS” and “secure” are different claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with kube-bench, on the nodes
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/aquasecurity/kube-bench" rel="noopener noreferrer"&gt;kube-bench&lt;/a&gt; is an open-source project from Aqua Security, and it does one blunt thing well: read the files and flags that exist on the host and compare them with the benchmark profile it knows about.&lt;/p&gt;

&lt;p&gt;The normal entry point is the &lt;code&gt;Job&lt;/code&gt; manifest in the root of the repo. On a self-managed cluster, apply it and read the logs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; https://raw.githubusercontent.com/aquasecurity/kube-bench/main/job.yaml
kubectl logs job/kube-bench
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every check comes back in one of four states. &lt;code&gt;PASS&lt;/code&gt; and &lt;code&gt;FAIL&lt;/code&gt; are what they sound like. &lt;code&gt;WARN&lt;/code&gt; means kube-bench could not make the call for you, either because the benchmark marks the check as manual or because the check is unscored. &lt;code&gt;INFO&lt;/code&gt; is a section header or a check you skipped. The habit to avoid is reading &lt;code&gt;WARN&lt;/code&gt; as “fine”. In the vanilla profile nothing in the policies section is scored, so nothing in it can ever show as &lt;code&gt;FAIL&lt;/code&gt;, and that section holds the RBAC and pod security checks you probably care about most.&lt;/p&gt;

&lt;p&gt;The Job needs &lt;code&gt;hostPID&lt;/code&gt; and that stack of read-only host mounts because kube-bench is looking at the node directly rather than asking the API server to summarise it. You can see the same assumption in the container one-liner from the &lt;a href="https://github.com/aquasecurity/kube-bench/blob/main/docs/running.md" rel="noopener noreferrer"&gt;upstream docs&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;--pid&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;host &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; /etc:/etc:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; /var:/var:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-t&lt;/span&gt; docker.io/aquasec/kube-bench:latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is how it catches things an API-only scanner cannot. It is also why the Job belongs in its own clearly named namespace, labelled &lt;code&gt;privileged&lt;/code&gt; for &lt;a href="https://dev.to/blog/podsecuritypolicy-replacement-pod-security-admission"&gt;Pod Security Admission&lt;/a&gt;, rather than squeezed under your strictest workload policy. One more practical note: to audit control-plane nodes the pod has to be scheduled onto them, which means adding a &lt;code&gt;nodeSelector&lt;/code&gt; and tolerations to the manifest first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmark lag is normal, so check which profile ran
&lt;/h2&gt;

&lt;p&gt;The easiest way to confuse yourself is to assume the latest kube-bench means the latest CIS revision. kube-bench maps your Kubernetes version to a benchmark profile using a table in its config, and the &lt;a href="https://github.com/aquasecurity/kube-bench/blob/main/docs/platforms.md" rel="noopener noreferrer"&gt;platforms page&lt;/a&gt; in the docs shows the same mapping. When your cluster is newer than anything in that table, kube-bench walks the minor version down until it finds a match and runs that profile instead. There is no error, and at the default log level there is no warning either.&lt;/p&gt;

&lt;p&gt;The JSON output is where you catch it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kube-bench run &lt;span class="nt"&gt;--targets&lt;/span&gt; node,policies &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each block carries a &lt;code&gt;version&lt;/code&gt; field, which is the profile that ran, and a &lt;code&gt;detected_version&lt;/code&gt; field, which is the Kubernetes version kube-bench found. Compare the pair with the platforms table before you write policy around the result.&lt;/p&gt;

&lt;p&gt;Once you know which profile you want, pin it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kube-bench run &lt;span class="nt"&gt;--benchmark&lt;/span&gt; &amp;lt;profile&amp;gt; &lt;span class="nt"&gt;--targets&lt;/span&gt; node,policies
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here &lt;code&gt;&amp;lt;profile&amp;gt;&lt;/code&gt; is a name from the “kube-bench config” column of that table. Pinning can’t give you a newer benchmark than the tool ships. What it buys is a result you can describe accurately, and a scheduled run that won’t switch profile underneath you on the day somebody bumps the image tag.&lt;/p&gt;

&lt;p&gt;The same table matters if you are not on vanilla Kubernetes. k3s, RKE2, OpenShift and the managed services each have their own profiles, because the vanilla one looks in paths those platforms do not use and rewards you with a page of false failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  One finding, end to end
&lt;/h2&gt;

&lt;p&gt;Reports are easy to admire from a distance, so take one finding all the way through. In the current vanilla profiles, check &lt;code&gt;5.2.3&lt;/code&gt; is “Minimize the admission of containers wishing to share the host process ID namespace”. Four questions get you from the line in the report to a decision:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Which check fired, and in which profile?&lt;/strong&gt; IDs move between profiles. The same control is &lt;code&gt;4.2.2&lt;/code&gt; in the EKS profile, which is one more reason to pin.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What did kube-bench run to decide?&lt;/strong&gt; For this check, a &lt;code&gt;kubectl&lt;/code&gt; loop over every pod in every namespace that reads &lt;code&gt;.spec.hostPID&lt;/code&gt;. It reports pods that share the host PID namespace today, and never asks whether a policy would stop the next one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Is this a real gap, an intentional exception, or the platform’s business?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What is the narrowest fix that changes the risk, not just the score?&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first answer matters more than it looks. The vanilla profile treats this check as unscored, so a failure prints as &lt;code&gt;WARN&lt;/code&gt;. The EKS profile treats it as automated and scored, so the same pods print as &lt;code&gt;FAIL&lt;/code&gt;. The headline changes with the profile; the risk does not.&lt;/p&gt;

&lt;p&gt;The benchmark’s remediation text is “Add policies to each namespace in the cluster which has user workloads to restrict the admission of &lt;code&gt;hostPID&lt;/code&gt; containers”, and on a current cluster the policy in question is a Pod Security Admission label. The &lt;code&gt;baseline&lt;/code&gt; level already forbids host namespaces, so you do not need &lt;code&gt;restricted&lt;/code&gt; to clear this finding. Ask the API server what would break before you enforce anything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl label &lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;server &lt;span class="nt"&gt;--overwrite&lt;/span&gt; ns &lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/enforce&lt;span class="o"&gt;=&lt;/span&gt;baseline
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A server-side dry run changes nothing, and it prints a warning for every namespace with pods that would violate the level (any &lt;code&gt;baseline&lt;/code&gt; violation, so host networking and &lt;code&gt;hostPath&lt;/code&gt; volumes show up too). The namespaces it names are your candidate exceptions: usually a node exporter, a logging or security agent, and kube-bench itself. From there the fix is a split model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;keep most namespaces on &lt;code&gt;baseline&lt;/code&gt; or &lt;code&gt;restricted&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;carve out one clearly named namespace for infrastructure that needs host access&lt;/li&gt;
&lt;li&gt;write down why that namespace exists and who can deploy into it&lt;/li&gt;
&lt;li&gt;re-run only the relevant check, so you know you changed the right thing
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kube-bench run &lt;span class="nt"&gt;--targets&lt;/span&gt; policies &lt;span class="nt"&gt;--check&lt;/span&gt; 5.2.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expect the check to stay at &lt;code&gt;WARN&lt;/code&gt;, because the pods in your exception namespace still have &lt;code&gt;hostPID: true&lt;/code&gt;. That is correct. What has changed is that you can now explain every pod behind it, and an explained exception is what an auditor is looking for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Managed clusters: part of the benchmark is out of reach
&lt;/h2&gt;

&lt;p&gt;The kube-bench docs are direct about this: “It is impossible to inspect the master nodes of managed clusters, e.g. GKE, EKS, AKS and ACK”. That sentence should change how you describe the result. On a managed control plane, you are auditing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the worker nodes you can reach&lt;/li&gt;
&lt;li&gt;the policy and workload checks that are visible from inside the cluster&lt;/li&gt;
&lt;li&gt;whatever managed-service profile kube-bench ships for your platform&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The project ships separate manifests such as &lt;code&gt;job-eks.yaml&lt;/code&gt; and &lt;code&gt;job-gke.yaml&lt;/code&gt; for exactly this reason. Running kube-bench on EKS, for example, you cannot schedule a pod onto the control-plane nodes, so the master checks are simply unavailable.&lt;/p&gt;

&lt;p&gt;Your compliance story on a managed platform is therefore partly technical and partly contractual. You verify what you can, and for the control plane you lean on the provider’s documentation, attestations and shared-responsibility model. Say so in the report: “out of our reach, covered by the provider’s attestation” is accurate, and easier to defend than a suspiciously complete set of green ticks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kubescape for the part that keeps running
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/kubescape/kubescape" rel="noopener noreferrer"&gt;Kubescape&lt;/a&gt; is a &lt;a href="https://www.cncf.io/projects/kubescape/" rel="noopener noreferrer"&gt;CNCF project&lt;/a&gt; and it now covers a lot more than compliance: misconfiguration scanning, image scanning, admission control and runtime detection all live under the same name. The slice that matters here is narrower: the API-side posture layer that sits next to kube-bench.&lt;/p&gt;

&lt;p&gt;The CLI is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubescape list frameworks
kubescape scan framework nsa
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run them in that order. The CIS framework IDs have the benchmark revision baked into the name, so an ID copied from an old tutorial may no longer exist in the version you installed. &lt;code&gt;nsa&lt;/code&gt; and &lt;code&gt;mitre&lt;/code&gt; are stable names; for CIS, list first and use what is actually there.&lt;/p&gt;

&lt;p&gt;For CI or a recurring check, the threshold flag is the practical one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubescape scan framework nsa &lt;span class="nt"&gt;--compliance-threshold&lt;/span&gt; 80
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Below the score you set, the command exits with code 1. That is usually enough to stop “we should probably look at this later” from becoming a permanent state. Pick the number from a real baseline run, though. A threshold the cluster can’t currently meet tends to get switched off rather than fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kubescape vs kube-bench: which question goes where
&lt;/h2&gt;

&lt;p&gt;Reach for &lt;strong&gt;kube-bench&lt;/strong&gt; when the question is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;are kubelet and control-plane settings aligned with the benchmark?&lt;/li&gt;
&lt;li&gt;do file permissions and host-level configuration match the guidance?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reach for &lt;strong&gt;Kubescape&lt;/strong&gt; when the question is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what is the posture of the workloads and resources in this cluster right now?&lt;/li&gt;
&lt;li&gt;can I run this in CI against manifests, before anything reaches a cluster?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The numbers won’t line up, because the two tools count different things. Where they do overlap, use it as a cross-check: if both are unhappy about the same class of policy gap, that finding goes to the top of the pile.&lt;/p&gt;

&lt;h2&gt;
  
  
  The in-cluster path is where Kubescape starts paying for itself
&lt;/h2&gt;

&lt;p&gt;The Kubescape operator is the natural next step once you are past one-off CLI scans, and the install is refreshingly ordinary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm repo add kubescape https://kubescape.github.io/helm-charts/
helm upgrade &lt;span class="nt"&gt;--install&lt;/span&gt; kubescape kubescape/kubescape-operator &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; kubescape &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--create-namespace&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details in the &lt;a href="https://github.com/kubescape/helm-charts/blob/main/charts/kubescape-operator/values.yaml" rel="noopener noreferrer"&gt;chart’s values&lt;/a&gt; are easy to miss. Out of the box, configuration scans run on a schedule; the change-driven rescans that most people picture when they hear “continuous” sit behind &lt;code&gt;capabilities.continuousScan&lt;/code&gt;, which is disabled by default. And the operator’s node agent mounts the host filesystem read-only, so the &lt;code&gt;kubescape&lt;/code&gt; namespace needs the same Pod Security exception as kube-bench. Your one documented exceptions namespace becomes two, which is still a list short enough to defend.&lt;/p&gt;

&lt;h2&gt;
  
  
  A minimal recurring setup, and where to start
&lt;/h2&gt;

&lt;p&gt;If you want the smallest loop that is still worth having, I would keep it to this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;kube-bench on a schedule&lt;/strong&gt;, as a &lt;code&gt;CronJob&lt;/code&gt; with the profile pinned, for the host-level checks only it can see&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kubescape in CI or in-cluster&lt;/strong&gt; for the API-side posture checks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;one threshold that fails something real&lt;/strong&gt;, instead of a dashboard with no owner&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;one exception path&lt;/strong&gt; for justified infrastructure privileges, documented and reviewed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As for where to start: run the kube-bench Job once, and before you fix anything, sort every &lt;code&gt;FAIL&lt;/code&gt; and &lt;code&gt;WARN&lt;/code&gt; into three piles: real fixes, documented exceptions, and the provider’s side of the line. That sorted list is a more useful deliverable than a score, and it is what the recurring scans then keep current.&lt;/p&gt;

</description>
      <category>security</category>
      <category>kubernetes</category>
      <category>cisbenchmark</category>
      <category>compliance</category>
    </item>
    <item>
      <title>Kubernetes is not a platform: it is a toolkit</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Wed, 09 Sep 2026 09:40:15 +0000</pubDate>
      <link>https://dev.to/coresolutions/kubernetes-is-not-a-platform-it-is-a-toolkit-16e2</link>
      <guid>https://dev.to/coresolutions/kubernetes-is-not-a-platform-it-is-a-toolkit-16e2</guid>
      <description>&lt;p&gt;A team says it has “built a platform” because it runs Kubernetes. Then a developer tries to ship a service and still has to assemble a &lt;code&gt;Deployment&lt;/code&gt;, &lt;code&gt;Service&lt;/code&gt;, &lt;code&gt;HorizontalPodAutoscaler&lt;/code&gt;, &lt;code&gt;NetworkPolicy&lt;/code&gt;, &lt;code&gt;PodDisruptionBudget&lt;/code&gt;, &lt;code&gt;ServiceAccount&lt;/code&gt;, and half a dozen bits of environment-specific glue by hand. That is the category error in one scene. &lt;strong&gt;Kubernetes is the substrate. The platform is still on the backlog.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Kubernetes project makes the same point in its &lt;a href="https://kubernetes.io/docs/concepts/overview/" rel="noopener noreferrer"&gt;own overview docs&lt;/a&gt;. Under the heading “What Kubernetes is not”, it says Kubernetes is “not a traditional, all-inclusive PaaS” and that it provides “the building blocks for building developer platforms”. Building blocks, from the people who maintain them. So when you hand product teams raw &lt;code&gt;kubectl&lt;/code&gt;, raw YAML and a wiki full of tribal knowledge, what you have handed over is a very capable control plane and a pile of homework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kubernetes is a very good substrate
&lt;/h2&gt;

&lt;p&gt;Kubernetes earns its place in all this. It gives you strong primitives and a programmable control plane, and that combination is hard to come by anywhere else.&lt;/p&gt;

&lt;p&gt;Concretely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;declarative desired state&lt;/li&gt;
&lt;li&gt;reconciliation loops&lt;/li&gt;
&lt;li&gt;scheduling and placement&lt;/li&gt;
&lt;li&gt;service discovery&lt;/li&gt;
&lt;li&gt;rollout machinery&lt;/li&gt;
&lt;li&gt;extensibility through CRDs and controllers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Which is a serious foundation, and it explains why Kubernetes keeps ending up in the middle of internal platforms.&lt;/p&gt;

&lt;p&gt;Foundations and finished products sit in different categories, though. Postgres is a foundation; the finance dashboard someone logs into every morning is the product built on it. A cluster running important workloads sits on the foundation side of that line, however well run it is.&lt;/p&gt;

&lt;p&gt;This is where a lot of platform work goes sideways. Infrastructure teams do the hard bit: secure the cluster, patch it, make it observable, get some governance around it. Somewhere in the following six months the organisation starts calling the cluster “the platform”, on the assumption that a working substrate produces a usable product on its own. It never has.&lt;/p&gt;

&lt;h2&gt;
  
  
  A platform is a product with users
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://tag-app-delivery.cncf.io/whitepapers/platforms/" rel="noopener noreferrer"&gt;CNCF Platforms white paper&lt;/a&gt; uses much stricter language than most teams do in everyday conversation. It defines a platform as an integrated collection of capabilities &lt;strong&gt;defined and presented according to the needs of the platform’s users&lt;/strong&gt;, and it calls out interfaces such as web portals, project templates, and self-service APIs.&lt;/p&gt;

&lt;p&gt;“We run Kubernetes” clears none of that bar. The platform is the opinionated layer above it, the one that turns raw technology into something people can use safely and repeatedly.&lt;/p&gt;

&lt;p&gt;Usually that means some combination of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;self-service interfaces&lt;/strong&gt; so teams can create or change what they need without waiting on tickets&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;paved roads&lt;/strong&gt; so the safe path is the easy path&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;policy and defaults&lt;/strong&gt; so teams do not need to rediscover every security and reliability lesson themselves&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;documentation and examples&lt;/strong&gt; that reflect how the organisation actually expects software to be shipped&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;support boundaries&lt;/strong&gt; so people know what is standard, what is allowed, and what will get help at 02:00&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last one does more work than it looks. A platform is partly software and partly promises. An interface where no one can tell you the support status is a hopeful abstraction: teams route around it the first time it surprises them, and now you maintain two paths instead of one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The category error shows up as cognitive rent
&lt;/h2&gt;

&lt;p&gt;When teams say “Kubernetes is our platform”, they usually mean “Kubernetes is where our applications run”. The second claim is easy to earn; the first one is a product commitment.&lt;/p&gt;

&lt;p&gt;There is a test for which one you have: &lt;strong&gt;what does a product engineer still need to understand to ship a normal service safely?&lt;/strong&gt; If the answer runs to “quite a lot of Kubernetes”, you are still handing over the substrate.&lt;/p&gt;

&lt;p&gt;That usually means developers still need working knowledge of things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;how requests and limits interact with scheduling&lt;/li&gt;
&lt;li&gt;when an &lt;code&gt;HPA&lt;/code&gt; will or will not behave the way they expect&lt;/li&gt;
&lt;li&gt;why a missing &lt;code&gt;PodDisruptionBudget&lt;/code&gt; only hurts during maintenance windows&lt;/li&gt;
&lt;li&gt;how &lt;code&gt;NetworkPolicy&lt;/code&gt; defaults work in their cluster&lt;/li&gt;
&lt;li&gt;which ingress, gateway, or certificate path is considered standard&lt;/li&gt;
&lt;li&gt;which bits of YAML are mandatory because the platform has no stronger abstraction yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Developers learning infrastructure is fine, and plenty of them should. The trouble starts when shipping an ordinary service requires them to learn the platform team’s entire problem space.&lt;/p&gt;

&lt;p&gt;That is the bill for confusing a toolkit with a platform. The complexity stays exactly where it was; it just gets paid for by every team that touches it, in small instalments, every time they deploy.&lt;/p&gt;

&lt;p&gt;This is why the Team Topologies framing still holds up. A platform is valuable in proportion to the cognitive load it takes off the teams consuming it. Ask developers to become part-time Kubernetes operators and you have moved that load rather than absorbed it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A thin platform still counts
&lt;/h2&gt;

&lt;p&gt;One reason teams overstate what Kubernetes gives them is that they imagine the alternative must be some giant internal developer portal project. If they are not ready to build Backstage templates, service catalogs, policy packs, and self-service workflows, they decide the cluster itself must count as the platform. Most of the useful work lives in the gap between those two options.&lt;/p&gt;

&lt;p&gt;Team Topologies has a saner idea: the &lt;strong&gt;thinnest viable platform&lt;/strong&gt;. On the Team Topologies TVP page, Matthew Skelton says that this can be “just a wiki page” if that is all you need at your current scale. That line keeps the argument grounded. What qualifies something as a platform is whether it gives your developers a clearer, safer, more repeatable way to get work done. Scale decides how thick it has to be; nothing else does.&lt;/p&gt;

&lt;p&gt;Sometimes the thinnest useful platform is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one golden path for a standard web service&lt;/li&gt;
&lt;li&gt;a starter template with the right defaults already wired in&lt;/li&gt;
&lt;li&gt;one supported CI path&lt;/li&gt;
&lt;li&gt;one documented way to expose a service&lt;/li&gt;
&lt;li&gt;one place to request a database or secret with known guardrails&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is still platform work, and at most organisations it is the right platform work to do first.&lt;/p&gt;

&lt;p&gt;Starting small is fine. The mistake is skipping the product layer entirely and calling the control plane beneath it finished.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a real platform adds on top of Kubernetes
&lt;/h2&gt;

&lt;p&gt;If Kubernetes is the substrate, what does the platform layer actually contribute?&lt;/p&gt;

&lt;p&gt;At minimum, it should remove repeated decisions that your teams do not benefit from making over and over.&lt;/p&gt;

&lt;p&gt;A decent internal platform usually adds:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;An opinionated service shape.&lt;/strong&gt; New services start from a template or scaffold that already includes the organisation’s baseline for probes, resources, identity, networking, and delivery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A supported path to production.&lt;/strong&gt; One default route that the platform team actively maintains, and that most services can take without needing a conversation first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails with teeth.&lt;/strong&gt; Policy, admission control, CI checks, and defaults that catch the boring mistakes before they get merged.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Self-service for common needs.&lt;/strong&gt; Databases, secrets, DNS, certificates, queues, environments, or delivery workflows that do not require a human to hand-hold every request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A product experience.&lt;/strong&gt; Good docs, clear ownership, visible support boundaries, and feedback loops so the thing improves instead of calcifying.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Notice what is missing from that list: “owning a cluster”. Owning a cluster is an operational capability, and platform engineering usually needs one. It buys you the right to start.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use the maturity model as a self-check
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://tag-app-delivery.cncf.io/whitepapers/platform-eng-maturity-model/" rel="noopener noreferrer"&gt;CNCF Platform Engineering Maturity Model&lt;/a&gt; is useful here because it stops the conversation from collapsing into slogans.&lt;/p&gt;

&lt;p&gt;It describes &lt;strong&gt;four levels&lt;/strong&gt; — &lt;code&gt;Provisional&lt;/code&gt;, &lt;code&gt;Operationalized&lt;/code&gt;, &lt;code&gt;Scalable&lt;/code&gt;, and &lt;code&gt;Optimizing&lt;/code&gt; — and looks across &lt;strong&gt;five aspects&lt;/strong&gt;: investment, adoption, interfaces, operations, and measurement.&lt;/p&gt;

&lt;p&gt;That works better as a diagnostic than asking whether you “have a platform”. Almost every team is somewhere in the middle of that grid, and the useful question is which cell they are in and which one they want to be in next.&lt;/p&gt;

&lt;p&gt;A few blunt checks help:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Interfaces:&lt;/strong&gt; do teams have a real self-service interface, or are they still stitching together YAML and Slack messages?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adoption:&lt;/strong&gt; do product teams voluntarily use the paved road because it is better, or because they were told to?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operations:&lt;/strong&gt; is there a supported day-two story behind the abstraction, or only a shiny day-one story?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measurement:&lt;/strong&gt; can you show that the platform shortens delivery loops or reduces support load, or is that still mostly assumed?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Investment:&lt;/strong&gt; is there sustained ownership, or is the “platform” just a side project orbiting the cluster team?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Answer those five properly and the picture resolves fast. A team with a well-run cluster, ad-hoc interfaces and adoption it has never measured is doing valuable infrastructure work. The platform sits a column to the right of where that team is standing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the finish line actually is
&lt;/h2&gt;

&lt;p&gt;The finish line is a state where most teams can ship a service without writing a production-grade manifest from scratch.&lt;/p&gt;

&lt;p&gt;You still want depth somewhere. Strong operational understanding in the platform team, and enough shared literacy that developers can reason about what runs underneath their code — that is a real asset, and it is a different thing from requiring it of everyone before they are allowed to deploy.&lt;/p&gt;

&lt;p&gt;Developers should be able to use the platform the way you hope customers use any good product — confidently, and without carrying a mental model of its internals around with them.&lt;/p&gt;

&lt;p&gt;Calling the substrate the platform has a specific cost. It lets you declare victory early, and after that every improvement gets spent on the cluster rather than on the people using it.&lt;/p&gt;

&lt;p&gt;If you want one thing to do this quarter: take the service shape your teams create most often and give it a single supported path — a template, a sane set of defaults, one documented route to production. Then time how long it takes somebody outside the platform team to get from an empty repository to a running service, before and after. If that number holds steady, you built the abstraction for your own convenience.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>platformengineering</category>
      <category>devrel</category>
      <category>internaldeveloperplatform</category>
    </item>
    <item>
      <title>Helm charts are technical debt: blame the templating model</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Mon, 31 Aug 2026 08:03:53 +0000</pubDate>
      <link>https://dev.to/coresolutions/helm-charts-are-technical-debt-blame-the-templating-model-19g9</link>
      <guid>https://dev.to/coresolutions/helm-charts-are-technical-debt-blame-the-templating-model-19g9</guid>
      <description>&lt;p&gt;The pull request is four lines of template, adding an optional &lt;code&gt;nodeSelector&lt;/code&gt; to an internal chart. CI has already done its job properly: &lt;code&gt;helm lint&lt;/code&gt; passed, the chart rendered against every environment's values file, &lt;code&gt;kubeconform&lt;/code&gt; validated each manifest that came out, and the snapshot diff of the rendered output is sitting in the check run waiting to be read.&lt;/p&gt;

&lt;p&gt;And you still cannot say, from the diff, what this change does in staging without going and reading three other templates first.&lt;/p&gt;

&lt;p&gt;Nobody files that as technical debt. It goes in the mental folder marked “charts are a bit fiddly”. But if your team maintains a dozen first-party charts, that tax gets levied on every review, and no amount of pipeline pays it off, because the pipeline was never the thing that was broken.&lt;/p&gt;

&lt;p&gt;This is not a “Helm was always bad” post. Helm solved a real problem, and most teams running Kubernetes are better off with it than without it. The rot is somewhere more specific: &lt;strong&gt;Helm renders Kubernetes manifests by running Go's &lt;code&gt;text/template&lt;/code&gt; over YAML as strings, and once you are authoring charts rather than just installing them, that model generates work forever.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That model has survived every major version so far, Helm 4 included. If your chart estate already feels brittle, the next release will not make it less so, and I will come back to why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Helm won for good reasons
&lt;/h2&gt;

&lt;p&gt;Helm is a CNCF graduated project and the default packaging story in Kubernetes, and it deserves to be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;install, upgrade, rollback and release history in a single tool&lt;/li&gt;
&lt;li&gt;a standard way for vendors and open-source projects to ship applications&lt;/li&gt;
&lt;li&gt;OCI registry support that slots into existing GitOps workflows&lt;/li&gt;
&lt;li&gt;ecosystem gravity, which by now matters nearly as much as technical merit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nobody sensible is still having the “should we use Helm at all?” argument. If you are installing Grafana, or &lt;a href="https://dev.to/blog/secrets-management-scale-external-secrets-operator"&gt;External Secrets Operator&lt;/a&gt;, or almost anything with a healthy community behind it, the chart is the path of least resistance and you should take it.&lt;/p&gt;

&lt;p&gt;The trouble is that the same tool that consumes packaged software so well also makes it very easy to start writing your own. One chart for the internal API. One for the batch worker. Then a shared library chart, because three of them had grown the same &lt;code&gt;_helpers.tpl&lt;/code&gt;. At some point you stopped being a Helm user and became the maintainer of a templating system. That is a perfectly reasonable thing to be, as long as it was a decision rather than a drift.&lt;/p&gt;

&lt;h2&gt;
  
  
  The debt is in the rendering model
&lt;/h2&gt;

&lt;p&gt;Helm templates know nothing about Kubernetes objects. They emit text that happens to be YAML.&lt;/p&gt;

&lt;p&gt;Here is the shape of it, lifted from roughly every &lt;code&gt;deployment.yaml&lt;/code&gt; in every internal chart:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# templates/deployment.yaml&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;{{&lt;/span&gt;&lt;span class="nv"&gt;- with .Values.nodeSelector&lt;/span&gt; &lt;span class="pi"&gt;}}&lt;/span&gt;
      &lt;span class="na"&gt;nodeSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;{{&lt;/span&gt;&lt;span class="nv"&gt;- toYaml . | nindent 8&lt;/span&gt; &lt;span class="pi"&gt;}}&lt;/span&gt;
      &lt;span class="pi"&gt;{{&lt;/span&gt;&lt;span class="nv"&gt;- end&lt;/span&gt; &lt;span class="pi"&gt;}}&lt;/span&gt;
      &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{{&lt;/span&gt; &lt;span class="nv"&gt;.Chart.Name&lt;/span&gt; &lt;span class="pi"&gt;}}&lt;/span&gt;
          &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{{&lt;/span&gt; &lt;span class="nv"&gt;.Values.image.repository&lt;/span&gt; &lt;span class="pi"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;:{{ .Values.image.tag }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;8&lt;/code&gt; is load-bearing. &lt;code&gt;nodeSelector:&lt;/code&gt; sits at column six, so its keys have to land at column eight, and &lt;a href="https://helm.sh/docs/chart_template_guide/function_list/" rel="noopener noreferrer"&gt;&lt;code&gt;nindent&lt;/code&gt;&lt;/a&gt; exists precisely because chart authors need to re-indent embedded blocks constantly. Write &lt;code&gt;nindent 6&lt;/code&gt; by mistake and you have not made a formatting error. You have set &lt;code&gt;nodeSelector&lt;/code&gt; to null and promoted its contents to siblings of &lt;code&gt;containers&lt;/code&gt;, and the render will print that quite happily, because it is still a well-formed YAML document. Nothing in the YAML layer objects. It takes schema validation downstream, &lt;code&gt;kubeconform -strict&lt;/code&gt; in the pipeline or a server-side dry run, before anything flags it. That is a lot of machinery to need in order to catch a two-space mistake.&lt;/p&gt;

&lt;p&gt;Two spaces. Helm's own &lt;a href="https://helm.sh/docs/chart_best_practices/templates/" rel="noopener noreferrer"&gt;template best practices&lt;/a&gt; page walks through indentation, whitespace chomping and generated-output formatting, and concedes the underlying problem in passing: “YAML is a whitespace-oriented language.” Full marks for honesty, but that is a peculiar property for an abstraction layer to have.&lt;/p&gt;

&lt;p&gt;All of which is the cheap version of the problem, and it is cheap precisely because a machine can express it. Schema validation catches a wrong shape, and it catches it in seconds. What no pipeline catches is a chart whose &lt;em&gt;right&lt;/em&gt; shapes have become impossible to hold in your head.&lt;/p&gt;

&lt;p&gt;Once your abstraction works on strings rather than typed objects, a few things follow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;indentation becomes part of correctness, not just tidiness&lt;/li&gt;
&lt;li&gt;a branch that renders the wrong shape still looks plausible in review&lt;/li&gt;
&lt;li&gt;values become an API surface long before anyone has designed them like one&lt;/li&gt;
&lt;li&gt;refactoring is hard, because the behaviour lives across helpers, partials and value conventions rather than in a schema your tooling understands&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Disciplined teams are not exempt from this. They often accumulate &lt;em&gt;more&lt;/em&gt; chart machinery, because they build helpers and layering and conventions to contain the complexity. That containment is genuinely useful, and it is also the tell: you are building scaffolding around a model that produces complexity faster than you can absorb it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it shows up on day two
&lt;/h2&gt;

&lt;p&gt;You rarely spot chart debt by reading a chart. You spot it in the changes that need far more scrutiny than their diff size suggests they should.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;values.yaml&lt;/code&gt; turns into an untyped product surface
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;values.yaml&lt;/code&gt; is the front door to your chart, which is convenient right up until you realise you have shipped a long-lived API with no type system behind it.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;values.schema.json&lt;/code&gt; belongs in every first-party chart, and Helm will enforce it on &lt;code&gt;install&lt;/code&gt;, &lt;code&gt;upgrade&lt;/code&gt;, &lt;code&gt;lint&lt;/code&gt; and &lt;code&gt;template&lt;/code&gt;. But the &lt;a href="https://helm.sh/docs/topics/charts/#schema-files" rel="noopener noreferrer"&gt;schema file is optional&lt;/a&gt;, so it is opt-in per chart and easy to leave behind on the one chart nobody has revisited, and it describes structure rather than behaviour. A value can be structurally perfect and still produce nonsense once it has threaded its way through conditionals, helper templates, defaults and subchart overrides.&lt;/p&gt;

&lt;p&gt;Then another team starts depending on a value name. Renaming a field stops being a cleanup and becomes a migration, complete with a deprecation window and a note in the release notes that nobody reads.&lt;/p&gt;

&lt;h3&gt;
  
  
  The source is harder to review than the output
&lt;/h3&gt;

&lt;p&gt;A rendered manifest can be entirely ordinary while the logic that produced it is anything but.&lt;/p&gt;

&lt;p&gt;Humans review the chart source. A pull request that flips a condition or threads one more flag through three templates is not asking “is this &lt;code&gt;Deployment&lt;/code&gt; valid?”. It is asking “which combinations of values, defaults and includes now produce a different document?”, and that question has no diff you can look at.&lt;/p&gt;

&lt;p&gt;Which is why the gates all end up pointed at the rendered output rather than the source. Those gates are worth building, and I will come back to ours. They are also a fair signal about the model: you test the output because nobody can reliably reason about the input.&lt;/p&gt;

&lt;h3&gt;
  
  
  Subcharts and value plumbing sprawl quietly
&lt;/h3&gt;

&lt;p&gt;Small Helm setups look tidy. Then the platform grows, one chart wraps five subcharts, globals appear, child values get overridden from the parent, and a flag in one place changes behaviour three templates away.&lt;/p&gt;

&lt;p&gt;Some of those values exist because users asked for them. Some exist because a subchart expects them. Some exist because somebody needed a workaround two years ago and nobody has ever felt brave enough to delete it. That is the point where a chart stops being a packaging unit and starts being institutional memory encoded as YAML conventions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Helm 4 improved Helm without changing this
&lt;/h2&gt;

&lt;p&gt;This is where lazy anti-Helm takes get sloppy, so let me be precise about it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/helm/helm/releases/tag/v4.0.0" rel="noopener noreferrer"&gt;Helm 4.0.0&lt;/a&gt; shipped in November 2025 and it is a genuine improvement: server-side apply, a redesigned plugin system with WebAssembly plugins, resource watching built on &lt;code&gt;kstatus&lt;/code&gt;, post-renderers as plugins, reproducible chart archives. If you run Helm heavily, upgrade.&lt;/p&gt;

&lt;p&gt;All of that is release machinery, though. Rendering still goes through Go's text templating, &lt;code&gt;values.yaml&lt;/code&gt; is still the primary interface, and the combinatorial growth of flags and helpers in your own charts remains entirely your problem. Helm 4 made Helm better at the job Helm is good at. Chart authoring was never on the list.&lt;/p&gt;

&lt;p&gt;A version bump on its own will not retire this argument, so here is what would: a release that swaps text templating for a typed object model, or one that makes schema validation mandatory instead of opt-in. If you are reading this a few releases later, that is the thing to go and check. Short of it, the version numbers in this post move and the point does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the alternatives actually change
&lt;/h2&gt;

&lt;p&gt;Different tools replace different &lt;em&gt;parts&lt;/em&gt; of Helm, which is where most “Helm alternatives” posts go wrong. There is no swap that hands you better authoring, better packaging, better release management and better ecosystem compatibility all at once.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;What it replaces&lt;/th&gt;
&lt;th&gt;What you still need&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Kustomize&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Templating, for the overlay case&lt;/td&gt;
&lt;td&gt;Packaging, versioning, release history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Timoni&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The authoring model, with CUE typing&lt;/td&gt;
&lt;td&gt;Ecosystem gravity, and patience while it is pre-1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;cdk8s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Templating, with a real programming language&lt;/td&gt;
&lt;td&gt;Packaging, plus a language runtime you now own&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Jsonnet / Tanka&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Templating, with a compact composition syntax&lt;/td&gt;
&lt;td&gt;Types, and a way to stop dynamic behaviour relocating complexity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;KCL&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Templating, with a typed config language&lt;/td&gt;
&lt;td&gt;Ecosystem gravity, and a team willing to learn a new language&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Score&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The developer-facing values surface&lt;/td&gt;
&lt;td&gt;A platform implementation underneath that satisfies the spec&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://kubectl.docs.kubernetes.io/references/kustomize/" rel="noopener noreferrer"&gt;Kustomize&lt;/a&gt; is the one I reach for most often outside Helm, largely because it already ships inside &lt;code&gt;kubectl&lt;/code&gt; and patches structured YAML instead of templating arbitrary strings. If your real problem is “I have plain manifests and need per-environment overlays”, it removes a surprising amount of cleverness for almost no adoption cost. Expecting it to do packaging too is how people end up disappointed by it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://timoni.sh/" rel="noopener noreferrer"&gt;Timoni&lt;/a&gt; is the most interesting answer to “what if we fixed the authoring model?”. CUE, real typing, modules distributed as OCI artifacts: conceptually most of what people are asking for. The caveat is maturity, and it is still on its 0.x line (&lt;code&gt;v0.29.0&lt;/code&gt; as I write this), so go and check where it has got to before committing. A 1.0 would change the calculation considerably more than another point release will.&lt;/p&gt;

&lt;p&gt;The code-generation family trades typing for operational surface. &lt;a href="https://www.cncf.io/projects/cdk-for-kubernetes-cdk8s/" rel="noopener noreferrer"&gt;cdk8s&lt;/a&gt; (joined the CNCF Sandbox in 2020) gives you real programming languages, Jsonnet with Tanka gives you a compact composition language, and &lt;a href="https://www.cncf.io/projects/kcl/" rel="noopener noreferrer"&gt;KCL&lt;/a&gt; (joined in September 2023) gives you an explicitly typed configuration language. In exchange you own a language runtime, a build step, and an onboarding cost for everyone who ever touches a manifest. Good trade for a platform team with a large estate, poor trade for a team with six charts.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.cncf.io/projects/score/" rel="noopener noreferrer"&gt;Score&lt;/a&gt; (joined in July 2024) sits one layer up. It is a workload specification, asking “what does this application need from the platform?” rather than “how do I render manifests for every deployment case?”. If your real complaint is that developers keep being handed an ever-growing values surface to fill in, that is the layer to attack. It still assumes a platform underneath that can satisfy it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How we actually run this
&lt;/h2&gt;

&lt;p&gt;We stayed on Helm, so it is only fair to say what containing the cost looks like day to day.&lt;/p&gt;

&lt;p&gt;Most of our platform components are wrapper charts, and that pattern carries a lot of the weight. &lt;code&gt;Chart.yaml&lt;/code&gt; declares a pinned dependency on the upstream chart, &lt;code&gt;Chart.lock&lt;/code&gt; freezes what actually resolved, and our own &lt;code&gt;templates/&lt;/code&gt; directory layers on what upstream does not ship: the NetworkPolicies, the PrometheusRules, the RBAC that upstream leaves to you. Renovate raises the upstream bumps, and they do not land until CI is green.&lt;/p&gt;

&lt;p&gt;Values sit in one place per chart. &lt;code&gt;values.yaml&lt;/code&gt; carries the defaults every environment shares, and a &lt;code&gt;values/&lt;/code&gt; directory holds the environment-specific overrides beside it, one file each. CI then lints the chart once against every one of those files rather than once overall, so a change that only breaks one environment cannot ride in on another's values.&lt;/p&gt;

&lt;p&gt;That catches everything Helm can be made to fail on. The quieter failure needs its own gate, because of a footgun worth internalising: Helm deep-merges maps, but it replaces lists outright. An environment file that overrides a list has to restate the entire list, so adding a component in one environment and forgetting another raises no error anywhere. Both files are valid. You just ship a shorter list to the environment nobody updated, and find out when something is missing. Nothing you can put in &lt;code&gt;values.yaml&lt;/code&gt; expresses “these lists must stay in step”, so a parity check in CI does it instead, and a red PR replaces copy-paste discipline. That is the scaffolding tax from earlier made concrete: a shell script exists because the type system does not.&lt;/p&gt;

&lt;p&gt;None of this makes the templating model typed. It makes the cost visible and bounded, which is the achievable goal rather than the ideal one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I would draw the line
&lt;/h2&gt;

&lt;p&gt;Helm is not technical debt because it exists in your stack, and authoring your own charts is not a mistake. We do both, and given the same constraints we would do both again. A third-party chart you install, version and occasionally override is packaging doing exactly the job it is good at, and a first-party chart is a reasonable way to ship your own software into a cluster where everything else already arrives as one.&lt;/p&gt;

&lt;p&gt;So I reach for Helm when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I am consuming third-party software that already ships a chart and I have no wish to become its packager&lt;/li&gt;
&lt;li&gt;I want release history, rollback and the install workflow Helm already gives me&lt;/li&gt;
&lt;li&gt;the surrounding ecosystem expects Helm artifacts, &lt;a href="https://dev.to/blog/gitops-for-kubernetes-with-argo-cd"&gt;GitOps tooling&lt;/a&gt; included&lt;/li&gt;
&lt;li&gt;I am packaging our own components for a platform where every other component already arrives as a chart&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What accrues interest is the estate you build afterwards, and how much of it you let grow without a gate on it. Our largest first-party chart carries a &lt;code&gt;values.yaml&lt;/code&gt; north of 800 lines. That is not automatically a failure, but it is a number I would want any team to be able to quote about themselves, because the moment nobody can tell you what a chart's values surface costs, you have stopped managing it. Same instinct as &lt;a href="https://dev.to/blog/stop-copying-big-tech-platform-architecture"&gt;right-sizing the rest of your platform&lt;/a&gt;: the tool is fine, and the real question is whether the estate around it still fits the team.&lt;/p&gt;

&lt;p&gt;The alternatives are worth knowing even when you stay. We have stayed, because packaging gravity and release history are worth more to us today than typed authoring would be. If that trade flips, if the estate outgrows its gates or Timoni reaches 1.0 with real ecosystem behind it, the honest move is to notice rather than to defend the original choice.&lt;/p&gt;

&lt;p&gt;You do not have to become anti-Helm to admit that the charts you maintain can rot. In most Kubernetes estates that is hardly a hot take. It is just what year two looks like.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>helm</category>
      <category>kustomize</category>
    </item>
    <item>
      <title>PodSecurityPolicy replacement: Pod Security Admission</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Mon, 31 Aug 2026 07:57:40 +0000</pubDate>
      <link>https://dev.to/coresolutions/podsecuritypolicy-replacement-pod-security-admission-45i3</link>
      <guid>https://dev.to/coresolutions/podsecuritypolicy-replacement-pod-security-admission-45i3</guid>
      <description>&lt;p&gt;If you still think of Pod Security Admission as the cut-down replacement for PodSecurityPolicy, that is exactly the mindset that makes rollouts harder than they need to be. It &lt;em&gt;is&lt;/em&gt; less powerful. That is also why it is practical. You get three fixed profiles, three modes, a handful of labels, and an admission controller that is already in the API server. No custom policy language, no mutation, no controller to install, and far fewer ways to surprise yourself mid-change.&lt;/p&gt;

&lt;p&gt;The trick is to use that simplicity properly. A safe rollout is not “turn on &lt;code&gt;restricted&lt;/code&gt; everywhere and hope for the best”. It is version-pinned labels, audit and warn first, then namespace-by-namespace enforcement once you have seen what will break. Done that way, you can move from the old PSP world to a built-in control that teams will actually leave enabled.&lt;/p&gt;

&lt;h2&gt;
  
  
  What replaced PodSecurityPolicy
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;PodSecurityPolicy&lt;/code&gt; was deprecated in Kubernetes &lt;code&gt;v1.21&lt;/code&gt; and removed in &lt;code&gt;v1.25&lt;/code&gt;. &lt;a href="https://kubernetes.io/docs/concepts/security/pod-security-admission/" rel="noopener noreferrer"&gt;Pod Security Admission&lt;/a&gt; became generally available in the same release, which matters because there is nothing extra to deploy: if your cluster is current enough, the mechanism is already there.&lt;/p&gt;

&lt;p&gt;The model is deliberately small:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Profiles&lt;/strong&gt;: &lt;code&gt;privileged&lt;/code&gt;, &lt;code&gt;baseline&lt;/code&gt;, &lt;code&gt;restricted&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Modes&lt;/strong&gt;: &lt;code&gt;enforce&lt;/code&gt;, &lt;code&gt;audit&lt;/code&gt;, &lt;code&gt;warn&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Version pinning&lt;/strong&gt;: &lt;code&gt;pod-security.kubernetes.io/&amp;lt;mode&amp;gt;-version&lt;/code&gt;, pinned to a Kubernetes minor version or &lt;code&gt;latest&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those labels live on namespaces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# pin policy to the minor version your cluster is running&lt;/span&gt;
&lt;span class="nv"&gt;pss&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;kubectl version &lt;span class="nt"&gt;-o&lt;/span&gt; json | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.serverVersion.gitVersion'&lt;/span&gt; | &lt;span class="nb"&gt;cut&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;-f1&lt;/span&gt;,2&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

kubectl label ns payments &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/enforce&lt;span class="o"&gt;=&lt;/span&gt;baseline &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/enforce-version&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$pss&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That simplicity is the point. PSP mixed policy with RBAC, allowed mutation, and had enough moving parts that many teams either avoided it or never quite trusted what a change would do. PSA is narrower. It sets a floor.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model: three profiles, three modes, one version pin
&lt;/h2&gt;

&lt;p&gt;You can think of PSA as a matrix:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;What happens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;enforce&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Rejects a non-compliant pod creation request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;audit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Allows it, but records an audit annotation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;warn&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Allows it, but shows a warning to the caller&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the profiles are ordered from least to most strict:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Profile&lt;/th&gt;
&lt;th&gt;What it is for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;privileged&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;System namespaces and workloads that genuinely need host access or privileged behaviour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;baseline&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A sensible minimum floor for ordinary application namespaces&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;restricted&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Current pod hardening best practice&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For most application namespaces, the safe starting pattern is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl label &lt;span class="nt"&gt;--overwrite&lt;/span&gt; ns payments &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/enforce&lt;span class="o"&gt;=&lt;/span&gt;baseline &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/enforce-version&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$pss&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/warn&lt;span class="o"&gt;=&lt;/span&gt;restricted &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/warn-version&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$pss&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/audit&lt;span class="o"&gt;=&lt;/span&gt;restricted &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/audit-version&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$pss&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gives you a firm minimum bar in &lt;code&gt;enforce&lt;/code&gt;, while &lt;code&gt;warn&lt;/code&gt; and &lt;code&gt;audit&lt;/code&gt; show you what stands between the namespace and &lt;code&gt;restricted&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The version pin is not optional bureaucracy. If you leave a namespace on &lt;code&gt;latest&lt;/code&gt;, any cluster upgrade can quietly tighten policy under your workloads. Pinning to the minor version you are actually running means policy changes happen when &lt;em&gt;you&lt;/em&gt; decide to move the label.&lt;/p&gt;

&lt;h2&gt;
  
  
  What &lt;code&gt;restricted&lt;/code&gt; actually checks
&lt;/h2&gt;

&lt;p&gt;The easiest way to lose credibility on this topic is to hand-wave the &lt;code&gt;restricted&lt;/code&gt; profile and get one detail wrong. Several of the checklists floating around in other posts are simply false — especially around &lt;code&gt;runAsUser&lt;/code&gt; and seccomp.&lt;/p&gt;

&lt;p&gt;Here is the practical checklist, straight from the &lt;a href="https://kubernetes.io/docs/concepts/security/pod-security-standards/" rel="noopener noreferrer"&gt;Pod Security Standards&lt;/a&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Control&lt;/th&gt;
&lt;th&gt;What &lt;code&gt;restricted&lt;/code&gt; expects&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Volume types&lt;/td&gt;
&lt;td&gt;Only &lt;code&gt;configMap&lt;/code&gt;, &lt;code&gt;csi&lt;/code&gt;, &lt;code&gt;downwardAPI&lt;/code&gt;, &lt;code&gt;emptyDir&lt;/code&gt;, &lt;code&gt;ephemeral&lt;/code&gt;, &lt;code&gt;persistentVolumeClaim&lt;/code&gt;, &lt;code&gt;projected&lt;/code&gt;, &lt;code&gt;secret&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Privilege escalation&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;allowPrivilegeEscalation: false&lt;/code&gt; on every container, including init and ephemeral containers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Non-root&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;runAsNonRoot: true&lt;/code&gt; at pod level or per container&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UID&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;runAsUser&lt;/code&gt; does &lt;strong&gt;not&lt;/strong&gt; need to be set, but if it is, it cannot be &lt;code&gt;0&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capabilities&lt;/td&gt;
&lt;td&gt;Drop &lt;code&gt;ALL&lt;/code&gt;; the only allowed add is &lt;code&gt;NET_BIND_SERVICE&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seccomp&lt;/td&gt;
&lt;td&gt;Must be explicitly &lt;code&gt;RuntimeDefault&lt;/code&gt; or &lt;code&gt;Localhost&lt;/code&gt;; leaving it unset is a violation under &lt;code&gt;restricted&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two of those rows are where the bad checklists go wrong. &lt;code&gt;runAsUser&lt;/code&gt; is not required — the policy forbids UID &lt;code&gt;0&lt;/code&gt;, it does not make you hard-code some other UID just to satisfy the admission check. And seccomp cuts the other way: under &lt;code&gt;baseline&lt;/code&gt; an unset profile is fine, but under &lt;code&gt;restricted&lt;/code&gt; it is a violation in its own right, which is one of the reasons plain demo manifests still bounce when teams first try this.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rollout that does not break production
&lt;/h2&gt;

&lt;p&gt;The safest PSA rollout starts with a dry run, not an enforcement change — the &lt;a href="https://kubernetes.io/docs/tasks/configure-pod-container/migrate-from-psp/" rel="noopener noreferrer"&gt;official migration guide&lt;/a&gt; is built around the same idea.&lt;/p&gt;

&lt;p&gt;First, test label application server-side across all namespaces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl label &lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;server &lt;span class="nt"&gt;--overwrite&lt;/span&gt; ns &lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/enforce&lt;span class="o"&gt;=&lt;/span&gt;baseline
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That does not change anything. It asks the API server what violations it would report if the label were applied.&lt;/p&gt;

&lt;p&gt;Then stage &lt;code&gt;audit&lt;/code&gt; and &lt;code&gt;warn&lt;/code&gt; cluster-wide:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl label &lt;span class="nt"&gt;--overwrite&lt;/span&gt; ns &lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/audit&lt;span class="o"&gt;=&lt;/span&gt;baseline &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/warn&lt;span class="o"&gt;=&lt;/span&gt;baseline
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or, if your end state is more ambitious, go straight to the pattern that tends to work well in practice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;enforce=baseline&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;warn=restricted&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;audit=restricted&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That gives you signal without immediately creating an outage because some forgotten sidecar, root-running image, or hostPath mount was hiding in a namespace nobody had looked at for months.&lt;/p&gt;

&lt;p&gt;You can also find namespaces that have no PSA labels yet:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get ns &lt;span class="nt"&gt;--selector&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'!pod-security.kubernetes.io/enforce'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the warnings are clean for a namespace, pin the version and turn on &lt;code&gt;enforce&lt;/code&gt; there. Namespace by namespace is slower than one giant switch, but slower is exactly what you want when admission control is involved.&lt;/p&gt;

&lt;p&gt;While the rollout is in flight, the API server's own metrics — &lt;code&gt;pod_security_evaluations_total&lt;/code&gt;, &lt;code&gt;pod_security_errors_total&lt;/code&gt; and &lt;code&gt;pod_security_exemptions_total&lt;/code&gt; — are worth scraping. They tell you whether evaluation is happening where you think it is, and whether exemptions are starting to sprawl.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a real rejection looks like
&lt;/h2&gt;

&lt;p&gt;The good news about PSA failures is that the error message is the fix list.&lt;/p&gt;

&lt;p&gt;Apply a plain pod into a namespace with &lt;code&gt;enforce=restricted&lt;/code&gt; and you get something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error from server (Forbidden): pods "nginx" is forbidden: violates PodSecurity "restricted:latest":
allowPrivilegeEscalation != false (container "nginx" must set securityContext.allowPrivilegeEscalation=false),
unrestricted capabilities (container "nginx" must set securityContext.capabilities.drop=["ALL"]),
runAsNonRoot != true (pod or container "nginx" must set securityContext.runAsNonRoot=true),
seccompProfile (pod or container "nginx" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That message is unusually helpful by Kubernetes standards. Every violation names the exact field to set and the value it wants, so you can work through it like a checklist.&lt;/p&gt;

&lt;p&gt;A minimal compliant pod looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;securityContext&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runAsNonRoot&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;seccompProfile&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RuntimeDefault&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
      &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginxinc/nginx-unprivileged:stable&lt;/span&gt;
      &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;containerPort&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8080&lt;/span&gt;
      &lt;span class="na"&gt;securityContext&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;allowPrivilegeEscalation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
        &lt;span class="na"&gt;capabilities&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;drop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;ALL&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two quiet fixes are hiding in that example: the container-level security context, and the image itself. A lot of “hello world” manifests still assume a root-running image on port &lt;code&gt;80&lt;/code&gt;, which is a fine way to discover that &lt;code&gt;restricted&lt;/code&gt; is not interested in your tutorial shortcuts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trap most teams hit once: Deployments apply, pods do not
&lt;/h2&gt;

&lt;p&gt;This is the operational detail that makes or breaks a rollout plan.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;warn&lt;/code&gt; and &lt;code&gt;audit&lt;/code&gt; evaluate workload resources such as &lt;code&gt;Deployment&lt;/code&gt; pod templates. &lt;code&gt;enforce&lt;/code&gt; does &lt;strong&gt;not&lt;/strong&gt;. &lt;code&gt;enforce&lt;/code&gt; applies only to the resulting pod objects.&lt;/p&gt;

&lt;p&gt;That means this can happen:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You apply a &lt;code&gt;Deployment&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Kubernetes accepts the &lt;code&gt;Deployment&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;ReplicaSet&lt;/code&gt; tries to create pods&lt;/li&gt;
&lt;li&gt;The pods are rejected by PSA&lt;/li&gt;
&lt;li&gt;You now have a deployment that looks “applied” but never becomes healthy&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When that happens, the signal is in the &lt;code&gt;ReplicaSet&lt;/code&gt; and the events, not in the initial &lt;code&gt;kubectl apply&lt;/code&gt; output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get rs &lt;span class="nt"&gt;-n&lt;/span&gt; payments
kubectl get events &lt;span class="nt"&gt;-n&lt;/span&gt; payments &lt;span class="nt"&gt;--sort-by&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;.metadata.creationTimestamp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you remember only one rollout lesson from this post, make it this one. Admission on pod creation and admission on workload templates are not symmetrical.&lt;/p&gt;

&lt;h2&gt;
  
  
  The namespaces that should not be &lt;code&gt;restricted&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;kube-system&lt;/code&gt; is the obvious example. CNIs, CSIs, node agents, and static pods routinely need things that &lt;code&gt;baseline&lt;/code&gt; or &lt;code&gt;restricted&lt;/code&gt; will reject: &lt;code&gt;hostPath&lt;/code&gt;, &lt;code&gt;hostNetwork&lt;/code&gt;, privileged containers, or other host-level access.&lt;/p&gt;

&lt;p&gt;Be explicit about that. Label those namespaces &lt;code&gt;privileged&lt;/code&gt; so the exception is documented in the cluster state rather than living as folklore.&lt;/p&gt;

&lt;p&gt;Istio adds another common surprise — relevant if you run &lt;a href="https://dev.to/blog/zero-trust-networking-kubernetes-istio-service-mesh"&gt;a mesh for zero-trust networking&lt;/a&gt;. With sidecar injection but without &lt;a href="https://istio.io/latest/docs/setup/additional-setup/cni/" rel="noopener noreferrer"&gt;Istio CNI&lt;/a&gt;, the injected &lt;code&gt;istio-init&lt;/code&gt; container needs &lt;code&gt;NET_ADMIN&lt;/code&gt; and &lt;code&gt;NET_RAW&lt;/code&gt; to set up traffic redirection. That is enough to fail &lt;strong&gt;baseline&lt;/strong&gt;, not just &lt;code&gt;restricted&lt;/code&gt;. The usual fix is to enable the CNI node agent and keep &lt;code&gt;istio-system&lt;/code&gt; itself &lt;code&gt;privileged&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why version pinning matters more than it looks
&lt;/h2&gt;

&lt;p&gt;A pinned PSA version is an operational control, not a documentation detail.&lt;/p&gt;

&lt;p&gt;Kubernetes &lt;code&gt;v1.34&lt;/code&gt; added a baseline control that blocks the &lt;code&gt;host&lt;/code&gt; field in &lt;code&gt;httpGet&lt;/code&gt; and &lt;code&gt;tcpSocket&lt;/code&gt; probes and lifecycle hooks. The reason is sensible: that field can be abused as an SSRF path through the kubelet. The operational consequence is that a cluster upgrade can start rejecting manifests that were fine the day before — &lt;em&gt;if&lt;/em&gt; your labels track &lt;code&gt;latest&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;With version pinning, you choose when that tighter rule lands. With &lt;code&gt;$pss&lt;/code&gt; still set to the minor version you were running before the upgrade:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl label &lt;span class="nt"&gt;--overwrite&lt;/span&gt; ns payments &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/enforce-version&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$pss&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/warn-version&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$pss&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  pod-security.kubernetes.io/audit-version&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$pss&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That makes upgrades much less dramatic. Upgrade the cluster first. Move the policy version when you are ready to deal with the findings.&lt;/p&gt;

&lt;h2&gt;
  
  
  The RBAC hole nobody should ignore
&lt;/h2&gt;

&lt;p&gt;PSA policy lives on namespace labels. That means anyone who can update namespace labels can also weaken or remove enforcement.&lt;/p&gt;

&lt;p&gt;Review that RBAC carefully. It is easy to focus on pod-creation rights and forget that &lt;code&gt;update&lt;/code&gt; on namespaces is effectively policy-admin access.&lt;/p&gt;

&lt;p&gt;There is a more central exemption mechanism through the API server &lt;code&gt;AdmissionConfiguration&lt;/code&gt;, where you can exempt namespaces, usernames, or runtime classes. That is useful on self-managed control planes. On managed services such as EKS, GKE, and AKS, you usually do not get to edit that file at all, which means namespace labels are the real policy surface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Pod Security Admission stops
&lt;/h2&gt;

&lt;p&gt;PSA is the built-in floor. That is its job, and it is a good one.&lt;/p&gt;

&lt;p&gt;There is no mutation, so no PSP-style defaulting is waiting to save a sloppy manifest. There is no custom rule language, and no way to express fine-grained per-workload exceptions inside one namespace. Once you want those things, you are into policy engines such as Kyverno or OPA Gatekeeper.&lt;/p&gt;

&lt;p&gt;That is not a weakness so much as a division of labour. Use PSA to guarantee the cluster-wide minimum. Use a policy engine above it when you need mutation, richer exceptions, or application-specific constraints.&lt;/p&gt;

&lt;p&gt;PodSecurityPolicy is gone. The sensible replacement is not to rebuild PSP in another tool and pretend nothing changed. It is to accept the smaller built-in model for what it is: predictable, cheap to operate, and strong enough to become the default floor across a cluster. That is a better trade than a perfect policy system nobody trusts enough to enable.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>security</category>
      <category>kubernetes</category>
    </item>
    <item>
      <title>Kubernetes cost allocation with OpenCost</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Wed, 05 Aug 2026 08:12:08 +0000</pubDate>
      <link>https://dev.to/bwlkr/kubernetes-cost-allocation-with-opencost-4dol</link>
      <guid>https://dev.to/bwlkr/kubernetes-cost-allocation-with-opencost-4dol</guid>
      <description>&lt;p&gt;You can have a perfectly healthy Kubernetes cluster and still have no idea who is driving the bill. The cloud invoice will tell you what the cluster cost. It will not tell you which namespace, workload, or team consumed that capacity, or how much of the spend was simply idle headroom nobody requested. OpenCost fills that gap. It gives you a vendor-neutral way to allocate Kubernetes cost down to cluster objects, query it over an API, and build showback without buying into a commercial platform first.&lt;/p&gt;

&lt;p&gt;As of v1.121.0 (July 2026), the current release, OpenCost is a CNCF Incubating project. The useful mental split is straightforward: OpenCost is the open-source engine and specification; Kubecost is the commercial product line in the same space. The OpenCost repository notes that the project was originally developed and open-sourced by Kubecost, while Kubecost's public site is now branded IBM Kubecost. For this post, we are staying on the OSS side: self-hosted OpenCost, standard Helm install, standard API, standard Prometheus queries.&lt;/p&gt;

&lt;h2&gt;
  
  
  The allocation model is the part to understand first
&lt;/h2&gt;

&lt;p&gt;If you only remember one detail from OpenCost, make it this one: workload cost is not just raw usage. In the OpenCost specification, workload costs are defined as &lt;code&gt;max(request, usage)&lt;/code&gt; for resources with allocation costs such as CPU and GPU. That matters because the bill follows reserved and allocatable capacity rather than whatever a container happened to burn in a five-minute slice.&lt;/p&gt;

&lt;p&gt;A pod that requests &lt;code&gt;2&lt;/code&gt; CPU and uses &lt;code&gt;200m&lt;/code&gt; is still consuming scheduling capacity someone else cannot have. OpenCost treats that as real cost, which is why its numbers are useful for showback and bill reconciliation instead of just efficiency charts.&lt;/p&gt;

&lt;p&gt;Idle cost is the second half of the picture. The specification defines cluster idle cost as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cluster Idle Cost = Cluster Asset Costs - Workload Costs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the spend you are carrying in the cluster but have not allocated to workloads. It is the cost of spare node capacity, overprovisioned requests, and all the other room you are paying for so the platform can absorb change.&lt;/p&gt;

&lt;p&gt;OpenCost also treats shared costs as first-class. The specification calls out three common ways to distribute them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;uniformly across tenants&lt;/li&gt;
&lt;li&gt;proportionally to a tenant's consumption&lt;/li&gt;
&lt;li&gt;by a custom metric such as network egress&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That means the honest answer to “what does this namespace cost?” is often “its direct workload cost, plus its share of idle and shared platform cost”. If you skip that second part, your showback report will look cleaner than the bill you actually pay.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install it with Helm against an existing Prometheus
&lt;/h2&gt;

&lt;p&gt;The current Helm docs assume you already have Prometheus. That is still the normal path.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm repo add opencost-charts https://opencost.github.io/opencost-helm-chart
helm repo update
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then point OpenCost at the Prometheus service it should read from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;opencost&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;prometheus&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;internal&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;namespaceName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;monitoring&lt;/span&gt;
      &lt;span class="na"&gt;serviceName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;prometheus-kube-prometheus-prometheus&lt;/span&gt;
      &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;9090&lt;/span&gt;
  &lt;span class="na"&gt;exporter&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;defaultClusterId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production-eks&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And install it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;helm &lt;span class="nb"&gt;install &lt;/span&gt;opencost opencost-charts/opencost &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; opencost &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--create-namespace&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-f&lt;/span&gt; values.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gives you the full OpenCost deployment plus the API and UI. The docs use &lt;code&gt;localhost:9003&lt;/code&gt; as the default API address, so a quick validation path is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl &lt;span class="nt"&gt;-n&lt;/span&gt; opencost port-forward deployment/opencost 9003 9090
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Port &lt;code&gt;9003&lt;/code&gt; is the API; &lt;code&gt;9090&lt;/code&gt; is the UI.&lt;/p&gt;

&lt;p&gt;There is a lighter path worth knowing about too. The Prometheus integration docs note that you do not need a Prometheus client just to emit cost metrics, and a "Promless" mode has been taking shape across releases since late 2025. Treat that as a fast-moving area rather than a settled default. If you want the most documented path today, use the Helm install against an existing Prometheus stack. If you want the leaner route, pin the version and read the release notes for the exact build you are adopting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the Allocation API, not a hand-rolled dashboard
&lt;/h2&gt;

&lt;p&gt;The fastest way to get useful showback is the Allocation API. It already applies the OpenCost cost model, which means you are querying the model's own output instead of rebuilding the logic yourself out of raw metrics.&lt;/p&gt;

&lt;p&gt;A seven-day namespace breakdown looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s1"&gt;'http://localhost:9003/allocation?window=7d&amp;amp;aggregate=namespace&amp;amp;shareIdle=true'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two parts of that query earn their keep:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;aggregate=namespace&lt;/code&gt;&lt;/strong&gt; gives you a shape platform teams can actually hand to someone. Most organisations do not want “cost by pod” as the first report; they want “what did team-a's namespace cost this week?” Namespace is a good default grain because it is close enough to ownership to start a useful conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;shareIdle=true&lt;/code&gt;&lt;/strong&gt; is what keeps the number honest. Without idle-cost sharing, namespace totals can look impressively low while a large chunk of the bill sits outside the report as unallocated capacity — fine for an efficiency drill-down, but weak showback.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you label workloads by team, you can roll the same data up at the ownership layer instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s1"&gt;'http://localhost:9003/allocation?window=7d&amp;amp;aggregate=label:team&amp;amp;shareIdle=true'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Allocation API accepts label-based aggregation, so you are not forced into namespace as the only billing dimension. That is useful when a single team owns several namespaces, or when shared namespaces are unavoidable and labels are the cleaner boundary.&lt;/p&gt;

&lt;p&gt;This is the point where OpenCost usually becomes operationally useful. You stop arguing about the cluster bill as one opaque number and start asking better questions: which teams carry the most idle share, which namespaces reserve far more than they use, and which platform defaults are making the bill harder to attribute than it should be.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prometheus is still useful, but use it for the right layer
&lt;/h2&gt;

&lt;p&gt;OpenCost also exposes metrics for Prometheus, which is what you want for dashboards and trend lines.&lt;/p&gt;

&lt;p&gt;For a simple monthly node-cost view, the documentation's example is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sum(node_total_hourly_cost) * 730
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is deliberately blunt. It tells you what the currently provisioned nodes cost over a rough 730-hour month. It is good for keeping an eye on baseline cluster spend, but it is not showback yet.&lt;/p&gt;

&lt;p&gt;For a namespace-level dashboard, you can build a query from OpenCost's allocation and node-cost metrics, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sum by (namespace) (
  container_cpu_allocation * on (node) group_left node_cpu_hourly_cost +
  container_memory_allocation_bytes * on (node) group_left node_ram_hourly_cost / (1024 * 1024 * 1024)
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is closer to what most teams expect to see in Grafana: cost broken down by namespace, derived from allocation metrics multiplied by node pricing. It is useful for trends, deltas, and “what changed this week?” views.&lt;/p&gt;

&lt;p&gt;The caveat is which layer you treat as authoritative. When you need a number to hand to finance or to another engineering team, start from the Allocation API. When you need a graph that shows whether cost by namespace is climbing, Prometheus is the right tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honesty box: what the numbers are, and what they are not
&lt;/h2&gt;

&lt;p&gt;Two assumptions trip people up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;That the allocation metrics are pure usage metrics.&lt;/strong&gt; They aren't. In the current code (v1.121.0) the synthetic &lt;code&gt;container_cpu_allocation&lt;/code&gt; and &lt;code&gt;container_memory_allocation_bytes&lt;/code&gt; metrics are built from &lt;code&gt;max(request, usage)&lt;/code&gt;, matching the specification — which is exactly why they beat a plain CPU-usage graph for cost allocation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;That pricing is magically exact out of the box.&lt;/strong&gt; OpenCost's own API docs describe the standard reporting path as on-demand list pricing, with cloud-provider cost-and-usage data available through integrations. The allocation model is useful immediately, but to reflect your committed discounts, private pricing, or billing exports, you have to wire that data in.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Neither of these is a flaw. They are just the boundaries of what the tool can do for you automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Showback first, chargeback if your organisation actually wants it
&lt;/h2&gt;

&lt;p&gt;FinOps discussions often make chargeback sound like the inevitable end state. In practice, most teams should start with showback.&lt;/p&gt;

&lt;p&gt;Showback means the spend is visible to the people who influence it. The namespace or team sees a monthly figure, understands what drove it, and can act on it. Chargeback adds the accounting decision on top: that cost is posted back to a budget, cost centre, or P&amp;amp;L.&lt;/p&gt;

&lt;p&gt;The FinOps Foundation's allocation guidance frames allocation as the foundation for either route. That is the useful way to think about OpenCost as well. It gives you a defensible allocation model and enough reporting surface to support showback immediately. Whether you turn that into formal chargeback is an organisational decision, not a Kubernetes feature flag.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical takeaway
&lt;/h2&gt;

&lt;p&gt;If your current Kubernetes cost story is “the cluster costs £X and nobody can explain why”, OpenCost is a solid place to start. Install it against the Prometheus stack you already run, use the Allocation API for the numbers you want people to trust, include idle cost when you report by namespace or team, and use Prometheus for dashboards rather than trying to rebuild the cost model yourself. That gets you from one undifferentiated cluster bill to a showback report people can actually argue with usefully, which is a much better place to be than arguing with the invoice alone.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>infrastructure</category>
      <category>kubernetes</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Nobody warns you about year two of platform engineering</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Thu, 30 Jul 2026 08:43:31 +0000</pubDate>
      <link>https://dev.to/bwlkr/nobody-warns-you-about-year-two-of-platform-engineering-fjo</link>
      <guid>https://dev.to/bwlkr/nobody-warns-you-about-year-two-of-platform-engineering-fjo</guid>
      <description>&lt;p&gt;Year one of a platform gets a budget, a roadmap, and a launch deck. Year two gets the support queue, the upgrade backlog, and the first awkward realisation that the golden path you shipped six months ago is already drifting away from the estate it was meant to simplify. &lt;strong&gt;That is the part nobody budgets for properly: a platform is a product, so shipping it once is the same mistake as shipping any product once.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is also why so many internal platforms feel promising in the demo and frustrating a year later. The visible work in year one is building: templates, guardrails, self-service workflows, maybe a portal, maybe a paved road people can actually use. The invisible work in year two is keeping the thing alive while the world underneath it keeps moving: Kubernetes releases, Helm chart churn, identity changes, cloud policy updates, framework upgrades, and the steady trickle of support requests from teams whose use case never quite fit the happy path.&lt;/p&gt;

&lt;p&gt;Platform engineering does not usually fail because the idea was wrong. It stalls because teams fund the build and underfund the maintenance.&lt;/p&gt;

&lt;h2&gt;
  
  
  The launch was the easy part
&lt;/h2&gt;

&lt;p&gt;There is a familiar shape to the first year of a platform effort.&lt;/p&gt;

&lt;p&gt;A team gets a mandate to improve developer experience, standardise delivery, or reduce ticket-driven toil. It has executive attention. It has a date to hit. It has a story people can understand: &lt;em&gt;we are building the internal platform&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;That phase is real work, but it has one advantage: the scope looks finite. You can imagine a first version. You can plan the launch. You can decide what the thinnest viable platform looks like and cut everything else.&lt;/p&gt;

&lt;p&gt;Year two is harder because the work stops looking like a project and starts looking like ownership.&lt;/p&gt;

&lt;p&gt;Now you are carrying things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;patching and upgrading the platform's own dependencies&lt;/li&gt;
&lt;li&gt;repaving templates and starter repos as upstream defaults change&lt;/li&gt;
&lt;li&gt;deciding which one-off exceptions should become product features&lt;/li&gt;
&lt;li&gt;pruning workflows nobody uses before they become permanent support burdens&lt;/li&gt;
&lt;li&gt;proving that adoption is still growing, rather than assuming it is&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that produces the same tidy before-and-after slide as the original launch. It is still the work that determines whether the platform becomes part of how the organisation ships software or just another ambitious internal system people route around.&lt;/p&gt;

&lt;p&gt;That is why &lt;a href="https://teamtopologies.com/platform-engineering" rel="noopener noreferrer"&gt;Team Topologies&lt;/a&gt; keeps insisting on &lt;strong&gt;platform as a product&lt;/strong&gt; rather than platform as a shared tools pile. The point is not that platforms need prettier language. The point is that products have ongoing discovery, maintenance, and evolution. A platform that is not continuously evolved is just a frozen opinion about how engineering used to work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Most of the field is only just getting to year two
&lt;/h2&gt;

&lt;p&gt;One reason there is so little useful day-two writing is that the discipline itself is still young.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://humanitec.com/whitepapers/state-of-platform-engineering-report-volume-3" rel="noopener noreferrer"&gt;2024 State of Platform Engineering report&lt;/a&gt; from Humanitec and Gitpod, based on a vendor-sponsored survey of platform teams, found that &lt;strong&gt;56% of respondent teams were less than two years old&lt;/strong&gt;, and only 13% had been doing platform engineering for more than five years. That is enough to explain why there is still so little hard-won writing about the maintenance phase: a lot of the field is only just arriving there.&lt;/p&gt;

&lt;p&gt;In other words, there is not a deep institutional playbook here. There are patterns. There are good principles. There is some solid research. But there is also a lot of collective improvisation.&lt;/p&gt;

&lt;p&gt;That matters because year-two problems do not look like year-one problems.&lt;/p&gt;

&lt;p&gt;The first year asks whether you can build a paved road. The second asks whether you can keep repaving it while traffic is already on it.&lt;/p&gt;

&lt;p&gt;The first year asks whether you can make self-service possible. The second asks whether the self-service path still fits the organisation as it changes.&lt;/p&gt;

&lt;p&gt;The first year asks whether the platform team can ship. The second asks whether it can avoid becoming the bottleneck it was created to remove.&lt;/p&gt;

&lt;h2&gt;
  
  
  Golden paths do not stay golden by themselves
&lt;/h2&gt;

&lt;p&gt;A platform's main promise is usually straightforward: give teams a safer, faster default path than stitching everything together from scratch.&lt;/p&gt;

&lt;p&gt;That is a good promise. It is also one with a permanent maintenance bill attached.&lt;/p&gt;

&lt;p&gt;The paved road starts decaying the moment upstream moves. A base chart changes shape. A cluster policy tightens. A framework introduces a new recommended default. A cloud provider deprecates some corner of the identity model you buried in a template months ago. A service team needs one capability the platform path did not anticipate, so it forks the template "just for now" and never really comes back.&lt;/p&gt;

&lt;p&gt;This is where platform credibility leaks away quietly. Not in one dramatic outage, but in a long series of tiny moments where engineers learn that the official path is slightly stale, slightly rigid, or slightly behind the reality of the stack. Once that happens often enough, people stop asking whether they should use the platform and start assuming they will need to work around it.&lt;/p&gt;

&lt;p&gt;That is also why the advice in &lt;a href="https://dev.to/blog/stop-copying-big-tech-platform-architecture/"&gt;Stop copying big-tech platform architecture: right-size to your scale&lt;/a&gt; matters even more in year two than in year one. Over-built platform machinery does not just cost more to launch. It costs more to keep current. Every extra control plane, abstraction layer, and bespoke workflow becomes another thing that needs tending after the excitement has worn off.&lt;/p&gt;

&lt;p&gt;A golden path without a gardener is just a future migration project.&lt;/p&gt;

&lt;h2&gt;
  
  
  The upgrade treadmill is not a side quest
&lt;/h2&gt;

&lt;p&gt;If you want one brutally practical example of year-two reality, use Kubernetes itself.&lt;/p&gt;

&lt;p&gt;Kubernetes now ships &lt;strong&gt;three minor releases a year&lt;/strong&gt;, and the project maintains patch support for only the most recent three minor versions. Under the &lt;a href="https://kubernetes.io/releases/patch-releases/" rel="noopener noreferrer"&gt;yearly support policy&lt;/a&gt; that works out to roughly a 14-month window per release: twelve months of standard patches, then two months of maintenance mode before end of life. If your platform sits on Kubernetes, the upgrade tax is not occasional. It is structural. You are budgeting for &lt;strong&gt;two to three cluster upgrades a year, forever&lt;/strong&gt;, before you even count the other components attached to the platform.&lt;/p&gt;

&lt;p&gt;And of course it is never just Kubernetes.&lt;/p&gt;

&lt;p&gt;There is ingress or Gateway API, cert management, policy tooling, observability agents, CI runners, secret syncing, identity plumbing, image build tooling, and whatever home-grown wrappers your platform created around those things to make them usable. One upstream change can ripple through templates, documentation, tests, and support expectations all at once.&lt;/p&gt;

&lt;p&gt;Teams often talk about this as if it were operational overhead adjacent to the roadmap. It is not. For a platform team, &lt;strong&gt;the upgrade treadmill is part of the roadmap&lt;/strong&gt;. If you do not plan for it explicitly, it eats the roadmap anyway.&lt;/p&gt;

&lt;p&gt;This is the part that breaks year-two planning. A team thinks it has funded one staff engineer to "keep the lights on" while everyone else moves on to the next strategic thing. Then the calendar turns, the supported version window narrows, and suddenly the strategic thing is an upgrade programme whether anyone likes it or not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Adoption stalls when nobody is listening
&lt;/h2&gt;

&lt;p&gt;There is a second year-two trap that looks softer but is just as destructive: the team assumes that shipping the platform was the same thing as winning adoption.&lt;/p&gt;

&lt;p&gt;It is not.&lt;/p&gt;

&lt;p&gt;DORA's &lt;a href="https://dora.dev/research/2024/dora-report/" rel="noopener noreferrer"&gt;2024 Accelerate State of DevOps report&lt;/a&gt; is useful here because it is more honest than most of the hype. It found that internal platforms were associated with better individual productivity and team performance, but also with weaker throughput and change stability where the implementation added complexity or reduced developer independence. DORA describes this as a &lt;strong&gt;J-curve&lt;/strong&gt;: you can see early gains, then hit a dip, and only recover if the platform matures in a &lt;a href="https://dora.dev/capabilities/platform-engineering/" rel="noopener noreferrer"&gt;user-centred way&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That finding lands because it matches what many teams feel in practice. A platform can absolutely remove toil. It can also create a new layer of process, abstraction, and waiting if the product discipline is weak.&lt;/p&gt;

&lt;p&gt;The warning signs are usually familiar:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;teams need tickets for routine changes the platform was meant to self-serve&lt;/li&gt;
&lt;li&gt;exceptions pile up because the happy path is too narrow&lt;/li&gt;
&lt;li&gt;the platform roadmap is driven by internal architecture debates rather than user pain&lt;/li&gt;
&lt;li&gt;the team measures launches and features, but not adoption or abandonment&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why Team Topologies talks about &lt;strong&gt;sensing&lt;/strong&gt; user needs rather than just building services for them. The &lt;a href="https://teamtopologies.com/key-concepts-content/what-is-a-thinnest-viable-platform-tvp" rel="noopener noreferrer"&gt;thinnest viable platform&lt;/a&gt; is only a starting point. If you stop there, the platform freezes while the organisation keeps changing around it.&lt;/p&gt;

&lt;p&gt;Puppet's &lt;a href="https://www.puppet.com/blog/state-devops-report-2024" rel="noopener noreferrer"&gt;2024 State of DevOps report&lt;/a&gt; makes the same point from a different angle: &lt;strong&gt;52% of respondents said a product manager was crucial to a platform team's success&lt;/strong&gt;, with a further 21% calling one nice to have. That is not a plea to add ceremony. It is a recognition that somebody needs to own adoption, prioritisation, and the discipline of learning what the users of the platform actually need next.&lt;/p&gt;

&lt;p&gt;Without that function, year two becomes a familiar kind of drift. The platform team keeps shipping things. The users keep needing different things. Everyone is technically busy, and the gap still widens.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is how the platform team becomes the bottleneck again
&lt;/h2&gt;

&lt;p&gt;Platform engineering is often sold as the way out of ticket Ops. Build the paved road, standardise the sharp edges, push routine work into self-service, and let product teams move without waiting for a central platform group.&lt;/p&gt;

&lt;p&gt;That is the right goal. But an under-staffed year-two platform team can relapse into the exact behaviour it was supposed to replace.&lt;/p&gt;

&lt;p&gt;You see it when every unusual request becomes a handoff. You see it when upgrades consume so much attention that user-facing platform work slips for a quarter. You see it when the team starts protecting itself with queues and process because it no longer has the slack to evolve the underlying system.&lt;/p&gt;

&lt;p&gt;At that point the platform still exists, but the interaction model has changed. The platform is no longer a capability. It is a department.&lt;/p&gt;

&lt;p&gt;And yes, that sounds harsh. It is still a useful test. If the main experience of the platform for most engineers is waiting for another team, something important has gone wrong even if the underlying architecture is perfectly respectable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What year-two funding actually looks like
&lt;/h2&gt;

&lt;p&gt;If the first-year mistake is treating the platform like a one-off delivery project, the year-two fix is not mysterious. It is just less glamorous than the original launch.&lt;/p&gt;

&lt;p&gt;A serious year-two platform budget needs to fund four things explicitly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;maintenance as first-class work&lt;/strong&gt;: upgrades, repaving, dependency churn, template hygiene, and pruning stale paths&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;product management&lt;/strong&gt;: somebody accountable for adoption, prioritisation, and the sensing loop with platform users&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;operational slack&lt;/strong&gt;: enough capacity that support and upgrades do not consume every sprint&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;evidence&lt;/strong&gt;: metrics for adoption, developer independence, and the support burden the platform is actually removing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point matters because year-two platform conversations often become strangely qualitative. People can tell that the platform feels helpful or painful, but the organisation stops short of measuring the signs that matter.&lt;/p&gt;

&lt;p&gt;A better year-two review asks questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;how many teams are still on the paved path versus a forked one?&lt;/li&gt;
&lt;li&gt;which support requests repeat often enough to deserve product work?&lt;/li&gt;
&lt;li&gt;how old are the core templates and platform dependencies?&lt;/li&gt;
&lt;li&gt;where does a platform user still need a human instead of self-service?&lt;/li&gt;
&lt;li&gt;did the last upgrade consume planned capacity, or did it blow up the quarter?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are product questions as much as technical ones. Which is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real work starts after the applause
&lt;/h2&gt;

&lt;p&gt;Platform engineering year two is where the discipline becomes real. The launch matters, but the launch is only the opening move. After that comes the unglamorous part: keeping the path current, paying the upgrade tax, listening hard enough to notice when adoption is flattening, and funding the platform team as owners of a product rather than custodians of a project.&lt;/p&gt;

&lt;p&gt;If you get that part right, the J-curve recovers and the platform becomes genuinely compounding infrastructure for the organisation. If you get it wrong, the platform slowly turns back into the queue it was meant to remove.&lt;/p&gt;

&lt;p&gt;Nobody warns you about year two because year one is easier to tell as a story. Year two is mostly maintenance, judgement, and sustained ownership. It is also the part that decides whether the platform was ever really a platform at all.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Observability in the age of LLMs: tracing what your AI actually did</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Tue, 28 Jul 2026 08:24:40 +0000</pubDate>
      <link>https://dev.to/bwlkr/observability-in-the-age-of-llms-tracing-what-your-ai-actually-did-42j9</link>
      <guid>https://dev.to/bwlkr/observability-in-the-age-of-llms-tracing-what-your-ai-actually-did-42j9</guid>
      <description>&lt;p&gt;A feature built on an LLM can hit every familiar SLO and still be failing in the way that matters. The API is up. The median latency looks fine. Error rate is flat. Meanwhile the model is taking the scenic route through three tool calls, burning tokens, and returning answers you would not want a customer to trust. The thing you need to see is not just whether the service is healthy. It is what the system actually did.&lt;/p&gt;

&lt;p&gt;That is why &lt;strong&gt;LLM observability&lt;/strong&gt; is best thought of as an extension of ordinary tracing, not a fresh religion with its own priesthood. An LLM call is a span. An agent run is a span tree. The useful questions are still familiar: where did the time go, which step failed, what changed, and how much did this request cost? What changes is the shape of the work. Now you also care about token usage, model choices, tool execution, and whether the answer was any good.&lt;/p&gt;

&lt;p&gt;If you already run OpenTelemetry, that is good news. You do not need to throw away your tracing model to observe AI features properly. You need to extend it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why classic observability stops short
&lt;/h2&gt;

&lt;p&gt;Traditional service telemetry assumes that the interesting failures are mostly about infrastructure or request handling: timeouts, retries, saturation, bad downstreams, deployment regressions. Those problems still exist in LLM-backed systems, but they are no longer the whole story.&lt;/p&gt;

&lt;p&gt;A generated answer can be wrong while every system metric stays green. A retrieval step can bring back unhelpful context without producing an obvious exception. An agent can choose a slow tool path, or call the same tool twice, or hand the model a bloated prompt that quietly doubles token spend. Keep the old graphs; just stop expecting them to catch this class of failure.&lt;/p&gt;

&lt;p&gt;OpenTelemetry made the same point in its &lt;a href="https://opentelemetry.io/blog/2026/genai-observability/" rel="noopener noreferrer"&gt;2026 GenAI observability post&lt;/a&gt;: every application call to an LLM hides "a chain of model calls, tool invocations, and token exchanges" behind the scenes. That is the part standard web-service monitoring does not naturally reveal. When a request takes 45 seconds, the question is no longer just "which service was slow?" but "which decision inside the trace made it slow?"&lt;/p&gt;

&lt;p&gt;This is the same reason &lt;a href="https://dev.to/blog/why-feedback-loops-matter-more-when-llms-write-your-code/"&gt;fast feedback loops matter more when LLMs write your code&lt;/a&gt;. When generation gets cheaper, validation becomes the scarce resource. In production, observability is one of those validation loops.&lt;/p&gt;

&lt;h2&gt;
  
  
  The debugging primitive is still the trace
&lt;/h2&gt;

&lt;p&gt;There is a mild industry temptation to treat LLM systems as too exotic for the tooling we already have. That is usually the wrong instinct.&lt;/p&gt;

&lt;p&gt;A useful mental model looks more like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;agent.run
├── llm.chat openai/gpt-5.6
├── tool.retrieval postgres/pgvector
├── llm.chat openai/gpt-5.6
├── tool.http billing-api
└── llm.chat openai/gpt-5.6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That tree is already enough to answer a lot of awkward production questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which step dominated latency&lt;/li&gt;
&lt;li&gt;whether the model asked for tools more than once&lt;/li&gt;
&lt;li&gt;whether retrieval happened before the answer changed course&lt;/li&gt;
&lt;li&gt;whether one provider or model variant behaved differently&lt;/li&gt;
&lt;li&gt;where token usage and therefore spend started to climb&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that needs a brand-new theory of observability. It just means treating model calls and tool calls as first-class trace events instead of leaving them hidden inside application logs or framework internals.&lt;/p&gt;

&lt;p&gt;That framing matters because it keeps you portable. If the trace is the source of truth, you can ship it to a general-purpose backend, to an AI-focused product, or to both. If the observability story starts and ends inside one vendor's SDK, you have made your debugging model harder to move.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the OpenTelemetry GenAI conventions give you
&lt;/h2&gt;

&lt;p&gt;The concrete standard here is the OpenTelemetry semantic conventions for Generative AI. As of July 2026 the &lt;code&gt;gen_ai&lt;/code&gt; work has moved into its own dedicated repository, &lt;a href="https://github.com/open-telemetry/semantic-conventions-genai" rel="noopener noreferrer"&gt;semantic-conventions-genai&lt;/a&gt;, and every document in it is still marked &lt;strong&gt;Development&lt;/strong&gt; rather than stable. The new repository does not even have a tagged release yet, so there is no version number to pin against; the honest options are pinning a commit, or recording the last main-repo release that carried the &lt;code&gt;gen_ai&lt;/code&gt; conventions before the move. That does not make them useless. It means you should adopt them with your eyes open and expect some movement.&lt;/p&gt;

&lt;p&gt;You do not need to memorise every field. The useful thing is the shape of the telemetry the conventions are trying to standardise.&lt;/p&gt;

&lt;p&gt;For spans and attributes, the useful core includes fields such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;gen_ai.operation.name&lt;/code&gt; for the kind of model operation&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gen_ai.provider.name&lt;/code&gt; for the provider&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gen_ai.request.model&lt;/code&gt; and &lt;code&gt;gen_ai.response.model&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gen_ai.usage.input_tokens&lt;/code&gt; and &lt;code&gt;gen_ai.usage.output_tokens&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gen_ai.response.finish_reasons&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gen_ai.tool.name&lt;/code&gt; for tool execution&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gen_ai.agent.name&lt;/code&gt; and &lt;code&gt;gen_ai.conversation.id&lt;/code&gt; when you need to group larger workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For metrics, the current GenAI docs define client-side measurements such as &lt;code&gt;gen_ai.client.token.usage&lt;/code&gt; and &lt;code&gt;gen_ai.client.operation.duration&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is enough to capture the parts of an LLM request you usually end up caring about first: how many tokens went in, how many came out, how long the operation took, which model handled it, and where tool execution sat in the middle.&lt;/p&gt;

&lt;p&gt;It is also worth being precise about one thing the conventions do &lt;strong&gt;not&lt;/strong&gt; settle for you: cost. Cost is still best treated as a derived metric rather than a magical standard field. In practice that means joining token counts to your model pricing table and calculating spend from there. If someone promises a universal &lt;code&gt;LLM cost&lt;/code&gt; number with no context, raise an eyebrow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument these first
&lt;/h2&gt;

&lt;p&gt;You can spend a lot of time arguing about evals, guardrails, and which agent framework is fashionable this month. Instrumentation is less glamorous, but it is the bit that keeps the rest honest.&lt;/p&gt;

&lt;p&gt;If you are tracing LLM-backed features today, start with these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One span per model call.&lt;/strong&gt; Capture provider, model, operation, latency, finish reason, and token counts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One span per tool call.&lt;/strong&gt; Retrieval, HTTP calls, database lookups, and internal helpers should sit in the same trace, not in a separate logging universe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt and response references, not reckless raw dumps.&lt;/strong&gt; Keep enough data to debug, but be deliberate about privacy, redaction, and storage costs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A derived cost metric.&lt;/strong&gt; Token counts are useful; token counts joined to pricing are useful to finance and engineering at the same time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality signals beside the trace.&lt;/strong&gt; User rating, evaluator score, moderation outcome, or task success should sit close enough to the trace that you can compare bad answers with the path that produced them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Getting the first span is cheaper than most teams expect. If your stack is Python calling OpenAI, the contrib instrumentation does it with zero code changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;opentelemetry-distro opentelemetry-instrumentation-openai-v2
&lt;span class="nv"&gt;OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true &lt;/span&gt;opentelemetry-instrument python app.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the environment variable. Prompt and completion content is off by default and opt-in, which is exactly the right default: you get spans, models, token counts, and finish reasons without capturing user text, and you turn on content capture deliberately, with redaction and retention already thought through.&lt;/p&gt;

&lt;p&gt;That is the practical core. You can add fancier workflow views later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Experimental is fine. Invisible is not.
&lt;/h2&gt;

&lt;p&gt;Some teams hesitate because the GenAI conventions are still moving. That is a real caveat, but it is not a good reason to keep model behaviour opaque.&lt;/p&gt;

&lt;p&gt;The honest trade-off is this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;if you wait for the conventions to feel perfectly settled, you stay blind longer&lt;/li&gt;
&lt;li&gt;if you adopt them now, you may need a migration later&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would take the migration.&lt;/p&gt;

&lt;p&gt;A moving standard is still much better than bespoke telemetry that only one SDK understands. The conventions already give you a common vocabulary for spans, token usage, and model metadata. Even if a field changes later, the shape of the problem is now visible in your traces instead of trapped inside ad hoc logs.&lt;/p&gt;

&lt;p&gt;The practical discipline is boring, which is usually how you know it is right:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;record the commit or snapshot of the conventions you instrumented against, since there is no tagged release to pin yet&lt;/li&gt;
&lt;li&gt;note that it is a Development-status schema&lt;/li&gt;
&lt;li&gt;keep your instrumentation behind a thin abstraction where you can&lt;/li&gt;
&lt;li&gt;re-check the docs before each major upgrade&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is manageable. Debugging a production agent system with no trace of its internal choices is much less manageable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI observability tools are mostly OpenTelemetry underneath
&lt;/h2&gt;

&lt;p&gt;Once you start looking around, the AI observability market can seem louder than the underlying technical differences justify.&lt;/p&gt;

&lt;p&gt;There are good tools here. Langfuse's own docs position its OpenTelemetry integration around ingesting OTLP traces for LLM observability. MLflow describes its tracing product as OpenTelemetry-compatible and aimed at LLM and agent observability. OpenInference, from the Arize ecosystem, describes itself more plainly still: OpenTelemetry instrumentation for AI observability.&lt;/p&gt;

&lt;p&gt;That is the important pattern.&lt;/p&gt;

&lt;p&gt;The credible tools are increasingly building on OpenTelemetry rather than asking you to abandon it. What they compete on is everything around the wire format:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;workflow visualisation&lt;/li&gt;
&lt;li&gt;evaluator and feedback loops&lt;/li&gt;
&lt;li&gt;prompt inspection&lt;/li&gt;
&lt;li&gt;dataset and experiment tooling&lt;/li&gt;
&lt;li&gt;agent-specific views&lt;/li&gt;
&lt;li&gt;hosted versus self-managed trade-offs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are meaningful differences. They just are not a reason to skip portable traces.&lt;/p&gt;

&lt;p&gt;If you already have an observability stack you trust, start there. Send the traces somewhere you can query them. Add a specialised tool when you have a concrete gap: better prompt inspection, easier annotation workflows, richer experiment tracking, or an agent timeline that your current backend does not give you.&lt;/p&gt;

&lt;p&gt;What I would not do is buy an AI observability platform before I had first proved that we were actually capturing the right spans.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent observability is where this gets interesting
&lt;/h2&gt;

&lt;p&gt;Plain LLM tracing is useful. Agent tracing is where the debugging payoff gets larger.&lt;/p&gt;

&lt;p&gt;The difficult production questions are rarely about one completion in isolation. They are about execution paths:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;why did this agent choose that tool&lt;/li&gt;
&lt;li&gt;why did it call the same tool twice&lt;/li&gt;
&lt;li&gt;why did it stop early&lt;/li&gt;
&lt;li&gt;why did one branch of the workflow take ten times longer than the others&lt;/li&gt;
&lt;li&gt;which step turned a reasonable user request into an expensive one&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is why the recent push toward agent timelines is worth paying attention to. Honeycomb launched its Agent Observability features on 12 May 2026, with an Agent Timeline view that is basically a more opinionated way to walk the same trace tree. That kind of view earns its keep because agent systems are not just slow or fast. They are branching, recursive, and occasionally a bit absurd.&lt;/p&gt;

&lt;p&gt;The frontier here is less about inventing new telemetry categories than about making traces easier for humans to read. Multi-agent systems create deeper trees, more spans, and more "why on earth did it do that?" moments. The backend that wins will be the one that makes that path legible under pressure.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do this week
&lt;/h2&gt;

&lt;p&gt;If you are shipping LLM-backed features and do not yet have a plan, the sensible order of work is refreshingly unromantic:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;instrument model calls as spans&lt;/li&gt;
&lt;li&gt;instrument tool calls in the same trace&lt;/li&gt;
&lt;li&gt;capture token counts and duration&lt;/li&gt;
&lt;li&gt;derive cost from usage and pricing&lt;/li&gt;
&lt;li&gt;attach one quality signal you trust&lt;/li&gt;
&lt;li&gt;only then decide whether you need a specialised platform&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That sequence gets you out of the worst failure mode, which is running a system that can think, call tools, and spend money without leaving behind an explanation.&lt;/p&gt;

&lt;p&gt;LLM observability is not separate from good observability practice. It is what happens when the service starts making choices on your behalf. Start with traces. Make token usage and cost visible. Keep the data portable. Then, if you need a shinier lens on top, add one later.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>monitoring</category>
      <category>performance</category>
    </item>
    <item>
      <title>Stop copying big-tech platform architecture: right-size to your scale</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Sat, 11 Jul 2026 07:57:12 +0000</pubDate>
      <link>https://dev.to/bwlkr/stop-copying-big-tech-platform-architecture-right-size-to-your-scale-3ejg</link>
      <guid>https://dev.to/bwlkr/stop-copying-big-tech-platform-architecture-right-size-to-your-scale-3ejg</guid>
      <description>&lt;p&gt;A lot of platform over-engineering starts with a diagram that looks serious. A small team has a few services, a modest growth curve, and perfectly ordinary operational needs, but the design doc still reaches for the full hyperscaler silhouette: Kubernetes everywhere, a service mesh, Kafka, multi-cluster, and a bespoke internal platform before anyone has proved they need one. The problem is not that those tools are bad. It is that they are answers to constraints most teams do not have. &lt;strong&gt;Right-sizing platform architecture means copying the principles behind big-tech systems, not the stack itself.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That becomes expensive quickly because complexity is not free. Every extra control plane becomes another upgrade path, another source of pages, another thing a new engineer has to understand before they can safely ship a change. Big companies take on that cost because, at their scale, it often pays for itself. Small and mid-sized teams usually take it on up front and hope the benefits show up later.&lt;/p&gt;

&lt;p&gt;They often do not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Big-tech architecture is the output of big-tech constraints
&lt;/h2&gt;

&lt;p&gt;It helps to say the quiet part out loud. Google, Amazon, and the rest did not end up with sprawling platform stacks because they enjoy making life difficult. They built those systems under pressures most teams will never see all at once:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;traffic and data volumes that make single-node or single-database assumptions collapse&lt;/li&gt;
&lt;li&gt;hundreds of teams that need to change software independently&lt;/li&gt;
&lt;li&gt;regulatory or blast-radius boundaries that force hard isolation&lt;/li&gt;
&lt;li&gt;dedicated platform and reliability teams to operate the machinery&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your team has three product engineers and one person who gets the on-call phone, you are operating under a different set of forces. Copying the architecture without copying the forces is how you end up with enterprise-grade overhead and startup-grade staffing.&lt;/p&gt;

&lt;p&gt;That is also the practical reading of &lt;a href="https://www.melconway.com/Home/Conways_Law.html" rel="noopener noreferrer"&gt;Conway’s Law&lt;/a&gt;. System design follows communication structure. A microservice estate with separate deployment pipelines, cross-cutting policy layers, and multiple control planes assumes an organisation shaped to own those boundaries. If you do not have that organisation, the elegant diagram turns into one team carrying all of the complexity by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  Copy the principles, not the implementation
&lt;/h2&gt;

&lt;p&gt;The useful parts of big-tech engineering are usually the ideas underneath the stack choice.&lt;/p&gt;

&lt;p&gt;You can borrow those almost anywhere:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;reduce cognitive load&lt;/strong&gt; so engineers can ship without learning every platform detail&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;improve feedback loops&lt;/strong&gt; so changes are cheap to validate and rollback&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;create clear boundaries&lt;/strong&gt; between services, teams, and operational ownership&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;standardise the paved road&lt;/strong&gt; so the safe path is also the easy path&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;treat cost as a design constraint&lt;/strong&gt;, not something to apologise for later&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What does not travel as cleanly is the implementation detail. Service mesh, multi-cluster, event streaming everywhere, and a bespoke internal developer platform are all legitimate responses to particular thresholds. They are poor defaults.&lt;/p&gt;

&lt;p&gt;This is where Dan McKinley’s &lt;a href="https://mcfunley.com/choose-boring-technology" rel="noopener noreferrer"&gt;Choose Boring Technology&lt;/a&gt; still lands so cleanly. His line about companies having roughly &lt;strong&gt;three innovation tokens&lt;/strong&gt; has aged well because the underlying point is still right: novelty is expensive, and you should spend it where it creates advantage rather than where it merely creates architecture. “Boring” in McKinley’s sense is not old or unfashionable. It is technology whose failure modes are already well understood.&lt;/p&gt;

&lt;p&gt;For platform teams, that is a strong default. If the thing you are building is not your differentiator, buying more operational complexity than you need is rarely a flex. It is a loan.&lt;/p&gt;

&lt;h2&gt;
  
  
  “You are not Google” is a useful design rule, not a slogan
&lt;/h2&gt;

&lt;p&gt;Oz Nova’s &lt;a href="https://blog.bradfieldcs.com/you-are-not-google-84912cf44afb" rel="noopener noreferrer"&gt;You Are Not Google&lt;/a&gt; remains one of the clearest antidotes to architecture cargo culting because it forces a more honest sequence of questions. What problem are you actually solving? Which candidate solutions exist? What historical pressure produced the famous answer everyone keeps citing? Does that answer fit your team, budget, and operational maturity now?&lt;/p&gt;

&lt;p&gt;Those questions are annoyingly effective because they make a lot of platform theatre disappear.&lt;/p&gt;

&lt;p&gt;A Kubernetes platform is a good example. Plenty of teams do need Kubernetes. Plenty also adopt it because it feels like the proper destination for anything serious. The real question is not “is Kubernetes modern?” It is “what operational problem does this solve for us today that simpler deployment machinery does not?” If the answer is mostly social proof, that is a warning light.&lt;/p&gt;

&lt;p&gt;The same applies higher up the stack. A bespoke internal platform can be the right move when many teams need a common workflow and the cost of hand-held support keeps rising. But Team Topologies’ &lt;a href="https://teamtopologies.com/key-concepts-content/what-is-a-thinnest-viable-platform-tvp" rel="noopener noreferrer"&gt;thinnest viable platform&lt;/a&gt; is a better starting point than an immediate platform product. Sometimes the right platform is a template, a few scripts, and good docs. It does not have to begin life as an ecosystem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Prime Video story is a correction, not a manifesto
&lt;/h2&gt;

&lt;p&gt;The most abused recent example in this space is Prime Video’s write-up on &lt;a href="https://www.primevideotech.com/video-infrastructure/scaling-up-the-prime-video-audio-video-monitoring-service-and-reducing-costs-by-90" rel="noopener noreferrer"&gt;reducing the cost of one monitoring service by more than 90%&lt;/a&gt;. The point of that post was not “Amazon abandoned microservices” and it certainly was not “monoliths always win”. It was narrower, and more useful.&lt;/p&gt;

&lt;p&gt;One team had built a distributed serverless design around Step Functions, Lambda, and S3. At the workload they were handling, the orchestration overhead had become the cost problem. They moved to a single process running on ECS and passed data in memory instead of via a chain of external services. The result was a &lt;strong&gt;cost reduction above 90%&lt;/strong&gt; for that tool.&lt;/p&gt;

&lt;p&gt;That matters because it shows what right-sizing actually looks like in practice. The original design was valid in the abstract. It was just mis-sized for the workload and economics in front of the team. Adrian Cockcroft’s corrective take on the reaction was useful here too: this was a refactoring decision inside one service, not a grand ideological retreat.&lt;/p&gt;

&lt;p&gt;The lesson is not “always build a monolith”. It is “distributed systems premiums only make sense when you are actually collecting the benefit”.&lt;/p&gt;

&lt;p&gt;Martin Fowler has been making the same point for years. &lt;a href="https://martinfowler.com/bliki/MonolithFirst.html" rel="noopener noreferrer"&gt;Monolith First&lt;/a&gt; is still one of the sanest defaults in software architecture because a modular monolith lets you prove boundaries before turning them into network calls. You can extract a service later. You cannot cheaply un-distribute a badly fragmented system that nobody really needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  A right-sizing table for platform teams
&lt;/h2&gt;

&lt;p&gt;If you want one decision aid from this post, make it this: treat big-tech patterns as threshold tools, not identity markers.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern people cargo-cult&lt;/th&gt;
&lt;th&gt;Right-sized default for most teams&lt;/th&gt;
&lt;th&gt;Real signal the bigger pattern may now pay off&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Service mesh&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Use ingress, &lt;code&gt;NetworkPolicy&lt;/code&gt;, and straightforward app-level telemetry first. Add mTLS or traffic shaping only when you can name the requirement.&lt;/td&gt;
&lt;td&gt;You need per-service identity, consistent mTLS, or progressive traffic control across many teams and many services.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Multi-cluster Kubernetes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Start with one well-run cluster and strong namespace, RBAC, quota, and environment boundaries.&lt;/td&gt;
&lt;td&gt;You have hard isolation needs: data residency, materially different blast radii, or control-plane scale limits.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Kafka for everything&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Start with a managed queue, an outbox pattern, or even Postgres when throughput and replay needs are modest.&lt;/td&gt;
&lt;td&gt;You genuinely need a durable, replayable event log with multiple independent consumers and sustained high throughput.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bespoke internal developer platform&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Start with a thinnest viable platform: templates, automation, docs, and one paved path that removes toil.&lt;/td&gt;
&lt;td&gt;Many teams are repeating the same workflow, support demand is compounding, and a platform product will clearly save more time than it costs to run.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Microservices from day one&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Start with a modular monolith and explicit internal boundaries.&lt;/td&gt;
&lt;td&gt;Teams need independent deployability, boundaries are already clear, and you have the operational discipline to support distributed systems well.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is also where the &lt;a href="https://dora.dev/" rel="noopener noreferrer"&gt;DORA research&lt;/a&gt; is often flattened into the wrong slogan. High-performing teams benefit from loose coupling and independent deployability. That does not automatically mean “split everything into microservices”. You can get plenty of that benefit from a well-structured monolith, a clean delivery pipeline, and sensible ownership boundaries.&lt;/p&gt;

&lt;p&gt;The label matters less than the change cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost, team shape, and on-call load should be first-class inputs
&lt;/h2&gt;

&lt;p&gt;Werner Vogels’ &lt;a href="https://thefrugalarchitect.com/" rel="noopener noreferrer"&gt;Frugal Architect&lt;/a&gt; material is useful beyond AWS because it frames architecture as a series of trade-offs rather than a hunt for the most sophisticated-looking answer. That sounds obvious, but teams forget it surprisingly easily when a design starts accumulating fashionable parts.&lt;/p&gt;

&lt;p&gt;A few questions cut through the theatre quickly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What does this add to the pager surface area?&lt;/li&gt;
&lt;li&gt;Who upgrades it, patches it, and debugs it at 02:00?&lt;/li&gt;
&lt;li&gt;What skill set does it require from every new engineer?&lt;/li&gt;
&lt;li&gt;What failure mode becomes easier because we adopted it?&lt;/li&gt;
&lt;li&gt;What failure mode becomes harder or more likely?&lt;/li&gt;
&lt;li&gt;Could a simpler design buy us another year before this complexity is justified?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If those questions make the design feel awkward, that is usually a useful signal rather than conservatism.&lt;/p&gt;

&lt;p&gt;This is especially true in platform engineering because platforms quietly multiply costs across the organisation. A hard-to-understand abstraction does not stay local to the team that built it. Every product engineer pays rent on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the bigger architecture is the right call
&lt;/h2&gt;

&lt;p&gt;None of this is an argument for permanent smallness or for avoiding ambitious systems. Some teams absolutely do cross the threshold where the heavier pattern is worth it.&lt;/p&gt;

&lt;p&gt;You are probably getting there when several of these are true at the same time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;multiple teams are blocked on each other because deployments cannot move independently&lt;/li&gt;
&lt;li&gt;reliability problems are coming from shared blast radius rather than application bugs&lt;/li&gt;
&lt;li&gt;platform support has turned into a queue because manual help does not scale&lt;/li&gt;
&lt;li&gt;data residency, compliance, or tenancy requirements demand hard technical boundaries&lt;/li&gt;
&lt;li&gt;workload volume has made the simpler option materially more expensive or less reliable&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At that point, more machinery can be the cheaper choice overall. But it is still better to adopt it deliberately than to inherit it by imitation.&lt;/p&gt;

&lt;p&gt;That is the part people miss when they turn “choose boring technology” into a bumper sticker. McKinley’s innovation tokens exist so you can spend them where they matter. The platform question is whether you are spending one on a real constraint or on architecture envy.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sophisticated choice is often the smaller one
&lt;/h2&gt;

&lt;p&gt;The cheapest architecture that meets your actual constraints is not the amateur option. Usually it is the disciplined one.&lt;/p&gt;

&lt;p&gt;Small teams do not fail because they lacked enough control planes. They usually fail because they made delivery too slow, operations too fragile, and ownership too muddy for their actual size. Big-tech architecture can solve those problems at big-tech scale. At smaller scale, it often creates them.&lt;/p&gt;

&lt;p&gt;So copy the habits that travel: clear interfaces, good defaults, fast feedback, cost awareness, and platforms that remove toil rather than advertise sophistication. Leave the rest until the numbers, the team shape, and the pager all agree that you have earned it.&lt;/p&gt;

</description>
      <category>architecture</category>
    </item>
    <item>
      <title>Progressive delivery on Kubernetes with Argo Rollouts</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Thu, 09 Jul 2026 19:05:06 +0000</pubDate>
      <link>https://dev.to/bwlkr/progressive-delivery-on-kubernetes-with-argo-rollouts-4l6a</link>
      <guid>https://dev.to/bwlkr/progressive-delivery-on-kubernetes-with-argo-rollouts-4l6a</guid>
      <description>&lt;p&gt;Most teams meet the limits of a plain Kubernetes &lt;code&gt;Deployment&lt;/code&gt; at the worst possible moment: mid-rollout, watching a new version march out to every pod while a graph slowly turns the wrong colour. A &lt;code&gt;Deployment&lt;/code&gt; rolls pods out carefully enough — it respects readiness, it surges and drains in order — but it has one stubborn blind spot. Once the new ReplicaSet is Ready, it keeps going unless something stops it. There is no built-in way to say “send 10% of traffic here, wait, check an SLO, then continue”, or “only promote this preview stack if the metrics hold up”.&lt;/p&gt;

&lt;p&gt;That missing control loop is exactly what Argo Rollouts adds. It keeps the &lt;code&gt;Deployment&lt;/code&gt; mental model you already know and layers canary, blue-green, analysis, and automated rollback on top. For platform teams, that quietly turns rollout safety from a runbook someone has to remember into a controller that does it the same way every time. As of Argo Rollouts v1.9.0, that controller can drive percentage-based canaries, blue-green cutovers, and metric-gated promotion without asking you to rip out the rest of your Kubernetes delivery stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap in a normal Deployment
&lt;/h2&gt;

&lt;p&gt;A &lt;code&gt;Deployment&lt;/code&gt; is genuinely good at one thing: swapping one ReplicaSet for another while keeping enough healthy pods around to serve traffic. Readiness probes, &lt;code&gt;maxUnavailable&lt;/code&gt;, and &lt;code&gt;maxSurge&lt;/code&gt; all pull their weight here. What none of them give you is a first-class sense of whether the new version is actually any good.&lt;/p&gt;

&lt;p&gt;That distinction matters most when “pods are Ready” and “this version is safe” quietly stop meaning the same thing. A service can start up cleanly and still come apart under real load. Error rate can sit flat through the first few requests and then climb once traffic arrives in earnest. Latency can drift at the 95th percentile while the average stays there looking reassuring. And a bad release often hurts just one path through the app — which is cold comfort when a straight rollout has already reached every user before anyone has the evidence to hit stop.&lt;/p&gt;

&lt;p&gt;Argo Rollouts lives precisely in that gap. It is not a replacement for the Kubernetes scheduler, and it is not a GitOps controller — it does not want either job. It is a rollout controller that takes over the update strategy for a single workload and adds the parts a &lt;code&gt;Deployment&lt;/code&gt; never had: promotion steps, pause points, analysis runs, and traffic-shaping integrations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model: Deployment behaviour, but with a promotion engine
&lt;/h2&gt;

&lt;p&gt;The easiest way to hold Argo Rollouts in your head is “a &lt;code&gt;Deployment&lt;/code&gt; with a state machine bolted on”. The &lt;code&gt;Rollout&lt;/code&gt; resource still owns ReplicaSets and still reacts to pod template changes, but on top of that it keeps track of where a release currently sits in its promotion flow.&lt;/p&gt;

&lt;p&gt;At a high level, the loop runs like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A new pod template lands in Git and is applied to the cluster.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;Rollout&lt;/code&gt; controller creates a new ReplicaSet.&lt;/li&gt;
&lt;li&gt;The strategy decides what happens next: shift some traffic, create a preview stack, pause, or run analysis.&lt;/li&gt;
&lt;li&gt;If the checks pass, the controller promotes the release.&lt;/li&gt;
&lt;li&gt;If the checks fail, the controller aborts and keeps the stable version serving traffic.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That separation is really the whole idea. Kubernetes carries on running the pods, exactly as it did before. Argo Rollouts is the part that decides whether a new version has earned more traffic yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pieces that matter
&lt;/h2&gt;

&lt;p&gt;As of v1.9.0, the CRDs worth genuinely understanding are &lt;code&gt;Rollout&lt;/code&gt;, &lt;code&gt;AnalysisTemplate&lt;/code&gt;, &lt;code&gt;ClusterAnalysisTemplate&lt;/code&gt;, &lt;code&gt;AnalysisRun&lt;/code&gt;, and &lt;code&gt;Experiment&lt;/code&gt;. That list looks longer than it needs to be, though — in day-to-day work most teams spend nearly all their time in the first three.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;Rollout&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;Rollout&lt;/code&gt; (&lt;code&gt;apiVersion: argoproj.io/v1alpha1&lt;/code&gt;, &lt;code&gt;kind: Rollout&lt;/code&gt;) sits at the centre of everything. It is deliberately built to be a drop-in replacement for a &lt;code&gt;Deployment&lt;/code&gt;, which is why migrating an existing workload is mostly a matter of changing the &lt;code&gt;apiVersion&lt;/code&gt;, the &lt;code&gt;kind&lt;/code&gt;, and the &lt;code&gt;strategy&lt;/code&gt; — the pod template underneath can stay much as it was.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;Rollout&lt;/code&gt; can either carry its own pod template or point at an existing &lt;code&gt;Deployment&lt;/code&gt; through &lt;code&gt;workloadRef&lt;/code&gt;. That second option is the gentle on-ramp: it lets you bring a live workload under Rollouts gradually, rather than rewriting the whole object in one nervous commit.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;AnalysisTemplate&lt;/code&gt; and &lt;code&gt;AnalysisRun&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;AnalysisTemplate&lt;/code&gt; describes the check; &lt;code&gt;AnalysisRun&lt;/code&gt; is that check actually running during a rollout. The template might ask Prometheus for a success rate, a latency figure, or some service-specific business metric that only your team would think to measure. The run takes the answer and decides whether the rollout carries on, pauses, or aborts.&lt;/p&gt;

&lt;p&gt;This is the part that makes Argo Rollouts something more than “rolling updates, but slower”. The controller is not simply counting down a timer between steps — it is waiting for evidence before it commits.&lt;/p&gt;

&lt;h3&gt;
  
  
  Canary vs blue-green
&lt;/h3&gt;

&lt;p&gt;Argo Rollouts supports both, and it helps to remember that they are answers to different questions.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Canary&lt;/strong&gt; is about gradual exposure. Raise the traffic a step at a time, watch what happens, then decide whether to keep going.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blue-green&lt;/strong&gt; is about a clean cutover. Stand the new version up beside the old one, satisfy yourself it works, then flip the active service across.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Neither is universally the right answer. Canary tends to be the better default for customer-facing services, where the whole point is for a bad version to touch as little traffic as possible before you catch it. Blue-green comes into its own when you want an obvious stable-versus-preview split and the reassurance of a fast switch back.&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete canary with analysis
&lt;/h2&gt;

&lt;p&gt;The heart of a real Argo Rollouts setup is rarely the sheer volume of YAML — it is the promotion logic hiding inside it. The example below is deliberately small: a canary that shifts traffic in stages and refuses to promote unless a Prometheus-backed success-rate check agrees.&lt;/p&gt;

&lt;p&gt;First, install the controller:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl create namespace argo-rollouts
kubectl apply &lt;span class="nt"&gt;-n&lt;/span&gt; argo-rollouts &lt;span class="nt"&gt;-f&lt;/span&gt; https://github.com/argoproj/argo-rollouts/releases/latest/download/install.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then define the analysis template:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argoproj.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AnalysisTemplate&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-success-rate&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service-name&lt;/span&gt;
  &lt;span class="na"&gt;metrics&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;success-rate&lt;/span&gt;
      &lt;span class="na"&gt;interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
      &lt;span class="na"&gt;successCondition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;result[0] &amp;gt;= &lt;/span&gt;&lt;span class="m"&gt;0.99&lt;/span&gt;
      &lt;span class="na"&gt;failureLimit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;
      &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;prometheus&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;address&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://prometheus.monitoring.svc.cluster.local:9090&lt;/span&gt;
          &lt;span class="na"&gt;query&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
            &lt;span class="s"&gt;sum(&lt;/span&gt;
              &lt;span class="s"&gt;rate(http_requests_total{service="{{args.service-name}}",status!~"5.."}[5m])&lt;/span&gt;
            &lt;span class="s"&gt;) /&lt;/span&gt;
            &lt;span class="s"&gt;sum(&lt;/span&gt;
              &lt;span class="s"&gt;rate(http_requests_total{service="{{args.service-name}}"}[5m])&lt;/span&gt;
            &lt;span class="s"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then wire it into the rollout:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;argoproj.io/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Rollout&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;replicas&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8&lt;/span&gt;
  &lt;span class="na"&gt;revisionHistoryLimit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
  &lt;span class="na"&gt;selector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;matchLabels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
          &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ghcr.io/example/checkout:v2&lt;/span&gt;
          &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;containerPort&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;8080&lt;/span&gt;
  &lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;canary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;setWeight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pause&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;analysis&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;templates&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;templateName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-success-rate&lt;/span&gt;
            &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service-name&lt;/span&gt;
                &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;setWeight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;30&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pause&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10m&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;setWeight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interesting part here is not the syntax — it is the sequence.&lt;/p&gt;

&lt;p&gt;At 10% traffic, the rollout deliberately pauses long enough to gather some real request data. The &lt;code&gt;AnalysisRun&lt;/code&gt; then measures the service against a threshold you wrote down in advance, and if that metric fails three times it aborts, rather than quietly waving the release on to 30% and then to everyone.&lt;/p&gt;

&lt;p&gt;There is one subtlety worth internalising early. Without a traffic-routing integration, &lt;code&gt;setWeight&lt;/code&gt; is an approximation drawn from ReplicaSet sizes rather than exact request-level shaping. At low replica counts that is often good enough, but it is not the same thing as a true 10% split. When you genuinely need that precision, wire Rollouts up to a traffic provider and give it &lt;code&gt;stableService&lt;/code&gt;, &lt;code&gt;canaryService&lt;/code&gt;, and &lt;code&gt;trafficRouting&lt;/code&gt;. That is the line between “best effort, by pod count” and “deliberate, by request path”.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where blue-green is simpler
&lt;/h2&gt;

&lt;p&gt;Blue-green is the model most teams find easiest to explain out loud, mostly because it has fewer moving parts. There is an active stack, a preview stack, and a single moment of promotion between them.&lt;/p&gt;

&lt;p&gt;In Rollouts, that shape becomes an &lt;code&gt;activeService&lt;/code&gt; carrying live traffic and, optionally, a &lt;code&gt;previewService&lt;/code&gt; pointed at the new ReplicaSet before it goes live. The part that makes it genuinely interesting is what you can hang either side of the cutover with &lt;code&gt;prePromotionAnalysis&lt;/code&gt; and &lt;code&gt;postPromotionAnalysis&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;prePromotionAnalysis&lt;/code&gt; has to be satisfied that the preview version is good enough before the switch happens at all.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;postPromotionAnalysis&lt;/code&gt; keeps watching once traffic has moved across, and can send it back to the old version if the new one starts to degrade.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The field I would commit to memory here is &lt;code&gt;scaleDownDelaySeconds&lt;/code&gt;. It defaults to &lt;strong&gt;30 seconds&lt;/strong&gt;, giving the service selectors and traffic infrastructure time to converge before the old ReplicaSet is scaled away. It sounds like a trivial setting, but it is exactly the sort of thing that separates a rollout which “worked” as far as the controller was concerned from one that quietly dropped traffic down in the dataplane.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this fits with GitOps
&lt;/h2&gt;

&lt;p&gt;Argo Rollouts and Argo CD get along well, but it is worth being clear that they are doing two different jobs.&lt;/p&gt;

&lt;p&gt;Argo CD reconciles the desired state in Git into the cluster. Argo Rollouts decides how a workload that has just changed should be promoted, once that change has landed. The boundary between them is clean, and that cleanness is the point.&lt;/p&gt;

&lt;p&gt;In practice it means you do not have to smuggle rollout logic into your CI pipelines or bend your GitOps repo layout around it. Git still holds the desired pod template, Argo CD still applies it, and Rollouts takes over only for the runtime question of whether the new version deserves to keep advancing.&lt;/p&gt;

&lt;p&gt;It also means rollback here is operational rather than declarative — a distinction that trips people up more often than you would expect. When a rollout aborts, Argo Rollouts does not go and rewrite Git on your behalf. It scales the release back to the stable version in the cluster, but the repository still holds the bad image tag until someone, or another controller, reverts it. That gap has teeth: if Argo CD is running with auto-sync and self-heal, it can read the aborted rollout as simply out of sync and try to re-apply the very version that just failed. Teams new to the model tend to walk straight past that boundary, and only notice it the first time the two controllers start pulling against each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production reality: the sharp edges
&lt;/h2&gt;

&lt;p&gt;Argo Rollouts demos beautifully and is surprisingly easy to misuse once real traffic is involved. The failure modes that actually bite are rarely exotic — they are the boring, avoidable ones.&lt;/p&gt;

&lt;h3&gt;
  
  
  Canary percentages are only exact if your traffic layer is exact
&lt;/h3&gt;

&lt;p&gt;A &lt;code&gt;setWeight: 10&lt;/code&gt; step reads as though it means precisely ten percent, but without traffic routing it is really just a pod-count approximation. If your service runs seven replicas, there is no way for the controller to carve traffic into a clean tenth. When exact exposure genuinely matters, reach for a supported traffic manager rather than trusting the arithmetic.&lt;/p&gt;

&lt;h3&gt;
  
  
  Traffic-routed canaries can overprovision
&lt;/h3&gt;

&lt;p&gt;The moment you add &lt;code&gt;trafficRouting&lt;/code&gt;, the replica arithmetic shifts under you. &lt;code&gt;maxSurge&lt;/code&gt; and &lt;code&gt;maxUnavailable&lt;/code&gt; are no longer the whole story, particularly once &lt;code&gt;setCanaryScale&lt;/code&gt; is in the mix, and it is easy to find yourself running more stable and canary pods at once than you had budgeted for during a promotion. That is usually the right trade for safety — just make sure it is a trade you planned for, rather than one you discover on the cluster.&lt;/p&gt;

&lt;h3&gt;
  
  
  Blue-green is simple, but not cheap
&lt;/h3&gt;

&lt;p&gt;Running blue and green side by side doubles parts of your footprint by design — that is the whole mechanism, not a bug. If you lean on preview environments heavily, keep them short-lived. Rollouts is built for promotion windows measured in minutes, not for quietly maintaining a second long-lived environment that drifts further from the first with every passing day.&lt;/p&gt;

&lt;h3&gt;
  
  
  HPA needs to target the &lt;code&gt;Rollout&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;If the workload is autoscaled, point the HPA at the &lt;code&gt;Rollout&lt;/code&gt; itself — not the &lt;code&gt;Deployment&lt;/code&gt; you migrated from, and certainly not an individual ReplicaSet. Get this wrong and you end up refereeing two controllers that hold different opinions about who owns the replica count, usually at the least convenient moment.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Argo Rollouts is the right call
&lt;/h2&gt;

&lt;p&gt;Reach for Argo Rollouts when the cost of a bad release is high enough that “the pods became Ready” simply is not a strong enough gate to stand behind. In practice that means customer-facing APIs, services with SLOs you actually measure, and teams who already have decent metrics but are still making the promotion call by hand, one deploy at a time.&lt;/p&gt;

&lt;p&gt;What it is not is a decoration. Adopting it to make a deployment diagram look more sophisticated is a poor reason on its own. If a service is low-risk, its traffic layer is simple, and rollback is already fast, a plain &lt;code&gt;Deployment&lt;/code&gt; may well remain the honest choice.&lt;/p&gt;

&lt;p&gt;The real win was never that Rollouts hands Kubernetes yet another CRD. It is that it gives your release process a proper control loop: expose a little, measure what happens, then decide. That single step — pausing to look before committing — is the one a plain &lt;code&gt;Deployment&lt;/code&gt; has never been able to take on its own.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>gitops</category>
      <category>kubernetes</category>
    </item>
    <item>
      <title>EKS node autoscaling with Karpenter: just-in-time capacity without node groups</title>
      <dc:creator>Billy Walker</dc:creator>
      <pubDate>Thu, 09 Jul 2026 19:01:51 +0000</pubDate>
      <link>https://dev.to/bwlkr/eks-node-autoscaling-with-karpenter-just-in-time-capacity-without-node-groups-pap</link>
      <guid>https://dev.to/bwlkr/eks-node-autoscaling-with-karpenter-just-in-time-capacity-without-node-groups-pap</guid>
      <description>&lt;p&gt;Most people arrive at Karpenter after a slow build-up of small frustrations rather than one dramatic outage. Managed node groups start out simple, and then one day you notice you are spending real time deciding which group should be a little bigger, which instance family to add next, and how much headroom to leave "just in case". It starts to feel less like running a cluster and more like tending a spreadsheet.&lt;/p&gt;

&lt;p&gt;Karpenter offers a gentler mental model. Instead of pre-sizing groups and hoping one of them fits, it watches the pods that cannot be scheduled, works out the capacity they actually need, and launches EC2 nodes to match. Two things shift as a result. Scaling up follows the workload rather than a template you wrote weeks ago, and scaling down becomes a policy you set once rather than a clean-up job you keep meaning to get to. On a busy cluster, that is often the difference between paying for a permanent buffer and only asking for capacity when there is real work waiting.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Karpenter changes EKS node autoscaling
&lt;/h2&gt;

&lt;p&gt;The simplest way to describe the change is this: Cluster Autoscaler works at the level of node groups, while Karpenter works at the level of individual pending pods. On paper that reads like a minor implementation detail. In day-to-day operation it turns out to be most of the story.&lt;/p&gt;

&lt;p&gt;With Cluster Autoscaler, you make the important decisions up front — which instance families exist, how much spare room each group keeps, and which group grows when pods start to queue. Karpenter takes those same decisions and turns them into guardrails, then chooses the actual node shape at launch time, once it can see what is really being asked for.&lt;/p&gt;

&lt;p&gt;That tends to suit clusters where the workloads do not all look alike:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;web services that spike during deployments&lt;/li&gt;
&lt;li&gt;queue workers whose demand follows traffic rather than CPU&lt;/li&gt;
&lt;li&gt;short-lived CI or batch pods&lt;/li&gt;
&lt;li&gt;teams mixing Spot and on-demand capacity in the same cluster&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this means node groups are a mistake. They are simply static by design, and that is perfectly fine right up until your workload mix changes often enough that static sizing starts to feel expensive or brittle. That is usually the moment Karpenter starts to earn its place.&lt;/p&gt;

&lt;h2&gt;
  
  
  NodePool and EC2NodeClass split policy from AWS plumbing
&lt;/h2&gt;

&lt;p&gt;On EKS, two Karpenter resources do most of the work: &lt;code&gt;NodePool&lt;/code&gt; and &lt;code&gt;EC2NodeClass&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;NodePool&lt;/code&gt; describes scheduling and disruption policy — what kinds of nodes are allowed, how aggressively Karpenter can consolidate, and which workloads are welcome to land there.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;EC2NodeClass&lt;/code&gt; describes the AWS side — subnets, security groups, AMI family, the IAM role or instance profile, storage, and the other EC2-specific details.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keeping these separate is more useful than it first appears, because it holds your capacity intent apart from the cloud plumbing underneath it. A single &lt;code&gt;EC2NodeClass&lt;/code&gt; can back several &lt;code&gt;NodePool&lt;/code&gt;s, and each &lt;code&gt;NodePool&lt;/code&gt; can express a different set of rules about the workloads it serves.&lt;/p&gt;

&lt;p&gt;A minimal setup looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.sh/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;NodePool&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;general&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;nodeClassRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;group&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.k8s.aws&lt;/span&gt;
        &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;EC2NodeClass&lt;/span&gt;
        &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
      &lt;span class="na"&gt;requirements&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.k8s.aws/instance-category&lt;/span&gt;
          &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
          &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;c'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;m'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;r'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;karpenter.sh/capacity-type&lt;/span&gt;
          &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
          &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;spot'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;on-demand'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubernetes.io/arch&lt;/span&gt;
          &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
          &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;amd64'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;disruption&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;consolidationPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;WhenEmptyOrUnderutilized&lt;/span&gt;
    &lt;span class="na"&gt;consolidateAfter&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10m&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things are worth remembering here, and the docs are clear about both. Karpenter will not do anything at all until at least one &lt;code&gt;NodePool&lt;/code&gt; exists — there is no implicit default waiting in the background. And a pod's own scheduling constraints are intersected with the &lt;code&gt;NodePool&lt;/code&gt; requirements, so if a workload asks for something the pool does not allow — an instance family, a zone, an architecture — Karpenter simply will not launch a node for it. It sounds obvious written down, but it is one of the more common ways for a pending pod to look like a capacity shortage when it is really a policy mismatch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why just-in-time provisioning beats a larger buffer
&lt;/h2&gt;

&lt;p&gt;The traditional habit on EKS is to keep enough spare capacity around that the next deployment or burst will probably fit. The trouble with "probably" is that it quietly shows up on the bill every month, whether or not you ever use it.&lt;/p&gt;

&lt;p&gt;Karpenter lets you run a little closer to the line, because it provisions from real pending demand rather than a guess. If a rollout suddenly needs six more replicas, it can pick capacity that fits those pods instead of inflating a general-purpose group and hoping the shape is close. When the burst passes, the same controller can quietly remove or replace whatever it over-provisioned.&lt;/p&gt;

&lt;p&gt;This helps most when your requests are uneven. A memory-heavy worker, an ARM-only build job, and a stateless frontend rollout do not really want to share a node. Node groups force you to approximate and accept some waste, whereas Karpenter can stay much closer to what each workload actually asked for.&lt;/p&gt;

&lt;p&gt;The catch is that your constraints have to be explicit. Instance family, architecture, capacity type, disruption tolerance — all of it needs to be written down properly. Karpenter takes away the provisioning ceremony, but it does not take away the thinking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consolidation, drift, and interruption handling are the real day-two features
&lt;/h2&gt;

&lt;p&gt;Fast provisioning is the headline, but the part you live with day to day is what happens after a node has landed.&lt;/p&gt;

&lt;p&gt;Karpenter's disruption controller looks after two cases that quietly save a lot of effort:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Consolidation&lt;/strong&gt; removes empty nodes, or swaps an underutilised one for a cheaper shape when the same pods can comfortably run elsewhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drift&lt;/strong&gt; replaces nodes that have fallen out of step with the current &lt;code&gt;NodePool&lt;/code&gt; or &lt;code&gt;EC2NodeClass&lt;/code&gt; — a new AMI selection, say, or changed launch requirements.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In a node-group-heavy setup, those two jobs — right-sizing and AMI rollovers — usually turn into separate pieces of scheduled maintenance. Karpenter keeps comparing what you asked for with what is actually running, so the fleet does not slowly drift out of shape between maintenance windows.&lt;/p&gt;

&lt;p&gt;There is an honest caveat here, and the Karpenter docs do not hide from it: consolidation is only ever as good as the disruption rules you allow. Preferred anti-affinity and topology spread can quietly stop nodes from consolidating even when the scheduler could place the pods elsewhere, and PodDisruptionBudgets can do the same. So when someone says Karpenter is not saving them any money, the controller is rarely the first place to look. More often it is the workload policy wrapped around the nodes it is trying to drain.&lt;/p&gt;

&lt;p&gt;Interruption handling matters too, especially if you lean on Spot. Karpenter can react to a Spot interruption warning, bring up replacement capacity ahead of time, and drain the old node in parallel — but all of that still depends on the workload tolerating disruption. Two minutes is not long when graceful shutdown, PDBs, and placement rules are all pulling in different directions.&lt;/p&gt;

&lt;p&gt;The cost story here is the unglamorous one: a cluster that already had plenty of capacity in aggregate, just trapped in the wrong shape. Consolidation and drift are what turn that from a tidy architecture diagram into an actual lower bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  Spot works well when the fallback path is real
&lt;/h2&gt;

&lt;p&gt;Karpenter and Spot are often mentioned in the same breath, and for good reason — it can choose from a far wider range of instance types than a hand-maintained node-group layout, which is exactly what makes Spot workable. The part worth dwelling on is that Spot only pays off when the fallback is genuinely there.&lt;/p&gt;

&lt;p&gt;A sensible EKS pattern usually looks something like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;prefer Spot for workloads that can tolerate interruption&lt;/li&gt;
&lt;li&gt;allow on-demand in the same &lt;code&gt;NodePool&lt;/code&gt; for when the service cannot wait for spare Spot capacity to appear&lt;/li&gt;
&lt;li&gt;keep topology spread and PDBs sensible, so replacing one node does not set off a second incident&lt;/li&gt;
&lt;li&gt;test your interruption handling before you actually need it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point carries more weight than it might seem. "We allow on-demand fallback" is a line of configuration; proving that a constrained workload can still schedule when a Spot pool dries up is a different thing entirely, and a good deal more reassuring when it happens at an awkward hour.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Karpenter is better than Cluster Autoscaler, and where it is not
&lt;/h2&gt;

&lt;p&gt;Karpenter tends to win out over Cluster Autoscaler when an EKS cluster carries a genuine mix of workloads and their shapes change often. It is good at turning pending work into right-sized nodes without asking you to pre-bake every possible capacity combination into managed groups ahead of time.&lt;/p&gt;

&lt;p&gt;It is a less obvious choice when the fleet is already settled — the node groups are well understood, the workload mix barely moves, and nothing about the current setup is causing pain. If what you have today is simple and cheap, Karpenter can end up being a sophisticated answer to a question nobody was really asking.&lt;/p&gt;

&lt;p&gt;So the useful question is not whether Karpenter is newer or cleverer, because it is comfortably both. It is whether managing node groups has quietly become a drag on your time. If it has, the added complexity is usually a fair trade. If it has not, a well-kept Cluster Autoscaler setup may still be all you need.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;Karpenter moves EKS node autoscaling away from managing node groups and towards provisioning that follows the workload. You describe the policy once through &lt;code&gt;NodePool&lt;/code&gt; and &lt;code&gt;EC2NodeClass&lt;/code&gt;, and from then on capacity is launched to match the pods actually in front of the scheduler, rather than the guess someone made last quarter.&lt;/p&gt;

&lt;p&gt;The real payoff is not only quicker scale-up. It is the combination of right-sized launches, consolidation, drift replacement, and Spot fallback that holds together under pressure. Between them, consolidation and Spot can take a meaningful bite out of node spend without making scale-up feel fragile, though the exact number will always depend on your particular workload mix.&lt;/p&gt;

&lt;p&gt;If you are already spending your week wrestling oversized node groups, slow scale-up, or a fleet that drifts out of shape between maintenance windows, Karpenter is well worth a serious look on EKS.&lt;/p&gt;

</description>
      <category>karpenter</category>
      <category>kubernetes</category>
      <category>eks</category>
      <category>aws</category>
    </item>
  </channel>
</rss>
