DEV Community

Cover image for Upgrade EKS Nodes Without Downtime: The Bridge Add-on Version Trick
Ajinkya
Ajinkya

Posted on

Upgrade EKS Nodes Without Downtime: The Bridge Add-on Version Trick

Move your add-ons to versions that work on both Kubernetes releases, and the node upgrade stops being the scary part.


The problem with "nodes first, add-ons after"

The AWS upgrade guide gives the order as: control plane → nodes → add-ons.

That works on paper, but it may leaves a gap. While your nodes roll from 1.35 to 1.36, you are running new nodes with old add-on versions, and those versions were only validated against the old Kubernetes version. The add-ons that hurt most are the ones every pod depends on:

  • vpc-cni hands out pod IPs on every node
  • CoreDNS resolves names for every workload
  • aws-ebs-csi-driver attaches volumes for stateful pods

If one of them misbehaves on a freshly rolled node, you find out in the middle of the rollout, with pods rescheduling all around you.

The idea: bridge versions

A bridge version is the newest add-on version that is supported on both your current and your target Kubernetes version.

If every add-on is on its bridge version before the nodes roll:

  • Old nodes (1.35) run an add-on version that supports 1.35.
  • New nodes (1.36) run the same add-on version, which also supports 1.36.
  • No node ever runs an add-on that wasn't validated for its Kubernetes version.

Because a bridge version is valid on both releases, you can apply it at any point before the node upgrade: at the very start, on 1.35, or right after the control plane moves to 1.36. Both work.

The order I use:

EKS control plane → bridge add-ons → nodes → (optional) default add-on versions.

This doesn't contradict the AWS order. It adds one safe step in front of the node roll.

Step 0: Find the bridge version for each add-on

EKS can list every version of an add-on for a given Kubernetes version. The bridge is the newest version that appears in both lists.

#!/usr/bin/env bash

OLD_K8S="1.35"
NEW_K8S="1.36"

get_versions() {
  aws eks describe-addon-versions \
    --addon-name "$1" \
    --kubernetes-version "$2" \
    --query 'addons[0].addonVersions[].addonVersion' \
    --output text 2>/dev/null | tr '\t' '\n' | tr -d '\r' | sort -V
}

for addon in vpc-cni coredns kube-proxy aws-ebs-csi-driver metrics-server; do
  old_versions=$(get_versions "$addon" "$OLD_K8S")
  new_versions=$(get_versions "$addon" "$NEW_K8S")

  echo "===== $addon ====="

  if [ -z "$old_versions" ] || [ -z "$new_versions" ]; then
    echo "WARNING: no versions returned - check addon name / region"
    echo
    continue
  fi

  old_latest=$(echo "$old_versions" | tail -1)
  new_latest=$(echo "$new_versions" | tail -1)

  common_version=$(grep -Fxf <(echo "$old_versions") <(echo "$new_versions") | sort -V | tail -1)

  echo "$OLD_K8S latest : $old_latest"
  echo "$NEW_K8S latest : $new_latest"

  if [ -z "$common_version" ]; then
    echo "No overlapping version found - you may need an intermediate hop"
  else
    echo "Bridge version (newest common to both): $common_version"
  fi
  echo
done
Enter fullscreen mode Exit fullscreen mode

Output from my cluster (1.35 → 1.36)

===== vpc-cni =====
1.35 latest : v1.23.1-eksbuild.1
1.36 latest : v1.23.1-eksbuild.1
Bridge version (newest common to both): v1.23.1-eksbuild.1

===== coredns =====
1.35 latest : v1.14.6-eksbuild.4
1.36 latest : v1.14.6-eksbuild.4
Bridge version (newest common to both): v1.14.6-eksbuild.4

===== kube-proxy =====
1.35 latest : v1.35.3-eksbuild.29
1.36 latest : v1.36.0-eksbuild.25
Bridge version (newest common to both): v1.35.3-eksbuild.29

===== aws-ebs-csi-driver =====
1.35 latest : v1.66.0-eksbuild.1
1.36 latest : v1.66.0-eksbuild.1
Bridge version (newest common to both): v1.66.0-eksbuild.1

===== metrics-server =====
1.35 latest : v0.9.0-eksbuild.11
1.36 latest : v0.9.0-eksbuild.11
Bridge version (newest common to both): v0.9.0-eksbuild.11
Enter fullscreen mode Exit fullscreen mode

How to read it:

  • For vpc-cni, coredns, aws-ebs-csi-driver and metrics-server, the latest version is identical on 1.35 and 1.36, so the bridge is simply the latest build.
  • kube-proxy is the exception. Its version tracks the Kubernetes minor, so the 1.36 latest (v1.36.0-eksbuild.25) is not available on 1.35, and the bridge is a 1.35.x build (v1.35.3-eksbuild.29). Keep it there until the nodes are on 1.36.

How the Terraform is wired

The EKS cluster (control plane)

The control plane version comes from a single variable:

resource "aws_eks_cluster" "eks" {
  name     = var.cluster_name
  role_arn = aws_iam_role.eks_cluster.arn
  version  = var.cluster_version

  vpc_config {
    # Pass both public and private subnets so the EKS control plane
    # ENIs are spread across AZs inside the VPC.
    subnet_ids              = concat(aws_subnet.public[*].id, aws_subnet.private[*].id)
    security_group_ids      = [aws_security_group.eks_cluster.id]
    endpoint_public_access  = true  # keeps kubectl working from your machine
    endpoint_private_access = true  # nodes in private subnets reach the API internally
  }

  access_config {
    authentication_mode                         = "API"
    bootstrap_cluster_creator_admin_permissions = true
  }

  tags = merge(var.tags, {
    Name = var.cluster_name
  })

  depends_on = [
    aws_iam_role_policy_attachment.eks_cluster_policy
  ]
}
Enter fullscreen mode Exit fullscreen mode

What matters for the upgrade:

  • version = var.cluster_version is the only line you change. Going from "1.35" to "1.36" makes Terraform update the control plane in place. EKS upgrades one minor version at a time and there is no rollback, so this is the step to plan carefully.
  • vpc_config and access_config stay the same during an upgrade. The public and private endpoints keep kubectl and the nodes talking to the API, and authentication_mode = "API" keeps your access entries working.
  • The nodes and add-ons don't change when this resource changes. That is why the control plane can be upgraded on its own with -target, as Step 1 shows.

The node group

The managed node group reads the cluster version and rolls when it changes:

resource "aws_eks_node_group" "ondemand-node" {
  cluster_name  = aws_eks_cluster.eks.name
  node_role_arn = aws_iam_role.node.arn
  subnet_ids    = aws_subnet.private[*].id
  capacity_type = "ON_DEMAND"
  ami_type      = var.ami_type

  version              = var.cluster_version  # ties node AMI to control plane version
  force_update_version = true                 # forces rolling update when version changes

  update_config {
    max_unavailable = 1 # only 1 of nodes can be unavailable at once
  }
  ...
}
Enter fullscreen mode Exit fullscreen mode

What each line does:

  • version = var.cluster_version sets the Kubernetes version of the nodes. With a managed node group and an ami_type, EKS picks the matching EKS-optimized AMI for that version. Change the variable and the node group rolls onto the new AMI.
  • force_update_version = true lets the update go through even when pods can't be drained because of a PodDisruptionBudget. Without it, a blocking PDB makes the update fail. With it, EKS eventually proceeds anyway, so pods can be evicted despite the PDB. That keeps upgrades from getting stuck, but it is also the one setting that can cost you availability, so make sure critical workloads have enough replicas.
  • max_unavailable = 1 rolls one node at a time, so the rest of the group keeps serving traffic.

The add-ons

Add-on versions come from a list variable, so bumping them is a data change:

resource "aws_eks_addon" "eks-addons" {
  for_each      = { for idx, addon in var.addons : idx => addon }
  cluster_name  = aws_eks_cluster.eks.name
  addon_name    = each.value.name
  addon_version = each.value.version

  # attach role ONLY for EBS CSI
  service_account_role_arn = each.value.name == "aws-ebs-csi-driver" ? aws_iam_role.ebs_csi.arn : null

  depends_on = [
    aws_eks_node_group.ondemand-node,
    aws_iam_role_policy_attachment.ebs_csi_policy
  ]
}
Enter fullscreen mode Exit fullscreen mode

Gotcha: -target pulls in dependencies

-target applies the resource you name plus everything it depends on. The add-on resource above has depends_on = [aws_eks_node_group.ondemand-node]. If the node group shares var.cluster_version with the control plane, then after you bump that variable, terraform apply -target=aws_eks_addon.eks-addons drags the node group in and rolls your nodes before the bridge add-ons land. That defeats the whole point.

The fix is a separate variable for the node group, so you can hold nodes on 1.35 while you upgrade the control plane and add-ons, then move them last (nodes are allowed to run behind the control plane):

  version              = var.node_group_version  # was var.cluster_version
  force_update_version = true
Enter fullscreen mode Exit fullscreen mode

The rest of this post assumes that split.

The upgrade, step by step with -target

Before each step, preview it with terraform plan using the same -target flags and check that only the resources you expect will change.

Step 1: Upgrade the EKS control plane

Set cluster_version = "1.36" and leave node_group_version = "1.35".

terraform plan  -target=aws_eks_cluster.eks
terraform apply -target=aws_eks_cluster.eks
Enter fullscreen mode Exit fullscreen mode

EKS only allows one minor version at a time and there is no rollback, so run your pre-checks first: deprecated APIs, PodDisruptionBudgets, and a supported Karpenter version for 1.36 if you use it.

Step 2: Apply the bridge add-on versions

Set var.addons to the bridge versions from Step 0:

addons = [
  { name = "vpc-cni",            version = "v1.23.1-eksbuild.1" },
  { name = "coredns",            version = "v1.14.6-eksbuild.4" },
  { name = "kube-proxy",         version = "v1.35.3-eksbuild.29" },
  { name = "aws-ebs-csi-driver", version = "v1.66.0-eksbuild.1" },
  { name = "metrics-server",     version = "v0.9.0-eksbuild.11" },
]
Enter fullscreen mode Exit fullscreen mode
terraform plan  -target=aws_eks_addon.eks-addons
terraform apply -target=aws_eks_addon.eks-addons
Enter fullscreen mode Exit fullscreen mode

You can also target a single add-on. Because of for_each, the address includes the key:

terraform apply -target='aws_eks_addon.eks-addons["1"]'   # just coredns, for example
Enter fullscreen mode Exit fullscreen mode

Check health before moving on:

kubectl get pods -n kube-system
Enter fullscreen mode Exit fullscreen mode

You don't have to wait for Step 1. Since bridge versions work on both releases, you can run this step first, on the old 1.35 control plane, and then upgrade the control plane. Steps 1 and 2 are interchangeable. The only rule is that the bridge add-ons land before the nodes.

Step 3: Roll the nodes

Set node_group_version = "1.36":

terraform plan  -target=aws_eks_node_group.ondemand-node
terraform apply -target=aws_eks_node_group.ondemand-node
Enter fullscreen mode Exit fullscreen mode

The node group rolls one node at a time. Every new node comes up on 1.36 running add-on versions that already support 1.36.

If you run Karpenter, its nodes are outside the managed node group. Check that your Karpenter release supports 1.36 before Step 1, then let drift replacement roll them onto the new AMI while respecting PDBs and disruption budgets. Watch progress with:

kubectl get nodes -o wide
Enter fullscreen mode Exit fullscreen mode

Step 4 (optional): Move to the default 1.36 add-on versions

Once all nodes are on 1.36, you can move add-ons to the versions EKS marks as default for 1.36. This command shows them:

K8S_VERSION="1.36"

for addon in vpc-cni coredns kube-proxy aws-ebs-csi-driver metrics-server; do
  echo "===== $addon ====="

  aws eks describe-addon-versions \
    --addon-name "$addon" \
    --kubernetes-version "$K8S_VERSION" \
    --query 'addons[0].addonVersions[?compatibilities[?defaultVersion==`true`]].{Version:addonVersion,Default:compatibilities[?defaultVersion==`true`].defaultVersion}' \
    --output table

  echo
done
Enter fullscreen mode Exit fullscreen mode

The --query keeps only the versions whose compatibility entry for 1.36 has defaultVersion set to true, and prints them as a table. The default is the version EKS would install if you created the add-on fresh on a 1.36 cluster. It can differ from the newest build, so compare it with the latest list from Step 0 before choosing.

Put the versions you pick into var.addons and apply:

terraform apply -target=aws_eks_addon.eks-addons
Enter fullscreen mode Exit fullscreen mode

This step is optional. A bridge version is already supported on 1.36, so you can stay on it and upgrade add-ons later on your own schedule. The one add-on to move promptly is kube-proxy, so it matches the node version again.

Step 5: A final plain apply

-target is meant for exceptional cases, and it can leave the rest of your state unreconciled. Finish with a normal run and confirm it reports nothing unexpected:

terraform plan
terraform apply
Enter fullscreen mode Exit fullscreen mode

The order at a glance

Step Action Result
0 Find bridge versions nothing changed
1 Upgrade control plane (-target=aws_eks_cluster.eks) cluster 1.36, nodes 1.35
2 Apply bridge add-ons (-target=aws_eks_addon.eks-addons) add-ons valid on 1.35 and 1.36
3 Roll nodes (-target=aws_eks_node_group.ondemand-node) nodes 1.36
4 Optional: default 1.36 add-on versions add-ons on EKS defaults for 1.36
5 Plain terraform apply state fully reconciled

Steps 1 and 2 are interchangeable.

Caveats

  • A bridge version reduces upgrade risk. It does not remove all of it. You still need multiple replicas for critical workloads, sensible PodDisruptionBudgets, and a check for deprecated APIs.
  • force_update_version = true can override a PDB during the node roll. It avoids stuck upgrades, but it also means the PDB is not an absolute guarantee.
  • If the script shows no overlapping version for an add-on, you'll need an intermediate hop for that one.
  • Read the release notes for each add-on. A version being listed as compatible is not a guarantee that nothing changed.

Takeaways

  • The AWS order (control plane, nodes, add-ons) leaves a window where new nodes run old add-ons.
  • A bridge version closes that window because it works on both Kubernetes versions, so it can go in before or after the control plane upgrade, as long as it lands before the nodes.
  • Use -target to control the order, but watch depends_on: give the node group its own version variable, and finish with a plain terraform apply.

Top comments (1)