Move your add-ons to versions that work on both Kubernetes releases, and the node upgrade stops being the scary part.
The problem with "nodes first, add-ons after"
The AWS upgrade guide gives the order as: control plane → nodes → add-ons.
That works on paper, but it may leaves a gap. While your nodes roll from 1.35 to 1.36, you are running new nodes with old add-on versions, and those versions were only validated against the old Kubernetes version. The add-ons that hurt most are the ones every pod depends on:
- vpc-cni hands out pod IPs on every node
- CoreDNS resolves names for every workload
- aws-ebs-csi-driver attaches volumes for stateful pods
If one of them misbehaves on a freshly rolled node, you find out in the middle of the rollout, with pods rescheduling all around you.
The idea: bridge versions
A bridge version is the newest add-on version that is supported on both your current and your target Kubernetes version.
If every add-on is on its bridge version before the nodes roll:
- Old nodes (1.35) run an add-on version that supports 1.35.
- New nodes (1.36) run the same add-on version, which also supports 1.36.
- No node ever runs an add-on that wasn't validated for its Kubernetes version.
Because a bridge version is valid on both releases, you can apply it at any point before the node upgrade: at the very start, on 1.35, or right after the control plane moves to 1.36. Both work.
The order I use:
EKS control plane → bridge add-ons → nodes → (optional) default add-on versions.
This doesn't contradict the AWS order. It adds one safe step in front of the node roll.
Step 0: Find the bridge version for each add-on
EKS can list every version of an add-on for a given Kubernetes version. The bridge is the newest version that appears in both lists.
#!/usr/bin/env bash
OLD_K8S="1.35"
NEW_K8S="1.36"
get_versions() {
aws eks describe-addon-versions \
--addon-name "$1" \
--kubernetes-version "$2" \
--query 'addons[0].addonVersions[].addonVersion' \
--output text 2>/dev/null | tr '\t' '\n' | tr -d '\r' | sort -V
}
for addon in vpc-cni coredns kube-proxy aws-ebs-csi-driver metrics-server; do
old_versions=$(get_versions "$addon" "$OLD_K8S")
new_versions=$(get_versions "$addon" "$NEW_K8S")
echo "===== $addon ====="
if [ -z "$old_versions" ] || [ -z "$new_versions" ]; then
echo "WARNING: no versions returned - check addon name / region"
echo
continue
fi
old_latest=$(echo "$old_versions" | tail -1)
new_latest=$(echo "$new_versions" | tail -1)
common_version=$(grep -Fxf <(echo "$old_versions") <(echo "$new_versions") | sort -V | tail -1)
echo "$OLD_K8S latest : $old_latest"
echo "$NEW_K8S latest : $new_latest"
if [ -z "$common_version" ]; then
echo "No overlapping version found - you may need an intermediate hop"
else
echo "Bridge version (newest common to both): $common_version"
fi
echo
done
Output from my cluster (1.35 → 1.36)
===== vpc-cni =====
1.35 latest : v1.23.1-eksbuild.1
1.36 latest : v1.23.1-eksbuild.1
Bridge version (newest common to both): v1.23.1-eksbuild.1
===== coredns =====
1.35 latest : v1.14.6-eksbuild.4
1.36 latest : v1.14.6-eksbuild.4
Bridge version (newest common to both): v1.14.6-eksbuild.4
===== kube-proxy =====
1.35 latest : v1.35.3-eksbuild.29
1.36 latest : v1.36.0-eksbuild.25
Bridge version (newest common to both): v1.35.3-eksbuild.29
===== aws-ebs-csi-driver =====
1.35 latest : v1.66.0-eksbuild.1
1.36 latest : v1.66.0-eksbuild.1
Bridge version (newest common to both): v1.66.0-eksbuild.1
===== metrics-server =====
1.35 latest : v0.9.0-eksbuild.11
1.36 latest : v0.9.0-eksbuild.11
Bridge version (newest common to both): v0.9.0-eksbuild.11
How to read it:
- For vpc-cni, coredns, aws-ebs-csi-driver and metrics-server, the latest version is identical on 1.35 and 1.36, so the bridge is simply the latest build.
-
kube-proxy is the exception. Its version tracks the Kubernetes minor, so the 1.36 latest (
v1.36.0-eksbuild.25) is not available on 1.35, and the bridge is a1.35.xbuild (v1.35.3-eksbuild.29). Keep it there until the nodes are on 1.36.
How the Terraform is wired
The EKS cluster (control plane)
The control plane version comes from a single variable:
resource "aws_eks_cluster" "eks" {
name = var.cluster_name
role_arn = aws_iam_role.eks_cluster.arn
version = var.cluster_version
vpc_config {
# Pass both public and private subnets so the EKS control plane
# ENIs are spread across AZs inside the VPC.
subnet_ids = concat(aws_subnet.public[*].id, aws_subnet.private[*].id)
security_group_ids = [aws_security_group.eks_cluster.id]
endpoint_public_access = true # keeps kubectl working from your machine
endpoint_private_access = true # nodes in private subnets reach the API internally
}
access_config {
authentication_mode = "API"
bootstrap_cluster_creator_admin_permissions = true
}
tags = merge(var.tags, {
Name = var.cluster_name
})
depends_on = [
aws_iam_role_policy_attachment.eks_cluster_policy
]
}
What matters for the upgrade:
-
version = var.cluster_versionis the only line you change. Going from"1.35"to"1.36"makes Terraform update the control plane in place. EKS upgrades one minor version at a time and there is no rollback, so this is the step to plan carefully. -
vpc_configandaccess_configstay the same during an upgrade. The public and private endpoints keepkubectland the nodes talking to the API, andauthentication_mode = "API"keeps your access entries working. - The nodes and add-ons don't change when this resource changes. That is why the control plane can be upgraded on its own with
-target, as Step 1 shows.
The node group
The managed node group reads the cluster version and rolls when it changes:
resource "aws_eks_node_group" "ondemand-node" {
cluster_name = aws_eks_cluster.eks.name
node_role_arn = aws_iam_role.node.arn
subnet_ids = aws_subnet.private[*].id
capacity_type = "ON_DEMAND"
ami_type = var.ami_type
version = var.cluster_version # ties node AMI to control plane version
force_update_version = true # forces rolling update when version changes
update_config {
max_unavailable = 1 # only 1 of nodes can be unavailable at once
}
...
}
What each line does:
-
version = var.cluster_versionsets the Kubernetes version of the nodes. With a managed node group and anami_type, EKS picks the matching EKS-optimized AMI for that version. Change the variable and the node group rolls onto the new AMI. -
force_update_version = truelets the update go through even when pods can't be drained because of a PodDisruptionBudget. Without it, a blocking PDB makes the update fail. With it, EKS eventually proceeds anyway, so pods can be evicted despite the PDB. That keeps upgrades from getting stuck, but it is also the one setting that can cost you availability, so make sure critical workloads have enough replicas. -
max_unavailable = 1rolls one node at a time, so the rest of the group keeps serving traffic.
The add-ons
Add-on versions come from a list variable, so bumping them is a data change:
resource "aws_eks_addon" "eks-addons" {
for_each = { for idx, addon in var.addons : idx => addon }
cluster_name = aws_eks_cluster.eks.name
addon_name = each.value.name
addon_version = each.value.version
# attach role ONLY for EBS CSI
service_account_role_arn = each.value.name == "aws-ebs-csi-driver" ? aws_iam_role.ebs_csi.arn : null
depends_on = [
aws_eks_node_group.ondemand-node,
aws_iam_role_policy_attachment.ebs_csi_policy
]
}
Gotcha: -target pulls in dependencies
-target applies the resource you name plus everything it depends on. The add-on resource above has depends_on = [aws_eks_node_group.ondemand-node]. If the node group shares var.cluster_version with the control plane, then after you bump that variable, terraform apply -target=aws_eks_addon.eks-addons drags the node group in and rolls your nodes before the bridge add-ons land. That defeats the whole point.
The fix is a separate variable for the node group, so you can hold nodes on 1.35 while you upgrade the control plane and add-ons, then move them last (nodes are allowed to run behind the control plane):
version = var.node_group_version # was var.cluster_version
force_update_version = true
The rest of this post assumes that split.
The upgrade, step by step with -target
Before each step, preview it with terraform plan using the same -target flags and check that only the resources you expect will change.
Step 1: Upgrade the EKS control plane
Set cluster_version = "1.36" and leave node_group_version = "1.35".
terraform plan -target=aws_eks_cluster.eks
terraform apply -target=aws_eks_cluster.eks
EKS only allows one minor version at a time and there is no rollback, so run your pre-checks first: deprecated APIs, PodDisruptionBudgets, and a supported Karpenter version for 1.36 if you use it.
Step 2: Apply the bridge add-on versions
Set var.addons to the bridge versions from Step 0:
addons = [
{ name = "vpc-cni", version = "v1.23.1-eksbuild.1" },
{ name = "coredns", version = "v1.14.6-eksbuild.4" },
{ name = "kube-proxy", version = "v1.35.3-eksbuild.29" },
{ name = "aws-ebs-csi-driver", version = "v1.66.0-eksbuild.1" },
{ name = "metrics-server", version = "v0.9.0-eksbuild.11" },
]
terraform plan -target=aws_eks_addon.eks-addons
terraform apply -target=aws_eks_addon.eks-addons
You can also target a single add-on. Because of for_each, the address includes the key:
terraform apply -target='aws_eks_addon.eks-addons["1"]' # just coredns, for example
Check health before moving on:
kubectl get pods -n kube-system
You don't have to wait for Step 1. Since bridge versions work on both releases, you can run this step first, on the old 1.35 control plane, and then upgrade the control plane. Steps 1 and 2 are interchangeable. The only rule is that the bridge add-ons land before the nodes.
Step 3: Roll the nodes
Set node_group_version = "1.36":
terraform plan -target=aws_eks_node_group.ondemand-node
terraform apply -target=aws_eks_node_group.ondemand-node
The node group rolls one node at a time. Every new node comes up on 1.36 running add-on versions that already support 1.36.
If you run Karpenter, its nodes are outside the managed node group. Check that your Karpenter release supports 1.36 before Step 1, then let drift replacement roll them onto the new AMI while respecting PDBs and disruption budgets. Watch progress with:
kubectl get nodes -o wide
Step 4 (optional): Move to the default 1.36 add-on versions
Once all nodes are on 1.36, you can move add-ons to the versions EKS marks as default for 1.36. This command shows them:
K8S_VERSION="1.36"
for addon in vpc-cni coredns kube-proxy aws-ebs-csi-driver metrics-server; do
echo "===== $addon ====="
aws eks describe-addon-versions \
--addon-name "$addon" \
--kubernetes-version "$K8S_VERSION" \
--query 'addons[0].addonVersions[?compatibilities[?defaultVersion==`true`]].{Version:addonVersion,Default:compatibilities[?defaultVersion==`true`].defaultVersion}' \
--output table
echo
done
The --query keeps only the versions whose compatibility entry for 1.36 has defaultVersion set to true, and prints them as a table. The default is the version EKS would install if you created the add-on fresh on a 1.36 cluster. It can differ from the newest build, so compare it with the latest list from Step 0 before choosing.
Put the versions you pick into var.addons and apply:
terraform apply -target=aws_eks_addon.eks-addons
This step is optional. A bridge version is already supported on 1.36, so you can stay on it and upgrade add-ons later on your own schedule. The one add-on to move promptly is kube-proxy, so it matches the node version again.
Step 5: A final plain apply
-target is meant for exceptional cases, and it can leave the rest of your state unreconciled. Finish with a normal run and confirm it reports nothing unexpected:
terraform plan
terraform apply
The order at a glance
| Step | Action | Result |
|---|---|---|
| 0 | Find bridge versions | nothing changed |
| 1 | Upgrade control plane (-target=aws_eks_cluster.eks) |
cluster 1.36, nodes 1.35 |
| 2 | Apply bridge add-ons (-target=aws_eks_addon.eks-addons) |
add-ons valid on 1.35 and 1.36 |
| 3 | Roll nodes (-target=aws_eks_node_group.ondemand-node) |
nodes 1.36 |
| 4 | Optional: default 1.36 add-on versions | add-ons on EKS defaults for 1.36 |
| 5 | Plain terraform apply
|
state fully reconciled |
Steps 1 and 2 are interchangeable.
Caveats
- A bridge version reduces upgrade risk. It does not remove all of it. You still need multiple replicas for critical workloads, sensible PodDisruptionBudgets, and a check for deprecated APIs.
-
force_update_version = truecan override a PDB during the node roll. It avoids stuck upgrades, but it also means the PDB is not an absolute guarantee. - If the script shows no overlapping version for an add-on, you'll need an intermediate hop for that one.
- Read the release notes for each add-on. A version being listed as compatible is not a guarantee that nothing changed.
Takeaways
- The AWS order (control plane, nodes, add-ons) leaves a window where new nodes run old add-ons.
- A bridge version closes that window because it works on both Kubernetes versions, so it can go in before or after the control plane upgrade, as long as it lands before the nodes.
- Use
-targetto control the order, but watchdepends_on: give the node group its own version variable, and finish with a plainterraform apply.
Top comments (1)