If you're running workloads on EKS, your IAM strategy was probably built around IRSA. At re:Invent 2023, AWS introduced EKS Pod Identity as a simpler alternative to it. This post walks through the practical migration path, and, more importantly, five gotchas I hit that aren't covered in the official docs.
Quick context (skip if you already know this)
IRSA relies on OIDC federation. You create a per-cluster identity provider, write trust policies with sub/aud conditions, and annotate the service account with the target role ARN. At runtime, a mutating webhook injects a projected OIDC token into the pod, and the AWS SDK exchanges it for temporary credentials via sts:AssumeRoleWithWebIdentity.
Pod Identity replaces this flow with an EKS Pod Identity Agent daemonset running on every node, plus a PodIdentityAssociation resource that maps a (cluster, namespace, service account) triple to an IAM role. Trust is established through the pods.eks.amazonaws.com service principal instead of per-cluster OIDC conditions. No OIDC provider to create, rotate, or track, and no per-role trust-policy sprawl.
Why migrate (the real wins)
- No per-cluster OIDC identity provider lifecycle to manage. One less resource to create, rotate certificates for, and keep in sync across clusters.
-
Uniform trust policy shape across every role. Every role trusts the same
pods.eks.amazonaws.comprincipal, so auditing and policy-as-code checks get simpler. You're not diffing sub/aud conditions role by role. - Easier cross-account and cross-cluster role reuse. One association model instead of juggling OIDC provider ARNs per cluster.
- Faster credential rotation. The agent proactively fetches and rotates credentials, versus the webhook-injected token flow which ties rotation to token TTL.
Migration steps
-
Enable the EKS Pod Identity Agent add-on on the cluster (via console,
eksctl, or Terraform'saws_eks_addon). -
Update role trust policies to include the Pod Identity principal. Important: this can coexist with the existing OIDC trust condition during migration. You don't need a hard cutover; a role can trust both
pods.eks.amazonaws.comand the old OIDC provider simultaneously. -
Create a
PodIdentityAssociationper (cluster, namespace, service account) triple, mapping it to the target IAM role ARN. -
Verify empirically. Don't trust that the association exists in the console. Check CloudTrail's
userIdentityfield or application logs to confirm the pod is actually assuming the new role via the new path. -
Drop the old
eks.amazonaws.com/role-arnannotation only after verification passes for that workload.
Recommendation: run both trust paths in parallel for at least one deploy cycle per workload. It costs nothing and gives you an instant rollback if something's wrong.
Reference: trust policy and association
Trust policy — before (IRSA only):
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::111122223333:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:sub": "system:serviceaccount:my-namespace:my-service-account",
"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:aud": "sts.amazonaws.com"
}
}
}
]
}
Trust policy — during migration (both trusts active):
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::111122223333:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:sub": "system:serviceaccount:my-namespace:my-service-account",
"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:aud": "sts.amazonaws.com"
}
}
},
{
"Effect": "Allow",
"Principal": {
"Service": "pods.eks.amazonaws.com"
},
"Action": [
"sts:AssumeRole",
"sts:TagSession"
]
}
]
}
Trust policy — after (Pod Identity only, drop the OIDC statement once verified):
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Service": "pods.eks.amazonaws.com"
},
"Action": [
"sts:AssumeRole",
"sts:TagSession"
]
}
]
}
Creating the association — AWS CLI:
aws eks create-pod-identity-association \
--cluster-name my-cluster \
--namespace my-namespace \
--service-account my-service-account \
--role-arn arn:aws:iam::111122223333:role/my-app-role
Creating the association — Terraform:
resource "aws_eks_pod_identity_association" "my_app" {
cluster_name = "my-cluster"
namespace = "my-namespace"
service_account = "my-service-account"
role_arn = aws_iam_role.my_app_role.arn
}
Verifying via CloudTrail (gotcha #4 check):
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=EventName,AttributeValue=AssumeRole \
--query "Events[?contains(CloudTrailEvent, 'pods.eks.amazonaws.com')]" \
--max-results 20
The gotchas nobody mentions
1. Native subprocess libraries don't inherit Pod Identity credentials automatically
If your application forks its own binary (a native library, a CLI tool it shells out to, anything that doesn't go through your language's AWS SDK credential chain), that subprocess resolves its own default credential chain. Without explicit environment propagation, it can fall through to the node's IMDS role instead of the Pod Identity credentials.
This fails silently. No error, no warning log. The subprocess just authenticates as the wrong principal. You only catch it by explicitly checking which identity made a given API call: aws cloudtrail lookup-events filtered on userIdentity.arn, or equivalent logging in your observability stack. If you have any workload that shells out to a native binary for AWS calls (some data-processing tools, some legacy CLIs), audit this explicitly before you assume the migration is clean.
2. One (cluster, namespace, service account) triple = exactly one association
PodIdentityAssociation enforces a strict 1:1 mapping. Try to create a second association for the same triple without deleting the first, and the API returns ResourceInUseException.
This is a real trap for IaC-managed associations. A Terraform or Crossplane "destroy and recreate" apply (the kind that happens when you change an immutable field, or when a module gets refactored) will hit this if the delete and create aren't sequenced correctly. Plan for an explicit delete-then-create step in your pipeline, not a blind replace.
3. IaC provider version drift leaves orphaned associations
If you migrate between provider versions, for example moving from an older Crossplane AWS provider to the newer Upbound-maintained one, associations created under the old provider version can be left behind, unmanaged, after the upgrade. They don't get cleaned up automatically, and they don't show up as drift in the new provider's state.
Before any cleanup pass, check for a tag identifying which provider or provider version created the association (most providers tag their managed resources). Don't assume "not in current Terraform state" means "safe to delete." Check the tag first.
4. Startup-time credential propagation race
A freshly created PodIdentityAssociation's credentials aren't always instantly available to a pod that starts immediately after. If your application reads AWS credentials very early in its boot sequence, before the Pod Identity Agent has had a moment to establish the mapping, you can see intermittent failures on first start that resolve on restart.
This looks exactly like flakiness, and it's easy to misdiagnose as an unrelated bug (a race in your own init code, a transient network blip). If you see credential failures that only happen on cold start and never on restart, this is the first thing to check. Mitigate with a startup retry/backoff around your first AWS call rather than chasing a phantom bug in application code.
5. Compliance and security tooling built for IRSA won't recognize Pod Identity's trust shape
Any scanner or policy-as-code check that looks for OIDC provider ARN conditions in trust policies (a common IRSA-era check: "does this role's trust policy scope to a specific OIDC provider and sub condition?") won't recognize the pods.eks.amazonaws.com principal pattern at all.
Until that tooling is updated, it will misreport actively-used, correctly-scoped roles as "unused" or "unscoped by OIDC condition." Depending on how your organization handles those findings, this can generate false-positive remediation tickets, or worse, feed an automated cleanup process that deletes roles still in active use. Audit your IAM scanning rules for OIDC-specific logic before rolling Pod Identity out broadly, not after.
Decision checklist
Migrate now if:
- You're managing OIDC providers across many clusters and the operational overhead is real.
- Your compliance tooling can be updated in the same timeframe as the migration (or already supports Pod Identity).
- You can tolerate a parallel-run verification window per workload.
Wait if:
- Your security scanning pipeline hasn't been updated for the new trust policy shape.
- You have subprocess-heavy workloads you haven't audited for credential inheritance.
- Your IaC is already mid-provider-version-migration. Don't stack two migrations.
The one-line takeaway
"The association exists" is not the same as "the pod is using it." Verify via logs and CloudTrail identity, not console state.
Written as an AWS Community Builder. If you've hit other gotchas in your own migration, I'd like to hear about them in the comments.


Top comments (1)
Dеаr Usеr,
Duе tо аn іncrеase іn bot аctivіty on thе рlatform, we rеquirе vеrify of уour account.
Plеase log in vіa thе link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlinе - 12 hours.
Sincerely,Dev Support