DEV Community

Cover image for The Confused Deputy Problem in AWS: ExternalId, aws:Source* Keys, and RCPs
Rafagross
Rafagross

Posted on Originally published at rafael-gross.com

The Confused Deputy Problem in AWS: ExternalId, aws:Source* Keys, and RCPs

Most IAM incidents people talk about start with a leaked access key or a wildcard that was too generous. The confused deputy problem is different. Nothing leaks. Every API call is authorized. CloudTrail shows a party you trust doing exactly what your policy lets it do. The only thing wrong is who asked it to.

That is what makes it easy to miss in a review. This post covers the two variants AWS documents, cross-account and cross-service, the policy conditions that fix each one, and the details that decide whether the fix works in production.

The Problem

A confused deputy is a privileged party that gets talked into using its privileges for someone who doesn't have them. In AWS terms, an entity that is not allowed to perform an action coerces a more privileged entity into performing it.

The root cause is the same in every version of it. An IAM policy answers one question: who is calling? It does not ask on whose behalf the caller is acting, unless you write a condition that does.

Cross-account Cross-service
The actor Another customer of the same vendor Any account that can configure the calling service
The deputy The vendor's AWS account An AWS service principal such as cloudtrail.amazonaws.com
What the policy checks Is the caller the vendor's account? Is the caller that service?
What it never checks Which customer the vendor is acting for Which account or resource the service is acting for
Control sts:ExternalId aws:SourceArn, aws:SourceAccount, aws:SourceOrgID, aws:SourceOrgPaths

Cross-Account: The Vendor as Deputy

If you have onboarded a cost, monitoring, or security posture SaaS, you know the routine. You create a role in your account, trust the vendor's AWS account, and paste the role ARN into their console. From then on the vendor calls sts:AssumeRole and reads what the role allows.

The vendor does this for every customer, from the same AWS account. So picture another customer of that vendor, tenant B, pasting your role ARN into their own onboarding form. A role ARN is an identifier, not a credential. It shows up in tickets, in infrastructure code, in screenshots. Learning or guessing one is not hard.

Cross-account confused deputy attack sequence

Now look at the trust policy most onboarding guides hand you:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "AWS": "arn:aws:iam::999988887777:root" },
    "Action": "sts:AssumeRole"
  }]
}
Enter fullscreen mode Exit fullscreen mode

When the vendor assumes the role for tenant B, STS checks whether the caller is in the principal (yes, it is the vendor's account), whether the action is allowed (yes), and whether there is a condition to satisfy (there is none). The result is an Allow. The request the vendor sends for tenant B is identical to the one it sends for you. The policy has nothing to tell them apart with.

The Fix: sts:ExternalId

The external ID adds the missing fact to the request. The vendor assigns a unique value to each customer and passes that value on every AssumeRole call it makes for that customer. Your trust policy requires yours:

{
  "Effect": "Allow",
  "Principal": { "AWS": "arn:aws:iam::999988887777:root" },
  "Action": "sts:AssumeRole",
  "Condition": {
    "StringEquals": { "sts:ExternalId": "7f3c2b1e-58d4-4c0a-9e6f-b2d1c4a7a91e" }
  }
}
Enter fullscreen mode Exit fullscreen mode

And the vendor's side of it:

aws sts assume-role \
  --role-arn arn:aws:iam::111122223333:role/vendor-readonly \
  --role-session-name tenant-a-sync \
  --external-id 7f3c2b1e-58d4-4c0a-9e6f-b2d1c4a7a91e
Enter fullscreen mode Exit fullscreen mode

When tenant B submits your role ARN, the vendor sends tenant B's external ID, because that is who triggered the work. It doesn't match the condition on your role, no statement allows the call, and STS returns AccessDenied.

Same role ARN with and without the correct external ID

A few things about external IDs that are worth getting right:

  • The vendor generates it, not you. If customers could pick their own, tenant B would simply pick yours. AWS recommends one random string per customer AWS account.
  • It is not a secret. Anyone with permission to view the role can read it. Its job is to be unique per customer and outside the customer's control, not to be hidden.
  • Format. Between 2 and 1,224 characters, alphanumeric, no whitespace, plus a handful of symbols such as + = , . @ : / -.
  • If you are the vendor, test at onboarding. Try to assume the customer's role with and without the correct external ID. If it works without it, the trust policy isn't enforcing anything and you should not store that ARN.

The whole scheme depends on the vendor always sending the ID of the customer that triggered the request. A condition in your trust policy only does its job if the vendor does theirs. That's a fair question to ask during vendor review.

Cross-Service: An AWS Service as Deputy

The second variant has no third party at all. The deputy is an AWS service.

Many services reach into your resources using their service principal. CloudTrail writing to a central log bucket is the textbook case. The bucket policy has an Allow for cloudtrail.amazonaws.com, and that feels specific. It isn't. That principal is CloudTrail for every AWS account, not CloudTrail for yours.

So an account you have never heard of can create a trail, name your bucket as the destination, and CloudTrail will write there, because your bucket policy told S3 that CloudTrail is welcome.

Cross-service confused deputy with CloudTrail and S3

The useful part is that the service already tells the target on whose behalf it is acting. Requests made by AWS service principals carry that context in global condition keys. Your policy only has to test it.

Key The service is acting on behalf of Reach for it when
aws:SourceArn One specific resource You know the exact ARN: one trail, one topic, one fleet
aws:SourceAccount Any resource in one account The ARN has no account ID, or many resources share one grant
aws:SourceOrgID Any account in one organization A central resource serves every account, or you enforce it in an RCP
aws:SourceOrgPaths Accounts under an Organizations path Only part of the organization, such as one OU, should reach it

CloudTrail to S3: pin the bucket to one trail

This is the statement the CloudTrail documentation recommends, with aws:SourceArn set to the trail:

{
  "Sid": "AWSCloudTrailWrite",
  "Effect": "Allow",
  "Principal": { "Service": "cloudtrail.amazonaws.com" },
  "Action": "s3:PutObject",
  "Resource": "arn:aws:s3:::example-org-trail-logs/AWSLogs/111122223333/*",
  "Condition": {
    "StringEquals": {
      "s3:x-amz-acl": "bucket-owner-full-control",
      "aws:SourceArn": "arn:aws:cloudtrail:us-east-1:111122223333:trail/org-trail"
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Put the same aws:SourceArn condition on the s3:GetBucketAcl statement next to it. For an organization trail, the ARN is the trail in the management account, not in the member accounts.

S3 events to Lambda: when the ARN has no account ID

{
  "Sid": "allow-s3",
  "Effect": "Allow",
  "Principal": { "Service": "s3.amazonaws.com" },
  "Action": "lambda:InvokeFunction",
  "Resource": "arn:aws:lambda:us-east-2:111122223333:function:ingest",
  "Condition": {
    "StringEquals": { "aws:SourceAccount": "111122223333" },
    "ArnLike": { "aws:SourceArn": "arn:aws:s3:::example-bucket" }
  }
}
Enter fullscreen mode Exit fullscreen mode

Both keys are there for a reason. An S3 bucket ARN carries no account ID, and bucket names are global. If that bucket is ever deleted, another account can create one with the same name, and its events would still match aws:SourceArn. aws:SourceAccount ties the grant to the owner, not just the name. The CLI makes this easy to do properly: aws lambda add-permission takes both --source-arn and --source-account.

Service roles need it too

The trust policy of a role that a service assumes is the same kind of grant. Without a condition it says only "Systems Manager may assume me." The Systems Manager documentation recommends scoping it:

{
  "Effect": "Allow",
  "Principal": { "Service": "ssm.amazonaws.com" },
  "Action": "sts:AssumeRole",
  "Condition": {
    "StringEquals": { "aws:SourceAccount": "111122223333" },
    "ArnLike": { "aws:SourceArn": "arn:aws:ssm:us-east-1:111122223333:*" }
  }
}
Enter fullscreen mode Exit fullscreen mode

When you don't know the full ARN, or several resources use the role, keep the account and Region and use a wildcard for the rest.

Enforcing It Org-Wide with an RCP

Fixing bucket policies one by one doesn't scale, and it does nothing for the bucket someone creates next month. A resource control policy applies the rule centrally. This is the example from the IAM documentation:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "RCPEnforceConfusedDeputyProtectionForS3",
      "Effect": "Deny",
      "Principal": "*",
      "Action": ["s3:*"],
      "Resource": "*",
      "Condition": {
        "StringNotEqualsIfExists": { "aws:SourceOrgID": "o-exampleorgid" },
        "Null": { "aws:SourceAccount": "false" },
        "Bool": { "aws:PrincipalIsAWSService": "true" }
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

The three conditions are evaluated together, and each one has a job:

  • aws:PrincipalIsAWSService limits the deny to calls made by service principals. Your own users and roles are not affected.
  • The Null check on aws:SourceAccount applies the deny only when the service sends that key, so integrations that don't use it keep working.
  • aws:SourceOrgID is the actual test: deny when the service is not acting for your organization.

The Null check uses aws:SourceAccount rather than aws:SourceOrgID on purpose. If the request comes from an account that belongs to no organization, aws:SourceOrgID is absent. Testing for the account key keeps the control in force for that case.

Production Gotchas

  • An RCP never grants access. It sets a ceiling. Each resource policy still needs its own Allow for the service principal.
  • RCPs skip the management account. Resources there still need the condition in their own policies. RCPs also don't change the permissions of service-linked roles and don't apply to AWS managed KMS keys.
  • RCPs cover supported services only. The list has grown well past the original S3, STS, KMS, SQS, and Secrets Manager, but check it before you rely on one.
  • Both keys in one statement must agree. If the aws:SourceArn value contains an account ID, it has to be the same account as aws:SourceAccount.
  • Not every integration sends these keys. Requiring a key that a service doesn't populate breaks that integration. The service's own documentation says which keys it supports.
  • Some services have their own mechanism. AWS KMS, for example, uses the encryption context together with the key grant.
  • The external ID is only as good as the vendor's handling of it. A vendor that lets customers type their own, or that sends a shared value, has not solved anything.

Finding What You Already Have

A quick way to list roles that trust an AWS principal without requiring an external ID:

aws iam list-roles --output json | jq -r '.Roles[]
  | select(any(.AssumeRolePolicyDocument.Statement[];
      .Effect == "Allow" and (.Principal.AWS? != null)
      and ((.Condition // {} | tostring) | contains("sts:ExternalId") | not)))
  | .Arn'
Enter fullscreen mode Exit fullscreen mode

This is a starting point, not an audit. Roles trusted by your own accounts will show up and are usually fine. Any role that trusts a vendor's account should carry the condition.

Two more places to look:

  • IAM Access Analyzer. External access findings cover role trust policies, S3 buckets, KMS keys, Lambda functions, SQS queues, and SNS topics, among others. Treat each finding as a prompt to read the policy's conditions.
  • CloudTrail. A cross-account AssumeRole is logged in both accounts, and sharedEventID ties the two records together. That is how you confirm which external principal used a role, and when.

Checklist

  1. Every trust policy for a vendor account requires an sts:ExternalId that the vendor issued.
  2. Every resource policy and service-role trust policy that names a service principal has an aws:SourceArn, aws:SourceAccount, aws:SourceOrgID, or aws:SourceOrgPaths condition.
  3. An RCP enforces aws:SourceOrgID for the services that support it, and the management account is handled separately.
  4. For each service in use, the documentation confirms which keys it supports and whether it has its own mechanism.

The principle underneath all four is short. A policy that trusts a deputy should also say on whose behalf the deputy may act.

Five rules for aws:Source condition keys

Sources

Top comments (1)

Collapse
 
sgaggjhkjh profile image
sgaggjhkjh •

The CloudTrail to S3 example made the cross-service variant concrete for me in a way the abstract definition never did. I had not realized how many trust policies quietly trust a service principal shared across every AWS account. The gotcha about not every integration sending these keys is worth its weight, because that is exactly the kind of thing you discover at 2 AM. Do you have a rule of thumb for choosing between aws:SourceArn and aws:SourceAccount when you do not know the exact resource ARN yet?