DEV Community

Cover image for Your AWS role can't tell a human from an agent anymore, part 1: the threat model and the identity problem

Your AWS role can't tell a human from an agent anymore, part 1: the threat model and the identity problem

Here's the honest thing nobody says at the "why hasn't AI coding taken over my org yet" keynote: the individual experience is genuinely great. Claude Code, Codex, OpenClaw — spin them up on a single machine and the productivity bump is real. You finish in an hour what used to take a day.

The step that dies is enterprise adoption.

Not because the tools stop working — but because a company running a proper SDLC can't just log in and let an agent have the run of the place at 9am for the next eight hours. And at the other extreme, you have the traditional banks and FSI shops that respond by locking everything behind a remote browser or a heavy VM sandbox — which works, but neuters the whole point of agentic tooling. The agent becomes a glorified browser that can't read your repo, can't run your tests, can't look at your infra. Useless for autonomous troubleshooting.

This series is written for the security or AI platform team on the receiving end of that request — the ones asked to hand out real account access to agents without betting the account on perfect model behavior. Over five posts I'll lay out four cheap layers that let agents do useful work inside a real AWS account, without either locking everything behind a remote browser or handing over the kingdom. Everything in this series was tested in a live account — I'll say so where it was.

This is not a series about agent harnesses, prompt engineering, or making a model smarter. It is about defense: how the security and infra team learns to protect an account when the same cloud identity can be wielded by a human or a machine. That is a genuinely new problem. For decades a principal was a person, and the only distinction IAM had to make was privilege level — senior engineer versus a fresh graduate. Same species, different permissions. Now the identity behind a single call can be you, or an agent running as you, and traditional RBAC has no word for that split.

There's a middle path. It isn't one silver bullet — it is a set of cheap layers that let agents do useful work inside a real cloud account without betting the account on perfect model behavior.

Layer 0 — The honest threat model

Before the tech, get the framing right. The question is never "can I trust the agent?". It's "what can the agent do even when I can't trust it?"

An AI agent on a developer laptop is, from the cloud's point of view, an extremely enthusiastic intern with a keyboard and your AWS session. It's not malicious — but it's also not careful. It guesses. It retries. It cleans up the wrong thing. And — critically — poisoned content will try to make it do things. Prompt injection into agentic CI is a real, growing attack surface: once an agent has write access, a crafted README or error message is a privilege escalation.

So the design principle is: defense in depth, where no single layer is trusted to save you. Every layer stops a different failure mode, and they should be cheap enough that you'll actually maintain them.

The identity chain: SSO to permission-set role to agent role with SourceIdentity, capped at one hour by role chaining

Layer 1 — Identity: stop handing out the long-lived keys

The most common mistake is the same one we made before AI agents existed, just accelerated: a long-lived aws_access_key_id sitting in ~/.aws/credentials, shared across the team, valid for months, with admin attached.

AI agents make this a single point of catastrophic failure — because they'll happily export it, paste it into a plan, or base64 it into a log. You cannot lose a key you never create.

What to do:

  • Use AWS IAM Identity Center (aws sso login) with temporary sessions instead of static keys. The credentials stop working when the configured session expires.
  • Scope the session the agent runs under — ideally a separate role with least privilege, not your own admin role.
  • Cap the agent role's maximum session duration. An hour is enough for a work session and short enough that a session copied off the host expires quickly; don't raise it toward the 12-hour ceiling just because you can.
  • Give the agent its own principal, not a shared human role. In the role's trust policy, allow sts:SetSourceIdentity and require the assume request to carry sts:SourceIdentity. Once the session exists, downstream requests expose it as aws:SourceIdentity, and CloudTrail records it. Session tags are separate: they require sts:TagSession, and only tags marked transitive survive role chaining.
  • If developers enter through an AWS IAM Identity Center permission set, have that SSO session assume a dedicated agent role you control. The generated permission-set role is the wrong place to build custom trust-policy machinery. That second assumption is role chaining, so the agent session is capped at one hour; make sure the credential provider can refresh it.
  • Never put real credentials in the repo, CI variables, or .env that an agent can read. The agent reads everything on the machine, remember.

A trust policy that does the source-identity part looks like this:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": {
      "AWS":
        "arn:aws:iam::111122223333:role/DeveloperTeamRole"
    },
    "Action": ["sts:AssumeRole", "sts:SetSourceIdentity"],
    "Condition": {
      "StringLike": { "sts:SourceIdentity": "claude-code-*" }
    }
  }]
}
Enter fullscreen mode Exit fullscreen mode

sts:SetSourceIdentity must be allowed on both sides — in the trust policy and in the permissions policy of the principal doing the assuming — or the assume call fails. And a source identity the agent's own wrapper chose is only as good as this condition: the trust policy is what makes it meaningful.

This removes the leaked long-lived key incident class, not credential theft. An agent can still copy a working session off the host. What you gain is a short, attributable blast window. If you need to cut live sessions, use IAM's Revoke active sessions operation: it attaches an AWSRevokeOlderSessions inline policy using aws:TokenIssueTime. That is per-role, which is another reason the agent needs its own role — and note it only kills sessions issued before the timestamp, so if the assume path is still alive the agent just re-authenticates. Revocation needs iam:PutRolePolicy, and removing the revocation policy needs iam:DeleteRolePolicy; the SCP in part 3 reserves those two calls for one protected incident-response role.

The safest useful pattern is broad metadata read access for troubleshooting, paired with narrowly scoped writes. Carve out the secret-bearing reads — pulling a secret value from Secrets Manager, decrypting SSM parameters, KMS decrypts, object reads on data buckets — none of these belong in a general troubleshooting role. Whatever the agent reads can leave the account in its model context.

Do not stop at the obvious read APIs. Interactive workload access can reach the same data indirectly when the workload role can read it: starting an SSM shell or sending a command, exec-ing into an ECS task, invoking a Lambda function, pushing an SSH key to an instance, or connecting to RDS with IAM auth. Each deserves an explicit, resource-scoped decision. Deny secret reads and these indirect paths by default in the agent role's identity policy, then allow only the exact troubleshooting workflows you intend — those decisions belong to the role, not the account-wide SCP, unless nobody in the whole sandbox account needs the capability.

Next up: the cheapest layer of all — making every agent call visible in CloudTrail, and how to roll that tag out across a fleet with MDM.

Sources

Top comments (0)