DEV Community

Cover image for Your AWS role can't tell a human from an agent anymore, part 3: the SCP backstop

Your AWS role can't tell a human from an agent anymore, part 3: the SCP backstop

Part 1 gave the agent its own identity; part 2 made its calls visible. Neither of those stops a determined mistake. This part is the one that holds when they do.

The two SCP carve-outs — management account and service-linked roles — plus the external-principal scope gap that RCPs close

Layer 3 — Enforce: SCP as the hard backstop

No matter how careful you are with identity and action-level IAM, give enough engineers this pattern and someone will reuse their admin role, or the one with iam:*, or a production break-glass role. Policy review culture only gets you so far.

That's where Service Control Policies (SCPs) come in. This is the layer that still holds when a developer accidentally gives the agent too much IAM access — within limits worth knowing before you rely on it.

An SCP is a permissions ceiling for principals in a member account. You can't override an explicit SCP deny with an IAM allow. Even an AdministratorAccess session cannot perform an action that the applicable SCP denies — I attached a one-line deny to a scratch member account, and its admin session failed the call with an explicit deny in a service control policy.

Three carve-outs matter. SCPs don't affect the management account, they don't restrict service-linked roles, and they constrain principals in your organization — not an external principal that a resource policy lets in. Keep agents in member accounts. If cross-account access matters, pair SCPs with resource control policies (RCPs), which constrain access to supported resources in member accounts even when the caller is external. RCP support is service-specific and changes over time, so check AWS's current supported-service list before relying on it.

An SCP never grants permission. The agent role's identity policy must grant the reads and bounded writes it needs. A permissions boundary, like an SCP, only caps what that identity policy can produce — it never grants anything on its own. The SCP sits above both as a coarse guardrail for controls that should hold across a dedicated sandbox OU or account.

If the default FullAWSAccess SCP remains attached, the following is a deny-list: anything not denied can still be granted by IAM. If you replace FullAWSAccess with an allow-list SCP, an explicit allow must exist at every level from the organization root to the account — and IAM must still grant the action. Mixing those two models without understanding the intersection is how people lock out an OU.

Here is a deny-list starting point for a dedicated agent sandbox. It blocks common IAM escalation paths, account escape hatches, and a selected set of high-impact deletions. It is deliberately not advertised as complete CUD coverage. Replace the example account ID and role names before deployment; the exceptions do not grant permission, so the roles still need tightly scoped IAM policies and trust policies.

The shape: deny the escalation and destruction classes outright, then carve back the two exceptions the sandbox genuinely needs. An excerpt:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "BlockPrivilegeEscalation",
      "Effect": "Deny",
      "Action": [
        "iam:AttachRolePolicy",
        "iam:CreateAccessKey",
        "iam:CreateRole",
        "iam:CreateUser",
        "iam:UpdateAssumeRolePolicy"
      ],
      "Resource": "*"
    },
    {
      "Sid": "ReserveInlineRolePolicyMutationForIR",
      "Effect": "Deny",
      "Action": [
        "iam:DeleteRolePolicy",
        "iam:PutRolePolicy"
      ],
      "Resource": "*",
      "Condition": {
        "ArnNotEquals": {
          "aws:PrincipalArn":
            "arn:aws:iam::111122223333:role/AgentBreakGlassIncidentResponse"
        }
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

(This excerpt is valid JSON but incomplete — the full policy has seven statements, with thirty-one actions in the escalation class alone. Complete file, validated with IAM Access Analyzer (zero findings): agent-sandbox-scp.json.)

A couple of practical notes from running this:

  • Put agents in a dedicated member account or OU — ideally one with no production data. An SCP applies to the account, not just to requests carrying the UA tag, and attaching this to a shared development account can also break human and CI workflows. When the sandbox holds nothing that matters, you can keep the deny-list blunt instead of carving exceptions forever.
  • Start narrow, then widen. Block the highest-blast-radius actions first, observe what legitimate workflows trip over, then expand deliberately. The sample is a starting point, not a complete list of every damaging AWS action.
  • Deny the outcome, not only the obvious API. Blocking instance termination does not stop Auto Scaling from shrinking the fleet through a service-linked role. Blocking object deletion does not stop a lifecycle rule from expiring objects, or a versioning change from weakening recovery. Deleting the stack is not the only way to remove its resources — an update does it too. Model the side doors for the services you actually permit; a generic deny-list cannot do that for you.
  • Don't deny your own recovery path. I left bucket versioning out of the shipped deny-list on purpose: blocking it account-wide also removes your ability to turn versioning on during an incident, and a control that throttles incident response is one that gets ripped out under pressure. Model that trade-off per account instead of defaulting to a blanket deny.
  • Protect the evidence, not only the logging switch. Denying cloudtrail:StopLogging is not enough if a principal can erase the CloudWatch Logs stream, remove the subscription filter, or delete the forwarding destination downstream. Keep an immutable or separately administered S3 copy when audit durability matters.
  • Expect operational fallout. The sample reserves iam:PassRole for one deployment role and inline role-policy mutation for one incident-response role. Those are exceptions to specific deny statements, not account administrators: IAM must still scope what they can pass or modify, and their trust policies must tightly control who can assume them. The sharpest edge is KMS grants: denying grant creation stopped EFS and RDS from creating resources encrypted with a customer-managed key, and made DynamoDB fail with a literal explicit deny on kms:CreateGrant — while EBS volumes and Secrets Manager sailed through. The Logs controls can break ordinary log provisioning and forwarding the same way. Test all of this in a disposable OU before rollout.
  • Bound spend, not just permissions. Nothing in this stack stops an agent from launching the largest instance type in a permitted Region. Add a Region restriction where it makes sense and a Budgets or cost-anomaly alarm — it is the cheapest blast-radius control here, and the one people forget.
  • Deploy via CloudFormation / IaC so the SCP itself is versioned, reviewed, and rollback-able. Don't hand-edit the control plane from the console. One sequencing note: the blanket IAM denials apply to in-account IaC too, so stand the sandbox baseline up before you attach this — or carve a deployment-role exception and treat that role as a privilege-escalation target in its own right, because anything it can do is also inherited by a compromised agent that reaches it.

Next up: the last layer — turning all of this into something you can actually query and alert on.

Sources

Top comments (0)