DEV Community

Vamshidher Reddy Jannapu Reddy
Vamshidher Reddy Jannapu Reddy

Posted on AI-assisted

My agent guardrail only lived on the laptop. So I compiled it into AWS.

For the past few weeks I've been building Aegis-DevOps, an open-source check that sits in front of the shell commands an AI coding agent runs. Before Claude Code, Codex, Copilot, Cursor or Gemini CLI runs kubectl, terraform or aws, a hook hands the command to Aegis, and Aegis blocks it if your signed policy says no. It works. It also has a hole I knew about from the start, and version 0.3.0 is my first attempt at closing it.

Disclosure: I drafted this article with help from an AI assistant and reviewed it myself. The commands and output below are from running aegis-devops 0.3.0 against its example policy.

The hole: a hook only sees commands

A pre-execution hook sees the command string the agent asks its shell tool to run. It doesn't see:

  • a Python script the agent writes and runs, which calls boto3 directly;
  • an SDK call from code the agent is "just testing";
  • a credential the agent found in a file and used from somewhere else.

In all three cases the destructive call reaches the AWS API without ever looking like aws s3 rb. Blocking on the client is the right first layer, because it's fast and it can explain itself to the model. It isn't the last layer.

The fix: let AWS refuse the call

AWS already has a control that sits above everything an identity inside an account can do: a Service Control Policy. An SCP attached to an account in an AWS Organization caps what any role in that account may do, and an agent with IAM rights inside the account can't detach it.

So 0.3.0 adds a compiler. It takes the same policy the hook enforces and writes it out as SCPs:

$ aegis compile aws --account 111122223333 --out build/aegis-aws
aegis: wrote 1 SCP(s) for account 111122223333 to build/aegis-aws
       (exact=2, not-applicable=27, partial=1)
aegis: agents.yaml is report-only: the SCPs are under report-only/ and are not for attaching;
       review coverage.md, then set enforcement: enforce
Enter fullscreen mode Exit fullscreen mode

The rule "never delete an S3 bucket or an RDS database" becomes an explicit Deny:

{
  "Effect": "Deny",
  "Action": ["rds:DeleteDBInstance", "rds:DeleteDBCluster", "rds:DeleteDBSnapshot",
             "rds:DeleteDBClusterSnapshot", "s3:DeleteBucket", "s3:DeleteObject",
             "s3:DeleteObjectVersion"],
  "Resource": ["*"],
  "Condition": {
    "ArnNotLike": { "aws:PrincipalArn": ["arn:aws:iam::111122223333:role/BreakGlass",
                                         "arn:aws:iam::111122223333:role/Deploy"] },
    "StringNotEqualsIfExists": { "aws:SourceIdentity": ["alice@example.com"] }
  }
}
Enter fullscreen mode Exit fullscreen mode

Now it doesn't matter whether the agent typed aws s3 rb, wrote a boto3 script, or found a key. If the call is made as an agent identity, AWS refuses it.

Who counts as an agent?

An API call carries an identity, not a hint about who was typing. So the compiler needs an identity model, in a signed agents.yaml. The default is deny-by-default: you list the identities you trust (people, your deploy role, CI bound to a specific workflow), plus at least one break-glass role, and everything else is treated as an agent.

I went with deny-by-default because the other way round, listing the agents, fails open. Forget one, or let an agent mint a new role, and the new one isn't restricted.

An agent will try to become trusted

If the only thing between an agent and DeleteDBInstance is "is this call from the Deploy role?", the obvious move is to become the Deploy role. So every compile also adds statements that stop exactly that:

  • no sts:AssumeRole into a trusted or break-glass role, and no iam:PassRole of one;
  • no edits to a trusted role's trust policy or permissions, and no new credentials for a trusted user;
  • no setting an exempt source identity on a session;
  • no organizations:LeaveOrganization, since an account that leaves the organization sheds its SCPs.

The part I care about most: saying what it doesn't cover

A guardrail that looks stronger than it is does more damage than no guardrail, because people stop checking. So every compile writes a coverage report, rule by rule:

### block-s3-bucket-delete: exact, mapping verified
- enforced: s3/bucket delete (s3:DeleteBucket, s3:DeleteObject, s3:DeleteObjectVersion)
- same effect, not covered: s3:PutLifecycleConfiguration: an expiration rule deletes every object
- same effect, not covered: s3:PutBucketPolicy: a deny-all bucket policy locks everyone out

### block-rds-delete: partial, mapping verified
- not enforced: 'rds/*' also matches resource types the action map does not know
- same effect, not covered: rds:ModifyDBInstance: BackupRetentionPeriod=0 deletes the automated backups
Enter fullscreen mode Exit fullscreen mode

"Same effect, not covered" is the list of other ways to reach the same outcome that the rule, as written, doesn't block. You can widen the rule, or accept the gap knowingly.

A few more choices along the same lines:

  • Report-only first. The first compile writes SCPs into report-only/ and tells you not to attach them. You switch agents.yaml to enforce (a signed change) once you've reviewed who would be restricted.
  • Aegis never holds cloud write credentials. It writes JSON, and you apply it with your own infrastructure as code.
  • Unverified rules never reach a compiled policy. The compiler starts from a verified snapshot, so a rule that was tampered with, forged, or written by someone not allowed to write it gets no vote, the same as on the client.
  • Every action mapping was tested in a sandbox AWS Organization, with live calls and the IAM policy simulator. A mapping added later starts out marked unverified.

What's still missing

It's a preview. AWS is in the 0.3.0 release; a Kubernetes target (ValidatingAdmissionPolicies scoped to agent identities) is merged and waiting for the next release, and GCP and Azure are designed but not built. SCPs don't apply to an organization's management account, so agents shouldn't run there at all. Rules with a time window or a rate limit stay client-side, because an SCP can't express them. And it's a young project: I'd much rather hear "this mapping is wrong" now than after someone relies on it.

Try it

pip install aegis-devops
aegis init .aegis
aegis compile aws --account <your-account-id> --out build/aegis-aws
Enter fullscreen mode Exit fullscreen mode

The client hook installs in one command for each agent (aegis install claude|codex|copilot|cursor|gemini|opencode), and there's a GitHub Action for Terraform and OpenTofu plans. Details on the compiler are in docs/server-side.md.

If you run agents with real AWS access, I'd like to know how you scope them today. Issues are open at github.com/moneytool/aegis-devops.

Top comments (1)

Collapse
 
mickyarun profile image
arun rajkumar •

Compiling the same policy down to an SCP is the right instinct, and deny-by-default on identities rather than listing the agents is the call I would have got wrong. Report-only first is also the correct default and most tools in this space do not ship it.

The hole I would put in the README is that both layers now come off the same file. The hook and the SCP share an author, so whatever the policy gets wrong, it gets wrong in both places in the same direction, and the second layer reports healthy while doing nothing. That is not defence in depth, it is one control deployed twice. It still buys you the thing the article is about, which is the boto3 path the hook cannot see, so the mechanism gain is real. The independence gain is zero and the compile step is what makes it zero.

The sharper version of that: an SCP written by someone who had not read the hook policy would occasionally disagree with it, and every disagreement would be a finding. Yours cannot disagree by construction.

The other thing I would look at is the SourceIdentity exemption. A human in the exempt list is the laundering route, because an agent running inside a session that already carries that human's source identity inherits the exemption. If your agents are ever invoked from a developer's own session, that condition hands them the bypass and the SCP still looks correct. Worth stating in the docs whether SourceIdentity is expected to survive into subprocesses in the setups you are targeting, because the answer decides whether the exemption is a convenience or a hole.

And one measurement, since the article's own premise is a guardrail that was not really running. Once an SCP is in enforce mode, the number that matters is how many calls it has actually denied, split between the deliberately tested ones and real ones. An SCP that has never denied anything real is indistinguishable from an SCP that was detached, and AWS will not tell you which. Zero real denies over a long window is the alarm, not the success metric.