DEV Community

Cover image for Your AWS Backup vault is locked. Who holds the key?
vadim albarov
vadim albarov

Posted on

Your AWS Backup vault is locked. Who holds the key?

AWS Backup Vault Lock in compliance mode is a strong guarantee. Once the cooling-off period ends, nobody can delete a recovery point early. Not an admin, not root, not AWS Support. For regulated workloads it is the closest thing to an offline tape you can get in the cloud.

We recently moved our RDS backups to this setup: an AWS Backup plan, a local vault, and a copy of every recovery point to a locked vault in a second region. Then a reviewer asked one question that undid the whole feeling of safety.

What protects the KMS key that encrypts the locked vault?

Nothing did. The vault key was a normal customer-managed key with the default policy. Any administrator in the account could disable it or schedule it for deletion. The recovery points would survive, locked and immutable, but nobody could decrypt them. A backup you cannot read is not a backup.

This post is about what we tried, what failed in a way the docs would have told us, and what we ended up with.

Step 1: Deny the destructive actions in the key policy

Three KMS actions can make a key unusable: kms:ScheduleKeyDeletion, kms:DisableKey, and kms:PutKeyPolicy. The third one matters because anyone who can rewrite the policy can remove the other two denies.

A key policy can deny these to everyone except a single break-glass principal:

statement {
  sid       = "DenyKeyDestruction"
  effect    = "Deny"
  actions   = ["kms:ScheduleKeyDeletion", "kms:DisableKey", "kms:PutKeyPolicy"]
  resources = ["*"]
  principals {
    type        = "AWS"
    identifiers = ["*"]
  }
  condition {
    test     = "StringNotEquals"
    variable = "aws:PrincipalArn"
    values   = [var.break_glass_role_arn]
  }
}
Enter fullscreen mode Exit fullscreen mode

Two details are easy to miss here.

First, this deny applies to the account root user too. The key policy is always evaluated, and an explicit deny in it beats the default root allow.

Second, KMS will refuse to create this key. The lockout safety check rejects any policy that denies PutKeyPolicy to the principal creating the key. You have to pass bypass_policy_lockout_safety_check = true. That flag exists for exactly this reason: if the exempt principal turns out to be unusable, the key becomes permanently unmanageable. Remember that sentence.

Step 2: Who is the exempt principal?

Our first exempt principal was the existing key-owners role in the same account. The reviewer rejected it in one line: any DevOps admin can assume that role. The deny adds one extra API call to an attack, not a wall.

Inside a single account, this is a hard problem. An admin with AdministratorAccess can edit any trust policy, attach any policy, and enrol a new MFA device on a stolen user. Whatever principal you exempt, the account can grant itself access to it.

So we wanted the exempt principal outside the account. A break-glass role in a separate, tightly controlled account, assumable only with MFA by two named people. Prod credentials could never reach it.

Step 3: The test that saved us

We did not apply this to the real key. We created a throwaway key with the same policy, plus an Allow statement granting the external break-glass role the key-management actions, and tried the role against it.

aws kms put-key-policy --key-id <test-key> --policy-name default --policy file://policy.json

An error occurred (AccessDeniedException) when calling the PutKeyPolicy operation:
This operation cannot be called cross account
Enter fullscreen mode Exit fullscreen mode

Same error for DisableKey, EnableKey, and GetKeyPolicy.

This is in the KMS documentation and has been for years. A key policy can grant principals in other accounts permission to use a key: encrypt, decrypt, generate data keys, describe it, create grants. It cannot let them administer it. Key-management operations are refused cross-account, whatever the key policy says.

Our throwaway key is now permanently unmanageable. Every principal in the account is denied by the key policy. The only exempt principal is in another account and cannot call the API. AWS Support is the documented way to recover such a key. It costs a dollar a month and encrypts nothing, so we left it as a monument.

Picture that on the real vault key, holding every copy of a production database.

Step 4: The design that works

The exempt principal must live in the same account as the key. The question becomes: can we make a same-account role that no same-account identity can assume?

Yes. Trust policies decide who can assume a role, and IAM permissions cannot override a trust policy. A role whose trust policy names only a principal from another account is unreachable from inside, no matter how many admin policies the caller holds.

resource "aws_iam_role" "dr_key_break_glass" {
  name = "${var.env}-dr-key-break-glass"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect    = "Allow"
      Action    = "sts:AssumeRole"
      Principal = { AWS = var.external_break_glass_role_arn }
    }]
  })
}

resource "aws_iam_role_policy" "dr_key_break_glass" {
  role = aws_iam_role.dr_key_break_glass.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect = "Allow"
      Action = [
        "kms:DescribeKey", "kms:GetKeyPolicy", "kms:PutKeyPolicy",
        "kms:DisableKey", "kms:EnableKey",
        "kms:ScheduleKeyDeletion", "kms:CancelKeyDeletion"
      ]
      Resource = aws_kms_key.dr_vault.arn
    }]
  })
}
Enter fullscreen mode Exit fullscreen mode

The key policy deny exempts this role's ARN. The creating principal is still denied PutKeyPolicy, so the bypass flag stays. The role is not listed as a principal anywhere in the key policy; the default root statement lets IAM policies take effect, and the inline policy above does the rest. That also avoids a race on first apply, where KMS rejects a key policy naming an IAM principal that was created seconds ago.

The legitimate path is two hops. A named user assumes the external break-glass role with MFA. That session assumes the in-account role. That session administers the key. Role chaining drops the MFA context, so the MFA condition lives on the first hop only.

We tested again on a second throwaway key:

Caller Action Result
Account admin AssumeRole into the in-account role AccessDenied
Account admin DisableKey, ScheduleKeyDeletion, PutKeyPolicy Explicit deny
Two-hop chain GetKeyPolicy, PutKeyPolicy OK
Two-hop chain DisableKey, then EnableKey OK
Two-hop chain ScheduleKeyDeletion OK, used for cleanup

What an attacker has left

One path. An admin in the account can edit the in-account role's trust policy to name themselves. That is an IAM write, iam:UpdateAssumeRolePolicy on a specific role name. IAM management events land in CloudTrail by default, so an EventBridge rule can page someone the moment it happens. Compare that with the original design, where the attack was a plain AssumeRole on a role the team uses every day.

To close that last path you need a control the account cannot undo. A Service Control Policy on the organizational unit, denying those IAM writes on the role and the three KMS actions on the key to every principal except the break-glass role, does it. An SCP mistake is reversible from the management account. A key policy mistake is not. A dedicated backup account that owns the vault and the key is the stronger end state. Both need the org management account, which for us meant a ticket to a partner. The in-account role shipped in the meantime.

Operational costs, stated plainly

  • Any Terraform apply that changes the key policy must run through the two-hop chain. Adding a new consumer does not, because the local admin role can still issue grants.
  • Destroying the module needs the chain too.
  • If the external break-glass role is deleted and recreated, the trust policy keeps the old principal ID. The next apply fixes it, but the chain is dead until then.
  • The deny matches the role ARN as a string. Recreating the in-account role with the same name keeps working. Renaming it freezes the policy.

Three habits

Read the current docs before designing an IAM control. The design came from Claude Code, which answered from memory. The cross-account rule was one paragraph away, on a page neither of us read.

Test anything that can lock a resource on something disposable. The lockout safety check is a warning, not an obstacle. If you bypass it on purpose, prove that the principal you trust can actually act, before the real key exists.

Ask what protects the thing that protects the thing. A locked vault, a versioned bucket, an immutable snapshot. Each one depends on a key, a policy, or a role that someone can still change. Follow the chain until you reach something the account itself cannot undo, and if you cannot reach it, know exactly which step is left open.

Top comments (0)