DEV Community

Cover image for How to prove a Kubernetes ServiceAccount is unused before deleting it
Muhammad Hassaan Javed for Infraforge

Posted on Originally published at infraforge.agency

How to prove a Kubernetes ServiceAccount is unused before deleting it

To prove a Kubernetes ServiceAccount is unused before you delete it, you need four sources rather than one: the API server audit log, CloudTrail for IRSA, CloudTrail for EKS Pod Identity, and the last-used label on any legacy token Secrets. The permission set says what the account could have done, which is why a wrong answer is expensive; it is not evidence about whether it was used. An ISO 27001 surveillance auditor sampled 8 privileged service accounts across four PCI-scope namespaces on a client's EKS 1.29 clusters. Four had been silent for 148 days and still held get, list, create, update and delete on Secrets at cluster scope. The quarterly access review had marked all four as active. It had checked the account name, not the use.

Problem signals:

  • kubectl auth can-i list secrets -n kube-system --as=system:serviceaccount:: returns yes for an account nobody can name an owner for
  • A ServiceAccount has no owning Deployment, CronJob or Job template, and no matching resource anywhere in the IaC repo
  • The access-review spreadsheet says status: retained, justification: active with no last-use timestamp column at all
  • Your CloudWatch log group for the cluster has 30-day retention but your access policy defines dormant as 90 days with no activity
  • Flux or Argo reports resources present in the cluster and absent from git, and the alert routes to a Slack channel nobody reads

Does your audit log actually cover the dormancy window you are claiming?

Before you promise a window you cannot prove

The precondition that decides whether this review is possible at all.

The workpaper asked for evidence that unused privileged accounts were identified and removed over an observation window of three months. That question is unanswerable unless the API server audit log covers the window. On EKS that is the only half you can check, and it is worth knowing why: the control plane exposes per-type enable and disable for api, audit, authenticator, controllerManager and scheduler, and nothing that lets you set the audit policy. You can read it, though not through the API: AWS publishes the policy in its EKS best-practices guide, and it drops some requests entirely, core-group Events among them, so an account whose only work is writing Events reads as silent. So the preconditions you control are enablement per cluster and the retention ceiling below. On a self-managed control plane the policy is yours, and there the level it records at is a second thing to check before you commit to a date on a call with an auditor. On EKS the control-plane audit log is off by default and is enabled per cluster; verify it is on for every cluster in scope, not just the one someone remembered. Then check the retention on the destination log group, because that number is the real ceiling on what you can prove.

We hit this on one of the six clusters in the estate. Five had 400-day retention on their log group. One, stood up eight months earlier by a different team from a copied module, had the default of never-expire on the group but audit logging enabled only two months prior. For that cluster we could prove 61 days of silence, and the internal policy defined dormant at 90. There is no clever query that recovers data that was never written. We told the auditor exactly that, in writing, and treated the accounts on that cluster as unproven rather than dormant. Unproven is a finding you can close with a compensating action. A dormancy claim you cannot evidence, discovered later, is a much worse conversation.

# 1. Confirm the control-plane AUDIT log is on, not merely that logging exists.
#    clusterLogging is a list of {types, enabled} groups and audit falls in
#    exactly one of them, so read that group's flag. Filtering on types != null
#    tests only that the field exists, which it always does.
aws eks describe-cluster --name <cluster> \
  --query 'cluster.logging.clusterLogging[?contains(types, `audit`)].enabled'
# -> [true] audit logging is on. [false] it is off. A wrong cluster name or
#    region does NOT land here: DescribeCluster fails with ResourceNotFoundException
#    (HTTP 404, and clusters are Region specific) before the query runs. [] means no
#    logging group listed audit at all, so read cluster.logging raw before concluding.

# 2. Find the retention ceiling on the destination log group
aws logs describe-log-groups \
  --log-group-name-prefix /aws/eks/<cluster>/cluster \
  --query 'logGroups[].{name:logGroupName,retentionDays:retentionInDays}' \
  --output json

# retentionDays absent means never expire, and that is NOT an unbounded
# window. The ceiling is min(retention, time since audit logging was
# switched on), and nothing above recovers the second half.

# 3. When did audit logging actually start? The earliest audit stream is the
#    cheapest proxy; the UpdateClusterConfig event in CloudTrail is the other.
aws logs describe-log-streams \
  --log-group-name /aws/eks/<cluster>/cluster \
  --log-stream-name-prefix kube-apiserver-audit \
  --query 'logStreams[].{stream:logStreamName,first:firstEventTimestamp}' \
  --output json
# A never-expire group whose audit logging was turned on two months ago
# evidences 61 days, not infinity. And the earliest stream proves a start
# date, not continuity: audit logging can be switched off and back on
# mid-window, so read every UpdateClusterConfig inside the window before you
# claim all of it.
Enter fullscreen mode Exit fullscreen mode

Retention is not a detail you check afterwards, and it is not the whole ceiling either. The longest silence you are entitled to claim is min(retention, time since audit logging was enabled), which is why step 3 exists.

One more precondition. If the estate runs more than one cluster, run the whole procedure against all of them before you report anything. The same ServiceAccount name exists in every cluster that shares a manifest repo, and an account that is genuinely dormant in staging can be carrying a nightly job in production. We check every cluster in scope and key the results on cluster plus namespace plus name, never on name alone.

Why 24 sampled rows fail and 100% enumeration passes

Sampling is what put you in the finding

The finding came from the auditor's eight accounts. The client's own quarterly access review, the one that had marked all four as active, sampled 24 ServiceAccounts out of 431 across 17 non-system namespaces. Roughly one in eighteen. Sampling is a defensible technique when the population is uniform and you are estimating a rate. It is the wrong technique when the control objective is that no unused privileged account exists, because a sample that misses the bad account reports the same clean result as a population that has none. Worse, the sampling frame in that script was a hardcoded list of eight namespaces, and nine namespaces had come online since the list was written. Those nine were not undersampled. They had never been in scope.

So the first move is to stop sampling privileged accounts and enumerate all of them. That is only tractable if privileged has a machine-readable definition, so write one down in a file and version it. We keep it as a single YAML document in the security baseline repo, and it is the input to both the review script and the admission policy that stops the next one appearing. One definition, two consumers. When the definition changes, the thing you review and the thing you enforce move together, which removes an entire class of argument about whether an account counts.

# security-baseline/privileged-rbac-definition.yaml
# A binding makes its subject privileged if it grants any of these.
# Everything here is DATA, deliberately. A rule that lives in a comment is a
# rule neither the review script nor the admission policy can apply, which
# defeats the point of the two reading the same file.
writeVerbsOnSensitiveResources:
  resources: [secrets, configmaps, serviceaccounts]
  verbs: [create, update, patch, delete, deletecollection]
readVerbsOnSecrets:
  resources: [secrets]
  verbs: [get, list, watch]
escalationPaths:
  resources: [pods/exec, pods/attach, pods/portforward]
  verbs: [create]
tokenMinting:
  # a subresource is its own resource name in RBAC, so a rule on
  # serviceaccounts does NOT cover this one
  resources: [serviceaccounts/token]
  verbs: [create]
impersonation:
  resources: [users, groups, serviceaccounts]
  verbs: [impersonate]
rbacEscalation:
  apiGroups: [rbac.authorization.k8s.io]
  resources: [roles, clusterroles, rolebindings, clusterrolebindings]
  verbs: [bind, escalate, create, update, patch]
nodeProxy:
  resources: [nodes/proxy]
  verbs: [get, create]
wildcards:
  resources: ['*']
  verbs: ['*']
clusterScope:
  bindingKind: ClusterRoleBinding
  subjectKind: ServiceAccount
  privilegedRegardlessOfVerb: true
  exemptNamespaces: [kube-system, flux-system, karpenter]
Enter fullscreen mode Exit fullscreen mode

The reviewer and the admission controller read the same file. That is the point of writing it down.

Now resolve what each account can actually do. Do not build this from a subject grep over RoleBindings and ClusterRoleBindings. That was the approach in the original review and it has a hole you will not notice until an auditor finds it: a binding whose subject is a Group, not a ServiceAccount. A ClusterRoleBinding pointing at the group system:serviceaccounts:settlement grants that ClusterRole to every account in that namespace, present and future, and a grep for the account name returns nothing. The same shape applies to system:authenticated, but not the same blast radius: that group is held by every authenticated principal in the cluster, human users and nodes included, so a binding against it is never a namespace-scoped problem and never has a per-account remedy. Aggregated ClusterRoles are the second hole: the control plane fills their rules from every ClusterRole matching a label selector and overwrites anything written there by hand, so the role in your git repo is not the role the account holds.

Use the API server's own authorizer instead. kubectl auth can-i --list --as=system:serviceaccount:: asks the same code path that will run at request time, so it sees group bindings, aggregation and wildcards without you reimplementing any of it. The output is a column-aligned table of resources, non-resource URLs, resource names and verbs, which is fine for a human and awkward to parse, so we drive the machine-readable pass with targeted boolean checks and keep the --list output as the human-readable attachment in the workpaper. A single check, whether the account can list Secrets across all namespaces with -A, answers the cluster-scope question in one word. A check against one namespace such as kube-system cannot, because a RoleBinding inside that namespace answers yes too.

# The boolean the reviewer attests to. Do NOT key a sweep on the printed
# string: a denial can carry the authorizer's reason with it, and it does not
# come back as a zero exit. kubectl documents -q/--quiet as "suppress output
# and just return the exit code", so across 431 accounts use that and read
# the status, rather than comparing stdout under set -e.
kubectl auth can-i list secrets -A \
  --as=system:serviceaccount:reconciler:ledger-backfill

# The human-readable attachment for the workpaper. --list evaluates ONE
# namespace at a time and defaults to your current context namespace, for
# every subject; there is no "the account's namespace" it recovers on its
# own. Pass the namespace you care about and run it once per namespace the
# account is bound in. Cluster-scoped grants appear in every namespace's
# output; namespaced ones only in their own, so one run is never the whole
# picture.
kubectl auth can-i --list -n reconciler \
  --as=system:serviceaccount:reconciler:ledger-backfill

# Note the --as flag needs impersonate permission on your own
# identity. Run it from a break-glass role, not from CI.
Enter fullscreen mode Exit fullscreen mode

The authorizer answers the question your jq join was guessing at, including group bindings a name grep never sees.

Enumerating 431 accounts across six clusters took the script 11 minutes. The honest cost of this recommendation is not machine time, it is reviewer time. The corrected review put 66 privileged accounts in front of a human instead of 24 skimmable rows, and the first cycle took most of a working day rather than the hour the spreadsheet used to take. That is the trade. A review that takes an hour and cannot survive a sample is worth less than no review, because it produces a signed attestation that is wrong.

How do you prove a ServiceAccount is dormant across four evidence sources?

Zero API calls is not the same as zero use

Here is the mistake that would have taken production down, and it is the reason we no longer accept a single-source dormancy claim. Twenty-four of the 66 privileged accounts, on the five clusters that could evidence the full 148-day window, showed zero API server activity across it. Three of those twenty-four were in daily use. They were annotated with eks.amazonaws.com/role-arn, so their projected tokens were being exchanged at STS for AWS credentials and never sent to the API server at all. The workload was reading from S3 and writing to DynamoDB with an identity that originates in Kubernetes and leaves no trace in the Kubernetes audit log. Delete the ServiceAccount and the pod stops getting credentials on its next token refresh.

So the last-use test is a join across sources, and the account is dormant only when all of them are silent. Source one is the API server audit log for direct Kubernetes calls. Source two is CloudTrail for AssumeRoleWithWebIdentity against the role ARN in the account's annotation. Source three is EKS Pod Identity, and leaving it out is how this procedure deletes a live account. Pod Identity maps a role to a namespace and service account through an association rather than an annotation, and per AWS it does not use OIDC identity providers, so credentials are assumed by the EKS Auth service and CloudTrail records AssumeRoleForPodIdentity instead of AssumeRoleWithWebIdentity. An account serving a Pod Identity workload therefore has no annotation to find, nothing in the source-two query, and nothing in the API server audit log if its work is entirely against AWS. AWS version-gates it only on 1.28, which needs platform version eks.4; every other Kubernetes version supports it on all platform versions, so on the 1.29 cluster this article assumes it is live. Enumerate with aws eks list-pod-identity-associations --cluster-name <cluster> before you conclude anything from the annotation being absent. Source four, for any legacy token Secrets still hanging around on a cluster that was upgraded through 1.24, is the last-used label Kubernetes stamps on those Secrets; treat it as corroboration rather than proof, because it is dated to the day and only appears on the legacy token type. And read a missing label correctly: it is written only for tokens actually used since the tracking feature was enabled, so its absence means there is no tracking data, not that the token went unused. That does not on its own make the account Unproven, because any use of a legacy token against the API server already shows up in source one; it means source four abstains rather than votes. It only corroborates once you know when tracking began, and the cluster records that directly: the ConfigMap kube-system/kube-apiserver-legacy-service-account-token-tracking holds the timestamp. One blind spot survives all four sources: a token presented to something that validates it without the API server seeing this account’s name. Vault’s Kubernetes auth, when it has its own reviewer token, and mesh control planes call TokenReview under their own identity, and EKS logs tokenreviews at Metadata, so the reviewed account is not in the log; a JWT validated offline against the cluster’s OIDC keys never reaches the API server at all. Check the projected-token audiences in pod specs, and the roles in whatever consumes them, before you call an account silent.

# Source 1: API server audit log, CloudWatch Logs Insights
fields @timestamp, user.username, verb, objectRef.resource
| filter user.username like /^system:serviceaccount:/
| stats count() as calls, latest(@timestamp) as lastSeen
    by user.username
| sort lastSeen asc
# READ THIS RESULT BACKWARDS. It lists the accounts WITH activity, one row
# each. A dormant account emits no events, so it produces NO ROW and can
# never appear here, and the sort surfaces the least recently active rather
# than the silent. Your candidates are the enumerated population MINUS the
# usernames this returns. An empty result is not an empty finding; it is
# every enumerated account being a candidate.

# Source 2: IRSA. Does this account carry an IAM role annotation at all?
kubectl get sa -A -o json | jq -r '
  .items[]
  | select(.metadata.annotations["eks.amazonaws.com/role-arn"])
  | "\(.metadata.namespace)/\(.metadata.name)\t\(.metadata.annotations["eks.amazonaws.com/role-arn"])"'

# Source 3: Pod Identity. This one has NO annotation to find, so an account
# using it looks identical to an unused one in every query above. Skip this
# block and the procedure deletes a live account.
aws eks list-pod-identity-associations --cluster-name <cluster> \
  --query 'associations[].[namespace,serviceAccount,associationArn]' --output text

# For each association above, look for EKS Auth activity in the window.
# The event is AssumeRoleForPodIdentity, NOT AssumeRoleWithWebIdentity, so
# the source 2 CloudTrail query cannot see it.
#   eventSource = eks-auth.amazonaws.com
#   eventName   = AssumeRoleForPodIdentity

# Source 4: legacy token Secrets and their last-used label. A missing label
# means no tracking data, not unused. It abstains rather than votes: it never
# decides dormancy on its own, and corroborates only after you have read when
# tracking began:
#   kubectl get cm -n kube-system kube-apiserver-legacy-service-account-token-tracking
kubectl get secrets -A -o json | jq -r '
  .items[]
  | select(.type=="kubernetes.io/service-account-token")
  | "\(.metadata.namespace)/\(.metadata.name)\t\(.metadata.labels["kubernetes.io/legacy-token-last-used"] // "never-recorded")"'
Enter fullscreen mode Exit fullscreen mode

The account is only dormant when every source is silent AND every source was capable of speaking. Apply that to CloudTrail as strictly as to the audit log: Event history keeps 90 days of management events, so a window longer than that is unanswerable without a trail delivering to S3, and the Athena path assumes a table and partitions that somebody had to create. An account whose trail predates nothing and whose objects have aged out reads exactly like an account nobody used. Tell the two apart. list-pod-identity-associations enumerates every association on the cluster, so a pair that is not in it genuinely does not use Pod Identity: that is a dispositive negative and it does not block the dormant branch. A missing last-used label is the other kind, meaning tracking recorded nothing, which is why it is corroboration and never the deciding source.

For source two there is a trap in the obvious command. aws cloudtrail lookup-events reads the management event history, and that history covers the trailing 90 days in the one Region you run it in. Run it anywhere but the Region the workload’s STS calls land in and a live account returns nothing. Our dormancy claim was 148 days. A lookup-events query returns nothing for the older window and nothing looks exactly like proof of non-use if you are in a hurry on day 4 of a 7-day response window. Query the trail's S3 destination through Athena for anything older than 90 days, filtering on eventName equal to AssumeRoleWithWebIdentity and the role ARN from the annotation, and then on the service account name CloudTrail records with the call, because one role’s trust policy can admit several accounts and a hit on the ARN alone does not say which one assumed it. If your organization trail does not reach back far enough either, say so and mark the account unproven, the same as with a short log retention.

Every path to deletion runs through a source that could still say the account is alive.

Every path to deletion runs through a source that could still say the account is alive.

The revocation order that turns a bad call into a readable 403

Unbind before you delete

Never delete the ServiceAccount first. If you are wrong about dormancy, deleting the account invalidates the bound tokens its pods are carrying, and the workload starts failing authentication. What lands in someone's alerting is a 401 from a client library, at whatever layer that library chose to surface it, often as a generic connection error two hops from the actual cause. Remove the binding instead. Authentication still succeeds, authorization denies, and the API server produces a message that names the account, the verb, the resource and the scope. That is a bug report someone can act on in thirty seconds.

$ kubectl get secrets --all-namespaces
Error from server (Forbidden): secrets is forbidden: User "system:serviceaccount:reconciler:ledger-backfill" cannot list resource "secrets" in API group "" at the cluster scope
Enter fullscreen mode Exit fullscreen mode

This is the failure mode you want if you got it wrong. It names the account, the verb and the scope.

The revocation runs in two stages with a soak between them, and the length of the soak is a judgment call about your slowest scheduled work. Thirty days catches monthly close jobs. If the platform runs anything quarterly, and a card-issuing or ledger platform usually does, the safe answer is a full quarter for accounts whose name suggests a batch or reconciliation role. We ran 30 days on 18 accounts and a deliberate 95 days on the three that looked seasonal. The commit that removes a binding is a one-line revert, so the cost of the longer soak is calendar time and nothing else.

  • Delete the ClusterRoleBinding or RoleBinding in the IaC repo, one account per commit, so a revert is surgical. This path only exists for a per-account binding. Where the privilege comes from a group subject such as system:serviceaccounts:settlement, there is nothing account-shaped to delete, and removing the group binding revokes the ClusterRole from every account in that namespace at once, so the denial watch in step 3 then fires for accounts that were never in scope. Narrow the group binding to named subjects instead. RBAC permissions are purely additive and there is no deny rule that can exempt one account from a group grant. Route that account out of the one-per-commit path. Let the GitOps controller apply it rather than doing it by hand, because a hand-applied deletion is the same out-of-band change that created the problem.
  • Verify immediately: the boolean can-i check that returned yes now returns no. That two-word before-and-after is the cleanest evidence line the workpaper can carry.
  • Watch the audit log for denials attributed to the account. Filter on the authorization decision annotation set to forbid and group by username, and alert on any non-zero result for accounts in the soak window.
  • After the soak, confirm nothing still names the account: no Pod, and no Deployment, StatefulSet, DaemonSet, Job or CronJob template with that serviceAccountName in that namespace. An account can make no calls at all and still be required, for example with automountServiceAccountToken set to false, and admission rejects any new pod that references a missing account, so the next rollout fails, and a node drain completes but its replacement pods are never created. Then delete the ServiceAccount. The token controller deletes the token Secrets tied to it, so confirm by name that they are gone rather than deleting them yourself.
  • Record the deletion commit SHA and the cluster list against the account in the review packet. The auditor's follow-up question is when and by what change, not whether.

Back-out is the revert of that one commit, and it is genuinely fast because the binding is the only thing you removed in stage one. If a 403 attributed to a soaking account shows up, do not restore the original cluster-wide binding and walk away. Restore it, then replace it in the same week with a namespace-local RoleBinding carrying only the verbs the denial showed the workload actually needed. Take a denial seriously even though the account showed no calls for the whole window. Anything that runs less often than your lookback, an annual close job for example, looks exactly like a dormant account. Two steps can still catch it: the pod-template check before delete, if the job is defined in the cluster, and a soak longer than the job’s interval, which a 30-day soak is not. A job that an outside scheduler launches ad hoc gets past both, so ask the owning team about anything seasonal before you trust either. The denial log is a free specification for the permission set the workload should have had all along.

Two things we will not do, both learned expensively. We do not backdate the defective workpaper to make the record look complete, because the certification body asks directly whether timestamps were altered and the only survivable answer is no. The old review stays on file, wrong, and the corrected review is dated to the day it ran and filed as remediation evidence. And we do not report only the accounts the auditor sampled. Disclosing that the methodology defect touched 66 accounts across 17 namespaces kept the finding at minor. Letting the auditor discover the wider pattern in a follow-up sample is how a minor becomes a major, and a major on a surveillance audit means a return visit. We have written more about the pattern in the infrastructure audit readiness guide.

Common questions on privileged ServiceAccount access reviews

What the auditor asks after you hand over the spreadsheet

These are the follow-ups that arrived within two days of filing the corrected review packet. Short answers, because the auditor wanted short answers.

Step What it does
Can I claim dormancy from workload absence instead of audit logs? No, and this is exactly what produced the finding. A ServiceAccount with no owning Deployment or CronJob is a strong signal that nobody owns it, and it is not evidence that its token was never used. Anyone holding an exported token can call the API from a laptop, and that call appears in the audit log and nowhere else. Absence of a workload goes in the packet as a supporting note, never as the last-use test.
Does this procedure work on managed control planes other than EKS? The Kubernetes half does. The enumeration, the authorizer checks and the unbind-then-delete order are the same on any 1.29 cluster. The cloud half changes: the second evidence source is whatever records your workload identity exchange, so on GKE that is the Workload Identity path and on AKS it is the workload identity federation events. Name the source your platform actually writes to, and check its retention the same way.
How do we stop the next batch appearing? An admission policy that enforces the same definition the review reads. Any new ClusterRoleBinding whose subject is a ServiceAccount in a workload namespace is rejected whatever role it names, because the clusterScope rule already calls that privileged. A RoleBinding is rejected when its roleRef names a privileged role. A ValidatingAdmissionPolicy cannot open the role a binding points at, so generate the privileged-role list into the policy’s parameters, and generate it from the live cluster on a schedule rather than from git, because aggregated ClusterRoles change without a commit. Test it by replaying the exact binding shape of one of the accounts you just deleted, and confirm the rejection message points at your runbook. Pair it with a nightly job that flags any privileged account with no use in 60 days and opens a ticket against the owning team.
What if the drift alerts were already firing and nobody looked? That is the common case. Our GitOps controller had been reporting resources present in the cluster and absent from git for two years into a Slack channel with no owner. Route those alerts to a ticket queue with an assignee instead. The first scan after we did that surfaced 23 resources across six clusters that existed only in live state, of which 19 were cruft and 4 were adopted into git.

The question underneath all four is whether the control is a document or a mechanism. A quarterly review that a human performs from a spreadsheet is a document. A definition file that feeds both an enumeration script and an admission policy, plus a nightly dormancy job that opens tickets, is a mechanism, and it is what moved this finding from open to closed at the next surveillance. If your release tooling is also the thing making out-of-band changes, that is a related problem and we cover it in Kubernetes and CI/CD stabilization.

When the sample has already been drawn and the clock is running

If the workpaper response is due this week

The hard part of this work is not the queries. It is deciding, under a 7-day response window, how wide to open the disclosure when you have just discovered the review methodology was wrong for two cycles, and doing that while the accounts in question still hold cluster-wide access to Secrets in a PCI-scoped namespace. Get that call wrong in either direction and you either escalate a minor into a major or you delete a ServiceAccount that was quietly holding an IAM role for a nightly settlement job.

We have sat on both sides of that call: building the corrected enumeration, and then being in the room when the lead auditor asks what else the defect touched. What we bring is the evidence pipeline that makes the answer defensible, the join across audit log and cloud trail that stops you revoking something live, and a read on which findings an auditor will accept as remediation versus which will read as backfill. If your sample has already been drawn and the response is due, book an infrastructure review and we will build the 100% privileged-account enumeration with your team inside the response window, starting the same week.


Originally published at https://infraforge.agency/insights/prove-kubernetes-serviceaccount-unused-before-delete/.

If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — see /review.

Top comments (0)