<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Muhammad Hassaan Javed</title>
    <description>The latest articles on DEV Community by Muhammad Hassaan Javed (@itxcrusher).</description>
    <link>https://dev.to/itxcrusher</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1254167%2F39e4646e-64c2-4e1c-a494-98a603ad7e41.jpeg</url>
      <title>DEV Community: Muhammad Hassaan Javed</title>
      <link>https://dev.to/itxcrusher</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/itxcrusher"/>
    <language>en</language>
    <item>
      <title>How to prove a Kubernetes ServiceAccount is unused before deleting it</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Fri, 25 Sep 2026 17:03:30 +0000</pubDate>
      <link>https://dev.to/infraforge/how-to-prove-a-kubernetes-serviceaccount-is-unused-before-deleting-it-1n9e</link>
      <guid>https://dev.to/infraforge/how-to-prove-a-kubernetes-serviceaccount-is-unused-before-deleting-it-1n9e</guid>
      <description>&lt;p&gt;To prove a Kubernetes ServiceAccount is unused before you delete it, you need four sources rather than one: the API server audit log, CloudTrail for IRSA, CloudTrail for EKS Pod Identity, and the last-used label on any legacy token Secrets. The permission set says what the account could have done, which is why a wrong answer is expensive; it is not evidence about whether it was used. An ISO 27001 surveillance auditor sampled 8 privileged service accounts across four PCI-scope namespaces on a client's EKS 1.29 clusters. Four had been silent for 148 days and still held get, list, create, update and delete on Secrets at cluster scope. The quarterly access review had marked all four as active. It had checked the account name, not the use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;kubectl auth can-i list secrets -n kube-system --as=system:serviceaccount:: returns yes for an account nobody can name an owner for&lt;/li&gt;
&lt;li&gt;A ServiceAccount has no owning Deployment, CronJob or Job template, and no matching resource anywhere in the IaC repo&lt;/li&gt;
&lt;li&gt;The access-review spreadsheet says status: retained, justification: active with no last-use timestamp column at all&lt;/li&gt;
&lt;li&gt;Your CloudWatch log group for the cluster has 30-day retention but your access policy defines dormant as 90 days with no activity&lt;/li&gt;
&lt;li&gt;Flux or Argo reports resources present in the cluster and absent from git, and the alert routes to a Slack channel nobody reads&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Does your audit log actually cover the dormancy window you are claiming?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Before you promise a window you cannot prove&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The precondition that decides whether this review is possible at all.&lt;/p&gt;

&lt;p&gt;The workpaper asked for evidence that unused privileged accounts were identified and removed over an observation window of three months. That question is unanswerable unless the API server audit log covers the window. On EKS that is the only half you can check, and it is worth knowing why: the control plane exposes per-type enable and disable for &lt;code&gt;api&lt;/code&gt;, &lt;code&gt;audit&lt;/code&gt;, &lt;code&gt;authenticator&lt;/code&gt;, &lt;code&gt;controllerManager&lt;/code&gt; and &lt;code&gt;scheduler&lt;/code&gt;, and nothing that lets you set the audit policy. You can read it, though not through the API: AWS publishes the policy in its EKS best-practices guide, and it drops some requests entirely, core-group Events among them, so an account whose only work is writing Events reads as silent. So the preconditions you control are enablement per cluster and the retention ceiling below. On a self-managed control plane the policy is yours, and there the level it records at is a second thing to check before you commit to a date on a call with an auditor. On EKS the control-plane audit log is off by default and is enabled per cluster; verify it is on for every cluster in scope, not just the one someone remembered. Then check the retention on the destination log group, because that number is the real ceiling on what you can prove.&lt;/p&gt;

&lt;p&gt;We hit this on one of the six clusters in the estate. Five had 400-day retention on their log group. One, stood up eight months earlier by a different team from a copied module, had the default of never-expire on the group but audit logging enabled only two months prior. For that cluster we could prove 61 days of silence, and the internal policy defined dormant at 90. There is no clever query that recovers data that was never written. We told the auditor exactly that, in writing, and treated the accounts on that cluster as unproven rather than dormant. Unproven is a finding you can close with a compensating action. A dormancy claim you cannot evidence, discovered later, is a much worse conversation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# 1. Confirm the control-plane AUDIT log is on, not merely that logging exists.
#    clusterLogging is a list of {types, enabled} groups and audit falls in
#    exactly one of them, so read that group's flag. Filtering on types != null
#    tests only that the field exists, which it always does.
aws eks describe-cluster --name &amp;lt;cluster&amp;gt; \
  --query 'cluster.logging.clusterLogging[?contains(types, `audit`)].enabled'
# -&amp;gt; [true] audit logging is on. [false] it is off. A wrong cluster name or
#    region does NOT land here: DescribeCluster fails with ResourceNotFoundException
#    (HTTP 404, and clusters are Region specific) before the query runs. [] means no
#    logging group listed audit at all, so read cluster.logging raw before concluding.

# 2. Find the retention ceiling on the destination log group
aws logs describe-log-groups \
  --log-group-name-prefix /aws/eks/&amp;lt;cluster&amp;gt;/cluster \
  --query 'logGroups[].{name:logGroupName,retentionDays:retentionInDays}' \
  --output json

# retentionDays absent means never expire, and that is NOT an unbounded
# window. The ceiling is min(retention, time since audit logging was
# switched on), and nothing above recovers the second half.

# 3. When did audit logging actually start? The earliest audit stream is the
#    cheapest proxy; the UpdateClusterConfig event in CloudTrail is the other.
aws logs describe-log-streams \
  --log-group-name /aws/eks/&amp;lt;cluster&amp;gt;/cluster \
  --log-stream-name-prefix kube-apiserver-audit \
  --query 'logStreams[].{stream:logStreamName,first:firstEventTimestamp}' \
  --output json
# A never-expire group whose audit logging was turned on two months ago
# evidences 61 days, not infinity. And the earliest stream proves a start
# date, not continuity: audit logging can be switched off and back on
# mid-window, so read every UpdateClusterConfig inside the window before you
# claim all of it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Retention is not a detail you check afterwards, and it is not the whole ceiling either. The longest silence you are entitled to claim is min(retention, time since audit logging was enabled), which is why step 3 exists.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One more precondition. If the estate runs more than one cluster, run the whole procedure against all of them before you report anything. The same ServiceAccount name exists in every cluster that shares a manifest repo, and an account that is genuinely dormant in staging can be carrying a nightly job in production. We check every cluster in scope and key the results on cluster plus namespace plus name, never on name alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 24 sampled rows fail and 100% enumeration passes
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Sampling is what put you in the finding&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The finding came from the auditor's eight accounts. The client's own quarterly access review, the one that had marked all four as active, sampled 24 ServiceAccounts out of 431 across 17 non-system namespaces. Roughly one in eighteen. Sampling is a defensible technique when the population is uniform and you are estimating a rate. It is the wrong technique when the control objective is that no unused privileged account exists, because a sample that misses the bad account reports the same clean result as a population that has none. Worse, the sampling frame in that script was a hardcoded list of eight namespaces, and nine namespaces had come online since the list was written. Those nine were not undersampled. They had never been in scope.&lt;/p&gt;

&lt;p&gt;So the first move is to stop sampling privileged accounts and enumerate all of them. That is only tractable if privileged has a machine-readable definition, so write one down in a file and version it. We keep it as a single YAML document in the security baseline repo, and it is the input to both the review script and the admission policy that stops the next one appearing. One definition, two consumers. When the definition changes, the thing you review and the thing you enforce move together, which removes an entire class of argument about whether an account counts.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# security-baseline/privileged-rbac-definition.yaml
# A binding makes its subject privileged if it grants any of these.
# Everything here is DATA, deliberately. A rule that lives in a comment is a
# rule neither the review script nor the admission policy can apply, which
# defeats the point of the two reading the same file.
writeVerbsOnSensitiveResources:
  resources: [secrets, configmaps, serviceaccounts]
  verbs: [create, update, patch, delete, deletecollection]
readVerbsOnSecrets:
  resources: [secrets]
  verbs: [get, list, watch]
escalationPaths:
  resources: [pods/exec, pods/attach, pods/portforward]
  verbs: [create]
tokenMinting:
  # a subresource is its own resource name in RBAC, so a rule on
  # serviceaccounts does NOT cover this one
  resources: [serviceaccounts/token]
  verbs: [create]
impersonation:
  resources: [users, groups, serviceaccounts]
  verbs: [impersonate]
rbacEscalation:
  apiGroups: [rbac.authorization.k8s.io]
  resources: [roles, clusterroles, rolebindings, clusterrolebindings]
  verbs: [bind, escalate, create, update, patch]
nodeProxy:
  resources: [nodes/proxy]
  verbs: [get, create]
wildcards:
  resources: ['*']
  verbs: ['*']
clusterScope:
  bindingKind: ClusterRoleBinding
  subjectKind: ServiceAccount
  privilegedRegardlessOfVerb: true
  exemptNamespaces: [kube-system, flux-system, karpenter]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The reviewer and the admission controller read the same file. That is the point of writing it down.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Now resolve what each account can actually do. Do not build this from a subject grep over RoleBindings and ClusterRoleBindings. That was the approach in the original review and it has a hole you will not notice until an auditor finds it: a binding whose subject is a Group, not a ServiceAccount. A ClusterRoleBinding pointing at the group system:serviceaccounts:settlement grants that ClusterRole to every account in that namespace, present and future, and a grep for the account name returns nothing. The same shape applies to &lt;code&gt;system:authenticated&lt;/code&gt;, but not the same blast radius: that group is held by every authenticated principal in the cluster, human users and nodes included, so a binding against it is never a namespace-scoped problem and never has a per-account remedy. Aggregated ClusterRoles are the second hole: the control plane fills their rules from every ClusterRole matching a label selector and overwrites anything written there by hand, so the role in your git repo is not the role the account holds.&lt;/p&gt;

&lt;p&gt;Use the API server's own authorizer instead. kubectl auth can-i --list --as=system:serviceaccount:: asks the same code path that will run at request time, so it sees group bindings, aggregation and wildcards without you reimplementing any of it. The output is a column-aligned table of resources, non-resource URLs, resource names and verbs, which is fine for a human and awkward to parse, so we drive the machine-readable pass with targeted boolean checks and keep the --list output as the human-readable attachment in the workpaper. A single check, whether the account can list Secrets across all namespaces with &lt;code&gt;-A&lt;/code&gt;, answers the cluster-scope question in one word. A check against one namespace such as kube-system cannot, because a RoleBinding inside that namespace answers yes too.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# The boolean the reviewer attests to. Do NOT key a sweep on the printed
# string: a denial can carry the authorizer's reason with it, and it does not
# come back as a zero exit. kubectl documents -q/--quiet as "suppress output
# and just return the exit code", so across 431 accounts use that and read
# the status, rather than comparing stdout under set -e.
kubectl auth can-i list secrets -A \
  --as=system:serviceaccount:reconciler:ledger-backfill

# The human-readable attachment for the workpaper. --list evaluates ONE
# namespace at a time and defaults to your current context namespace, for
# every subject; there is no "the account's namespace" it recovers on its
# own. Pass the namespace you care about and run it once per namespace the
# account is bound in. Cluster-scoped grants appear in every namespace's
# output; namespaced ones only in their own, so one run is never the whole
# picture.
kubectl auth can-i --list -n reconciler \
  --as=system:serviceaccount:reconciler:ledger-backfill

# Note the --as flag needs impersonate permission on your own
# identity. Run it from a break-glass role, not from CI.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The authorizer answers the question your jq join was guessing at, including group bindings a name grep never sees.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Enumerating 431 accounts across six clusters took the script 11 minutes. The honest cost of this recommendation is not machine time, it is reviewer time. The corrected review put 66 privileged accounts in front of a human instead of 24 skimmable rows, and the first cycle took most of a working day rather than the hour the spreadsheet used to take. That is the trade. A review that takes an hour and cannot survive a sample is worth less than no review, because it produces a signed attestation that is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you prove a ServiceAccount is dormant across four evidence sources?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Zero API calls is not the same as zero use&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here is the mistake that would have taken production down, and it is the reason we no longer accept a single-source dormancy claim. Twenty-four of the 66 privileged accounts, on the five clusters that could evidence the full 148-day window, showed zero API server activity across it. Three of those twenty-four were in daily use. They were annotated with eks.amazonaws.com/role-arn, so their projected tokens were being exchanged at STS for AWS credentials and never sent to the API server at all. The workload was reading from S3 and writing to DynamoDB with an identity that originates in Kubernetes and leaves no trace in the Kubernetes audit log. Delete the ServiceAccount and the pod stops getting credentials on its next token refresh.&lt;/p&gt;

&lt;p&gt;So the last-use test is a join across sources, and the account is dormant only when all of them are silent. Source one is the API server audit log for direct Kubernetes calls. Source two is CloudTrail for AssumeRoleWithWebIdentity against the role ARN in the account's annotation. Source three is EKS Pod Identity, and leaving it out is how this procedure deletes a live account. Pod Identity maps a role to a namespace and service account through an association rather than an annotation, and per AWS it does not use OIDC identity providers, so credentials are assumed by the EKS Auth service and CloudTrail records &lt;code&gt;AssumeRoleForPodIdentity&lt;/code&gt; instead of &lt;code&gt;AssumeRoleWithWebIdentity&lt;/code&gt;. An account serving a Pod Identity workload therefore has no annotation to find, nothing in the source-two query, and nothing in the API server audit log if its work is entirely against AWS. AWS version-gates it only on 1.28, which needs platform version &lt;code&gt;eks.4&lt;/code&gt;; every other Kubernetes version supports it on all platform versions, so on the 1.29 cluster this article assumes it is live. Enumerate with &lt;code&gt;aws eks list-pod-identity-associations --cluster-name &amp;lt;cluster&amp;gt;&lt;/code&gt; before you conclude anything from the annotation being absent. Source four, for any legacy token Secrets still hanging around on a cluster that was upgraded through 1.24, is the last-used label Kubernetes stamps on those Secrets; treat it as corroboration rather than proof, because it is dated to the day and only appears on the legacy token type. And read a missing label correctly: it is written only for tokens actually used since the tracking feature was enabled, so its absence means there is no tracking data, not that the token went unused. That does not on its own make the account Unproven, because any use of a legacy token against the API server already shows up in source one; it means source four abstains rather than votes. It only corroborates once you know when tracking began, and the cluster records that directly: the ConfigMap &lt;code&gt;kube-system/kube-apiserver-legacy-service-account-token-tracking&lt;/code&gt; holds the timestamp. One blind spot survives all four sources: a token presented to something that validates it without the API server seeing this account’s name. Vault’s Kubernetes auth, when it has its own reviewer token, and mesh control planes call TokenReview under their own identity, and EKS logs tokenreviews at Metadata, so the reviewed account is not in the log; a JWT validated offline against the cluster’s OIDC keys never reaches the API server at all. Check the projected-token audiences in pod specs, and the roles in whatever consumes them, before you call an account silent.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Source 1: API server audit log, CloudWatch Logs Insights
fields @timestamp, user.username, verb, objectRef.resource
| filter user.username like /^system:serviceaccount:/
| stats count() as calls, latest(@timestamp) as lastSeen
    by user.username
| sort lastSeen asc
# READ THIS RESULT BACKWARDS. It lists the accounts WITH activity, one row
# each. A dormant account emits no events, so it produces NO ROW and can
# never appear here, and the sort surfaces the least recently active rather
# than the silent. Your candidates are the enumerated population MINUS the
# usernames this returns. An empty result is not an empty finding; it is
# every enumerated account being a candidate.

# Source 2: IRSA. Does this account carry an IAM role annotation at all?
kubectl get sa -A -o json | jq -r '
  .items[]
  | select(.metadata.annotations["eks.amazonaws.com/role-arn"])
  | "\(.metadata.namespace)/\(.metadata.name)\t\(.metadata.annotations["eks.amazonaws.com/role-arn"])"'

# Source 3: Pod Identity. This one has NO annotation to find, so an account
# using it looks identical to an unused one in every query above. Skip this
# block and the procedure deletes a live account.
aws eks list-pod-identity-associations --cluster-name &amp;lt;cluster&amp;gt; \
  --query 'associations[].[namespace,serviceAccount,associationArn]' --output text

# For each association above, look for EKS Auth activity in the window.
# The event is AssumeRoleForPodIdentity, NOT AssumeRoleWithWebIdentity, so
# the source 2 CloudTrail query cannot see it.
#   eventSource = eks-auth.amazonaws.com
#   eventName   = AssumeRoleForPodIdentity

# Source 4: legacy token Secrets and their last-used label. A missing label
# means no tracking data, not unused. It abstains rather than votes: it never
# decides dormancy on its own, and corroborates only after you have read when
# tracking began:
#   kubectl get cm -n kube-system kube-apiserver-legacy-service-account-token-tracking
kubectl get secrets -A -o json | jq -r '
  .items[]
  | select(.type=="kubernetes.io/service-account-token")
  | "\(.metadata.namespace)/\(.metadata.name)\t\(.metadata.labels["kubernetes.io/legacy-token-last-used"] // "never-recorded")"'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The account is only dormant when every source is silent AND every source was capable of speaking. Apply that to CloudTrail as strictly as to the audit log: Event history keeps 90 days of management events, so a window longer than that is unanswerable without a trail delivering to S3, and the Athena path assumes a table and partitions that somebody had to create. An account whose trail predates nothing and whose objects have aged out reads exactly like an account nobody used. Tell the two apart. &lt;code&gt;list-pod-identity-associations&lt;/code&gt; enumerates every association on the cluster, so a pair that is not in it genuinely does not use Pod Identity: that is a dispositive negative and it does not block the dormant branch. A missing last-used label is the other kind, meaning tracking recorded nothing, which is why it is corroboration and never the deciding source.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For source two there is a trap in the obvious command. aws cloudtrail lookup-events reads the management event history, and that history covers the trailing 90 days in the one Region you run it in. Run it anywhere but the Region the workload’s STS calls land in and a live account returns nothing. Our dormancy claim was 148 days. A lookup-events query returns nothing for the older window and nothing looks exactly like proof of non-use if you are in a hurry on day 4 of a 7-day response window. Query the trail's S3 destination through Athena for anything older than 90 days, filtering on eventName equal to AssumeRoleWithWebIdentity and the role ARN from the annotation, and then on the service account name CloudTrail records with the call, because one role’s trust policy can admit several accounts and a hit on the ARN alone does not say which one assumed it. If your organization trail does not reach back far enough either, say so and mark the account unproven, the same as with a short log retention.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJx9VdFu4zYQfM9XDA7oPSVpcumTU1yh2Mkl6F3i2E58heEHWlxb7FFcgVzZEZz-e0HKkZ1D0TfJpGZnZmfXS8ubvFBeMBkcAdls6M3aWFqRxjjD0nOJ87OzX0CuLskrMezmODn5jKttX1kbYByy4R0C-TV5qFobgeXV7wv_6-eNkcI4SEHYGKd588c_R8BV_P61ofCK0WxEoozrQYlQEOS19-QEdaB5d9XxK_rbzDkWJaQRYaEc7rJv8GwJ2eg-Iff3yINtFkJd0ogtTY0UU1rcaXJipImU-5ZrPfHK2ER0bRQyKcipYziOAvhHXZ3QmpyEA-jIZLgdskYHpkLg3CRjsGSf4KQwAU6VFCqVE5TTyR8Tn_OcaycJc3BgRPeaShwBw_3h44GWG_ZD1v8jhWMbouNBle9sH3bw19sBUzhoC5ZGYFwwmnD1MLmFJ4kF2CXIyD9eTs09IacWlsrYJa2EEvbjeyWP-1JHwHX39jR7cpXnNbkePFXsBSog1HlxnAppTu7nVpkSmn2pnMw7hAQ_abnvRUPlUitrG_Qfnq9Hh2Fre9FeWtExxhewZkl5k1s6hkQZqJQXE5WGqGOyp9q9pbLP224w2qFQWBinjVtBCiWpUmo4Jrd343ddfu4wp9tpYfICK891lc6m6exDaIJQ2dtlZPdx6H20cunCx5VcfnjF9Hx2ZVUQeKVNHWBCqryP2Wkica-8501yoW_rIORjaq7euDJUunb4I_sEohHqxd-USzjFN1630ijdHmfgWlCScgGeck9K4nlKDZRjKcjvPehCX7YwtCbfYMP-h2WldxOiBJ6W5MnlFGCkB4XSrNoV006hciBtUgLeG6XqOKti8rgOojef_sMbwvT24es1-l-fxpPrUS_VbakUdalSDd3yHGdRRzr7SUXBVkd2p7hnVORPds0BvZi26_RiQrRsUhDY2QaeStJN5OCpsio_NPItM2mD0UtlTW7ENj_5nwSfpy37JT5-6h6f94m8mQ3aAUGunDZxFFuJnkpeUzLgkPBb6aXxIVl606Jufzu7gCZnlA1xC3uzqIV2XWKkVdb6c3EGrZp2HX7ZE7mdjSgIe9qXiAOyMtIOdci5ohgVzRsXt4yQ0vMOIw7G923mGsS1yh4VawiVlVXSmhbEWNvm_I1O4vB9z-Gv2Z9EVdI8zo5RuwXXTrcTcWN264s3jjwWtIxcNVmKGe7t4khw9CLx78TGqG-4thpLZey8qxSZ3s0G8cNY5hI5u6XxJYyEnV0_yGEc50MCNnFDpkYbmf8LhyeH6Q" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJx9VdFu4zYQfM9XDA7oPSVpcumTU1yh2Mkl6F3i2E58heEHWlxb7FFcgVzZEZz-e0HKkZ1D0TfJpGZnZmfXS8ubvFBeMBkcAdls6M3aWFqRxjjD0nOJ87OzX0CuLskrMezmODn5jKttX1kbYByy4R0C-TV5qFobgeXV7wv_6-eNkcI4SEHYGKd588c_R8BV_P61ofCK0WxEoozrQYlQEOS19-QEdaB5d9XxK_rbzDkWJaQRYaEc7rJv8GwJ2eg-Iff3yINtFkJd0ogtTY0UU1rcaXJipImU-5ZrPfHK2ER0bRQyKcipYziOAvhHXZ3QmpyEA-jIZLgdskYHpkLg3CRjsGSf4KQwAU6VFCqVE5TTyR8Tn_OcaycJc3BgRPeaShwBw_3h44GWG_ZD1v8jhWMbouNBle9sH3bw19sBUzhoC5ZGYFwwmnD1MLmFJ4kF2CXIyD9eTs09IacWlsrYJa2EEvbjeyWP-1JHwHX39jR7cpXnNbkePFXsBSog1HlxnAppTu7nVpkSmn2pnMw7hAQ_abnvRUPlUitrG_Qfnq9Hh2Fre9FeWtExxhewZkl5k1s6hkQZqJQXE5WGqGOyp9q9pbLP224w2qFQWBinjVtBCiWpUmo4Jrd343ddfu4wp9tpYfICK891lc6m6exDaIJQ2dtlZPdx6H20cunCx5VcfnjF9Hx2ZVUQeKVNHWBCqryP2Wkica-8501yoW_rIORjaq7euDJUunb4I_sEohHqxd-USzjFN1630ijdHmfgWlCScgGeck9K4nlKDZRjKcjvPehCX7YwtCbfYMP-h2WldxOiBJ6W5MnlFGCkB4XSrNoV006hciBtUgLeG6XqOKti8rgOojef_sMbwvT24es1-l-fxpPrUS_VbakUdalSDd3yHGdRRzr7SUXBVkd2p7hnVORPds0BvZi26_RiQrRsUhDY2QaeStJN5OCpsio_NPItM2mD0UtlTW7ENj_5nwSfpy37JT5-6h6f94m8mQ3aAUGunDZxFFuJnkpeUzLgkPBb6aXxIVl606Jufzu7gCZnlA1xC3uzqIV2XWKkVdb6c3EGrZp2HX7ZE7mdjSgIe9qXiAOyMtIOdci5ohgVzRsXt4yQ0vMOIw7G923mGsS1yh4VawiVlVXSmhbEWNvm_I1O4vB9z-Gv2Z9EVdI8zo5RuwXXTrcTcWN264s3jjwWtIxcNVmKGe7t4khw9CLx78TGqG-4thpLZey8qxSZ3s0G8cNY5hI5u6XxJYyEnV0_yGEc50MCNnFDpkYbmf8LhyeH6Q" alt="Every path to deletion runs through a source that could still say the account is alive." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Every path to deletion runs through a source that could still say the account is alive.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The revocation order that turns a bad call into a readable 403
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Unbind before you delete&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Never delete the ServiceAccount first. If you are wrong about dormancy, deleting the account invalidates the bound tokens its pods are carrying, and the workload starts failing authentication. What lands in someone's alerting is a 401 from a client library, at whatever layer that library chose to surface it, often as a generic connection error two hops from the actual cause. Remove the binding instead. Authentication still succeeds, authorization denies, and the API server produces a message that names the account, the verb, the resource and the scope. That is a bug report someone can act on in thirty seconds.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;$&lt;/span&gt; &lt;span class="nx"&gt;kubectl&lt;/span&gt; &lt;span class="nx"&gt;get&lt;/span&gt; &lt;span class="nx"&gt;secrets&lt;/span&gt; &lt;span class="nx"&gt;--all-namespaces&lt;/span&gt;
&lt;span class="nx"&gt;Error&lt;/span&gt; &lt;span class="nx"&gt;from&lt;/span&gt; &lt;span class="nx"&gt;server&lt;/span&gt; &lt;span class="err"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;Forbidden&lt;/span&gt;&lt;span class="err"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;secrets&lt;/span&gt; &lt;span class="nx"&gt;is&lt;/span&gt; &lt;span class="nx"&gt;forbidden&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;User&lt;/span&gt; &lt;span class="s2"&gt;"system:serviceaccount:reconciler:ledger-backfill"&lt;/span&gt; &lt;span class="nx"&gt;cannot&lt;/span&gt; &lt;span class="nx"&gt;list&lt;/span&gt; &lt;span class="nx"&gt;resource&lt;/span&gt; &lt;span class="s2"&gt;"secrets"&lt;/span&gt; &lt;span class="nx"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;API&lt;/span&gt; &lt;span class="nx"&gt;group&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt; &lt;span class="nx"&gt;at&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="nx"&gt;cluster&lt;/span&gt; &lt;span class="nx"&gt;scope&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;This is the failure mode you want if you got it wrong. It names the account, the verb and the scope.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The revocation runs in two stages with a soak between them, and the length of the soak is a judgment call about your slowest scheduled work. Thirty days catches monthly close jobs. If the platform runs anything quarterly, and a card-issuing or ledger platform usually does, the safe answer is a full quarter for accounts whose name suggests a batch or reconciliation role. We ran 30 days on 18 accounts and a deliberate 95 days on the three that looked seasonal. The commit that removes a binding is a one-line revert, so the cost of the longer soak is calendar time and nothing else.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Delete the ClusterRoleBinding or RoleBinding in the IaC repo, one account per commit, so a revert is surgical. This path only exists for a per-account binding. Where the privilege comes from a group subject such as &lt;code&gt;system:serviceaccounts:settlement&lt;/code&gt;, there is nothing account-shaped to delete, and removing the group binding revokes the ClusterRole from every account in that namespace at once, so the denial watch in step 3 then fires for accounts that were never in scope. Narrow the group binding to named subjects instead. RBAC permissions are purely additive and there is no deny rule that can exempt one account from a group grant. Route that account out of the one-per-commit path. Let the GitOps controller apply it rather than doing it by hand, because a hand-applied deletion is the same out-of-band change that created the problem.&lt;/li&gt;
&lt;li&gt;Verify immediately: the boolean can-i check that returned yes now returns no. That two-word before-and-after is the cleanest evidence line the workpaper can carry.&lt;/li&gt;
&lt;li&gt;Watch the audit log for denials attributed to the account. Filter on the authorization decision annotation set to forbid and group by username, and alert on any non-zero result for accounts in the soak window.&lt;/li&gt;
&lt;li&gt;After the soak, confirm nothing still names the account: no Pod, and no Deployment, StatefulSet, DaemonSet, Job or CronJob template with that serviceAccountName in that namespace. An account can make no calls at all and still be required, for example with automountServiceAccountToken set to false, and admission rejects any new pod that references a missing account, so the next rollout fails, and a node drain completes but its replacement pods are never created. Then delete the ServiceAccount. The token controller deletes the token Secrets tied to it, so confirm by name that they are gone rather than deleting them yourself.&lt;/li&gt;
&lt;li&gt;Record the deletion commit SHA and the cluster list against the account in the review packet. The auditor's follow-up question is when and by what change, not whether.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Back-out is the revert of that one commit, and it is genuinely fast because the binding is the only thing you removed in stage one. If a 403 attributed to a soaking account shows up, do not restore the original cluster-wide binding and walk away. Restore it, then replace it in the same week with a namespace-local RoleBinding carrying only the verbs the denial showed the workload actually needed. Take a denial seriously even though the account showed no calls for the whole window. Anything that runs less often than your lookback, an annual close job for example, looks exactly like a dormant account. Two steps can still catch it: the pod-template check before delete, if the job is defined in the cluster, and a soak longer than the job’s interval, which a 30-day soak is not. A job that an outside scheduler launches ad hoc gets past both, so ask the owning team about anything seasonal before you trust either. The denial log is a free specification for the permission set the workload should have had all along.&lt;/p&gt;

&lt;p&gt;Two things we will not do, both learned expensively. We do not backdate the defective workpaper to make the record look complete, because the certification body asks directly whether timestamps were altered and the only survivable answer is no. The old review stays on file, wrong, and the corrected review is dated to the day it ran and filed as remediation evidence. And we do not report only the accounts the auditor sampled. Disclosing that the methodology defect touched 66 accounts across 17 namespaces kept the finding at minor. Letting the auditor discover the wider pattern in a follow-up sample is how a minor becomes a major, and a major on a surveillance audit means a return visit. We have written more about the pattern in &lt;a href="https://infraforge.agency/infrastructure-audit-readiness/" rel="noopener noreferrer"&gt;the infrastructure audit readiness guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common questions on privileged ServiceAccount access reviews
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;What the auditor asks after you hand over the spreadsheet&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;These are the follow-ups that arrived within two days of filing the corrected review packet. Short answers, because the auditor wanted short answers.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Can I claim dormancy from workload absence instead of audit logs?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No, and this is exactly what produced the finding. A ServiceAccount with no owning Deployment or CronJob is a strong signal that nobody owns it, and it is not evidence that its token was never used. Anyone holding an exported token can call the API from a laptop, and that call appears in the audit log and nowhere else. Absence of a workload goes in the packet as a supporting note, never as the last-use test.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Does this procedure work on managed control planes other than EKS?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The Kubernetes half does. The enumeration, the authorizer checks and the unbind-then-delete order are the same on any 1.29 cluster. The cloud half changes: the second evidence source is whatever records your workload identity exchange, so on GKE that is the Workload Identity path and on AKS it is the workload identity federation events. Name the source your platform actually writes to, and check its retention the same way.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;How do we stop the next batch appearing?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;An admission policy that enforces the same definition the review reads. Any new ClusterRoleBinding whose subject is a ServiceAccount in a workload namespace is rejected whatever role it names, because the clusterScope rule already calls that privileged. A RoleBinding is rejected when its roleRef names a privileged role. A ValidatingAdmissionPolicy cannot open the role a binding points at, so generate the privileged-role list into the policy’s parameters, and generate it from the live cluster on a schedule rather than from git, because aggregated ClusterRoles change without a commit. Test it by replaying the exact binding shape of one of the accounts you just deleted, and confirm the rejection message points at your runbook. Pair it with a nightly job that flags any privileged account with no use in 60 days and opens a ticket against the owning team.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;What if the drift alerts were already firing and nobody looked?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;That is the common case. Our GitOps controller had been reporting resources present in the cluster and absent from git for two years into a Slack channel with no owner. Route those alerts to a ticket queue with an assignee instead. The first scan after we did that surfaced 23 resources across six clusters that existed only in live state, of which 19 were cruft and 4 were adopted into git.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The question underneath all four is whether the control is a document or a mechanism. A quarterly review that a human performs from a spreadsheet is a document. A definition file that feeds both an enumeration script and an admission policy, plus a nightly dormancy job that opens tickets, is a mechanism, and it is what moved this finding from open to closed at the next surveillance. If your release tooling is also the thing making out-of-band changes, that is a related problem and we cover it in &lt;a href="https://infraforge.agency/kubernetes-cicd/" rel="noopener noreferrer"&gt;Kubernetes and CI/CD stabilization&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the sample has already been drawn and the clock is running
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If the workpaper response is due this week&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hard part of this work is not the queries. It is deciding, under a 7-day response window, how wide to open the disclosure when you have just discovered the review methodology was wrong for two cycles, and doing that while the accounts in question still hold cluster-wide access to Secrets in a PCI-scoped namespace. Get that call wrong in either direction and you either escalate a minor into a major or you delete a ServiceAccount that was quietly holding an IAM role for a nightly settlement job.&lt;/p&gt;

&lt;p&gt;We have sat on both sides of that call: building the corrected enumeration, and then being in the room when the lead auditor asks what else the defect touched. What we bring is the evidence pipeline that makes the answer defensible, the join across audit log and cloud trail that stops you revoking something live, and a read on which findings an auditor will accept as remediation versus which will read as backfill. If your sample has already been drawn and the response is due, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review&lt;/a&gt; and we will build the 100% privileged-account enumeration with your team inside the response window, starting the same week.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/prove-kubernetes-serviceaccount-unused-before-delete/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/prove-kubernetes-serviceaccount-unused-before-delete/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>audit</category>
      <category>readiness</category>
      <category>auditreadiness</category>
    </item>
    <item>
      <title>Every terraform state command and when each is safe to run</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Thu, 24 Sep 2026 04:39:23 +0000</pubDate>
      <link>https://dev.to/infraforge/every-terraform-state-command-and-when-each-is-safe-to-run-al8</link>
      <guid>https://dev.to/infraforge/every-terraform-state-command-and-when-each-is-safe-to-run-al8</guid>
      <description>&lt;p&gt;Two engineers ran terraform state list against what they both called the same workspace and got 23 resources and 19. Neither laptop was broken. One was on a checkout predating the &lt;code&gt;backend "s3"&lt;/code&gt; block, so its configuration named no backend at all, Terraform used the local one, and the committed terraform.tfstate was the state; the other was on current main and initialized against S3. Check the commit before you check the state, because the same command answers to a different authority depending on it. The terraform state commands below are the ones you reach for when that happens, and the only column that matters is which of them touch live infrastructure, which touch only the state file, and which quietly do neither. This is the reference we keep open during a state split.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;terraform state list returns a different resource count on two laptops, and the checkouts are not on the same commit&lt;/li&gt;
&lt;li&gt;Error: Error acquiring the state lock with a Lock Info block naming a Path you do not recognise&lt;/li&gt;
&lt;li&gt;apply fails with Error 409: The resource ... already exists, alreadyExists on something you never removed&lt;/li&gt;
&lt;li&gt;git ls-files shows a tracked terraform.tfstate next to your backend block&lt;/li&gt;
&lt;li&gt;plan proposes to create a resource you can see running in the cloud console&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What a four-resource gap in terraform state list actually means
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;23 resources on one laptop, 19 on the other&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first thing to establish in a state split is not what the state contains. It is which state you are reading. Terraform records the backend it resolved for a working directory in &lt;code&gt;.terraform/terraform.tfstate&lt;/code&gt;, which is a state-shaped file with a &lt;code&gt;backend&lt;/code&gt; key, and it is not the same file as the root &lt;code&gt;terraform.tfstate&lt;/code&gt; that holds your resources. Be precise about when that produces a split, because Terraform does not quietly fall back. If the configuration declares &lt;code&gt;backend "s3"&lt;/code&gt; and the cached backend is missing or records something else, it refuses to run anything at all, &lt;code&gt;terraform state list&lt;/code&gt; included, with &lt;code&gt;Initial configuration of the requested backend&lt;/code&gt; or &lt;code&gt;Backend configuration block has changed&lt;/code&gt; and a demand that you run &lt;code&gt;terraform init&lt;/code&gt;. A directory treats a committed &lt;code&gt;terraform.tfstate&lt;/code&gt; as authoritative only when the backend it resolves is the local one at the default path. In this story that means BOTH sides agree there is no backend: the configuration declares none, and &lt;code&gt;.terraform/terraform.tfstate&lt;/code&gt; records none either. That is a checkout predating the block in a directory that never initialised against S3. Get one of the two wrong and the refusal is symmetric: a config that declares a backend the cache does not know gives you &lt;code&gt;Initial configuration of the requested backend&lt;/code&gt;, and a cache holding a backend the config no longer declares gives you &lt;code&gt;Unsetting the previously set backend "s3"&lt;/code&gt;. Both demand an init, and the init that clears them can copy state in the direction you may not want. Update that checkout and the same directory stops being quiet: the configuration now declares &lt;code&gt;backend "s3"&lt;/code&gt; while the directory still has no backend cache at all, which is the first of those two refusals, and Terraform will not run until you init.&lt;/p&gt;

&lt;p&gt;So a four-resource gap has three silent shapes, and the cheapest to rule out is the checkout. A working directory sitting on a commit that predates the &lt;code&gt;backend&lt;/code&gt; block resolves the local backend and prints the committed file's resources without complaint, which is the case in the lede; &lt;code&gt;git log -1&lt;/code&gt; and &lt;code&gt;grep -r 'backend "' *.tf&lt;/code&gt; settle it before you touch anything else. The other two need both directories properly initialised: a different selected workspace, or a different &lt;code&gt;-backend-config&lt;/code&gt; key or bucket. Resist adding credentials to that list, however tempting, because two engineers always do have different identities and it is the entry most likely to look confirmed. On this backend credentials decide whether the read SUCCEEDS, not which object is read: the object is identified by the bucket and key in the backend configuration. Different credentials against the same configured bucket either read the same object or fail loudly on access. Their real damage is in &lt;code&gt;plan&lt;/code&gt;, where Terraform reads a different account’s infrastructure, not in &lt;code&gt;state list&lt;/code&gt;. The loud case is the mirror of the first. Pull that stale checkout forward past the backend block and the directory stops reading anything at all: the config now names a backend this directory has never initialised, Terraform demands an init, and that init offers to copy the committed file into the backend without being asked to. The danger moves to that prompt, below.&lt;/p&gt;

&lt;p&gt;So the diagnostic is three commands, not one, and you run them on both machines before anyone argues about who is right. The backend cache tells you where you are pointed, and the selected workspace tells you which object you get once you are pointed there, which the cache cannot: a &lt;code&gt;workspace_key_prefix&lt;/code&gt; setup shows the same key on both laptops while &lt;code&gt;state pull&lt;/code&gt; reads two different objects. &lt;code&gt;terraform show -json&lt;/code&gt; against an explicit file path tells you what a specific state file on disk contains, without going near the network and without asking Terraform to decide which backend it prefers. It does need the provider plugins installed locally, and when they are missing it fails rather than printing an empty list, which only helps if the script checks.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Which backend did THIS working directory resolve to? An init with no backend&lt;/span&gt;
&lt;span class="c"&gt;# block writes no backend cache at all, so ANY recorded type, local included,&lt;/span&gt;
&lt;span class="c"&gt;# means a backend block was declared at some point. With the block still there&lt;/span&gt;
&lt;span class="c"&gt;# the cache is expected; with it gone, Terraform answers "Unsetting the&lt;/span&gt;
&lt;span class="c"&gt;# previously set backend" and refuses. No cache and no block is the healthy&lt;/span&gt;
&lt;span class="c"&gt;# case this article's lede is about.&lt;/span&gt;
&lt;span class="c"&gt;# declared() only reads *.tf here. It misses *.tf.json and a cloud {} block, so&lt;/span&gt;
&lt;span class="c"&gt;# if either is in use, read the configuration by eye before trusting a branch.&lt;/span&gt;
declared&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-qE&lt;/span&gt; &lt;span class="s1"&gt;'^[[:space:]]*backend[[:space:]]+"'&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt;.tf 2&amp;gt;/dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="nv"&gt;cached_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; .terraform/terraform.tfstate &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.backend.type // "none"'&lt;/span&gt; .terraform/terraform.tfstate &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo &lt;/span&gt;none &lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cached_type&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; none &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; declared&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;jq &lt;span class="s1"&gt;'.backend.type, .backend.config.bucket, .backend.config.key'&lt;/span&gt; .terraform/terraform.tfstate
&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cached_type&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; none &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Cache records a backend (&lt;/span&gt;&lt;span class="nv"&gt;$cached_type&lt;/span&gt;&lt;span class="s2"&gt;) the config NO LONGER declares. Terraform"&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'answers Unsetting the previously set backend and refuses until init. Stop.'&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;elif &lt;/span&gt;declared&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"A backend is DECLARED and none is resolved. Terraform answers Initial"&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"configuration of the requested backend and refuses until init. Stop here."&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;else
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"no cache and no backend block: this directory resolved the LOCAL backend"&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Both refusals are stops, not warnings. Past either one the pull below fails,&lt;/span&gt;
&lt;span class="c"&gt;# leaves an empty file, and the diff reads as an empty remote, which is the&lt;/span&gt;
&lt;span class="c"&gt;# reading that sends someone to push local state over a good remote one.&lt;/span&gt;

&lt;span class="c"&gt;# Which workspace is selected? The backend cache does NOT record this, and with&lt;/span&gt;
&lt;span class="c"&gt;# workspace_key_prefix the cached config.key is identical across every workspace&lt;/span&gt;
&lt;span class="c"&gt;# of the same configuration, so two laptops can print the same bucket and key&lt;/span&gt;
&lt;span class="c"&gt;# and still pull two different objects.&lt;/span&gt;
terraform workspace show 2&amp;gt;/dev/null &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;cat&lt;/span&gt; .terraform/environment 2&amp;gt;/dev/null &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo &lt;/span&gt;default

&lt;span class="c"&gt;# Render both sides through the same function, or the diff is fiction:&lt;/span&gt;
&lt;span class="c"&gt;# terraform show -json emits full resource ADDRESSES, and recurse() is what&lt;/span&gt;
&lt;span class="c"&gt;# stops it silently dropping everything that lives inside a module.&lt;/span&gt;
&lt;span class="c"&gt;# show -json needs the provider plugins installed locally, and when it fails jq&lt;/span&gt;
&lt;span class="c"&gt;# still exits 0 on empty input, so check the exit status or a failure reads as&lt;/span&gt;
&lt;span class="c"&gt;# a state with nothing in it.&lt;/span&gt;
addrs&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nb"&gt;local &lt;/span&gt;j
  &lt;span class="nv"&gt;j&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;terraform show &lt;span class="nt"&gt;-json&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"show -json failed on &lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;. Stop."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;return &lt;/span&gt;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$j&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'[.values.root_module | recurse(.child_modules[]?) | .resources[]?.address] | sort[]'&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;# What does the committed file on disk contain, independent of any backend?&lt;/span&gt;
addrs terraform.tfstate &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /tmp/local.txt &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1

&lt;span class="c"&gt;# And what does the CONFIGURED backend hold? Not necessarily the remote one:&lt;/span&gt;
&lt;span class="c"&gt;# state pull reads whatever backend line 1 just printed. If that says "local",&lt;/span&gt;
&lt;span class="c"&gt;# both sides of this diff come from the same file and it returns a false clean.&lt;/span&gt;
terraform state pull &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /tmp/remote.tfstate &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1
addrs /tmp/remote.tfstate &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /tmp/remote.txt &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1

diff /tmp/local.txt /tmp/remote.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Neither engineer was wrong about what they saw. They were reading two different files, and only the backend cache says which. Read line 1 before you trust the diff: &lt;code&gt;state pull&lt;/code&gt; follows the configured backend, so on the machine whose cache says local, or has no cache at all, you are comparing a file with itself. Build both sides with the same function too: reading &lt;code&gt;.values.root_module.resources[]&lt;/code&gt; on its own returns root-module resources only, so on any workspace with modules the diff invents a gap that was never there.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In the case that pushed us to write this down, the diff came back with five resources present only in S3 and one &lt;code&gt;google_sql_database_instance&lt;/code&gt; present only locally, which is where the four comes from. Four of the five were ordinary additions merged since that checkout diverged. The fifth was a &lt;code&gt;google_compute_firewall&lt;/code&gt; created by hand in the console during an outage the previous weekend and imported to the remote state by whoever fixed it. The SQL instance existed in the local file because the engineer working from the stale local backend had applied it. Both of those changes were real, and each was recorded in exactly one of the two states, which is the part a resource count cannot show you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJyNkdtu2kAQhu95iv8BONyjKhGkqIpaiARIqLK4WNZjs8LeQbNDLAR998pj4rR3uVt7d775D0XFjT86UWy_D4BZpiTiCpYaIQZFiHBoWE4hlsiD7DEaPWF-y5kS9EjwHItQfjvI5CknXzkhOBycP1HMn_8MgHk7cY98x-o27umT_jTWIqlTMoSQZ8kTZqvfH5AhKvauQoi-uuTUQVc9dJH9suvH6-lDVF0HVcoNWoSK8LqxG1u17wlXSnese5X2-fJ1mS1xM1su_jP80mtbZ9s-TaHikiiB3kmuJtDFfGy0V4XFuXrbonBV5wXKnfFW7NpSX2bbI0FCOiF0u7uKFDm1tDQ1XNAEz-crzsL1WeFdRMn2vhGOJRp33fc6zfLullxN1nM6O09wMUf7y4Cjh7tR1zVOdJ0cLv5Ean53vd9NtgkVRUU6V0Gn0IYhVLM-ck9D40UGibDAKVxlBnefWt5-ZnPWIyiWIRJJgpDrirQCW6F9iwsL5ke2bfhzYGiL3UWPLEHDx1qO_zhshzfd8ABYdqe_whQIFA" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJyNkdtu2kAQhu95iv8BONyjKhGkqIpaiARIqLK4WNZjs8LeQbNDLAR998pj4rR3uVt7d775D0XFjT86UWy_D4BZpiTiCpYaIQZFiHBoWE4hlsiD7DEaPWF-y5kS9EjwHItQfjvI5CknXzkhOBycP1HMn_8MgHk7cY98x-o27umT_jTWIqlTMoSQZ8kTZqvfH5AhKvauQoi-uuTUQVc9dJH9suvH6-lDVF0HVcoNWoSK8LqxG1u17wlXSnese5X2-fJ1mS1xM1su_jP80mtbZ9s-TaHikiiB3kmuJtDFfGy0V4XFuXrbonBV5wXKnfFW7NpSX2bbI0FCOiF0u7uKFDm1tDQ1XNAEz-crzsL1WeFdRMn2vhGOJRp33fc6zfLullxN1nM6O09wMUf7y4Cjh7tR1zVOdJ0cLv5Ean53vd9NtgkVRUU6V0Gn0IYhVLM-ck9D40UGibDAKVxlBnefWt5-ZnPWIyiWIRJJgpDrirQCW6F9iwsL5ke2bfhzYGiL3UWPLEHDx1qO_zhshzfd8ABYdqe_whQIFA" alt="The committed state file only wins when the working directory resolves the local backend at the default path, which in a split like this one means the configuration declares no backend and  raw `.terraform/terraform.tfstate` endraw  records none. That is the branch to rule out first." width="1167" height="1466"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The committed state file only wins when the working directory resolves the local backend at the default path, which in a split like this one means the configuration declares no backend and &lt;code&gt;.terraform/terraform.tfstate&lt;/code&gt; records none. That is the branch to rule out first.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  terraform init flags: which one moves state and which one only re-points it
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The five init flags that change where writes land&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;These five flags get confused constantly, and the confusion is expensive because three of them decide whether an existing state gets copied to the new location or abandoned there. We assume Terraform CLI open source, 1.9.x, with an S3 backend and DynamoDB locking. Current versions of the S3 backend can lock with a lockfile in the bucket itself (&lt;code&gt;use_lockfile&lt;/code&gt;), and HashiCorp now documents DynamoDB locking as deprecated, so on a newer version the lock you are reading may be an S3 object rather than a table item. Terraform Cloud and Enterprise workspaces have their own migration surface and different prompts, so do not carry these across.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FLAG                       WHAT IT DOES                                  SAFE WHEN                          DAMAGE WHEN NOT
-------------------------  --------------------------------------------  ---------------------------------  --------------------------------------
-migrate-state             Offers to copy state from the previously      You are moving a workspace from    You accept the prompt while the source
                           configured backend into the newly             one backend to another and want    is the stale side of a split. The stale
                           configured one, with a yes/no prompt.         the contents carried over.         contents become the remote authority.

-reconfigure               Re-initializes the backend and ignores the    The directory holds no local       Used as a shortcut past a migration
                           previous one: nothing is copied FROM it.      terraform.tfstate, and the         prompt, it leaves the old backend with
                           But a non-empty local terraform.tfstate in    backend already holds the state    the only copy. And any non-empty local
                           the directory still takes the migration       you want.                          terraform.tfstate, tracked or not,
                           path, usually with a yes/no prompt.                                              turns it into a migration. Yes copies
                                                                                                            that file over the remote state even
                                                                                                            when the remote is newer; no leaves
                                                                                                            the remote alone. Either answer, or no
                                                                                                            prompt at all, can leave the local
                                                                                                            file empty, with its old contents in
                                                                                                            terraform.tfstate.backup.

-backend-config=KV|FILE    Supplies backend settings (bucket, key,       Values live outside the config,    A wrong `key` adopts a different
                           region) at init time.                         e.g. per-environment .hcl files.   workspace's state. From another
                                                                                                            configuration, plan proposes to CREATE
                                                                                                            all of yours and DESTROY all of theirs.
                                                                                                            From another environment of the SAME
                                                                                                            configuration the addresses match, and
                                                                                                            plan shows updates and replacements of
                                                                                                            their real resources instead, which is
                                                                                                            much harder to spot.

-backend=false             Initializes providers and modules only.       You want a provider install or a   You expect it to protect you from
                           Leaves the backend selection untouched.       validate pass without touching     writes generally. It does not gate
                                                                         backend credentials.               apply in a later, normal init.

-force-copy                Suppresses the migration prompts and          You are scripting a migration you  It performs the copy that the
                           answers yes to all of them. Turning it        have already proven by hand, and   -migrate-state row tells you to stop
                           on also enables -migrate-state.               the copy direction is settled.     and read, with no prompt to read.
                                                                                                            Check the CI init line before you
                                                                                                            trust a pipeline: it is often there.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;-migrate-state offers to copy from whatever the directory used before. -reconfigure ignores the old backend but usually still offers to copy a local terraform.tfstate. Picking the wrong one at 2am is how the wrong side of a split becomes the official one.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The specific trap is the copy prompt. Plain &lt;code&gt;init&lt;/code&gt; can raise it the first time a directory holding a local &lt;code&gt;terraform.tfstate&lt;/code&gt; meets a backend block, &lt;code&gt;-reconfigure&lt;/code&gt; again whenever that file is non-empty, and &lt;code&gt;-migrate-state&lt;/code&gt; when you move between configured backends. During an incident it reads like a formality. It is not. It is Terraform asking which of your two states you want to keep, and it will happily copy the four-resource-short file over a correct remote one, newer serial or not. Read the source and destination in the prompt text out loud before typing yes. Nor is no a safe answer when the source is that local file, on the first init after a backend block appears, with or without &lt;code&gt;-migrate-state&lt;/code&gt;, or under &lt;code&gt;-reconfigure&lt;/code&gt;: whatever you answer, Terraform can empty the local file afterwards, leaving its old contents in &lt;code&gt;terraform.tfstate.backup&lt;/code&gt;, and it can do that without asking at all when the file already matches the backend's copy. And &lt;code&gt;-force-copy&lt;/code&gt; answers yes to every prompt, so read the init line you inherited before you rely on reading anything. So take the backup before the init and outside Terraform, because &lt;code&gt;terraform state pull&lt;/code&gt; cannot give you one here: before &lt;code&gt;init&lt;/code&gt; it fails with the same backend error as every other command, and after it reads the newly configured destination, not the source you meant to keep. Copy the source directly, &lt;code&gt;cp terraform.tfstate backup-$(git rev-parse --short HEAD).tfstate&lt;/code&gt; where the source is the local file, or run &lt;code&gt;terraform state pull&lt;/code&gt; against the OLD backend configuration before you re-point it. If an init has already run on one of those paths, check whether the local file is now empty, and if it is, copy &lt;code&gt;terraform.tfstate.backup&lt;/code&gt; somewhere safe before anything else: it holds what the file held, and a later command that writes local state there can overwrite it. &lt;code&gt;git checkout -- terraform.tfstate&lt;/code&gt; restores only the last committed version, which can predate applies nobody committed.&lt;/p&gt;

&lt;p&gt;The other habit worth building is that a tracked state file is a repo problem, not a Terraform problem. &lt;code&gt;git ls-files -- '*.tfstate' '*.tfstate.backup'&lt;/code&gt; takes half a second and answers the question that took our client's team an hour of laptop-to-laptop comparison. If it returns anything, you have a second authority in the repo regardless of what your backend block says, and every fresh clone is a coin flip. We wrote up the wider version of this failure in &lt;a href="https://infraforge.agency/terraform-state-recovery/" rel="noopener noreferrer"&gt;the Terraform state recovery playbook&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  terraform state subcommands: what each one touches and the damage when misused
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Ranked by what they can destroy, not by how often you type them&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here is the table itself. The column that decides everything is the middle one: whether the command touches only the state file, only live infrastructure, or both. Most state accidents come from an operator who believed a state-only command was going to change the cloud, or believed a cloud-touching command was going to be recorded.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;COMMAND                    TOUCHES              SAFE WHEN                                  SPECIFIC DAMAGE WHEN NOT
-------------------------  -------------------  -----------------------------------------  ----------------------------------------------
terraform state list       nothing (read)       Always, and the cheapest thing here. Run    None. Note it reads the CONFIGURED backend,
                                                it after the pull, not instead of it.       so it inherits whatever split you are debugging.

terraform state show ADDR  nothing (read)       Always. Shows recorded attributes for one   None, but the values are what state believes,
                                                address.                                    not what the cloud currently holds.

terraform state pull       nothing (read)       Always, before any of the four commands     None. Take it and keep it. What any of these
                                                below that write state.                     commands writes automatically depends on the
                                                                                            backend and the version, `state push` has no
                                                                                            `-backup` option at all, and none of that is worth
                                                                                            working out during an incident. The pull is the
                                                                                            copy you control, and the rule is the same for
                                                                                            `push`, `rm`, `mv` and `replace-provider`.

terraform state push FILE  state only, and      You are restoring a known-good snapshot     CheckValidImport refuses three things: an
                           wholesale            and you have read the three guards below,   unrelated lineage, a serial lower than the
                                                because the restore case trips one of       destination, and an EQUAL serial whose contents
                                                them.                                       differ. Restoring after something wrote to
                                                                                            state is the second, so you meet a refusal
                                                                                            rather than an overwrite. The quiet danger is a
                                                                                            hand-edited pull with the serial bumped: that
                                                                                            reads as newer, passes all three, and
                                                                                            overwrites wholesale. -force removes the
                                                                                            checks.

terraform state rm ADDR    state only           You are about to re-import the same         The resource keeps running and is now
                                                resource under a different address in      unmanaged. Use it thinking it deletes and you
                                                the same change window.                    are paying for an orphan nobody plans against.

terraform state mv SRC DST state only           A resource moved address because of a      A typo'd destination silently creates a state
                                                rename or a module refactor and the        entry that matches no config block, and the
                                                remote object is unchanged.                next plan proposes to destroy the real object.

terraform state            state only           A provider source address changed, as in    Rewrites EVERY resource using the from-provider
  replace-provider                              a registry namespace move, and the          in one pass, so a wrong FQN moves the whole set
                                                resource schemas are compatible.            at once. Point it at an incompatible provider
                                                                                            and state holds entries the new one cannot
                                                                                            decode. Take the pull first and do not count on
                                                                                            an automatic backup here.

terraform import ADDR ID   state only           A live object exists and a matching        A wrong ID binds a real object to that address
                           (writes the object   config block already exists for it.        with no check that it is the one you meant, and
                           into state)                                                     the next plan reconfigures or replaces it. A
                                                                                           matching resource block is required before you
                                                                                           import, so create it first rather than treating
                                                                                           import as the step that finds what is missing.

terraform force-unlock ID  the lock record      You have confirmed the holder named in     You break a lock held by a run that is still
                           only                 the Lock Info block is dead.               mid-apply. Two writers, one state file, and a
                                                                                           serial you cannot reconstruct.

terraform apply            state only           Credentials and account/project resolve to  Wrong-scoped credentials either fail the
  -refresh-only            (rewrites from live  the same place the state was written from,  refresh or, where the provider reads the
                           attributes)          and attributes drifted for a resource       API's answer as not-found, read resources
                                                already in state. Check that first.         as gone, and applying that plan forgets
                                                                                            them all rather than one object. It also
                                                                                            adopts nothing new: a resource created
                                                                                            outside Terraform stays outside. And if an
                                                                                            object in state really was DELETED outside
                                                                                            Terraform, applying drops it and the next
                                                                                            apply proposes to create it. Read the plan;
                                                                                            this writes state.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;state rm does not delete. refresh-only does not adopt. Those two sentences are most of what goes wrong.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two rows deserve more than a table cell. &lt;code&gt;terraform state push&lt;/code&gt; is the only command here that can lose resources in a single keystroke, because it replaces rather than merges. Terraform does check lineage and serial and will refuse a push it considers a regression, and there is a force option that skips those checks, which means the guardrail is exactly one flag away from being off. And the checks only run against a destination that already holds state. Push into an empty one, which is what declining a migration into a new, empty backend leaves behind, and none of them apply, with no flag at all. Treat any push as a restore operation with a change record behind it, not as an edit.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;terraform force-unlock&lt;/code&gt; is the one people reach for fastest and should reach for slowest. The lock error prints who holds it, and that block is the entire decision.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: Error acquiring the state lock

Error message: operation error DynamoDB: PutItem, https response error
StatusCode: 400, RequestID: 7Q2G5M0N8LDH3S4K6V1R9B2C0TJ5UQAEMVJF66Q9ASUAAJG4KQ9X,
ConditionalCheckFailedException: The conditional request failed
Lock Info:
  ID:        4d1f9b0c-6a37-4d2e-9c11-2f8a3e5b7d41
  Path:      tf-state-prod/data-layer/terraform.tfstate
  Operation: OperationTypeApply
  Who:       ci-runner@runner-7
  Version:   1.9.5
  Created:   2026-08-24 15:31:07.882119 +0000 UTC
  Info:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Operation: OperationTypeApply and a Created stamp two minutes old is a live writer. Force-unlocking that is how you get a state file nobody can reconstruct.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The rule we give teams: force-unlock only after you have found the holder in the other system. If &lt;code&gt;Who&lt;/code&gt; names a CI runner, open that job and confirm it exited. If it names a person, message them. A client that reports a held lock while the DynamoDB console shows no lock item usually means the console is looking at a different table, region or account from the one the backend uses. Find the one the backend names before anything else. An IAM or region mismatch on the lock table itself fails fast, with an access or not-found error, not a lock you cannot see. And with &lt;code&gt;use_lockfile&lt;/code&gt; there may be no table item to find at all: the lock is a &lt;code&gt;.tflock&lt;/code&gt; object beside the state in the bucket.&lt;/p&gt;

&lt;h2&gt;
  
  
  Adopting a console hotfix: why refresh-only is the wrong tool and import needs config first
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The refresh-only pass that adopted nothing&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When we reached the reconciliation step, the first plan looked correct and was not. A second firewall rule created by hand during that outage had never been imported into either state, and it needed to come under management. &lt;code&gt;terraform apply -refresh-only&lt;/code&gt; is the command everyone suggests. It ran, it reported no changes to state, and the next normal plan still proposed to create a firewall rule that already existed. Refresh-only updates recorded attributes for resources Terraform already tracks. An object that has never been in state is invisible to it. That distinction cost about forty minutes.&lt;/p&gt;

&lt;p&gt;The other common wrong turn is importing before writing the config block. Import will not run at all until a &lt;code&gt;resource&lt;/code&gt; block matches the address you name. It stops with &lt;code&gt;Error: resource address "google_compute_firewall.allow_health_checks" does not exist in the configuration.&lt;/code&gt; and writes nothing, so what the mistake costs you is the change window, not the resource. The order is fixed for that reason: write the block, then import, then plan and expect zero changes. The destroy proposal people half-remember comes from the other direction, deleting a config block after a successful import. On 1.5 and later you can declare &lt;code&gt;import {}&lt;/code&gt; blocks in configuration and see the adoption in the plan output before anything is written, which is what we now default to, because it makes the import reviewable in a PR instead of being a thing someone typed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Snapshot first. Always.&lt;/span&gt;
terraform state pull &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; pre-reconcile.tfstate

&lt;span class="c"&gt;# 2. Config block exists for the console-created rule, then declare the import.&lt;/span&gt;
&lt;span class="c"&gt;#    main.tf&lt;/span&gt;
resource &lt;span class="s2"&gt;"google_compute_firewall"&lt;/span&gt; &lt;span class="s2"&gt;"allow_health_checks"&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  name    &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"allow-health-checks"&lt;/span&gt;
  network &lt;span class="o"&gt;=&lt;/span&gt; google_compute_network.core.id
  &lt;span class="c"&gt;# ... attributes matched to the live object&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

import &lt;span class="o"&gt;{&lt;/span&gt;
  to &lt;span class="o"&gt;=&lt;/span&gt; google_compute_firewall.allow_health_checks
  &lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"projects/analytics-prod/global/firewalls/allow-health-checks"&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;# 3. Review the adoption in the plan BEFORE it is written to state.&lt;/span&gt;
terraform plan &lt;span class="nt"&gt;-out&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;reconcile.tfplan

&lt;span class="c"&gt;# 4. Only then.&lt;/span&gt;
terraform apply reconcile.tfplan
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The import block turns state adoption into something a reviewer can see in a PR diff. The CLI import command writes first and shows you afterwards.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The other half of the reconciliation was the reverse case: the SQL instance that existed only in the local file, because it had been applied from the stale side of the split. The remote state had never recorded it, so it took the same path as the firewall rule: merge its resource block into main, declare the import, and plan for zero changes. &lt;code&gt;terraform state rm&lt;/code&gt; followed by an import belongs to a different case, an address in state bound to the wrong object, and there the ordering matters for a reason people miss. The CLI &lt;code&gt;terraform import&lt;/code&gt; refuses an address that is already in state, and an &lt;code&gt;import {}&lt;/code&gt; block aimed at one is skipped without a word: the plan then compares your config with the wrongly bound object, showing no changes only if they happen to match, and otherwise changes to that object, up to a replacement that destroys it. Remove the wrong binding first either way. The object keeps running through both commands, which is the point of &lt;code&gt;state rm&lt;/code&gt; and also exactly why it is dangerous when someone reaches for it expecting a delete.&lt;/p&gt;

&lt;p&gt;The tell that you are done is a plan with no changes and no refresh noise, after an adopting plan whose summary line counted the imports you expected. A skipped &lt;code&gt;import {}&lt;/code&gt; block is not counted, and when nothing else differs the plan just says No changes, so check for the import line rather than for silence. If plan still proposes to create something you can see running, read what it pairs that create with. A destroy of an address you recognise means an address mismatch, and the fix is &lt;code&gt;state mv&lt;/code&gt; or a &lt;code&gt;moved {}&lt;/code&gt; block, not another import. A create with no matching destroy means the object is not in state at all, and it does need the import. Teams stuck in that loop for more than an hour usually have a module refactor tangled into the same change; we cover unwinding those separately under &lt;a href="https://infraforge.agency/terraform-iac-debt/" rel="noopener noreferrer"&gt;Terraform and IaC debt&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The changes that stuck afterward were small. An 11-line pre-commit hook that rejects any staged path matching &lt;code&gt;*.tfstate&lt;/code&gt; or &lt;code&gt;*.tfstate.backup&lt;/code&gt;, which has fired twice since on branches nobody would have checked. A CI step that runs &lt;code&gt;jq -r '.backend.type' .terraform/terraform.tfstate&lt;/code&gt; after init and fails the job unless it prints &lt;code&gt;s3&lt;/code&gt;, which catches a job that initialised without the backend, or against a different one, before it plans against the wrong state. And a rule that any console change during an incident gets an import block opened in the same hour, not the same week, because state drift you know about is a ticket and state drift you forget about is an outage with a four-hour tail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common questions about terraform state commands and backend migration
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The questions that come next&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;These are the follow-ups we field most often after a state split, answered for Terraform CLI open source with an S3 backend.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Does terraform state rm delete the resource?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No. It removes the entry from state and leaves the object running in the cloud, unmanaged. That is what you want immediately before re-importing it at a corrected address, and it is a bill you keep paying if you meant to destroy something. Use terraform destroy with a -target for that case, after a plan you have read.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Is terraform state pull safe on production?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes. It reads the configured backend and writes to stdout. Redirect it to a file before every other command in this reference. It is the only free insurance in the whole list, and the teams who skip it are the ones with no way back after a bad push.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Can I force-unlock when DynamoDB shows no lock item?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;There is nothing to unlock in that case, so look at the connection instead. Check that the lock table name, region and IAM permissions in your backend config match the table you are inspecting in the console, and confirm the credentials the CLI resolved are the ones you think they are.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Does -migrate-state work between any two backends?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;It offers to copy state from the previously configured backend to the newly configured one, including local to S3 and back. Read the source and destination in the prompt before confirming. If either side might be stale, back up the source and diff the two before you run init, not at the prompt. Declining is not always neutral: on the first init after a backend block appears, or with -reconfigure, answering no can still empty a local terraform.tfstate, leaving its old contents in terraform.tfstate.backup.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Will apply -refresh-only pick up a resource I built in the console?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No. It refreshes attributes for resources already in state. Adopting something Terraform has never seen requires an import, and the config block for it has to exist first, or the CLI import refuses and writes nothing.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Getting a second pair of eyes on a split state before you push anything
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;When both states look plausible and applies are blocked&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hard part of a state split is never the commands. It is the twenty minutes where two states each look defensible, applies are blocked, and the fastest-looking move is a &lt;code&gt;state push&lt;/code&gt; that quietly drops four resources into being unmanaged. Nobody wants to be the person who made that call alone at the end of a long day, and the decision is genuinely hard: it depends on which side has the console hotfix, which side CI last wrote to, and whether the lock you are staring at belongs to a job that is still running.&lt;/p&gt;

&lt;p&gt;We do this reconciliation work with platform teams often enough that we walk in with the diff commands already written and a snapshot taken before anything else happens. What we bring is mostly the discipline to not push until the two states have been diffed resource by resource against the cloud, plus enough scar tissue to recognise which flavour of drift you have from the first plan output. It is unglamorous and it is the difference between a four hour incident and a two day one.&lt;/p&gt;

&lt;p&gt;If you are looking at two state files right now and cannot tell which one is authoritative, do not push either. Take a &lt;code&gt;terraform state pull&lt;/code&gt; from the remote backend, keep the local file, and &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review&lt;/a&gt;; we will sit on a call the same day and pick the authority with you before anything writes.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/terraform-state-commands-safe-reference/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/terraform-state-commands-safe-reference/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>iac</category>
      <category>recovery</category>
      <category>terraformstate</category>
    </item>
    <item>
      <title>Why S3 traffic goes through NAT: the missing endpoint route</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Wed, 23 Sep 2026 17:29:44 +0000</pubDate>
      <link>https://dev.to/infraforge/why-s3-traffic-goes-through-nat-the-missing-endpoint-route-d00</link>
      <guid>https://dev.to/infraforge/why-s3-traffic-goes-through-nat-the-missing-endpoint-route-d00</guid>
      <description>&lt;p&gt;NAT gateway hours were flat while NAT gateway data processing was up 4.6x, and that pair was the whole diagnosis. Something inside the VPC had started pushing bulk traffic through NAT that should never have touched it. It was same-region S3 reads from a new EKS node group whose subnets were never added to the S3 gateway endpoint's route tables. The nightly rollup job moved 2.7 TB a day down that path across two shards, 1.9 TB of it on the largest, for 22 days before anyone looked, at about $2,860 above baseline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;NAT gateway data-processing bytes climb 4.6x while NAT gateway hours stay perfectly flat&lt;/li&gt;
&lt;li&gt;The EC2 - Other line on the monthly preview runs several times its forecast with no new instances and no new AZ&lt;/li&gt;
&lt;li&gt;A VPC Flow Logs top-talker query puts one node IP at 1.9 TB in a single day against the NAT ENIs&lt;/li&gt;
&lt;li&gt;A private subnet’s route table ID is missing from the S3 gateway endpoint’s RouteTableIds in describe-vpc-endpoints&lt;/li&gt;
&lt;li&gt;A CronJob that never appeared in a cost report starts appearing right after a nodeSelector change&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where does a 4.6x NAT gateway data processing spike usually come from?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The CRM sync was innocent, and so was the retry loop&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There was no page. There was a FinOps close-out preview on 2026-07-28 showing EC2 - Other running at about $5,050 a month against a baseline near $1,150. The analyst drilled the sub-lines and it all sat in one usage type. Data processed by the NAT gateways had gone from roughly 760 GB/day to roughly 3,470 GB/day, starting 2026-07-06, part of a day at first because the rollup only moved that evening, then holding at full height every day since. Gateway hours had not moved by a minute.&lt;/p&gt;

&lt;p&gt;Flat hours means the same NAT gateways ran for the same number of seconds. Nobody added a gateway, nobody added an AZ, nobody changed the topology. Only the volume per gateway changed. That one fact killed a whole class of explanation before we opened a terminal against the cluster, and it is the first thing we check now when a bill moves and the architecture did not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws ce get-cost-and-usage \
  --time-period Start=2026-06-25,End=2026-07-28 \
  --granularity DAILY \
  --metrics UnblendedCost UsageQuantity \
  --filter '{"Dimensions":{"Key":"SERVICE","Values":["EC2 - Other"]}}' \
  --group-by Type=DIMENSION,Key=USAGE_TYPE \
  --output json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Two usage types dominate the result: gateway hours, dead flat across 33 days, and gateway bytes, a step that starts on 2026-07-06 and squares up from the 7th. The hours ruled out half our hypotheses for the price of one API call.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The platform lead's first read was a new external integration behaving badly. Two candidates had shipped inside the window: a CRM sync service on 2026-07-04 and an SMS verification service on 2026-07-11. The SMS service was the wrong date, five days late for an anomaly that started on the 6th. The CRM sync was documented at around 200 MB/day of payload, four orders of magnitude short of a 2.7 TB/day delta. Both vendor dashboards agreed with their own docs. Dead end.&lt;/p&gt;

&lt;p&gt;Second guess was a runaway HTTP client, because a retry loop hammering a third-party API and dragging response bodies back is the classic version of this bill. The on-call SRE grepped two weeks of application logs in Loki for elevated retry counters and for 429 and 503 responses. Two services came back hot. One was the CRM sync, already accounted for. The other was an internal service calling an in-cluster endpoint, which never touches NAT at all. Dead end.&lt;/p&gt;

&lt;p&gt;Third guess was a DNS leak: a service resolving an internal name against a public resolver, getting a public address back, and hairpinning out through NAT and back in. CoreDNS forward counts to the upstream resolver were flat and there was no NXDOMAIN storm in the logs. Dead end, and about two hours gone.&lt;/p&gt;

&lt;p&gt;All three guesses shared an assumption. Each one took for granted that expensive NAT bytes were bytes that genuinely belonged outside the VPC. Nobody had yet considered that the traffic was going to a bucket 15 milliseconds away in the same region. That inversion, where the cost spike is in the path rather than the payload, is the shape we hit most often in &lt;a href="https://infraforge.agency/problems/cloud-cost-spikes/" rel="noopener noreferrer"&gt;cloud cost spike work&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you find the top talker through a NAT gateway in VPC Flow Logs?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;One node moved 1.9 TB to read its own region's bucket&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;At 11:20 the SRE stopped interrogating the cluster and started interrogating the network. VPC Flow Logs had been enabled at the VPC level since the VPC was built, landing in S3 on a 90-day lifecycle, queried by Athena approximately never. They were already paid for, so the marginal cost of the answer was one Athena scan, and they were the only place with a per-source byte count.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SELECT
  CASE WHEN srcaddr IN ('10.42.0.219', '10.42.32.87', '10.42.64.140')
       THEN dstaddr ELSE srcaddr END AS node_ip,
  sum(bytes) / 1e9 AS gb,
  count(*)         AS flows
FROM vpc_flow_logs
WHERE day = '2026/07/27'   -- slash form: the projected partition
                             -- format is yyyy/MM/dd, and a hyphenated
                             -- value matches nothing and returns zero
                             -- rows without erroring
  AND (srcaddr IN ('10.42.0.219', '10.42.32.87', '10.42.64.140')
       OR dstaddr IN ('10.42.0.219', '10.42.32.87', '10.42.64.140'))
  -- Keep only the node-to-NAT legs, where both addresses are VPC-private.
  -- That drops the NAT-to-internet hop, which carries the same bytes again.
  AND srcaddr LIKE '10.42.%'
  AND dstaddr LIKE '10.42.%'
GROUP BY 1
ORDER BY gb DESC
LIMIT 10;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The NAT ENI private addresses come from describe-nat-gateways. The CASE is what makes this work: on a NAT ENI the download leg arrives as srcaddr = the NAT address and dstaddr = the node, so a plain GROUP BY srcaddr tops out at an S3 public address and the node never appears. Keying on whichever side is not the NAT sums both directions against the workload that caused them. It is the node rather than the pod because the VPC CNI SNATs pod traffic to the node primary ENI for anything outside the VPC CIDR.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The top row was 10.42.174.83 at 1.9 TB for the day, and almost all of it inbound. Read that address correctly, because the obvious reading wastes an hour: it is a NODE, not a pod. The VPC CNI translates a pod address to the primary private address of its node primary ENI for any destination outside the VPC CIDR, and S3 prefixes are outside it, so what reaches a NAT gateway ENI, and therefore the flow log, is the node. Grep a pod listing for it and you find nothing. Only a cluster running &lt;code&gt;AWS_VPC_K8S_CNI_EXTERNALSNAT=true&lt;/code&gt; hands SNAT to the NAT gateway and preserves pod addresses that far, and this one was not. So &lt;code&gt;kubectl get nodes -o wide&lt;/code&gt; first, which placed it in the new memory group, and then the pods scheduled on that node, which gave a nightly rollup pod in the reporting namespace, one shard of a CronJob that had been running daily for months and had never shown up in a cost conversation. Second row, 800 GB, was another node in the same group running a sibling shard. Together the two shards account for 2,700 of the 2,710 GB/day the meter had gained, and the last 10 GB is image pulls on the new pool (more on that below). Third row, 400 GB, was a general-pool node whose traffic had always been there and sat inside the old baseline.&lt;/p&gt;

&lt;p&gt;The rollup job reads parquet from a bucket in the same region and the same account, and writes aggregates to the warehouse. That traffic should never be metered by a NAT gateway, and for months it had not been. Something had changed underneath a job whose own manifest had not been touched in eleven weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does a new EKS node group lose the S3 gateway endpoint?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Nine subnets, six route tables on the endpoint&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;S3 traffic stays inside the region's gateway endpoint only if the subnet's route table carries a route for it. This account had had an S3 gateway endpoint since the VPC was built in 2024, and it had always just worked, which is exactly why nobody had a mental model of it. Describing the endpoint returned six associated route tables. Those six covered every subnet the platform team had provisioned since 2024: three private application subnets and three private data subnets. The three public subnets, one per zone and each holding a NAT gateway, never needed it.&lt;/p&gt;

&lt;p&gt;Then we mapped node private IPs to subnet CIDRs. The rollup pod's node sat in 10.42.174.0/24, which belonged to a route table created on 2026-07-05 and named for a memory-optimised pool. It was not one of the six.&lt;/p&gt;

&lt;p&gt;Git history on the EKS module and a Slack thread from the same afternoon filled in the rest. On 2026-07-05 the model team asked for r6i-class nodes because their feature-store rebuild kept getting OOM-killed on the general pool. The platform team was heads-down on a certificate rotation, so a model-team engineer wrote the PR themselves: new node group, a new subnet per AZ, new route tables, default route to the NAT gateway. A peer on the same team reviewed it. It merged, and it was correct as far as it went.&lt;/p&gt;

&lt;p&gt;The subnet module had never been built to touch the gateway endpoint. The endpoint's route table associations live in a separate networking module with its own state file, owned by a different team, and nothing in the PR diff pointed at it. For one day it cost nothing, because nothing in the new subnets read S3. On 2026-07-06 the model team's validation job pulled a slice of the analytics data. Those were the first S3 bytes to leave via NAT.&lt;/p&gt;

&lt;p&gt;The second change landed that same evening, 2026-07-06, and did the real damage. The platform team let the rollup CronJob schedule onto the memory pool as well, because those nodes sat idle at night while the model jobs were dormant. One line in a nodeSelector and tolerations block, no ticket, a rubber-stamp review from a lead who was still mid certificate rotation. From that night on, a job reading 1.9 TB of parquet a day landed on nodes in a subnet with no route to the endpoint.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxdjzFrwzAQhff-ioemFuw0abYOBaeULE0ItuliMsj2xRWVpUOS4wb844ucZmg5uOXe93HvpO3YfEoX8J7fAYdVJZzVemCwbROwU2cZCJIZfqgNBXFEmr4gL7NKuJBK5nQlxTGyT39ZQyN66q27_Ed3M9pTf0PzMounSTg7BHpGsQY7OqlvaOUDgsWZGxITPg6vb5Uo1uhkoFFeQKZlq0xIYCyYXLrdIPbp6Fe8u4qNjdJZn8BbLBfzPC6jfJ-VYoq7EvusvLkT1EpraqMW2w3Y2Ya8p3Y2x0_mOsW6uo8f1UPzRSGBlz3BUaesEQ8xGZXX4A9P7m_n" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxdjzFrwzAQhff-ioemFuw0abYOBaeULE0ItuliMsj2xRWVpUOS4wb844ucZmg5uOXe93HvpO3YfEoX8J7fAYdVJZzVemCwbROwU2cZCJIZfqgNBXFEmr4gL7NKuJBK5nQlxTGyT39ZQyN66q27_Ed3M9pTf0PzMounSTg7BHpGsQY7OqlvaOUDgsWZGxITPg6vb5Uo1uhkoFFeQKZlq0xIYCyYXLrdIPbp6Fe8u4qNjdJZn8BbLBfzPC6jfJ-VYoq7EvusvLkT1EpraqMW2w3Y2Ya8p3Y2x0_mOsW6uo8f1UPzRSGBlz3BUaesEQ8xGZXX4A9P7m_n" alt="Same pod, same bucket, same region, same account. One line in a nodeSelector decided which of these two paths it took, and one of them has a meter on it." width="1203" height="222"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Same pod, same bucket, same region, same account. One line in a nodeSelector decided which of these two paths it took, and one of them has a meter on it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Neither change was wrong on its own. The PR was small and the reviewer was competent. Bin-packing a nightly batch job onto idle memory nodes is advice we would give. The failure is that two Terraform state files present subnet creation and endpoint association as unrelated concerns, so the coupling between them exists only in the head of whoever built the VPC. That seam shows up constantly in &lt;a href="https://infraforge.agency/migrations/" rel="noopener noreferrer"&gt;migration recovery&lt;/a&gt; work, where a landing zone gets stood up by one team and the endpoint policy by another, and nothing surfaces the gap until a bill or an outage finds it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually proves S3 traffic is using the gateway endpoint?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Checking the remote IP told us nothing&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;By 14:15 both engineers were on, with roughly three days before the billing cycle closed. There were two ways out and only one of them was the fix.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Pin the job back to the old pool&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One line, in place inside 20 minutes, though the bleeding would not have stopped until that night: the job is nightly, so the next run is the earliest either fix could show anything. We rejected it. The model team's validation job had already read S3 from those same subnets, and every future workload scheduled onto that pool would have re-armed the trap. Scheduling was where the cost showed up, not where the bug lived.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Add the three route tables to the endpoint&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Three new route-table associations in the networking module the platform lead owned. The plan read 3 to add, 0 to change, 0 to destroy, which is the diff you want to see at 14:30 with three days of the cycle still to run. We applied that one.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;We applied at 14:40, deliberately between rollup runs. Changing the path a multi-terabyte transfer is currently using is not an experiment worth running during an incident, and waiting cost us nothing we were not already spending.&lt;/p&gt;

&lt;p&gt;Then we nearly declared victory on a check that proved nothing. Someone shelled into a node on the new pool, made a request against the regional S3 endpoint, and read back the remote address. It came back a public S3 address, and it comes back as a public S3 address either way. A gateway endpoint does not hand S3 a private address. It installs a route for the S3 managed prefix list into your route table and points that route at the endpoint. The destination on the packet is identical on both paths, so the remote IP is not evidence of anything.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# 1. which route tables does the gateway endpoint actually cover?
aws ec2 describe-vpc-endpoints \
  --filters Name=service-name,Values=com.amazonaws.eu-west-1.s3 \
  --query 'VpcEndpoints[].{Id:VpcEndpointId,Type:VpcEndpointType,RouteTables:RouteTableIds}'

# 2. does THIS route table carry the prefix-list route to the endpoint?
aws ec2 describe-route-tables \
  --route-table-ids rtb-0c9a1f4e7b2d5a613 \
  --query 'RouteTables[0].Routes[].{Dest:DestinationCidrBlock,Prefix:DestinationPrefixListId,Gateway:GatewayId,Nat:NatGatewayId}'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;On a covered route table the second command returns a route whose prefix-list ID is the S3 managed list with the vpce ID as its target. Project both target fields or the output lies to you: a NAT route reports under &lt;code&gt;NatGatewayId&lt;/code&gt; and leaves &lt;code&gt;GatewayId&lt;/code&gt; null, so a query that reads only GatewayId blanks the target on exactly the subnet you are diagnosing. Before the fix, the memory subnets returned one usable route, 0.0.0.0/0 with the NAT gateway under Nat and no prefix-list row at all.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The confirming evidence was the meter, not the address. We reran the top-talker query that same afternoon and both rollup shards were gone from the top ten, and on its own that proved nothing. The job is nightly. An afternoon window is empty whether or not the route landed, which makes it exactly the kind of check that cannot fail. The one that could was the next night's run window, and the meter behind it. Overnight the NAT gateways settled at about 750 GB/day against a pre-incident baseline near 760. Nothing further accrued before the cycle closed. The damage was already done: 22 days at about $130/day above baseline, roughly $2,860 of pure waste, every dollar of it avoidable by three route-table associations in a module nobody knew to open.&lt;/p&gt;

&lt;p&gt;Three controls went in that week, and they are not equally important. The load-bearing one took two attempts, and the first is worth describing because it is the version most teams reach for. We wrote roughly 40 lines of rego against the Terraform plan JSON, in CI, on every PR touching the networking directory, failing when a route table attached to a private subnet was absent from a gateway-type endpoint's associations. It would never have fired. The PR that caused this created subnets and route tables in the EKS module and never touched the networking directory, so the check does not run on the one PR shape that matters. And the two concerns live in separate state files, so a plan JSON for either workspace cannot see the other side: the networking plan holds the endpoint's association list and no knowledge of the new route tables, the subnet plan holds the route tables and no endpoint. The predicate has no left-hand side and passes on an empty set.&lt;/p&gt;

&lt;p&gt;What went in instead reads the account rather than a plan. A scheduled conformance check calls &lt;code&gt;describe-route-tables&lt;/code&gt; and &lt;code&gt;describe-vpc-endpoints&lt;/code&gt; and fails when a route table associated with a private subnet is missing from a gateway-type endpoint's associations. Because it queries the account it sees both state files' resources, and because it runs on a timer it does not care which directory a PR touched. The cost is real and worth stating: it catches the gap after the merge rather than before it, within an hour rather than at review time. An hour of NAT charges is about $5.40 at this volume. The plan-time version would have caught nothing at all.&lt;/p&gt;

&lt;p&gt;The second control is a tighter cost alert on NAT data processing, and getting there needs one step people skip. Cost Anomaly Detection cannot watch a usage type: a monitor's dimension is one of &lt;code&gt;SERVICE&lt;/code&gt;, &lt;code&gt;LINKED_ACCOUNT&lt;/code&gt;, &lt;code&gt;TAG&lt;/code&gt; or &lt;code&gt;COST_CATEGORY&lt;/code&gt;, and nothing else. So we created a Cost Category whose rule matches the NAT data-processing usage type and pointed a &lt;code&gt;COST_CATEGORY&lt;/code&gt; monitor at that, with a subscription threshold well under the old one. Match it with &lt;code&gt;CONTAINS&lt;/code&gt; on &lt;code&gt;NatGateway-Bytes&lt;/code&gt; rather than an exact string. Usage types carry a Region billing code, and only us-east-1 line items appear bare, so in this account (eu-west-1) the value is &lt;code&gt;EU-NatGateway-Bytes&lt;/code&gt;, and Ireland uses the legacy &lt;code&gt;EU&lt;/code&gt; code rather than the &lt;code&gt;EUW1&lt;/code&gt; you would guess from the Region name. An exact rule on the unprefixed form categorises nothing and the replacement monitor never fires, which is the original failure wearing a new hat. An AWS Budget filtered on &lt;code&gt;UsageType&lt;/code&gt; works too and is quicker to stand up. Be careful about why the old monitor stayed quiet, because the obvious reading is wrong and it changes how you size the new one. A Cost Anomaly Detection threshold is not a daily rate: it is the anomaly’s total cost impact, actual minus expected spend accumulated over the anomaly’s whole duration, so a $500 threshold against a $130/day drift is four days away, not unreachable. So why did nothing fire? Not because the step was small: EC2 - Other sat near $1,150 a month, about $38 a day, and $130 a day on top of that is more than triple the pool, which is the 4.6x the preview showed. A service-level monitor is exactly what catches that. It did not fire because there was no subscription on that monitor anyone read: the alerts were addressed to a shared mailbox the platform team had stopped opening when it filled with RI-expiry notices. The lesson is not about thresholds at all, it is that an unread alert and no alert cost the same. Scoped to the category, the same pattern surfaces inside about three days. The third is CODEOWNERS on the node group and subnet modules, routing review to the platform engineer who owns the VPC rather than to any platform engineer who is free. The change looked EKS-shaped. The risk was VPC-shaped. We also updated the runbook, which is not a control at all: documentation explains why, it does not block a merge. If you are auditing this class of coupling across your own modules, it sits in the same family as the &lt;a href="https://infraforge.agency/terraform-iac-debt/" rel="noopener noreferrer"&gt;Terraform and IaC debt&lt;/a&gt; problems that only surface on the invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ: S3 gateway endpoints, NAT charges and new subnets
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The four questions the retro kept circling&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does an S3 gateway endpoint cost anything to run? Gateway endpoints carry no hourly and no per-GB charge, which is why the gap between the two paths in this story is the entire NAT data-processing bill. Interface endpoints are the ones with hourly plus per-GB pricing, and they are a different decision.&lt;/li&gt;
&lt;li&gt;We attached the endpoint but S3 still resolves to a public address. Did it work? Yes, that is expected. The gateway endpoint changes routing, not addressing. Verify by reading the route table for a prefix-list route targeting the vpce, then watch the NAT gateway bytes metric fall. The remote IP will look the same on both paths.&lt;/li&gt;
&lt;li&gt;Does this hit ECR image pulls too? Yes. Layer downloads come from S3-backed storage, so pods cold-starting in a subnet without the route pull their images through NAT. On the new pool that was about 20 cold starts a day at roughly 500 MB an image, near 10 GB/day. Small next to 3.5 TB, but it is the signal that arrives before your batch job does.&lt;/li&gt;
&lt;li&gt;Do interface endpoints have the same route table problem? Different failure mode. An interface endpoint is an ENI placed in specific subnets plus a private DNS name, so it does not depend on a route table association. Create a new subnet and leave it off the endpoint's subnet list, and resolution still points at an ENI in another AZ. It works, and it will not show on the bill: since April 2022 AWS does not charge inter-AZ transfer for traffic through an interface endpoint. What you get instead is every call from that subnet crossing a zone boundary, which is added latency and a dependency on another zone that nobody designed. Check it whenever you add an AZ.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When the bill moves and nothing in the cluster did
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If your EC2 - Other line just doubled&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hard part of this class of incident is that every signal you normally trust reads clean. Node count normal, pod count normal, error rates normal, mesh RPS normal. The expensive traffic is going somewhere your dashboards consider boring, and the evidence sits in flow logs nobody queries and endpoint associations nobody reads, one Terraform state file away from the change that caused it. Teams lose weeks to it because there is nothing to debug, only something to notice.&lt;/p&gt;

&lt;p&gt;We come at this from the network side rather than the cluster side. Route tables, endpoint associations and flow logs first, then the workloads: which of your subnets are paying per gigabyte for traffic that should be free, and which module owns the line that has to change. The finding is usually smaller than the invoice that prompted the call.&lt;/p&gt;

&lt;p&gt;If your NAT data processing has stepped up and the cluster looks healthy, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review&lt;/a&gt; and we will start on your flow logs and gateway endpoint associations the same business day.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/s3-traffic-through-nat-missing-endpoint-route/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/s3-traffic-through-nat-missing-endpoint-route/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>networking</category>
      <category>recovery</category>
      <category>cloudnetworking</category>
    </item>
    <item>
      <title>How to recover a Helm 3 release stuck in pending-upgrade</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Thu, 03 Sep 2026 16:43:46 +0000</pubDate>
      <link>https://dev.to/infraforge/how-to-recover-a-helm-3-release-stuck-in-pending-upgrade-123j</link>
      <guid>https://dev.to/infraforge/how-to-recover-a-helm-3-release-stuck-in-pending-upgrade-123j</guid>
      <description>&lt;p&gt;A Helm release stuck in pending-upgrade blocks &lt;code&gt;helm upgrade&lt;/code&gt; with &lt;code&gt;another operation (install/upgrade/rollback) is in progress&lt;/code&gt;, and the useful surprise is that it does not block &lt;code&gt;helm rollback&lt;/code&gt;. The pending check lives in &lt;code&gt;prepareUpgrade&lt;/code&gt; and guards concurrent upgrades; &lt;code&gt;pkg/action/rollback.go&lt;/code&gt; has no equivalent. So try &lt;code&gt;helm rollback&lt;/code&gt; to the last good revision FIRST. Deleting a release-history Secret is a real mutation of Helm's storage and it is almost never the answer. A stuck &lt;code&gt;pending-install&lt;/code&gt; on revision 1 has no earlier revision to roll back to, but &lt;code&gt;helm uninstall&lt;/code&gt; has no pending check either and clears the abandoned objects with it, so that case is an uninstall and a reinstall. On a payments platform we work with, that was revision 47 sitting pending for 22 minutes after the CI runner was killed mid-upgrade, with a chart bump that had added a &lt;code&gt;values.schema.json&lt;/code&gt; the production values file no longer satisfied waiting to break the retry. This guide is the order we run it in, with the two moves that will cost you a resource if you get them backwards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;helm upgrade exits immediately with &lt;code&gt;Error: UPGRADE FAILED: another operation (install/upgrade/rollback) is in progress&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;helm history -o json shows the newest revision as pending-upgrade with no deployed revision above it&lt;/li&gt;
&lt;li&gt;The release does not appear in &lt;code&gt;helm list -n &amp;lt;ns&amp;gt;&lt;/code&gt; at all, only in &lt;code&gt;helm list -a -n &amp;lt;ns&amp;gt;&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;A Secret created by the half-finished upgrade exists but its key decodes to 0 bytes, and the pods CrashLoop on auth&lt;/li&gt;
&lt;li&gt;helm upgrade with the existing production values file returns &lt;code&gt;values don't meet the specifications of the schema(s) in the following chart(s)&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why does helm upgrade say another operation is in progress?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Read the history before you type rollback&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The lock is a stored revision record, not a process holding a mutex.&lt;/p&gt;

&lt;p&gt;This guide assumes Helm 3 driving the chart directly, from CI or from a laptop, with the default &lt;code&gt;secret&lt;/code&gt; storage driver, against EKS 1.29. If Argo CD renders your chart and applies the manifests itself, none of what follows applies, because in that mode there is no Helm release history to inspect and the recovery is a git and sync-status problem instead.&lt;/p&gt;

&lt;p&gt;Helm does not hold a lock in memory. It writes a Secret per revision, named &lt;code&gt;sh.helm.release.v1.&amp;lt;release&amp;gt;.v&amp;lt;n&amp;gt;&lt;/code&gt;, and every operation reads the newest one first. If that newest record says &lt;code&gt;pending-upgrade&lt;/code&gt;, &lt;code&gt;helm upgrade&lt;/code&gt; concludes an operation is still running somewhere and refuses to start another. The CI job that wrote it may have been killed 22 minutes ago. Helm has no way to know that, so the release stays wedged until you change the record.&lt;/p&gt;

&lt;p&gt;The first surprise is that the release can look like it does not exist. &lt;code&gt;helm list -n settlement&lt;/code&gt; filters to deployed and failed releases, so a pending one is simply absent from the output, which sends people looking for a deleted release that is sitting right there. Run &lt;code&gt;helm list -a -n settlement&lt;/code&gt; instead and it appears.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# The two cheap reads. Do both before you touch anything.
helm history payments-worker -n settlement -o json \
  | jq -r '.[] | "\(.revision)\t\(.status)\t\(.description)"'

# Same answer straight from the storage layer, no Helm binary involved:
kubectl get secret -n settlement \
  -l owner=helm,name=payments-worker \
  -L version,status --sort-by=.metadata.creationTimestamp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;helm history renders an aligned table by default, which is useless in a ticket. The -o json form pipes cleanly, and the kubectl form works when the Helm binary is arguing with you.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In our case the JSON came back with revision 46 still marked &lt;code&gt;deployed&lt;/code&gt; and revision 47 as &lt;code&gt;pending-upgrade&lt;/code&gt;, description &lt;code&gt;Preparing upgrade&lt;/code&gt;. That is the shape you are looking for: a head revision in a pending state, with the last good revision still &lt;code&gt;deployed&lt;/code&gt; beneath it. Helm demotes the previous revision to &lt;code&gt;superseded&lt;/code&gt; only on the success path, so a run killed mid-apply leaves 46 exactly where it was. If instead the head revision says &lt;code&gt;failed&lt;/code&gt;, you are in a different and much easier situation, because Helm will happily accept a &lt;code&gt;helm rollback&lt;/code&gt; against a failed release without any of the surgery below.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJyNkk1u2zAQhfc-xTuAFaAXaFHHQdwEKAy0WQRGFrQ4lhjTHGGGjCFYuXtAyaIX3XQnCjPf-8EcPJ_r1kjE3_UC-LlryZ_QOo0sPSrGu3J4Q1V9x-qi0cSk4ANaMhZCH04dh88FsMojg6XOc092wP3uN8fWhQYaU328wysnQSe893SCU5BXOrck9Fa2D8b5vLueTAh7vzf1EZHhjUY0zDfR21pHwbrQVKlrxFgCC-ZfM2LA88UpYtaDCSAj3pEU2I_Pf2guaDTeg0OZwrcB24t1NoMwD9RCJhJ4_051VPSccEoacSTqRu7zyO1JBzz-b7JpJ_CA7QLYltfLBEjhqr7MVkLxYhrjxv3tTfNht8pSqYOQJ6OEP1QLRV3CkqdsPfh-jHTNDqGaxS4hNIPPLraoqmiOVPE5kGjruiz0OJ7G5lIikQiL5toMTk514iknqWmsY3Oz9mu3nhxkce8-5havqQqzxNqUIp6mIq4XKdSxRMV8f3n2ZXT2tAAepq8vXjn_vg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJyNkk1u2zAQhfc-xTuAFaAXaFHHQdwEKAy0WQRGFrQ4lhjTHGGGjCFYuXtAyaIX3XQnCjPf-8EcPJ_r1kjE3_UC-LlryZ_QOo0sPSrGu3J4Q1V9x-qi0cSk4ANaMhZCH04dh88FsMojg6XOc092wP3uN8fWhQYaU328wysnQSe893SCU5BXOrck9Fa2D8b5vLueTAh7vzf1EZHhjUY0zDfR21pHwbrQVKlrxFgCC-ZfM2LA88UpYtaDCSAj3pEU2I_Pf2guaDTeg0OZwrcB24t1NoMwD9RCJhJ4_051VPSccEoacSTqRu7zyO1JBzz-b7JpJ_CA7QLYltfLBEjhqr7MVkLxYhrjxv3tTfNht8pSqYOQJ6OEP1QLRV3CkqdsPfh-jHTNDqGaxS4hNIPPLraoqmiOVPE5kGjruiz0OJ7G5lIikQiL5toMTk514iknqWmsY3Oz9mu3nhxkce8-5havqQqzxNqUIp6mIq4XKdSxRMV8f3n2ZXT2tAAepq8vXjn_vg" alt="The rollback is the first branch. A stuck pending-install on revision 1 has no earlier revision to roll back to, so it goes to helm uninstall, not to a hand-edit of Helm storage." width="1336" height="1545"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The rollback is the first branch. A stuck pending-install on revision 1 has no earlier revision to roll back to, so it goes to helm uninstall, not to a hand-edit of Helm storage.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the rollback first, then clear the record only if you must
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Back up every release Secret first&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Start with the rollback, because the pending guard does not apply to it. &lt;code&gt;helm rollback payments-worker 46 -n settlement --wait --timeout 5m&lt;/code&gt; writes a new revision rather than destroying one. If it returns cleanly, you are done.&lt;/p&gt;

&lt;p&gt;We did not start there. On the night this happened we went straight to clearing the record by hand, and the rollback we ran afterwards is the one that fixed it. Reading &lt;code&gt;pkg/action/rollback.go&lt;/code&gt; later made the point uncomfortable: there is no pending check in the rollback path, so that rollback would have worked on its own and the Secret delete bought us nothing. It is worth being plain about that, because the procedure below circulates as folklore and most of the people running it do not need to.&lt;/p&gt;

&lt;p&gt;You need it in one situation, and it is narrower than the folklore suggests. When the head revision is &lt;code&gt;pending-install&lt;/code&gt; on revision 1 there is no earlier revision to roll back to, but that does not make a hand-edit of storage the answer. &lt;code&gt;helm uninstall&lt;/code&gt; carries no pending check either, &lt;code&gt;uninstall.go&lt;/code&gt; tests only for an already-uninstalled release, and it clears both the release record and the objects the half-finished install left behind. Uninstall, then install again. Deleting the Secret by hand is defensible only when those objects have to stay in place, and then you still owe Helm an adoption step: &lt;code&gt;--take-ownership&lt;/code&gt; on the next install or upgrade, which landed in Helm 3.17. A failed rollback is NOT that situation. &lt;code&gt;Rollback.Run&lt;/code&gt; calls &lt;code&gt;Releases.Create&lt;/code&gt; with &lt;code&gt;pending-rollback&lt;/code&gt; before it touches any resource, and on failure records that same revision as &lt;code&gt;failed&lt;/code&gt;, so after a rollback that errors your head is a failed rollback revision and the fix is at the resource level, not in storage. Re-read &lt;code&gt;helm history&lt;/code&gt; before you conclude otherwise.&lt;/p&gt;

&lt;p&gt;Before deleting a revision record, take the whole history to a file. This is the one step that turns an irreversible mistake into an inconvenient one,   and it costs six seconds.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get secret &lt;span class="nt"&gt;-n&lt;/span&gt; settlement &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-l&lt;/span&gt; &lt;span class="nv"&gt;owner&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;helm,name&lt;span class="o"&gt;=&lt;/span&gt;payments-worker &lt;span class="nt"&gt;-o&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="s1"&gt;'del(.items[].metadata.resourceVersion, .items[].metadata.uid,
         .items[].metadata.creationTimestamp, .items[].metadata.managedFields)'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /tmp/payments-worker-helm-history.json

&lt;span class="c"&gt;# Confirm the file has every revision that still exists, not just the ones you remember.&lt;/span&gt;
&lt;span class="c"&gt;# With the default --history-max 10 this lists about ten, not one per revision.&lt;/span&gt;
jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.items[].metadata.name'&lt;/span&gt; /tmp/payments-worker-helm-history.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The jq filter is the point. A plain get -o yaml carries resourceVersion, which the API server refuses outright on a create, plus uid and creationTimestamp that it would silently overwrite, so the untouched dump restores nothing at the moment you need it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Now confirm which revision is actually stuck, and confirm it from the payload rather than the label. The release body inside that Secret is not JSON and you cannot patch it as JSON. Helm gzips the release JSON, base64 encodes the result itself, and then Kubernetes base64 encodes the whole thing again into &lt;code&gt;data.release&lt;/code&gt;. Open the Secret in an editor and you get an opaque blob. Anyone who tells you to flip the status field from &lt;code&gt;pending-upgrade&lt;/code&gt; to &lt;code&gt;failed&lt;/code&gt; with a &lt;code&gt;kubectl patch&lt;/code&gt; has not tried it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get secret sh.helm.release.v1.payments-worker.v47 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-n&lt;/span&gt; settlement &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;jsonpath&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{.data.release}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;base64&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; | &lt;span class="nb"&gt;base64&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; | &lt;span class="nb"&gt;gunzip&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.info.status, .info.description'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Two base64 decodes, then gunzip. One decode gives you more base64 and people assume the data is corrupt.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That printed &lt;code&gt;pending-upgrade&lt;/code&gt; and &lt;code&gt;Preparing upgrade&lt;/code&gt;, which matched the label. Read what follows as the record of what we ran that night, not as the step to copy. We deleted that single Secret, which made revision 46 the head of the history again and let Helm treat the release as one it could operate on. With 46 sitting there &lt;code&gt;deployed&lt;/code&gt;, the rollback on its own would have done the same work, so the delete bought us nothing. It is written out because these are the commands that circulate as folklore, and the two judgment calls buried in them are worth having before you are somewhere you need them. If your own head revision is pending and an earlier revision is beneath it, stop at the rollback.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl delete secret sh.helm.release.v1.payments-worker.v47 &lt;span class="nt"&gt;-n&lt;/span&gt; settlement

helm rollback payments-worker 46 &lt;span class="nt"&gt;-n&lt;/span&gt; settlement &lt;span class="nt"&gt;--wait&lt;/span&gt; &lt;span class="nt"&gt;--timeout&lt;/span&gt; 5m
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Delete the record, not the release: &lt;code&gt;helm uninstall&lt;/code&gt; here would take the workload down with it, which is why the uninstall route belongs to a stuck &lt;code&gt;pending-install&lt;/code&gt; on revision 1 and not to this. This pair is what we ran, not what this case needs. A pending head with an earlier revision under it stops at the rollback.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two judgment calls in that pair of commands. Delete exactly one Secret, the pending one, and name it in full. A label selector delete against &lt;code&gt;owner=helm,name=payments-worker&lt;/code&gt; removes the entire history including the revision you are about to roll back to, and then your only route home is the backup file you just wrote. Second, name the target revision explicitly. &lt;code&gt;helm rollback payments-worker -n settlement&lt;/code&gt; with no revision argument rolls back to the revision before the head. That is the trap after a delete: with 47 gone the head IS 46, so a bare rollback targets 45 and quietly skips the revision you were aiming for. Type the number.&lt;/p&gt;

&lt;p&gt;We use &lt;code&gt;--wait --timeout 5m&lt;/code&gt; rather than a bare rollback because a rollback that returns success while pods are still terminating tells you nothing. With &lt;code&gt;--wait&lt;/code&gt;, Helm returns only after the Deployment reports its expected replicas ready, so a non-zero exit is real information. The cost is honest: for as long as the rollback actually runs, up to that five minute ceiling, no &lt;code&gt;helm upgrade&lt;/code&gt; against that release will start. A second rollback or an uninstall still would, since neither carries the pending guard.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 'no ConfigMap with the name X found' means during a rollback
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;When rollback fails on a live resource Helm has no record of&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The second failure people hit is a rollback that gets past the lock and then dies on a specific resource. The message reads like Helm is confused about what exists. It is not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: no ConfigMap with the name "payments-worker-broker" found
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Helm 3 emits this from its update path when the object exists in the cluster but has no entry in the release record it is diffing against. A rollback prints it bare, as here, while the same condition reached through &lt;code&gt;helm upgrade&lt;/code&gt; arrives prefixed with &lt;code&gt;UPGRADE FAILED:&lt;/code&gt;. Match on the &lt;code&gt;no &amp;lt;Kind&amp;gt; with the name "&amp;lt;name&amp;gt;" found&lt;/code&gt; body, which is stable across both, rather than on the prefix. Rollback records its own &lt;code&gt;Rollback "payments-worker" failed: ...&lt;/code&gt; as the revision description, not on stderr.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Get the direction right, because getting it backwards sends you looking in the wrong place for an hour. &lt;code&gt;Update&lt;/code&gt; in Helm 3's &lt;code&gt;pkg/kube/client.go&lt;/code&gt; takes the previous release's manifest and the target manifest, and walks the target. It looks each resource up live first, and a miss there is harmless: Helm creates the object and carries on. The error comes one step later, when it looks the same resource up in the previous release's manifest and finds nothing. So the object is in the cluster, and in the manifest being applied, and absent from the record Helm is comparing against. That is what a rollback to revision 46 runs into when something outside the release created the object, or an earlier half-finished upgrade left it behind without recording it.&lt;/p&gt;

&lt;p&gt;The recovery is to delete the live object and run the rollback again. Deleting it makes Helm's live lookup miss, and a miss is the harmless path: Helm creates the resource from revision 46's manifest and records it properly this time. Read what revision 46 expects first with &lt;code&gt;helm get manifest payments-worker --revision 46 -n settlement&lt;/code&gt;, so you know what is about to be recreated and can confirm nothing else depends on the object's current contents. Do not reach for &lt;code&gt;--force&lt;/code&gt; on the rollback, and the reason is sharper than "it is risky": on this path it does nothing at all. &lt;code&gt;force&lt;/code&gt; is only consulted inside &lt;code&gt;updateResource&lt;/code&gt;, and this error never gets that far: the visitor checks the cluster with &lt;code&gt;helper.Get&lt;/code&gt; first, then looks the resource up in the previous release's manifest, and returns &lt;code&gt;no %s with the name %q found&lt;/code&gt; from that second lookup. The flag sits downstream of the point where it fails. Where &lt;code&gt;--force&lt;/code&gt; does apply it sends a full replace (&lt;code&gt;helper.Replace&lt;/code&gt;, a PUT) instead of patching, which discards fields another controller owns, an HPA-managed &lt;code&gt;replicas&lt;/code&gt; being the usual casualty, and fails outright on immutable resources such as a &lt;code&gt;Job&lt;/code&gt; or a &lt;code&gt;Service&lt;/code&gt; clusterIP.&lt;/p&gt;

&lt;p&gt;There is a related trap on the workload itself. A half-finished upgrade frequently creates a Secret from a template whose input value never rendered, so the Secret exists, the Deployment mounts it, and the key inside is an empty string. Kubernetes is perfectly happy with that. The pods are not, and you get a CrashLoopBackOff whose logs blame the broker rather than the chart. Check the length, not the presence.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Presence proves nothing. Length does.&lt;/span&gt;
kubectl get secret payments-worker-broker-auth &lt;span class="nt"&gt;-n&lt;/span&gt; settlement &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;jsonpath&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{.data.password}'&lt;/span&gt; | &lt;span class="nb"&gt;base64&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt;
&lt;span class="c"&gt;# 0&lt;/span&gt;

&lt;span class="c"&gt;# After a clean rollback or upgrade, the same command returns the real byte count.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;A Secret that decodes to 0 bytes passes every existence check and fails every connection.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Verification is four reads and they should all agree. &lt;code&gt;helm status payments-worker -n settlement -o json | jq -r '.info.status'&lt;/code&gt; returns &lt;code&gt;deployed&lt;/code&gt;. &lt;code&gt;helm history payments-worker -n settlement -o json | jq -r '.[-1].status'&lt;/code&gt; returns &lt;code&gt;deployed&lt;/code&gt; for the head revision. The old pending entry is still listed if you got here by rolling back, because a rollback appends a revision rather than removing one; it is gone only if you deleted its Secret. &lt;code&gt;kubectl get pods -n settlement -l app.kubernetes.io/name=payments-worker&lt;/code&gt; shows every pod Running with a restart count that stops climbing over the next few minutes. And &lt;code&gt;helm get manifest&lt;/code&gt; against the head revision matches what is live, which is the check that catches a rollback that succeeded on paper while something else was quietly reconciling the cluster back.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you fix a chart schema error without deleting the schema?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Getting the values file past values.schema.json&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The release is deployable again, and now you still have the original problem: the chart version you were upgrading to ships a &lt;code&gt;values.schema.json&lt;/code&gt; that your production values file does not satisfy. Helm validates values against that schema in &lt;code&gt;prepareUpgrade&lt;/code&gt;, before it writes a revision or touches the API server. That ordering matters more than it looks: a run that fails schema validation creates no release record at all, so a schema error can never be what left you in &lt;code&gt;pending-upgrade&lt;/code&gt;. Revision 47 wedged because the runner was killed; the schema was waiting to break the retry, which is exactly what it did.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: UPGRADE FAILED: values don't meet the specifications of the schema(s) in the following chart(s):
payments-worker:
- broker.pool.maxIdle: Invalid type. Expected: integer, given: string
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The chart did not change what the field means. It started enforcing a type that was previously accepted as free text.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That is the whole class of failure. A field that was unstructured for two years held &lt;code&gt;"25"&lt;/code&gt; in quotes because someone templated it out of a CI variable, and the new schema declares it an integer. Nothing about the running workload was wrong. The schema simply started looking.&lt;/p&gt;

&lt;p&gt;Fix it locally before you go near the cluster. &lt;code&gt;helm lint ./charts/payments-worker -f values/production.yaml&lt;/code&gt; runs the same schema validation without a Kubernetes connection, which turns a ten minute deploy-and-fail cycle into a two second one. Once lint is clean, &lt;code&gt;helm upgrade payments-worker ./charts/payments-worker -n settlement -f values/production.yaml --dry-run=server&lt;/code&gt; renders with a cluster connection, so &lt;code&gt;lookup&lt;/code&gt; functions resolve and &lt;code&gt;.Capabilities.APIVersions&lt;/code&gt; reflects what the cluster actually serves instead of Helm's built-in defaults. That catches a template reaching for an API version the cluster has dropped. It does not submit anything for admission review, so a Gatekeeper, Kyverno or pod-security rejection is still waiting for you on the real upgrade; to cover that, run &lt;code&gt;helm template ... | kubectl apply --server-side --dry-run=server -f -&lt;/code&gt;. The server-side form of &lt;code&gt;--dry-run&lt;/code&gt; arrived in Helm 3.13, so on older clients you get the client-side render only.&lt;/p&gt;

&lt;p&gt;We do not delete or blank out the &lt;code&gt;values.schema.json&lt;/code&gt; in a vendored copy of the chart to make the error go away. We have inherited two clusters where someone did exactly that, and in both the next chart bump reintroduced the schema and the same incident happened again with a different on-call engineer and no memory of the first one. Fix the values file. If the schema itself is genuinely wrong for your use, pin the chart version, open the issue upstream, and write the pin's reason in the values file where the next person will read it.&lt;/p&gt;

&lt;p&gt;For prevention, &lt;code&gt;helm upgrade --atomic&lt;/code&gt; helps, but not with the failure in this story. It implies &lt;code&gt;--wait&lt;/code&gt; and rolls the release back if the upgrade does not converge, and that rollback is issued by the Helm client process itself. Kill that process and nothing is left to issue it, so a hard-killed CI job still leaves the release pending. What guards against this specific case is the opposite ordering to the one people reach for: Helm's &lt;code&gt;--timeout&lt;/code&gt; must expire BEFORE the CI runner's own job timeout. Then Helm gives up on its own terms, unwinds, and exits. Set the runner's limit shorter and it kills Helm partway through, which is precisely how the release ends up pending. On Helm 4 the flag is &lt;code&gt;--rollback-on-failure&lt;/code&gt;; &lt;code&gt;--atomic&lt;/code&gt; is deprecated on &lt;code&gt;helm upgrade&lt;/code&gt; and is an unknown flag on &lt;code&gt;helm install&lt;/code&gt;. The tradeoff is real and we tell clients about it up front: &lt;code&gt;--atomic&lt;/code&gt; holds the release for the entire timeout window, so on a rollout that takes four minutes with a ten minute timeout, a failed deploy blocks the next one for ten minutes rather than failing fast. That is usually the right trade for a payments path and usually the wrong one for a batch worker that deploys thirty times a day. We walk through where that line sits per service in &lt;a href="https://infraforge.agency/kubernetes-cicd/" rel="noopener noreferrer"&gt;our Kubernetes and CI/CD stabilization work&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common questions about clearing a stuck Helm release
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;What people ask after the first rollback lands&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Will helm rollback work while the release is still in pending-upgrade?&lt;/strong&gt; Yes, and it is the first thing to try. The guard is upgrade-only: &lt;code&gt;prepareUpgrade&lt;/code&gt; returns &lt;code&gt;errPending&lt;/code&gt; when &lt;code&gt;lastRelease.Info.Status.IsPending()&lt;/code&gt;, and its own comment scopes it to concurrent upgrades acting as a pessimistic lock. &lt;code&gt;pkg/action/rollback.go&lt;/code&gt; carries no such check, so a rollback to a known-good revision runs and leaves a new deployed revision as the head. The pending row keeps its &lt;code&gt;pending-upgrade&lt;/code&gt; status in the history, because &lt;code&gt;performRollback&lt;/code&gt; supersedes only revisions that were already &lt;code&gt;deployed&lt;/code&gt;, and it blocks nothing once the guard reads a deployed head. When there is no prior revision at all, a pending-install on revision 1, the answer is still not storage surgery: &lt;code&gt;helm uninstall&lt;/code&gt; has no pending check either and takes the abandoned objects with it, so uninstall and install again. A rollback that ERRORS is not that case: it has already written its own failed revision, so re-read the history before touching storage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is deleting a sh.helm.release.v1 Secret safe?&lt;/strong&gt; Deleting one revision record removes Helm's memory of that revision. It does not touch a single running resource. The danger is scope, not the act: delete by full name, never by label selector, and keep the backup file until &lt;code&gt;helm status&lt;/code&gt; reports deployed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can I just patch the status field inside the release Secret?&lt;/strong&gt; The payload is gzipped JSON, base64 encoded by Helm and then base64 encoded again by Kubernetes, so there is no JSON there to patch. If you want to inspect it, use the decode round trip above. To change the release state, use the supported route: clear the pending revision, then roll back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does any of this apply if Argo CD manages the app?&lt;/strong&gt; No. When Argo CD renders the chart and applies the manifests itself, there is no Helm release history in the cluster to repair. The failure looks similar and the fix is entirely different.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My rollback succeeded but the pods still CrashLoop.&lt;/strong&gt; Check the Secret and ConfigMap byte lengths, not their existence. A key that decodes to 0 bytes satisfies every reference check the Deployment does and fails at connection time, which is why the logs blame a downstream service.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When the stuck release is on the payments path and nobody wants to type delete
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If the release has been pending for an hour&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hard part of this procedure is not the commands. It is that the correct move is &lt;code&gt;kubectl delete secret&lt;/code&gt; against a revision record on a production release, with an engineering lead watching, at the point in an incident where confidence is lowest. The failure modes are unforgiving in a specific way: a label-selector delete that takes the whole history, a rollback to an implicit revision that lands one further back than you meant, a &lt;code&gt;--force&lt;/code&gt; that replaces a StatefulSet and drops fields you needed left alone. Every one of those is recoverable if the backup exists and unrecoverable if it does not.&lt;/p&gt;

&lt;p&gt;We do this work with teams who deploy through Helm from CI and have never had to open a release Secret before. Our part is being the second pair of eyes on the delete, reading the history with you before anything is typed, and then leaving behind the &lt;code&gt;--atomic&lt;/code&gt; and timeout settings that stop the same job from wedging the release next quarter. Most of these calls run under an hour once we can see &lt;code&gt;helm history&lt;/code&gt; output.&lt;/p&gt;

&lt;p&gt;If a release is pending right now and you would rather not be the one running the delete, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review&lt;/a&gt; and we will get on a call and work through the history with you the same day. If it is already recovered and you want the CI path hardened so it does not recur, that is the same conversation with less adrenaline, and it is the shape of work described in &lt;a href="https://infraforge.agency/problems/kubernetes-release-failures/" rel="noopener noreferrer"&gt;our Kubernetes release failure playbook&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/helm-rollback-failed-release-recovery/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/helm-rollback-failed-release-recovery/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>k8s</category>
      <category>howto</category>
      <category>kubernetescicd</category>
    </item>
    <item>
      <title>ServerlessDatabaseCapacity pinned: Aurora v2 will not scale down</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Tue, 01 Sep 2026 17:02:09 +0000</pubDate>
      <link>https://dev.to/infraforge/serverlessdatabasecapacity-pinned-aurora-v2-will-not-scale-down-55di</link>
      <guid>https://dev.to/infraforge/serverlessdatabasecapacity-pinned-aurora-v2-will-not-scale-down-55di</guid>
      <description>&lt;p&gt;ServerlessDatabaseCapacity sitting flat near max_capacity while pg_stat_activity is clean almost always means one of two things: a min_capacity or max_capacity change someone made in the console and never reverted, or a logical replication slot whose restart_lsn has stopped advancing. It is almost never a runaway query. Check the scaling configuration first, because that is one API call and it rules out half the problem. Then read pg_replication_slots, before you restart or drop anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ServerlessDatabaseCapacity is a flat line for days where it used to be a sawtooth, with no matching rise in DatabaseConnections&lt;/li&gt;
&lt;li&gt;Cost Explorer shows Aurora:ServerlessV2Usage stepping up on one calendar day and holding, while every other RDS usage type is flat&lt;/li&gt;
&lt;li&gt;Capacity at 04:00 UTC is the same number as capacity at 14:00 UTC&lt;/li&gt;
&lt;li&gt;pg_stat_activity has nothing older than 60 seconds and pg_stat_progress_vacuum returns zero rows&lt;/li&gt;
&lt;li&gt;pg_replication_slots shows an active slot with hundreds of GB between restart_lsn and pg_current_wal_lsn()&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What ServerlessDatabaseCapacity flat near max_capacity actually costs
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;218 hours at 148 ACU and not one page&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A claims-analytics platform we work with runs one Aurora PostgreSQL Serverless v2 cluster in us-east-1: a writer and a single reader, roughly 15 services in front of it. For 218 hours the pair sat at a combined 148 ACU. The writer floated between 96 and 128, the reader between 32 and 48. Nothing paged. Latency was fine, the on-call rotation had a quiet week, and the only artifact of the whole thing was a CloudWatch chart that had gone flat where it used to be a sawtooth.&lt;/p&gt;

&lt;p&gt;At the us-east-1 list rate of $0.12 per ACU-hour, 148 ACU-hours per hour for 218 hours is $3,872 of capacity. The same window at that cluster's normal 6.4 ACU average would have been $167. The gap was found by a monthly close-out preview, not by anything in the observability stack, which is the part worth sitting with. This article assumes Aurora PostgreSQL 15.x on Serverless v2, AWS CLI v2, and that you have psql access to the writer.&lt;/p&gt;

&lt;p&gt;The flat line is most of the diagnosis. Aurora Serverless v2 scales up in seconds and comes down gradually, because shrinking capacity means giving back memory that the buffer cache is holding. A cluster with a genuinely quiet overnight window draws a sawtooth. A cluster that draws a flat line at or near its ceiling is either being told to stay there, or being kept from leaving. Those are two different bugs with two different fixes, and they are frequently both present at once.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws ce get-cost-and-usage \
  --time-period Start=2026-07-15,End=2026-08-01 \
  --granularity DAILY \
  --metrics UnblendedCost UsageQuantity \
  --filter '{"Dimensions":{"Key":"SERVICE","Values":["Amazon Relational Database Service"]}}' \
  --group-by Type=DIMENSION,Key=USAGE_TYPE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The dollar total is smoothed by every other RDS instance in the account. The Aurora:ServerlessV2Usage line is not. In regions other than us-east-1 the usage type carries a region prefix.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Run that before you run anything on the database, because it dates the step change to a single day. Knowing the cluster went from 6 to 148 ACU-hours per hour on a Tuesday morning and never came back is worth more than an hour of query analysis. It converts an open-ended performance question into "what happened on that day", and somebody's memory usually answers it in about four minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which three causes pin Aurora Serverless v2 at high capacity?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;pg_stat_activity was clean, which is the useful part&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We rank these by how often we actually find them, not by how interesting they are. The first one accounts for more of these calls than the other two combined, and it is the one nobody wants to check because it feels too dumb to be the answer.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. A scaling change made in the console and never reverted&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Someone raised max_capacity for a load test, a migration cutover, or a seasonal peak, and raised min_capacity at the same time so the cluster would not warm up between iterations. The work finished, the revert became a mental note, and the mental note lost to an afternoon incident. On the cluster above, min_capacity had gone from 2 to 16 on both instances. That floor alone bills $80.64 a day whether anyone queries the database or not.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. A logical replication slot whose restart_lsn stopped advancing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A CDC connector, a DMS task, or a hand-rolled pglogical consumer holds a slot open permanently. That is normal. What is not normal is restart_lsn standing still: the cluster retains every WAL segment behind it, the walsender stays attached, and the writer never reaches the quiet state that scale-down needs. This one hides well because the consumer's own health check is green.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. A resident working set plus a job that never lets it go idle&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A rollup that runs every five minutes and briefly touches a large table will hold the buffer cache warm and reset the scale-down evaluation before it completes. Capacity ratchets up on each burst and never gets a long enough gap to come down. You see this as a capacity floor that is high but not at the ceiling, with a saw-tooth of a few ACU riding on top of it.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two things teams reach for instead are a runaway query and a connection leak. Both are real failure modes and both are visible in a single query, so spend the ninety seconds and rule them out: if nothing in pg_stat_activity is older than 60 seconds, if the connection count matches last month, and if pg_stat_progress_vacuum is empty, stop looking at the workload. The other reflex is to blame the service, as in "Serverless v2 just does not scale down properly". It scales down. Something on this cluster is asking it not to, and the next section tells you which.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check that separates a console change from a stuck replication slot
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Two calls, ninety seconds, and the branch is decided&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# 1. What is the ceiling, and what is the floor you pay for around the clock?
aws rds describe-db-clusters \
  --db-cluster-identifier claims-analytics-prod \
  --query 'DBClusters[0].ServerlessV2ScalingConfiguration'

# 2. Is anything holding WAL? Run this on the writer.
SELECT slot_name,
       plugin,
       slot_type,
       active,
       active_pid,
       restart_lsn,
       pg_size_pretty(
         pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)
       ) AS retained_wal
FROM pg_replication_slots
ORDER BY pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) DESC;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;One call returns the ceiling you set. One query returns the thing that ignores it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first call returns a small object with MinCapacity and MaxCapacity. If MinCapacity reads 16 and your Terraform says 2, you are done with half the investigation: that is your floor, you have been paying it every hour since it changed, and CloudTrail will tell you who set it and when. Do not stop there, though. A raised floor explains a floor. It does not explain a writer sitting at 128.&lt;/p&gt;

&lt;p&gt;The second query is the one that decides the rest. On this cluster it returned exactly one row: a logical slot on the pgoutput plugin, active true, with a live active_pid, and retained_wal of 241 GB. The number by itself is not proof of anything, because a healthy consumer can be legitimately behind during a backfill. Run the query twice, sixty seconds apart, and compare &lt;code&gt;restart_lsn&lt;/code&gt; itself rather than the size sitting beside it. &lt;code&gt;retained_wal&lt;/code&gt; is a distance from a moving target: &lt;code&gt;pg_current_wal_lsn()&lt;/code&gt; advances with every write, so a constant retained_wal means &lt;code&gt;restart_lsn&lt;/code&gt; is advancing at exactly the write rate, which is a consumer holding a steady lag, not a stuck one. The stuck condition is &lt;code&gt;restart_lsn&lt;/code&gt; that has not moved at all between the two samples, and retained_wal growing is its confirming symptom. Falling means the consumer is catching up and what you have is a throughput problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxVksFqGzEQhu95iv8Bsq0LJYdQUmI7TgPtqYEeFmNmpbFXtVbjzmjtGNN3LyuviXvR5Z__49Mw6ygH15JmvM5vgMf6J-ueNbLZnDI1ZDyjHbmQj1hHykhMio7elqiqB0xPns1paLjyTeVib5nV7r80-vGhCwlSZuFaShv2X__eAFNUFY5spT-rZ5JMIo8jSLxnhQ5vZv-hgL7LgRWOQwxpczuGWEcRxT4QXlmV1qLd8oJPUujz026zUt7F4CgHSSuLkkc7ZcukeRUtoU-jYEnIqZjhbgJjJ8lb0Z5faz_V13XLvdueVRfhDbllOEnWd6y3IO_RMmlumLItL6TRcHF6MUw-308mcOOaC4f_9BSRBZ_-y4rJ4trkuf4lug1pA-MMZQueUz7L_OBO9Fg10icPcxS58nJIywtjdPhWP8Jcy76P7PFbGgQbSJxzSJtCGr70DoCL4rYDZlYAL_VUcotGKbmWDaQMWWdOyNpzAUgqazHqGOOVDP2nc_8f4WTXAg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxVksFqGzEQhu95iv8Bsq0LJYdQUmI7TgPtqYEeFmNmpbFXtVbjzmjtGNN3LyuviXvR5Z__49Mw6ygH15JmvM5vgMf6J-ueNbLZnDI1ZDyjHbmQj1hHykhMio7elqiqB0xPns1paLjyTeVib5nV7r80-vGhCwlSZuFaShv2X__eAFNUFY5spT-rZ5JMIo8jSLxnhQ5vZv-hgL7LgRWOQwxpczuGWEcRxT4QXlmV1qLd8oJPUujz026zUt7F4CgHSSuLkkc7ZcukeRUtoU-jYEnIqZjhbgJjJ8lb0Z5faz_V13XLvdueVRfhDbllOEnWd6y3IO_RMmlumLItL6TRcHF6MUw-308mcOOaC4f_9BSRBZ_-y4rJ4trkuf4lug1pA-MMZQueUz7L_OBO9Fg10icPcxS58nJIywtjdPhWP8Jcy76P7PFbGgQbSJxzSJtCGr70DoCL4rYDZlYAL_VUcotGKbmWDaQMWWdOyNpzAUgqazHqGOOVDP2nc_8f4WTXAg" alt="Two branches, and the common case is that both fire. Fixing only the scaling config leaves the cluster pinned and makes it look like the fix failed." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Two branches, and the common case is that both fire. Fixing only the scaling config leaves the cluster pinned and makes it look like the fix failed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here is the part a generic answer misses. A slot can be stuck while the connector is entirely healthy, because restart_lsn only advances when the consumer commits an offset for a change it captured. If the tables in the publication are quiet while the rest of the database is hot, the connector has nothing to commit, so it never acknowledges any position, so the cluster keeps every WAL segment written since the last real event. Every dashboard is green. Every health check passes. The slot has been standing still for nine days. We check the captured tables' write rate against the cluster's total write rate whenever the two look decoupled, and it is the fastest way to catch this.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you bring capacity down without forcing a CDC re-snapshot?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Lower the ceiling before you go anywhere near the slot&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Capture before you change. A writer reboot or a failover clears the buffer cache and resets pg_stat_statements, which destroys the memory-pressure evidence and the query history you will want if the capacity climbs back. Dropping the slot destroys the lag measurement that proves what happened. Before any mutation, save the ServerlessDatabaseCapacity series for the past 14 days, the full output of the slot query, a snapshot of pg_stat_statements ordered by total_exec_time (the column is total_exec_time on PostgreSQL 13 and later), and the CloudTrail event for the scaling change if there is one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws cloudwatch get-metric-statistics \
  --namespace AWS/RDS \
  --metric-name ServerlessDatabaseCapacity \
  --dimensions Name=DBInstanceIdentifier,Value=claims-analytics-prod-writer \
  --start-time 2026-07-07T00:00:00Z \
  --end-time 2026-07-21T00:00:00Z \
  --period 3600 \
  --statistics Average Maximum \
  --output json &amp;gt; capacity-before.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Two weeks of hourly capacity, on disk, before anything changes. Note where the window ENDS: on the step-change day Cost Explorer already handed you, 2026-07-21 here. Run it through to today instead and the Maximum series hands back the pinned value you are trying to remove, 128, rather than the ceiling the cluster genuinely used before the change.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The scaling configuration comes down first, because it is the cheapest and most reversible move you have. It is a soft limit that takes effect at the next capacity evaluation and restarts nothing. Set max_capacity to a number the cluster actually reached before the change, which is the Maximum in capacity-before.json now that the window stops at the step change, and not to a guess. Here that Maximum read 32, and capacity fell from 112 to 32 within a few minutes with no query errors and a small, settled rise in p99.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws rds modify-db-cluster \
  --db-cluster-identifier claims-analytics-prod \
  --serverless-v2-scaling-configuration MinCapacity=2,MaxCapacity=32 \
  --apply-immediately

# Then land the same values in code so the next console edit reads as drift.
resource "aws_rds_cluster" "claims_analytics" {
  # ...
  serverlessv2_scaling_configuration {
    min_capacity = 2
    max_capacity = 32
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The CLI stops the bleeding in one call. The Terraform block is what stops the same person doing it again next quarter.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The slot is where people cause real damage, so go slowly. Postgres will refuse to drop a slot a consumer is attached to, and that refusal is a favour: dropping it forces most CDC connectors to re-snapshot the source tables, which on a 180 GB dataset is a multi-hour outage for everything downstream.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR:  replication slot "cdc_claim_events" is active for PID 24817
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Read this as a warning, not an obstacle. The slot is load-bearing for a pipeline somebody owns.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Work in this order instead. Confirm who owns the consumer and whether it is running. If it is running and behind, let it catch up and re-measure; a backfill that finishes takes the pressure off by itself. If it is running but &lt;code&gt;restart_lsn&lt;/code&gt; is not moving, the offset commit is the problem, and the fix is to make the connector emit something to acknowledge even when the captured tables are quiet. Debezium's PostgreSQL connector exposes a heartbeat interval and a heartbeat action query for exactly this (&lt;code&gt;heartbeat.interval.ms&lt;/code&gt; and &lt;code&gt;heartbeat.action.query&lt;/code&gt;); check the property names against the connector version you run, and point the action query at a small scratch table. One requirement decides whether this works at all: on &lt;code&gt;pgoutput&lt;/code&gt;, decoding only emits changes for tables in the publication, so a scratch table outside it generates WAL and produces no decoding output, the connector still has nothing to acknowledge, and the slot does not move. Add it first, &lt;code&gt;ALTER PUBLICATION &amp;lt;name&amp;gt; ADD TABLE &amp;lt;scratch&amp;gt;&lt;/code&gt;, then watch restart_lsn advance. Skip that and the fix looks like it failed when it was never wired up. A controlled pause and resume, after verifying the connector has flushed its offsets, is the safe way to nudge a slot that has drifted. Dropping the slot is the last option, taken deliberately, with the re-snapshot scheduled.&lt;/p&gt;

&lt;p&gt;Confirm the fix on the shape of the curve, not on a single reading. Watch one full traffic cycle: the number you care about is the off-peak floor, not the peak. On this cluster the writer settled to 3 ACU at 04:00 UTC and rode up to 12 ACU during the business-hours ramp, which is a sawtooth again, and the day's total came to $19. Check that against the baseline rather than against relief: with the reader back on its restored floor of 2, the pair averages about 6.6 ACU, which is the 6.4 this cluster ran at before any of this started. A day that lands at $47 is 16 ACU average and still nearly three times baseline, and it will feel like success because it is so much better than $426. If the off-peak floor still matches the daytime peak, one of the two causes is still live and you have fixed the visible one.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Alarm on ServerlessDatabaseCapacity itself: average above roughly twice your known peak for 6 consecutive hourly datapoints. On this cluster that threshold would have paged inside a day instead of at monthly close.&lt;/li&gt;
&lt;li&gt;Watch the scaling configuration, not just the capacity. An EventBridge rule matching CloudTrail's ModifyDBCluster events from aws.rds, routed to a channel with the calling IAM identity in the message, turns a silent console edit into a named change in minutes rather than at the next nightly plan.&lt;/li&gt;
&lt;li&gt;Export retained WAL bytes per slot from the writer as a custom metric on a one-minute schedule. It is the single number that predicts this bill, and it is cheap to ship.&lt;/li&gt;
&lt;li&gt;PostgreSQL 13 and later expose max_slot_wal_keep_size, which invalidates a slot once its retained WAL passes the limit. Confirm whether your Aurora parameter group exposes it before you plan around it, and be clear with the pipeline owner that invalidation means a re-snapshot.&lt;/li&gt;
&lt;li&gt;Give temporary parameter overrides an expiry. A short-lived branch with an expiry date and a scheduled job that opens the revert PR the next day costs almost nothing and closes the exact hole that produced this.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One control worth more than the rest: if the drift alert exists and nobody sees it, it does not exist. On this engagement the nightly drift plan had been running the whole time and posting into a channel that had been muted weeks earlier after a noisy false-positive run. Muting a channel to silence one bad alert is how a working detection system becomes decoration. When we audit a cost incident we now check which alerting channels have received zero human reads recently, and it turns up more than it should. We have written about the wider pattern in &lt;a href="https://infraforge.agency/problems/cloud-cost-spikes/" rel="noopener noreferrer"&gt;our work on cloud cost spikes&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Aurora Serverless v2 scale-down questions we get asked next
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;What people type into search right after this one&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Short answers to the follow-ups that arrive within an hour of the first fix.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does a high max_capacity cost anything if the cluster never reaches it? You are billed for ACUs in use, with min_capacity as the floor, so the ceiling itself is not the charge. The risk is that a high ceiling removes the guardrail that would have capped the climb, which is precisely what happened here.&lt;/li&gt;
&lt;li&gt;Can I just drop the replication slot to make capacity fall? Yes, and capacity will fall. Most CDC consumers will then re-snapshot the source tables, so treat it as a scheduled outage for the downstream pipeline rather than a quick fix. Confirm the owner first.&lt;/li&gt;
&lt;li&gt;Will a failover or a writer reboot clear this? It will reset capacity for a while and it will clear your buffer cache and your pg_stat_statements history, which is evidence you want. The slot backlog is not cleared by a failover, so if the slot is the cause, the capacity climbs straight back.&lt;/li&gt;
&lt;li&gt;Does any of this apply to Aurora MySQL Serverless v2? The scaling behaviour and the console-drift failure mode carry over directly. The replication mechanism does not: on MySQL the equivalent thing to look at is binlog retention and whoever is consuming it.&lt;/li&gt;
&lt;li&gt;What min_capacity should we set? We have stopped setting min_capacity above 4 ACU on clusters with a genuinely quiet overnight window. The honest cost of that position: the first minute of the morning ramp runs against a cold cache, and on this cluster that read as roughly 23ms of extra p99 before it settled. Against $80.64 a day for a floor nobody uses, we take the 23ms.&lt;/li&gt;
&lt;li&gt;Why did capacity stay high even after the connector caught up? Scale-down is gradual by design, and it is bounded by memory that the buffer cache is still holding. Give it a full off-peak window before you conclude the fix did not work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you arrived here because a CDC pipeline you inherited is doing something you cannot explain, the replication slot is usually the least documented part of the whole system. We cover that ground in &lt;a href="https://infraforge.agency/migrations/" rel="noopener noreferrer"&gt;our migration recovery work&lt;/a&gt;, where a half-finished cutover leaving a live slot behind is one of the more common things we walk into.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the ACU line is flat and the invoice is not waiting for you
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If the cycle closes before the capacity comes down&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hard version of this is not the diagnosis. It is the moment you have found a 241 GB slot, the pipeline that owns it belongs to a team that is mid-sprint on something else, the billing cycle closes in 36 hours, and the safe fix and the fast fix are not the same fix. Getting that call right needs someone who has seen both outcomes: the controlled restart that costs 78 seconds, and the slot drop that costs a six-hour re-snapshot and an apology to three downstream consumers.&lt;/p&gt;

&lt;p&gt;We do this work on live Aurora clusters with the pipeline owners in the room, and we care as much about the guardrail you land afterwards as about the number on the dashboard tonight. Most of these engagements end with two alarms, one Terraform PR, and a slot-lag metric that nobody had before. If your capacity chart has been flat for days and the close is coming, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review&lt;/a&gt; and we will be on a call with you the same day to work the two branches in order.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/aurora-serverless-v2-not-scaling-down/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/aurora-serverless-v2-not-scaling-down/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>aurora</category>
      <category>cost</category>
      <category>troubleshooting</category>
      <category>database</category>
    </item>
    <item>
      <title>How a do-not-disrupt annotation broke Karpenter consolidation</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Tue, 18 Aug 2026 11:36:38 +0000</pubDate>
      <link>https://dev.to/infraforge/how-a-do-not-disrupt-annotation-broke-karpenter-consolidation-56fi</link>
      <guid>https://dev.to/infraforge/how-a-do-not-disrupt-annotation-broke-karpenter-consolidation-56fi</guid>
      <description>&lt;p&gt;Our Karpenter cluster went from 108 nodes to 198 over one weekend while pod count barely moved, consolidation fell to about an eighth of its normal rate, and our EC2 on-demand bill jumped 2.2x. The cause was one line added to a shared internal Helm chart on Friday evening: karpenter.sh/do-not-disrupt: true, applied to every pod template the chart rendered. Renovate auto-merged the bump across 34 tenant repos over the weekend, and by Monday about 1,450 pods carried the annotation. Karpenter will not consolidate a node while any running pod on it carries that annotation, so roughly 85% of nodes became unconsolidatable while rollouts and spot replacements kept adding nodes. Here is how we spotted it and unpinned the fleet without an emergency restart of every workload the chart owned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;karpenter_disruption_nodes_disrupted_total{method="consolidation"} drops to a fraction of its usual rate while node count climbs&lt;/li&gt;
&lt;li&gt;EC2 on-demand line items on the two largest instance sizes double while spot line items stay flat&lt;/li&gt;
&lt;li&gt;Node count climbs steadily through a quiet weekend with no matching increase in pod count or RPS&lt;/li&gt;
&lt;li&gt;kubectl finds karpenter.sh/do-not-disrupt: true on a majority of pods even though only a handful of workloads legitimately need it&lt;/li&gt;
&lt;li&gt;On-demand share of total node hours drifts from ~30% to &amp;gt;60% even though the NodePool allows spot&lt;/li&gt;
&lt;li&gt;kubectl get events shows DisruptionBlocked: Cannot disrupt Node: Pod "/" has "karpenter.sh/do-not-disrupt" annotation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The metric that redirected us away from HPAs
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Node count up 83%, pod count flat, consolidation down to an eighth&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Our first read was wrong. The Monday FinOps digest fired at 08:15 UTC showing EC2 on-demand up 2.2x for the trailing 72 hours, and the platform lead's instinct was 'someone shipped an HPA that went sideways over the weekend.' We checked requests per second on the top 20 services; nothing moved more than 8% week over week. Total pod count was up about 40 pods against a fleet of 2,400. Not an HPA problem. Something was keeping nodes without a matching demand signal, and the extra capacity was mostly on-demand even though the NodePool allowed spot, which Karpenter prefers when it can choose.&lt;/p&gt;

&lt;p&gt;The evidence that pointed us at the right layer came from Karpenter's own counters. We were on v0.37, and we read them for the 60-hour incident window against the same window a week earlier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight prometheus"&gt;&lt;code&gt;&lt;span class="c"&gt;# Prometheus, Karpenter v0.37 metric names. Run each as written, then again&lt;/span&gt;
&lt;span class="c"&gt;# with "offset 7d" after the [60h] for the comparison week.&lt;/span&gt;
&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;increase&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;karpenter_nodes_created&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;60h&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;increase&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;karpenter_disruption_nodes_disrupted_total&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"consolidation"&lt;/span&gt;&lt;span class="p"&gt;}[&lt;/span&gt;&lt;span class="mi"&gt;60h&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;

                  &lt;span class="n"&gt;incident&lt;/span&gt;   &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;week&lt;/span&gt; &lt;span class="n"&gt;earlier&lt;/span&gt;
&lt;span class="n"&gt;nodes&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt;        &lt;span class="mi"&gt;196&lt;/span&gt;          &lt;span class="mi"&gt;412&lt;/span&gt;
&lt;span class="n"&gt;consolidated&lt;/span&gt;          &lt;span class="mi"&gt;41&lt;/span&gt;          &lt;span class="mi"&gt;347&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Launches were not the signal. They fell too, since part of a normal week's launches are the replacement nodes consolidation itself starts. What collapsed was the rate at which Karpenter took nodes away.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The provisioning path was healthy. There were no &lt;code&gt;failed provisioning&lt;/code&gt; or &lt;code&gt;unschedulable&lt;/code&gt; events. Karpenter launched 196 nodes against 412 in the comparison week, but it removed far fewer: 41 through consolidation against 347, plus the usual 65 or so lost to spot interruptions. Over 60 hours that gap accumulated into 90 extra nodes, most of them on-demand, sitting quietly and costing money. The on-demand skew is consolidation's absence too: on a normal weekend Karpenter replaces the on-demand nodes it launches during a spot shortfall once spot capacity returns, and that replacement is itself a consolidation.&lt;/p&gt;

&lt;h2&gt;
  
  
  One line in a shared chart, 34 tenants downstream
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The chart shipped Friday at 19:47 UTC and Renovate did the rest&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Once the question was 'why isn't consolidation firing', the answer came out fast. Karpenter's consolidation logic evaluates candidate nodes by checking whether every pod on the node can be safely rescheduled. Any node running a pod with karpenter.sh/do-not-disrupt: true is disqualified as a candidate for as long as that pod runs there, and Karpenter says so in a DisruptionBlocked event on the node, which names the pod. The annotation exists for legitimate reasons: Kafka Streams state, long-running batches with expensive warmup, workloads that lose data on SIGTERM. In our cluster, four workloads owned it and were the only ones that should have had it.&lt;/p&gt;

&lt;p&gt;The reality was different:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get events &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="nt"&gt;--field-selector&lt;/span&gt; &lt;span class="nv"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;DisruptionBlocked,involvedObject.kind&lt;span class="o"&gt;=&lt;/span&gt;Node &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;jsonpath&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{range .items[*]}{.message}{"\n"}{end}'&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-3&lt;/span&gt;

Cannot disrupt Node: Pod &lt;span class="s2"&gt;"payments/checkout-api-6f9c7d5b8-x2k4q"&lt;/span&gt; has &lt;span class="s2"&gt;"karpenter.sh/do-not-disrupt"&lt;/span&gt; annotation
Cannot disrupt Node: Pod &lt;span class="s2"&gt;"search/indexer-5b7d9c6f4-p8v2m"&lt;/span&gt; has &lt;span class="s2"&gt;"karpenter.sh/do-not-disrupt"&lt;/span&gt; annotation
Cannot disrupt Node: Pod &lt;span class="s2"&gt;"notifications/dispatcher-7c8f6d9b5-j3n6r"&lt;/span&gt; has &lt;span class="s2"&gt;"karpenter.sh/do-not-disrupt"&lt;/span&gt; annotation

kubectl get pods &lt;span class="nt"&gt;--all-namespaces&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'[.items[] | select(.metadata.annotations["karpenter.sh/do-not-disrupt"] == "true")] | length'&lt;/span&gt;

1450

kubectl get pods &lt;span class="nt"&gt;--all-namespaces&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'[.items[] | select(.metadata.annotations["karpenter.sh/do-not-disrupt"] == "true") | .metadata.namespace] | unique | length'&lt;/span&gt;

38
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;1,450 pods across 38 namespaces carried the annotation. With pods distributed by the default scheduler, that pinned roughly 85% of nodes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Our first suspicion was a mutating admission webhook. &lt;code&gt;kubectl get mutatingwebhookconfigurations&lt;/code&gt; returned nothing new. The annotation was in the pod specs at the source. We picked one pod, walked back to its ReplicaSet, saw the annotation in the pod template, and ran &lt;code&gt;git blame&lt;/code&gt; on the tenant's Deployment manifest. The blame line pointed at an internal chart bump: &lt;code&gt;internal-charts/base-deployment&lt;/code&gt; moved from 2.7.4 to 2.7.5 on Friday at 19:47 UTC. The diff was two lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;+ annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;span class="na"&gt;+   karpenter.sh/do-not-disrupt&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Commit message: 'add do-not-disrupt to prevent midday restarts during batch runs.'&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The chart owner had one tenant whose Kafka Streams StatefulSet was losing ~90 seconds of state on every consolidation event. The narrow fix was to annotate that one StatefulSet, which required coordinating with the tenant. The wide fix was to put the annotation in the shared chart, which required coordinating with no one. They picked the wide fix. Our Renovate config auto-merged minor and patch bumps from internal charts, and 2.7.4 to 2.7.5 was a patch. Thirty-four tenant repos consumed base-deployment. Over the weekend, thirty-four PRs opened, thirty-four PRs merged, thirty-four ArgoCD syncs rolled deployments, and every rolled pod inherited the annotation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJwtjk1PwkAQhu_8ivfiTRAURI0hgfLhwRjT4GnDYdsdaJPtTLO7hfTibzcMHid5nmfeo5dLWdmQ8JkPgKXJ9HgczUczNBRO5N6L8LDYhhqT17fpHD_77IDhcIGVyYnlbBNBWuKo3NMUidhywnceDwNgpWxmll2SoQYhrGhTswRYdmhtKisUXdOqkqmyNstwkmyN2HN5iwfxnhwctV76hjgpvlZ8Yyb309kYrbgIlosKpQ2hR6oIllmSTbXwVdmosjW_L7M7yBEsjm4vOi6Fo_ja2WQLT1d6q_TO5OK9dCnq5thKQqDW25J0iurWuVsMLKmq-YRAjZxJh-6082G-xBFK6TiByQbfw0lX-P8FcqaA5zEq6UI8_AGaKoeY" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJwtjk1PwkAQhu_8ivfiTRAURI0hgfLhwRjT4GnDYdsdaJPtTLO7hfTibzcMHid5nmfeo5dLWdmQ8JkPgKXJ9HgczUczNBRO5N6L8LDYhhqT17fpHD_77IDhcIGVyYnlbBNBWuKo3NMUidhywnceDwNgpWxmll2SoQYhrGhTswRYdmhtKisUXdOqkqmyNstwkmyN2HN5iwfxnhwctV76hjgpvlZ8Yyb309kYrbgIlosKpQ2hR6oIllmSTbXwVdmosjW_L7M7yBEsjm4vOi6Fo_ja2WQLT1d6q_TO5OK9dCnq5thKQqDW25J0iurWuVsMLKmq-YRAjZxJh-6082G-xBFK6TiByQbfw0lX-P8FcqaA5zEq6UI8_AGaKoeY" alt="A single opinionated chart plus a permissive auto-merge is the shape of the whole incident. Nothing in the loop was malicious. Every step was doing what it was configured to do." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A single opinionated chart plus a permissive auto-merge is the shape of the whole incident. Nothing in the loop was malicious. Every step was doing what it was configured to do.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The right recovery is on the pods, not the templates
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The patch we almost ran would have rolled 1,450 pods&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;By 09:00 UTC Monday we had the diagnosis and a bad plan. The obvious move was to patch the Deployment template in each of the 34 tenants, removing the annotation at the source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl patch deployment &amp;lt;name&amp;gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &amp;lt;ns&amp;gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s1"&gt;'{"spec":{"template":{"metadata":{"annotations":{"karpenter.sh/do-not-disrupt":null}}}}}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The command that looked surgical and was not.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We almost ran the loop. Then one of us asked the question that matters: does changing a template annotation trigger a rollout? It does. The Deployment controller's ComputeHash runs DeepHashObject over the entire PodTemplateSpec, and &lt;code&gt;metadata.annotations&lt;/code&gt; is part of the template. Change any pod-template annotation and the template hash changes, which creates a new ReplicaSet, which rolls every pod. The cleanest proof of this is that &lt;code&gt;kubectl rollout restart&lt;/code&gt; triggers a rollout precisely by stamping &lt;code&gt;kubectl.kubernetes.io/restartedAt&lt;/code&gt; into &lt;code&gt;spec.template.metadata.annotations&lt;/code&gt;. If template annotations were excluded from the hash, &lt;code&gt;rollout restart&lt;/code&gt; could not work.&lt;/p&gt;

&lt;p&gt;If we had run the patch loop across 34 tenants at 09:00 UTC on a Monday, we would have rolled about 1,450 pods in flight, cascaded PDB blocks into deploy pipelines, and (worst) triggered even more Karpenter provisioning as the rollouts churned. The recovery would have looked identical to a second, larger incident.&lt;/p&gt;

&lt;p&gt;The correct move was to strip the annotation from the &lt;em&gt;running pods&lt;/em&gt;, not from the template. A ReplicaSet reconciles pod count and pod ownership, not annotation drift on pods that already exist. So mutating a live pod's annotations is safe and does not trigger replacement. Karpenter's disruption loop runs every 10 seconds, and it reconsiders consolidation when pod placement or nodes change (an annotation edit does not count) and at least once every five minutes, so the newly-eligible nodes start draining within a few minutes, on their own.&lt;/p&gt;

&lt;p&gt;Here is what we ran. It strips the annotation only from pods outside an explicit keep list, the four workloads that own the annotation plus the fraud-scoring batch, so the pods that need it keep it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Pods that keep the annotation: the four workloads that own it, and the&lt;/span&gt;
&lt;span class="c"&gt;# fraud-scoring batch mid-run. Keyed on namespace and workload-name prefix,&lt;/span&gt;
&lt;span class="c"&gt;# which every pod carries. A lookalike name is kept too, so check what was&lt;/span&gt;
&lt;span class="c"&gt;# kept by running the same pipeline with grep -E instead of grep -Ev.&lt;/span&gt;
&lt;span class="nv"&gt;KEEP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'^(streams/orders-enricher-|ledger/settlement-writer-|ml/feature-backfill-|media/transcoder-|risk/fraud-scoring-)'&lt;/span&gt;

kubectl get pods &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.items[]
      | select(.metadata.annotations["karpenter.sh/do-not-disrupt"] == "true")
      | "\(.metadata.namespace)/\(.metadata.name)"'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-Ev&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$KEEP&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt;/ &lt;span class="s1"&gt;'{print $1, $2}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | xargs &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="nt"&gt;-n2&lt;/span&gt; &lt;span class="nt"&gt;-P8&lt;/span&gt; sh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'kubectl annotate pod -n "$0" "$1" karpenter.sh/do-not-disrupt-'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The trailing dash on the annotation key removes it. 1,426 pods, eight calls at a time, took just under three minutes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Verification was two commands, one for the pod count and one for Karpenter's response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get pods &lt;span class="nt"&gt;--all-namespaces&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'[.items[] | select(.metadata.annotations["karpenter.sh/do-not-disrupt"] == "true")] | length'&lt;/span&gt;

24

&lt;span class="c"&gt;# The chart logs JSON and runs two replicas: read both pods, lift kubectl's&lt;/span&gt;
&lt;span class="c"&gt;# 10-line default for selectors with --tail=-1, and parse.&lt;/span&gt;
kubectl &lt;span class="nt"&gt;-n&lt;/span&gt; karpenter logs &lt;span class="nt"&gt;-l&lt;/span&gt; app.kubernetes.io/name&lt;span class="o"&gt;=&lt;/span&gt;karpenter &lt;span class="nt"&gt;--since&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10m &lt;span class="nt"&gt;--tail&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nt"&gt;-1&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-Rr&lt;/span&gt; &lt;span class="s1"&gt;'fromjson? | select((.message // "") | startswith("disrupting via consolidation")) | .message'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="s1"&gt;'s/ pods).*/ pods)/'&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-5&lt;/span&gt;

disrupting via consolidation delete, terminating 1 nodes &lt;span class="o"&gt;(&lt;/span&gt;6 pods&lt;span class="o"&gt;)&lt;/span&gt;
disrupting via consolidation delete, terminating 1 nodes &lt;span class="o"&gt;(&lt;/span&gt;4 pods&lt;span class="o"&gt;)&lt;/span&gt;
disrupting via consolidation replace, terminating 1 nodes &lt;span class="o"&gt;(&lt;/span&gt;11 pods&lt;span class="o"&gt;)&lt;/span&gt;
disrupting via consolidation delete, terminating 2 nodes &lt;span class="o"&gt;(&lt;/span&gt;9 pods&lt;span class="o"&gt;)&lt;/span&gt;
disrupting via consolidation delete, terminating 1 nodes &lt;span class="o"&gt;(&lt;/span&gt;3 pods&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Down to 24 annotated pods (the four legitimate workloads plus a 20-pod fraud-scoring batch six hours into its run) and Karpenter firing consolidation events within four minutes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The strip did not stay done. The 34 Deployment templates still carried the annotation until the chart fix shipped, so every pod a ReplicaSet created afterwards came back annotated, including the pods evicted from the nodes consolidation was now removing, and each one re-pinned the node it landed on. We re-ran the strip every 15 minutes until base-deployment 2.7.6 was out, and the annotated count crept back from 24 to around 60 between passes. Over the next 90 minutes node count dropped from 198 to 141 and plateaued. The 24 kept pods were pinning about 22 nodes: 20 fraud-batch pods spread across 18 nodes, and 4 legitimate pods on 4 more. Consolidation could not touch those. We coordinated a restart of the fraud batch with its owning team. It checkpointed cleanly and restarted onto 5 nodes, tight-packed by Karpenter's provisioning-time bin packing, which freed 13 nodes and took us to 128. Over the next 90 minutes the repeated strips and natural pod churn freed another 20 nodes, and Karpenter consolidated those too. Node count landed at 108 by 15:00 UTC, back inside baseline range.&lt;/p&gt;

&lt;p&gt;Then we rolled out the fixed chart. base-deployment 2.7.6 removed the blanket annotation and gated it behind an explicit &lt;code&gt;values.yaml&lt;/code&gt; opt-in (&lt;code&gt;karpenter.doNotDisrupt: false&lt;/code&gt; default, &lt;code&gt;true&lt;/code&gt; for the four legitimate tenants). This rollout DID replace pods, because it changed the template hash, and we scheduled it deliberately. Renovate opened the PRs, but we had switched off auto-merge for this chart for the rollout. The platform team merged them in batches of five with a ten-minute soak between batches so any interaction with PDBs or startup probes surfaced before the next wave. Six hours end to end, no SLO impact.&lt;/p&gt;

&lt;h2&gt;
  
  
  The changes that landed in the two weeks after
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Three controls that would have caught this on Saturday&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The postmortem produced three durable controls. Each one closes a specific step in the causal chain above.&lt;/p&gt;

&lt;p&gt;The first is a Kyverno ClusterPolicy that only permits the annotation on pods that explicitly opt in via a label. It ran in Audit mode for two weeks, we watched the drift metric go to zero, then we flipped it to Enforce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kyverno.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterPolicy&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;restrict-karpenter-do-not-disrupt&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;validationFailureAction&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Enforce&lt;/span&gt;
  &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;require-opt-in-label&lt;/span&gt;
    &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;any&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;kinds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Pod&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;preconditions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;all&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;request.object.metadata.annotations.&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;karpenter.sh/do-not-disrupt&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;||&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;''&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;}}"&lt;/span&gt;
        &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Equals&lt;/span&gt;
        &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true"&lt;/span&gt;
    &lt;span class="na"&gt;validate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;karpenter.sh/do-not-disrupt&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;requires&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;label&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;karpenter.platform/opt-in=true"&lt;/span&gt;
      &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;karpenter.platform/opt-in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;A future chart change that adds the annotation without the matching label is rejected at admission and never lands.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The second is an SLO alert on Karpenter's own consolidation metric. &lt;code&gt;karpenter_disruption_nodes_disrupted_total&lt;/code&gt; was already in our Prometheus scrape config; nobody had put it on a dashboard. We added a Grafana panel showing nodes consolidated per hour, filtered to &lt;code&gt;method="consolidation"&lt;/code&gt; (the &lt;code&gt;action&lt;/code&gt; label holds &lt;code&gt;delete&lt;/code&gt;, &lt;code&gt;replace&lt;/code&gt; or &lt;code&gt;no-op&lt;/code&gt;, never &lt;code&gt;consolidation&lt;/code&gt;), and paged if the rate stayed below 20% of the 7-day rolling median for more than two hours, at any hour. Backtesting against the incident, that alert would have fired around 04:00 UTC Saturday, roughly 8 hours in, instead of 60 hours in. On Karpenter v1 those metric names are gone; alert on &lt;code&gt;karpenter_voluntary_disruption_decisions_total{reason=~"underutilized|empty"}&lt;/code&gt; instead, which leaves drift out. It counts decisions, not nodes, so a multi-node consolidation counts once.&lt;/p&gt;

&lt;p&gt;The third is a check in the shared chart's own CI pipeline. Any change, minor or patch, that touches a scheduling-related annotation key (&lt;code&gt;karpenter.sh/*&lt;/code&gt;, &lt;code&gt;karpenter.k8s.aws/*&lt;/code&gt;, &lt;code&gt;scheduler.alpha.kubernetes.io/*&lt;/code&gt;, &lt;code&gt;node.kubernetes.io/*&lt;/code&gt;, &lt;code&gt;cluster-autoscaler.kubernetes.io/*&lt;/code&gt;) fails the chart's release job until the platform team approves it, so a version like 2.7.5 cannot be published, and never reaches Renovate, without that review. Non-scheduling changes still flow through the previous auto-merge path unchanged. The tradeoff we now pay is about one platform-team review per quarter on this chart; the incident that check would have caught cost us roughly $3,500 in on-demand overage across the 60-hour window.&lt;/p&gt;

&lt;p&gt;We considered tightening the NodePool's &lt;code&gt;disruption.budgets&lt;/code&gt; and did not. The root cause was not consolidation being too eager; consolidation was fine when it was allowed to fire. Tightening budgets would slow legitimate consolidation without preventing another annotation-blast. We put a comment in the NodePool YAML pointing to the postmortem so a future engineer does not try to 'fix' the budgets in response to reading it.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ: Karpenter consolidation and the do-not-disrupt annotation
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Common questions we get about this pattern&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does &lt;code&gt;karpenter.sh/do-not-disrupt&lt;/code&gt; pin just the pod or the whole node?&lt;/strong&gt; It effectively pins the node. Consolidation eligibility is evaluated per-node, and any node running at least one annotated pod is disqualified until that pod is gone or loses the annotation. One annotated pod on a 32-vCPU node keeps the node from consolidating for as long as that pod runs there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Would tightening PodDisruptionBudgets have prevented this?&lt;/strong&gt; No. Karpenter checks PDBs at the same gate as the annotation: a node whose pods a PDB would refuse to evict is also kept out of consolidation, with its own DisruptionBlocked event. Tighter PDBs add another way to pin nodes; they do nothing about a chart stamping the annotation on every pod.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can we detect this via CloudWatch or Cost Explorer alone?&lt;/strong&gt; Not fast enough. Cost Explorer reports daily unless you enable hourly granularity, and our FinOps digest ran weekly. The signal you want is Karpenter's own consolidation rate from &lt;code&gt;/metrics&lt;/code&gt;, which moves within minutes of the problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this pattern apply to Karpenter v1?&lt;/strong&gt; Yes. In v1 a running pod with &lt;code&gt;karpenter.sh/do-not-disrupt: "true"&lt;/code&gt; still blocks Karpenter from voluntarily disrupting its node. Three things differ. From v1.12 the annotation also accepts a duration such as &lt;code&gt;"30m"&lt;/code&gt;, which protects a pod only for that long after it starts running; earlier releases honour only &lt;code&gt;"true"&lt;/code&gt;. A NodePool with &lt;code&gt;terminationGracePeriod&lt;/code&gt; set may disrupt a node through drift even with such a pod on it. And the annotation does not exclude a node from the forceful methods, expiration, interruption and manual deletion among them, so &lt;code&gt;expireAfter&lt;/code&gt; without &lt;code&gt;terminationGracePeriod&lt;/code&gt; can leave an annotated pod blocking a node's drain indefinitely. A NodePool that sets neither field gets exactly that: &lt;code&gt;expireAfter&lt;/code&gt; defaults to 720h and &lt;code&gt;terminationGracePeriod&lt;/code&gt; has no default. The Deployment controller's PodTemplateSpec hash behavior is a Kubernetes property, not a Karpenter property, and it does not change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if we want every pod in a Deployment to be do-not-disrupt by design?&lt;/strong&gt; Put the annotation in the template and accept that changing it triggers a rollout. Use &lt;code&gt;maxSurge&lt;/code&gt; and PDBs to control that rollout the way you would for any other template change. The mistake in our incident was not that a template had the annotation; it was that 34 templates got it without the owning teams knowing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Karpenter regressions get stuck
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If your on-demand line jumped and no one shipped anything&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The awkward thing about this class of Karpenter regression is that every piece looks correct in isolation. The annotation was a real feature intended for real workloads. The chart bump was a legitimate response to a real production pain. The Renovate auto-merge policy was tuned to reduce toil on well-behaved chart bumps. The NodePool was configured the way the docs recommend. It took the combination, plus a weekend, plus a weekly cost digest, to produce a nearly doubled node count and a mid-four-figure overage. The diagnostic move that mattered was reading Karpenter's consolidation rate against its own baseline, then its DisruptionBlocked events, which name the pod in the way. The recovery move that mattered was knowing which kubectl mutations trigger a rollout and which do not.&lt;/p&gt;

&lt;p&gt;We have written more on this shape of failure in the &lt;a href="https://infraforge.agency/kubernetes-cicd/" rel="noopener noreferrer"&gt;Kubernetes and CI/CD stabilization pillar&lt;/a&gt;, and the specific pattern of a shared chart change cascading through GitOps sits in the &lt;a href="https://infraforge.agency/argocd-gitops-recovery/" rel="noopener noreferrer"&gt;ArgoCD and GitOps recovery cluster&lt;/a&gt;. If your Karpenter cluster is growing and you cannot see why, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review&lt;/a&gt; and we will pull the consolidation counters and DisruptionBlocked events together, read them the same way we read ours, and get you a diagnosis inside a single working day.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/karpenter-consolidation-broken-by-do-not-disrupt-annotation/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/karpenter-consolidation-broken-by-do-not-disrupt-annotation/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>recovery</category>
      <category>kubernetescicd</category>
    </item>
    <item>
      <title>Worker drops jobs intermittently: startup race, env drift, schema skew</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Mon, 10 Aug 2026 03:15:00 +0000</pubDate>
      <link>https://dev.to/infraforge/worker-drops-jobs-intermittently-startup-race-env-drift-schema-skew-3ngb</link>
      <guid>https://dev.to/infraforge/worker-drops-jobs-intermittently-startup-race-env-drift-schema-skew-3ngb</guid>
      <description>&lt;p&gt;If your worker is dropping jobs intermittently, sometimes crashing at boot with 'Error 111 connecting to redis:6379. Connection refused' and sometimes producing an empty result.json with no error at all, the cause is almost never a Redis performance problem. In most of the cases we get called into, it is a startup race between the worker and its dependency, hiding two quieter bugs behind it: an environment variable that silently falls back to localhost, and a job schema that got renamed on one side of the wire. This guide walks the three causes in the order they actually occur, gives you the one log grep that separates them, and shows the fix sequence that does not destroy the evidence you still need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Some runs crash at boot with &lt;code&gt;redis.exceptions.ConnectionError: Error 111 connecting to redis:6379. Connection refused.&lt;/code&gt; and others start cleanly&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;result.json&lt;/code&gt; is written on some runs, missing on others, sometimes present but with 0 processed jobs and no error in the log&lt;/li&gt;
&lt;li&gt;Restarting the worker alone (without the backend) makes the problem go away for a while, then it comes back after the next deploy&lt;/li&gt;
&lt;li&gt;A grep for &lt;code&gt;KeyError&lt;/code&gt; shows sporadic &lt;code&gt;KeyError: 'job_id'&lt;/code&gt; traces that are being caught by a broad except block&lt;/li&gt;
&lt;li&gt;The worker container's uid is 1000, &lt;code&gt;/app/output/&lt;/code&gt; is owned by root:root with mode 0755, and nobody remembers why&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The symptom, and what it usually is not
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Three failure modes wearing one costume&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The reported symptom is almost always the same sentence: 'the worker is flaky, sometimes it processes jobs and sometimes it doesn't, and the logs don't really say why.' That sentence hides three separate bugs that compound. We rank them by the order we actually find them, not by how loud they are in the log:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1. &lt;strong&gt;Startup race.&lt;/strong&gt; Worker container comes up before Redis (or the backend that populates Redis) is accepting connections. Retries are missing or set to a value that gives up in under 2 seconds. Roughly 60% of the incidents we see.&lt;/li&gt;
&lt;li&gt;2. &lt;strong&gt;Environment variable drift.&lt;/strong&gt; The backend sets &lt;code&gt;REDIS_URL&lt;/code&gt;, the worker reads &lt;code&gt;CACHE_URL&lt;/code&gt;, and the worker's client library silently defaults to &lt;code&gt;redis://localhost:6379&lt;/code&gt;. It 'connects' to nothing and hangs or reads empty queues. Roughly 25%.&lt;/li&gt;
&lt;li&gt;3. &lt;strong&gt;Schema skew after a rename.&lt;/strong&gt; Someone renamed &lt;code&gt;job_id&lt;/code&gt; to &lt;code&gt;id&lt;/code&gt; on the producer side and missed one &lt;code&gt;.get('job_id')&lt;/code&gt; on the consumer. &lt;code&gt;.get&lt;/code&gt; returns &lt;code&gt;None&lt;/code&gt;, the worker's outer &lt;code&gt;except Exception&lt;/code&gt; swallows the downstream &lt;code&gt;KeyError&lt;/code&gt;, the job is skipped, no error is logged. Roughly 15%, and the hardest to see.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cause everyone blames first, and which is almost never it: 'Redis is slow' or 'the worker needs more memory'. We have not once found this to be the actual cause in this failure shape. If you are already sizing up the Redis instance, stop and read the next section first.&lt;/p&gt;

&lt;p&gt;There is also a fourth cause that shows up as a hard &lt;code&gt;PermissionError: [Errno 13] Permission denied: '/app/output/result.json'&lt;/code&gt; when the worker finally does try to write. It is real, but it is loud. This guide is about the quiet failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  The discriminating check: read the logs before you touch anything
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;One grep that ranks the three&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Before you restart, redeploy, or scale anything, capture the current state. Restarting the worker throws away the evidence that tells you which of the three you are dealing with. On docker compose, snapshot both services' logs to disk. On Kubernetes, capture &lt;code&gt;kubectl logs --previous&lt;/code&gt; for the last crashed worker AND &lt;code&gt;kubectl describe pod&lt;/code&gt; for its events. Do this first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# capture before you touch anything
docker compose logs --no-color --timestamps worker  &amp;gt; /tmp/worker.log
docker compose logs --no-color --timestamps backend &amp;gt; /tmp/backend.log

# k8s equivalent
kubectl logs -n jobs worker-7c9d8f5b6-x2k4m --previous &amp;gt; /tmp/worker.log
kubectl describe pod -n jobs worker-7c9d8f5b6-x2k4m       &amp;gt; /tmp/worker-describe.txt

# then run the discriminating grep
grep -E 'Connection refused|CACHE_URL|localhost:6379|KeyError' /tmp/worker.log | head -40
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Snapshot first, grep second. Anything that recreates the container (&lt;code&gt;docker compose up --force-recreate worker&lt;/code&gt;, or a &lt;code&gt;down&lt;/code&gt; followed by an &lt;code&gt;up&lt;/code&gt;) discards the previous container's stdout under the default json-file driver. A plain &lt;code&gt;restart&lt;/code&gt; keeps the log file, but it costs you the live process state that made the failure reproducible.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The grep output tells you which cause you have, in this order:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you see &lt;code&gt;Error 111 connecting to redis:6379. Connection refused&lt;/code&gt; in the first 5 seconds of the worker's log and never again after that, it is the &lt;strong&gt;startup race&lt;/strong&gt;. The worker gave up before Redis was ready.&lt;/li&gt;
&lt;li&gt;If you see &lt;code&gt;Connecting to redis://localhost:6379&lt;/code&gt; (note: localhost, not the service name), or you see NO connection log at all and the queue depth reads always return 0, it is &lt;strong&gt;env drift&lt;/strong&gt;. The worker never got the right URL and its client defaulted.&lt;/li&gt;
&lt;li&gt;If the connection logs are clean, the worker is clearly consuming messages, but the count of 'processed' log lines is lower than the count of 'received' lines, and you can find any &lt;code&gt;KeyError&lt;/code&gt; (even one, even caught) in the traces, it is &lt;strong&gt;schema skew&lt;/strong&gt;. The worker is silently dropping jobs whose payload does not match its expected shape.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxNjcFKw0AQhu99iv8BDB5ERZGKbVo9eKvgIeSw7s42S9KdODNJkeK7i6GBXuf75_tix0ffODF8lAvgpfpkaUnQ8V7hXW-DUKhRFEusTmvOmbwlzhCKg1JAyohJ1HCrz78LYIWiwA_p9LGudubEhh7iPNUzzjzRcvYpjNGxd13Dao93N_cPT19yvWT5n34PNBBMXIzJT43ysrGpNnnE6ARBUrR65ufI9iTkKY0U4HnIhuWk7oU9qc7XSbu91L5WO9_QwUFbOl5BU0fZEIT7ep6eC2_VO3MLZ-hJDkk1cVa4HNA7a7T-A7-8cSA" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxNjcFKw0AQhu99iv8BDB5ERZGKbVo9eKvgIeSw7s42S9KdODNJkeK7i6GBXuf75_tix0ffODF8lAvgpfpkaUnQ8V7hXW-DUKhRFEusTmvOmbwlzhCKg1JAyohJ1HCrz78LYIWiwA_p9LGudubEhh7iPNUzzjzRcvYpjNGxd13Dao93N_cPT19yvWT5n34PNBBMXIzJT43ysrGpNnnE6ARBUrR65ufI9iTkKY0U4HnIhuWk7oU9qc7XSbu91L5WO9_QwUFbOl5BU0fZEIT7ep6eC2_VO3MLZ-hJDkk1cVa4HNA7a7T-A7-8cSA" alt="Order matters. The race hides the drift, and the drift hides the schema skew. Fix them in the order they surface." width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Order matters. The race hides the drift, and the drift hides the schema skew. Fix them in the order they surface.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The safe fix sequence for each cause
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Fix least-destructive first&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Fix each cause with the smallest change that discriminates against the others. Do not batch the three fixes into one PR; you will not know which one worked, and the next incident will look identical. Ship them separately, verify each one, then move on.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Startup race&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Add a real readiness gate, not a sleep. On docker compose, use &lt;code&gt;depends_on&lt;/code&gt; with &lt;code&gt;condition: service_healthy&lt;/code&gt; and a healthcheck on the backend/Redis. On Kubernetes, use an initContainer that runs &lt;code&gt;nc -z redis 6379&lt;/code&gt; in a loop with a bounded timeout (60s), and set the worker's client to retry with exponential backoff for at least 30s after boot. &lt;code&gt;sleep 10&lt;/code&gt; in the entrypoint is the fix everyone reaches for; it papers over the problem until the day Redis takes 11 seconds.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Env var drift&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rename one side to match the other in a single PR. Do NOT add a fallback like &lt;code&gt;CACHE_URL or REDIS_URL&lt;/code&gt;; that is how you got here. Then add a boot-time assertion: if the resolved URL is &lt;code&gt;localhost&lt;/code&gt; or empty, the worker exits with a clear error instead of connecting to nothing. This is the single change that pays for itself the fastest.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. Schema skew&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Find every &lt;code&gt;.get('job_id')&lt;/code&gt; and &lt;code&gt;.get('id')&lt;/code&gt; across producer and consumer. Pick one name (we prefer &lt;code&gt;id&lt;/code&gt; because it is what most queue libraries default to). Add a schema validation step at the consumer boundary that raises loudly on a missing key, and remove any &lt;code&gt;except Exception: pass&lt;/code&gt; you find on the way. The broad except is what turned this into a silent bug.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4. Output permissions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;If &lt;code&gt;PermissionError&lt;/code&gt; on &lt;code&gt;/app/output/result.json&lt;/code&gt; is in your logs, the fix is a Dockerfile line: &lt;code&gt;RUN mkdir -p /app/output &amp;amp;&amp;amp; chown -R 1000:1000 /app/output&lt;/code&gt; before the &lt;code&gt;USER 1000&lt;/code&gt; directive. Do not &lt;code&gt;chmod 777&lt;/code&gt;. If the directory is a mounted volume, set &lt;code&gt;fsGroup: 1000&lt;/code&gt; in the pod's securityContext instead.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One thing to name explicitly: the tempting single-line fix of adding &lt;code&gt;sleep 15&lt;/code&gt; to the worker's entrypoint 'solves' the startup race in staging and hides all three bugs in production. We have watched this exact patch get merged, celebrated, and then paged the same team six weeks later when Redis restarted during a maintenance window and the sleep was not long enough. The cost of the right fix (a healthcheck plus retry) is roughly 40 lines of yaml and one afternoon. Pay it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# docker-compose.yml, the readiness gate that actually works
services:
  redis:
    image: redis:7.2-alpine
    healthcheck:
      test: ["CMD", "redis-cli", "ping"]
      interval: 2s
      timeout: 1s
      retries: 15

  backend:
    depends_on:
      redis:
        condition: service_healthy

  worker:
    depends_on:
      backend:
        condition: service_started
      redis:
        condition: service_healthy
    environment:
      REDIS_URL: redis://redis:6379/0    # one name, both sides
    # no sleep, no CACHE_URL fallback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The gate is &lt;code&gt;condition: service_healthy&lt;/code&gt; plus a real healthcheck on the dependency. &lt;code&gt;condition: service_started&lt;/code&gt; alone only waits for the container to exist, not to be ready.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Confirm the fix by running the stack from cold at least ten times in a row and checking that the processed count equals the received count on every run. If nine out of ten pass, you have not fixed it; you have improved it. Race conditions do not get 90% fixed. For deeper patterns on this kind of intermittent K8s failure, we have written up &lt;a href="https://infraforge.agency/problems/kubernetes-release-failures/" rel="noopener noreferrer"&gt;Kubernetes release failure recovery&lt;/a&gt; separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ: variants readers usually ask right after
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The questions that come next&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is it safe to just add &lt;code&gt;restart: always&lt;/code&gt; and let the worker crash-loop until Redis is up?&lt;/strong&gt; It works, but it makes your logs noisier, it costs you real seconds on every boot, and it hides the underlying dependency graph from anyone reading the compose file later. Use a healthcheck. Save &lt;code&gt;restart: always&lt;/code&gt; for actual transient failures in steady state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does this apply to RabbitMQ, NATS, or Kafka workers too?&lt;/strong&gt; Yes, the shape is identical. The verbatim error string changes (&lt;code&gt;Connection refused&lt;/code&gt; on RabbitMQ, &lt;code&gt;dial tcp: connection refused&lt;/code&gt; on NATS, &lt;code&gt;NoBrokersAvailable&lt;/code&gt; on Kafka), but the three causes and the fix order are the same. The env var drift bug is especially common on Kafka clients because &lt;code&gt;KAFKA_BROKERS&lt;/code&gt; vs &lt;code&gt;BOOTSTRAP_SERVERS&lt;/code&gt; is a coin flip in the ecosystem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I skip the healthcheck if I use Kubernetes with readiness probes?&lt;/strong&gt; No, readiness probes gate traffic to a pod, not startup order between pods. You still need an initContainer or an application-level retry loop for the worker to wait on its dependency. Readiness alone will not save you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not just fix the broad &lt;code&gt;except Exception: pass&lt;/code&gt; and call it a day?&lt;/strong&gt; Because removing it in isolation will surface the schema skew as a hard crash on production traffic, and if you have not fixed the startup race first, you will not be able to tell which crash is which. Fix in the order the causes appear at boot: connectivity, config, contract.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I stop this from recurring?&lt;/strong&gt; Three things, in order of ROI: (1) a boot-time assertion in every service that fails fast if a required env var is missing or resolves to localhost, (2) a schema check at the consumer boundary that rejects malformed messages loudly instead of silently, (3) a pre-merge integration test that starts the stack from cold and asserts processed == received on 100 test jobs. That third one is the single highest-value test we recommend for job-processing systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you are staring at an empty result.json and the clock is running
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;When the worker is dropping jobs right now&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;What makes this class of failure genuinely hard is not any single one of the three bugs. It is that they compound: the race gives you enough noise in the log that the drift looks like a symptom of the race, and by the time you fix the race, the drift has been silently corrupting queue state for hours, and the schema skew is dropping the recovery jobs you are firing to catch up. Untangling that under time pressure is where teams get stuck at 2 in the morning.&lt;/p&gt;

&lt;p&gt;We have spent a lot of engineering hours in exactly this shape of incident, on Redis, RabbitMQ, and Kafka, across docker compose and Kubernetes. The pattern above is what we run on the call: capture logs first, grep to rank the causes, fix in dependency order, verify with cold-start runs. If you want a second set of eyes on it while it is happening, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book a same-day infrastructure review&lt;/a&gt; and we will get on a bridge with your on-call engineer inside a few hours and work the sequence together. If the fires are already out and you want the pre-merge integration test and boot-time assertions in place before the next one, our &lt;a href="https://infraforge.agency/services/" rel="noopener noreferrer"&gt;platform reliability engagements&lt;/a&gt; cover exactly that.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/worker-drops-jobs-intermittently-startup-race-env-drift/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/worker-drops-jobs-intermittently-startup-race-env-drift/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>platform</category>
      <category>troubleshooting</category>
      <category>linuxsysadmin</category>
    </item>
    <item>
      <title>How a missing S3 gateway endpoint route quintupled our NAT bill</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Thu, 16 Jul 2026 03:32:55 +0000</pubDate>
      <link>https://dev.to/infraforge/how-a-missing-s3-gateway-endpoint-route-quadrupled-our-nat-bill-4i9</link>
      <guid>https://dev.to/infraforge/how-a-missing-s3-gateway-endpoint-route-quadrupled-our-nat-bill-4i9</guid>
      <description>&lt;p&gt;The finance lead asked why AWS charged us $2,100 for NAT gateway data processing last month. Our normal was around $400. Nothing in the release calendar explained it: no new services, no traffic bump in our own metrics, no scaling events. The bill just quintupled. Then someone opened VPC Flow Logs and filtered on the NAT's ENI. Roughly 71% of the bytes had destinations in 52.216.0.0/15 and 3.5.0.0/16, which is S3. We had an S3 gateway endpoint on the VPC. It was supposed to be handling that traffic on the AWS backbone for free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;NAT gateway data-processing charges (NatGateway-Bytes) climb week over week with no code or workload change&lt;/li&gt;
&lt;li&gt;VPC Flow Logs show heavy traffic to S3 IP ranges (52.216.0.0/15, 3.5.0.0/16) hitting the NAT's ENI instead of the gateway endpoint&lt;/li&gt;
&lt;li&gt;aws ec2 describe-route-tables on a private subnet returns no route for the S3 managed prefix list (pl-xxx)&lt;/li&gt;
&lt;li&gt;NAT gateway BytesOutToDestination stays high while VPC Flow Logs show traffic to the S3 prefix list still routing through NAT (the S3 gateway endpoint itself emits no CloudWatch metrics to check)&lt;/li&gt;
&lt;li&gt;The S3 gateway endpoint exists in the VPC, but the bill still shows $0.045/GB on service-to-service S3 traffic&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  $412 to $2,103 on a flat workload, all of it NatGateway-Bytes
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The bill line item that should not have existed&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The NatGateway-Bytes cost for us-east-1 climbed from $412 in March to $2,103 in April. Nothing else on the bill moved. Egress was flat. EC2 was flat. RDS was flat. The whole delta was NAT data processing at $0.045 per GB.&lt;/p&gt;

&lt;p&gt;Cost Explorer's usage-type breakdown confirmed the shape. The entire spike was NAT-Bytes on one VPC, in one region, over a six-week ramp. No cliff, no sudden jump, just a slow gradient upward that nobody watched because 'NAT costs a few hundred bucks' was our mental default.&lt;/p&gt;

&lt;p&gt;So we mapped the destinations. AWS publishes the IP ranges for S3, DynamoDB, and every other service in &lt;a href="https://ip-ranges.amazonaws.com/ip-ranges.json" rel="noopener noreferrer"&gt;ip-ranges.json&lt;/a&gt;. We pulled an hour of VPC Flow Logs, joined destination IPs against those ranges, and found the answer. S3 was 71% of the NAT-processed bytes. That should have been impossible. We had a gateway endpoint for S3, and gateway endpoints are free. Their entire point is to keep S3 traffic off NAT.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross-region buckets, then a hardcoded public endpoint, then the truth
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;What we thought first, and why the first two theories died fast&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first theory was that a new batch job was writing to a bucket in a different region. Cross-region S3 traffic does not hit the same-region gateway endpoint; it goes out through NAT. Reasonable theory, wrong theory. We grepped Terraform for cross-region bucket references and found nothing. We checked CloudTrail for recent bucket creations. Nothing new. All our S3 traffic was to buckets in the same region as the workload.&lt;/p&gt;

&lt;p&gt;The second theory was that an application was talking to S3 via the public global endpoint URL instead of the regional one. That is a real failure mode: SDKs pointed at a hardcoded &lt;a href="https://s3.amazonaws.com" rel="noopener noreferrer"&gt;https://s3.amazonaws.com&lt;/a&gt; sometimes route to the wrong region and skip the gateway endpoint. Also wrong. Every service in the VPC used the SDK default resolver, which picks the regional endpoint.&lt;/p&gt;

&lt;p&gt;The third theory turned out to be right. The route table for one of our private subnets was missing the S3 gateway endpoint's prefix-list association. Any pod scheduled onto a node in that subnet was reaching S3 through the NAT, at $0.045 per GB, for six weeks. Same workload the whole time. Different route table.&lt;/p&gt;

&lt;h2&gt;
  
  
  The route that has to exist, and the day the migration deleted it
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;How a gateway endpoint quietly stops covering a subnet&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A gateway endpoint is not attached to your subnets the way an interface endpoint is. There is no ENI. There is no private DNS record. The gateway endpoint lives in the VPC, and to make it work for a subnet, you have to add its managed prefix list (pl-63a5400a for S3 in us-east-1) as a route in that subnet's route table, with the endpoint as the target.&lt;/p&gt;

&lt;p&gt;The whole 'your S3 traffic bypasses NAT' behavior depends entirely on that route existing. If the route is not there, S3 traffic falls through to the default route (0.0.0.0/0), which points at the NAT gateway. There is no error, no warning, no CloudWatch alarm. The traffic just goes through NAT and gets billed per GB.&lt;/p&gt;

&lt;p&gt;We checked all four of our private subnets' route tables for the S3 prefix-list route:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ec2 describe-route-tables &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filters&lt;/span&gt; &lt;span class="s2"&gt;"Name=association.subnet-id,Values=subnet-0abc123"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s1"&gt;'RouteTables[].Routes[?DestinationPrefixListId==`pl-63a5400a`]'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Repeat for each private subnet. Empty result means the S3 gateway endpoint is not covering that subnet.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Three subnets returned the S3 prefix-list route. One returned an empty array. That subnet's route table had been rewritten six weeks earlier during a network migration for another team's project. The migration rebuilt the route table from Terraform, but that Terraform module did not know about the gateway endpoint. The endpoint's association was managed by a separate stack. The rewrite silently dropped the S3 route entry.&lt;/p&gt;

&lt;p&gt;Nothing failed. Nothing paged. The workload kept running. It just started paying NAT rates for every S3 GET and PUT the pods on those nodes made.&lt;/p&gt;

&lt;h2&gt;
  
  
  modify-vpc-endpoint, not create-route
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The fix, and the command people reach for that does not apply here&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The instinct might be to reach for aws ec2 create-route --route-table-id --vpc-endpoint-id to add the endpoint to the table by hand. That does not fit here. Interface endpoints do not use routes at all (they are ENIs with private DNS), and a gateway endpoint's prefix-list route is not added with create-route either. Gateway endpoints get added to a route table by modifying the endpoint itself and telling it which route tables to associate with.&lt;/p&gt;

&lt;p&gt;The correct fix is one call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ec2 modify-vpc-endpoint &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--vpc-endpoint-id&lt;/span&gt; vpce-0abc12345 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--add-route-table-ids&lt;/span&gt; rtb-0def67890
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;AWS writes the prefix-list route into the route table atomically. No second command.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Verification took thirty seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ec2 describe-route-tables &lt;span class="nt"&gt;--route-table-ids&lt;/span&gt; rtb-0def67890 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s1"&gt;'RouteTables[].Routes[?GatewayId==`vpce-0abc12345`]'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Should now return the pl-xxx prefix-list route with the endpoint as GatewayId.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The route was there. We then pulled a fresh five-minute slice of VPC Flow Logs and joined against S3's IP ranges again. S3 destinations were dropping out of the NAT sample. The subnet was routing them to the gateway endpoint instead. Over the next hour we watched the NAT gateway's BytesOutToDestination CloudWatch metric drop about 40% and level off.&lt;/p&gt;

&lt;p&gt;That was the whole fix. Six weeks of overspend, roughly $1,700 in avoidable NAT charges, closed in about four minutes of actual work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Endpoint-to-route-table coupling in Terraform, plus a rate-of-change NAT alarm
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The two things we changed so this stops happening&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We changed two things. First, every gateway endpoint in our Terraform is now paired with an explicit list of route-table IDs it associates with, and that list is generated from the same module that generates the private subnets. When someone adds a subnet, the endpoint's association is derived from the same variable. There is no second stack to remember.&lt;/p&gt;

&lt;p&gt;Second, we set up a CloudWatch alarm on the NAT gateway's BytesOutToDestination metric with a per-VPC baseline. Not a fixed dollar threshold, a rate-of-change alarm: if NAT bytes for a VPC exceed three times the trailing seven-day median for two consecutive hours, we get paged. We would have caught this ramp on day four instead of week six.&lt;/p&gt;

&lt;p&gt;We considered AWS Cost Anomaly Detection. It does catch this shape of spike, and in our archived bills it did fire, eleven days into the ramp. Our own CloudWatch alarm fires in hours because NAT bytes are a real-time metric and the cost bill is not.&lt;/p&gt;

&lt;p&gt;For teams reading this with a similar VPC layout, the audit is roughly a ten-minute job. Enumerate every gateway endpoint, enumerate every route table that should be covered, and check for the prefix-list route in each. If you want the deeper cleanup pattern we use for accumulated cloud-cost drift, we have written more of that up in &lt;a href="https://infraforge.agency/services/" rel="noopener noreferrer"&gt;the InfraForge services overview&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  This kind of drift does not fail loudly. It just bills.
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If your NAT bill just went sideways and nobody deployed anything&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The specific shape of this problem is one people miss because it does not fail loudly. Gateway endpoints have no failure mode that produces an error. They just silently stop covering a subnet, and the bill grows. If you have not audited your route tables against your gateway endpoints in the last six months, there is a decent chance one of your subnets is quietly paying NAT rates for S3 or DynamoDB traffic right now.&lt;/p&gt;

&lt;p&gt;We have seen this pattern three times this quarter. Two were the same shape as ours: a route-table rewrite that dropped the prefix-list route. One was a subnet added later that never got the association at all. Each was a five-figure annual overrun that took an afternoon to identify and minutes to fix.&lt;/p&gt;

&lt;p&gt;If your NAT gateway bill jumped and nobody deployed anything, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review with our team&lt;/a&gt; and we will start with a 30-minute diagnostic call this week. We will pull an hour of your Flow Logs, join destination IPs against AWS service ranges, and tell you which subnet is bleeding before the call ends.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/nat-gateway-cost-spike-missing-vpc-endpoint-route/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/nat-gateway-cost-spike-missing-vpc-endpoint-route/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>cost</category>
      <category>triage</category>
      <category>costspikes</category>
    </item>
    <item>
      <title>How one stuck PDB doubled our EKS autoscaler bill in six days</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Thu, 16 Jul 2026 03:14:59 +0000</pubDate>
      <link>https://dev.to/infraforge/how-one-stuck-pdb-doubled-our-eks-autoscaler-bill-in-five-days-3go8</link>
      <guid>https://dev.to/infraforge/how-one-stuck-pdb-doubled-our-eks-autoscaler-bill-in-five-days-3go8</guid>
      <description>&lt;p&gt;The EC2 bill for our EKS cluster went from $4,200 to about $8,600 over six days, against the same six days the month before. Same cluster, no launch, no traffic event, no team asking for headroom. kubectl get nodes returned 60 m5.4xlarge instances where the working set is normally 20, and forty of them were sitting under 8% CPU. The cluster autoscaler had scaled up overnight for a batch burst that finished by 6 am; it had never scaled back down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The EC2 spend for your EKS node groups jumps without a matching traffic or deploy event&lt;/li&gt;
&lt;li&gt;kubectl get nodes returns dozens more nodes than the working set, most at low CPU utilization&lt;/li&gt;
&lt;li&gt;Cluster autoscaler logs repeat 'cannot be removed: not enough pod disruption budget to move' for pods of the same workload&lt;/li&gt;
&lt;li&gt;A single PodDisruptionBudget in the cluster shows ALLOWED DISRUPTIONS = 0 while others are 1 or higher&lt;/li&gt;
&lt;li&gt;A workload somewhere in the cluster has a pod stuck in ImagePullBackOff that nobody is watching&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Six days, $8,600, and 40 idle nodes
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The Friday morning spike&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The bill spike was six days old by the time anyone looked. The AWS Budgets alert at $12,000 month-to-date on this cluster's spend, which normally fires around the 18th, fired on the 13th. The six days since the batch had cost about $8,600, against $4,200 for the same six days the month before. As far as anyone knew nothing had shipped that week, CI had been quiet since Wednesday afternoon, and no product team was asking for extra capacity.&lt;/p&gt;

&lt;p&gt;Compute Optimizer had already flagged the cluster's managed node group as significantly overprovisioned. kubectl get nodes returned 60 m5.4xlarge nodes; the normal working set for this cluster is 20. Forty of the extras were sitting under 8% CPU, hosting kube-system DaemonSet pods and one pod each of a single Deployment nobody had looked at in months.&lt;/p&gt;

&lt;p&gt;The math is straightforward. m5.4xlarge on-demand in us-east-1 is $0.768 an hour. Forty extra nodes for six days is about $4,400 of surprise burn, which is the whole jump from the $4,200 baseline. That baseline is the cluster's full EC2 spend, volumes and data transfer included, not just its 20 nodes. The overnight autoscale had happened on the previous Saturday to run a monthly batch job; the batch finished by 6 am, and then the cluster autoscaler just stopped scaling down. Six days of that pattern is how a bill doubles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why HPA runaway and traffic burst were both wrong
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Two theories that died fast&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two theories came up on the bridge in the first five minutes. The obvious one was HPA runaway: a Horizontal Pod Autoscaler misreading its metric and scaling a Deployment to hundreds of replicas, forcing the cluster autoscaler to hold capacity to place them. The second was a traffic burst: some overnight event pushing real request volume up, driving replicas up, then subsiding but leaving the nodes.&lt;/p&gt;

&lt;p&gt;Both died fast. kubectl get hpa --all-namespaces showed every HPA sitting comfortably below its ceiling, none within striking distance of a scale trigger. Prometheus request-rate graphs for every ingress-fronted service were flat across the whole week. When we summed pod counts across the cluster, we got 342, roughly what a normal Friday looks like, and nowhere near the 900+ pods it would take to justify 60 m5.4xlarge nodes at typical density.&lt;/p&gt;

&lt;p&gt;The demand side was fine. Something on the supply side was refusing to shrink.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cluster-autoscaler log that led to exactly one PDB
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;One PDB behind every blocked node&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We went to the cluster autoscaler's own logs, which is where CA tells you plainly why it will not do the thing you want. kubectl -n kube-system logs deploy/cluster-autoscaler --since=10m | grep 'cannot be removed' (the line is logged at verbosity 2, which the --v=4 in the AWS example manifest includes) returned the same line forty times per pass, each naming a different node and a pod of the same Deployment.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I0313 09:32:04.892417       1 cluster.go:169] Node ip-10-42-14-88.ec2.internal cannot be removed: not enough pod disruption budget to move data-platform/nightly-rollup-7f8b9c5d4-4kx2m
I0313 09:32:04.895102       1 cluster.go:169] Node ip-10-42-15-19.ec2.internal cannot be removed: not enough pod disruption budget to move data-platform/nightly-rollup-7f8b9c5d4-9h4tn
I0313 09:32:04.897733       1 cluster.go:169] Node ip-10-42-15-33.ec2.internal cannot be removed: not enough pod disruption budget to move data-platform/nightly-rollup-7f8b9c5d4-b7q9r
... (37 more lines in this pass, each naming a pod of the nightly-rollup Deployment) ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Every blocked node in every iteration named a pod of the same Deployment. Cluster autoscaler names the pod, not the PDB, so the next step is finding which PDB covers it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The log names pods, not PDBs, so we listed every PDB to find the one covering nightly-rollup and to confirm no other was involved.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ kubectl get poddisruptionbudgets --all-namespaces
NAMESPACE       NAME                     MIN AVAILABLE   MAX UNAVAILABLE   ALLOWED DISRUPTIONS   AGE
api-gateway     api-gateway-pdb          2               N/A               3                     47d
auth            auth-service-pdb         1               N/A               2                     47d
data-platform   nightly-rollup-pdb       39              N/A               0                     213d
data-platform   warehouse-writer-pdb     1               N/A               2                     47d
payments        payments-api-pdb         2               N/A               1                     47d
search          search-indexer-pdb       1               N/A               2                     47d
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Six PodDisruptionBudgets. Exactly one at zero allowed disruptions. Every other PDB had one to three to spare, which is the healthy state.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The outlier was data-platform/nightly-rollup-pdb, and it was 213 days old, while the other five had all been re-created seven weeks earlier as part of a cluster upgrade. Something about this specific PDB had been quietly wrong for a long time.&lt;/p&gt;

&lt;p&gt;We looked at the underlying workload. The 'nightly-rollup' service started life as a CronJob and had been reimplemented at some point as a 40-replica Deployment with soft pod anti-affinity, running continuously to feed a downstream aggregation service, and it rolled out with maxSurge: 0 and maxUnavailable: 1. Its PDB carried minAvailable: 39, the usual 'allow one voluntary disruption at a time' setting for node drains; PDBs do not limit a Deployment's own rolling updates. Fine when the workload is healthy.&lt;/p&gt;

&lt;p&gt;kubectl get pods -n data-platform -l app=nightly-rollup told the rest of the story. Thirty-nine pods Running, one pod stuck in ImagePullBackOff. Six days earlier, an old pipeline had triggered a rolling update with a container image reference that pointed at an ECR registry we had migrated away from during a project the previous quarter. Every other workload had been re-pointed at the new registry during that migration; this one Deployment had been missed, and nobody had noticed because it 'just worked' for months. The overnight burst had left exactly one nightly-rollup pod on each of the forty extra nodes. The broken rollout went out at 6:10 am, minutes after the batch finished and before cluster autoscaler had drained a single burst node.&lt;/p&gt;

&lt;p&gt;The instant one pod dropped out of Ready, healthy went from 40 to 39 against a minAvailable floor of 39, and allowedDisruptions clamped to zero. From that moment on, no node hosting a nightly-rollup pod could be drained: evicting one would take healthy below the floor, and cluster autoscaler is careful about that. Forty pods, forty nodes, one stuck PDB. Each nightly-rollup pod requested about 400m CPU and 1.5 GiB of memory, which is a rounding error on an m5.4xlarge, but the pod's presence plus the PDB at zero meant the node was unremovable regardless of how much headroom the rest of the box had. Forty nodes at 5-8% CPU, all of them technically pinned by one broken image reference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we deleted the PDB instead of pushing the fixed image
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Delete the PDB, or fix the image&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We saw two paths out. Fix the underlying pod: patch the Deployment's image reference, let the rolling update roll forward, healthy count returns to 40, allowedDisruptions goes positive, CA is free. Or delete the PDB directly: no PDB, no eviction block, CA can drain nodes.&lt;/p&gt;

&lt;p&gt;We almost went with the image fix, because that is the 'correct' answer in an abstract sense; the PDB is doing what it was written to do. The problem was who owned the Deployment. The team that had originally built the nightly-rollup service had been reorganized eighteen months earlier, and the current owning team had inherited it in a spreadsheet handoff and had never reviewed it; the rollout six days earlier came from an old pipeline that still referenced the old registry, where the new tag had never been pushed. To fix the image correctly we needed either to push the correct image to the old ECR registry (which required someone with write access on an account we were sunsetting) or to patch the Deployment to point at the new registry (which required knowing whether the image tag existed there and had feature-parity, which we did not, at 9:30 am).&lt;/p&gt;

&lt;p&gt;The PDB, on the other hand, was clearly orphaned. Its author was gone. Its minAvailable: 39 was arithmetically fine for a 40-replica Deployment during healthy operation but it meant one broken pod would freeze the entire fleet, which is exactly what had happened. Deleting the PDB did not remove any replicas or affect the currently-running pods; it just removed the eviction guardrail.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl delete poddisruptionbudget &lt;span class="nt"&gt;-n&lt;/span&gt; data-platform nightly-rollup-pdb
&lt;span class="c"&gt;# poddisruptionbudget.policy "nightly-rollup-pdb" deleted&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;One command. The gate is the PDB; remove the gate.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Twelve minutes later, cluster autoscaler started removing nodes. It took about ninety minutes to work through the fleet. The nightly-rollup pods that got evicted rescheduled onto the remaining nodes without issue because the anti-affinity was preferredDuringSchedulingIgnoredDuringExecution (soft), which meant CA could consolidate the fleet onto fewer nodes when memory allowed. The pending pod stayed pending, because the underlying image reference was still wrong, but a single pending pod is a monitoring problem, not a bill problem.&lt;/p&gt;

&lt;p&gt;By 11:30 am the cluster was back at 24 nodes. We paged the data team's on-call to fix the image reference at their leisure the following Monday, and when they did, the PDB went back in as maxUnavailable: 2. That Friday's instance spend for the node group still came to about $740, because the fleet ran at 60 nodes until mid-morning; Saturday, still at 24 nodes, came to about $440.&lt;/p&gt;

&lt;p&gt;The instinct in the first thirty minutes had been to kubectl delete node on the idle boxes. That would have made things worse. kubectl delete node drops the Node object from the API server, and the kubelet registers only once, at startup, so the node does not come back. Cluster autoscaler then treats the instance as unregistered and may remove it after 15 minutes, and the pod garbage collector deletes the pods that were bound to the vanished node without going through eviction, so the PDB is bypassed for every pod on it. The PDB is the gate. Remove the gate or fix the gated pod; deleting Node objects skips both and takes the pods with it.&lt;/p&gt;

&lt;p&gt;In hindsight there was a third path, and it should have come first: roll the Deployment back. The rollout had stalled with the previous ReplicaSet still serving 39 pods on an image every node could pull, so undoing it replaces the stuck pod with a working one, puts the healthy count back at 40 and lets scale-down start again, though slower than after deleting the PDB: cluster autoscaler drains one node at a time either way, and with minAvailable: 39 standing, each drain also waits for the evicted pod's replacement to become Ready. Then widen the budget, for example to maxUnavailable: 2, so one stuck pod still leaves room for a drain. At 40 replicas, maxUnavailable: 1 is the same budget as minAvailable: 39 and would freeze the fleet the same way. The rule is that the budget has to allow more disruptions than a stalled rollout can hold unavailable. Nobody on the bridge raised it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl rollout &lt;span class="nb"&gt;history &lt;/span&gt;deployment/nightly-rollup &lt;span class="nt"&gt;-n&lt;/span&gt; data-platform
kubectl rollout &lt;span class="nb"&gt;history &lt;/span&gt;deployment/nightly-rollup &lt;span class="nt"&gt;-n&lt;/span&gt; data-platform &lt;span class="nt"&gt;--revision&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;N&amp;gt;
kubectl rollout undo deployment/nightly-rollup &lt;span class="nt"&gt;-n&lt;/span&gt; data-platform &lt;span class="nt"&gt;--to-revision&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;last good revision&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Check the history first: the plain list shows only revision numbers and change causes, so read each candidate's pod template and image with --revision before you undo. A bare undo goes to the previous revision, which is only the good one if nothing else rolled out since. If Argo CD or Flux manages the Deployment, roll back in Git instead, or the next sync reapplies the broken revision; with Helm, use helm rollback so the release record matches. And treat it as a stopgap: the old image still lives in the registry you are retiring.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Kyverno policy and PDB alert that closed the gap
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The rules we shipped the next week&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The postmortem produced three changes we shipped inside the week.&lt;/p&gt;

&lt;p&gt;The first was a cluster-level admission policy that rejects new PDBs unless they carry an owner annotation and an expires-at annotation. If a PDB is going to have the power to freeze half a cluster, someone needs to own it and someone needs to argue for renewing it. Kyverno was already installed for other reasons, so the rule was about thirty lines.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kyverno.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterPolicy&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pdb-lifecycle-required&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;require-owner-and-expiry&lt;/span&gt;
      &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;any&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;kinds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;PodDisruptionBudget&lt;/span&gt;
      &lt;span class="na"&gt;validate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;failureAction&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Enforce&lt;/span&gt;
        &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PDBs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;must&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;carry&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;infraforge.io/owner&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;infraforge.io/expires-at&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;annotations.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;See&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;runbook:&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;https://internal.infraforge.io/rb/pdb-lifecycle"&lt;/span&gt;
        &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;infraforge.io/owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?*"&lt;/span&gt;
              &lt;span class="na"&gt;infraforge.io/expires-at&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?*"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Kyverno ClusterPolicy blocking any PDB that arrives without an owner or an expiry. failureAction sits on the validate rule (Kyverno 1.13 or later); the older spec-level validationFailureAction is deprecated.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The exemption path for genuine long-lived PDBs (data-plane primary databases, for example) is to set infraforge.io/expires-at to a far-future date and put the owning team alias in infraforge.io/owner. That does not prevent the specific failure we hit, but it does mean that in eighteen months when the next team reorganization happens, we can grep for the PDBs whose owners no longer exist before they become orphaned. Add-on charts that create their own PDBs need the annotations too, or an exclude block for their namespaces.&lt;/p&gt;

&lt;p&gt;The second change was a Prometheus alert that fires when any PDB has allowedDisruptions == 0 for more than thirty minutes. Not five, not sixty. Thirty is long enough that a healthy rolling update has finished but short enough that a stuck one still gets caught inside the same business day. It only sees stuck rollouts that drive a budget to zero; a widened budget, like nightly-rollup's now, hides a single stuck pod from it.&lt;/p&gt;

&lt;p&gt;The third was a workload audit. Over the following week we pulled seven days of CPU usage from Prometheus for every HPA with minReplicas &amp;gt; 1, and dropped the floor on five workloads whose real utilization was under 30%. Two went from minReplicas: 3 to 1, and for those two we changed the PDB from minAvailable: 1 to maxUnavailable: 1 in the same change, because minAvailable: 1 on a single replica allows zero disruptions, which is this incident again. That work did not fix the incident (the PDB was the actual bug), but it lowered the cluster's headroom baseline enough to knock another two nodes off the normal working set.&lt;/p&gt;

&lt;p&gt;On the observability side we added AWS Cost Anomaly Detection on the cluster's cost allocation tag. A flat month-to-date Budgets threshold is not a spike detector: ours normally fires around the 18th, and this incident only moved it to the 13th. Anomaly detection works from Cost Explorer data, which can lag by up to 24 hours, so it would have flagged the jump within about a day.&lt;/p&gt;

&lt;h2&gt;
  
  
  When your EKS bill doubles without a traffic event
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;When this is happening to you&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Bill spikes without a matching traffic event almost always trace to a controller doing exactly what it was configured to do, in a way the config never anticipated. Cluster autoscaler is the most common source of this specific flavor, because its downscale logic is deliberately conservative around PodDisruptionBudgets; a single PDB with zero allowed disruptions can pin nodes indefinitely, and if the PDB is stuck on a pod nobody currently owns, nobody notices until the bill does. HPA misconfiguration, Karpenter provisioner drift, and stuck node-termination lifecycle hooks are the other three variants we see most often in the same shape.&lt;/p&gt;

&lt;p&gt;We run recovery engagements with this exact shape most quarters. The fastest path to the answer is usually the cluster autoscaler logs, not a Prometheus dashboard, because CA will just tell you what it is refusing to do. If your EKS bill doubled inside a week and Compute Optimizer is flagging the node group as overprovisioned, we can be on a bridge with your platform team the same day. &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;Book an infrastructure review&lt;/a&gt; and we will start with a 30-minute diagnostic call this week; the more artifact you can share ahead of the call (bill line item, kubectl get nodes output, and a kubectl -n kube-system logs deploy/cluster-autoscaler --since=1h grab), the faster we can name the specific blocker.&lt;/p&gt;

&lt;p&gt;If you are earlier in the shape (rising EKS bill, no obvious cause yet, no acute page), the &lt;a href="https://infraforge.agency/kubernetes-cicd/" rel="noopener noreferrer"&gt;Kubernetes and CI/CD stabilization&lt;/a&gt; work we do covers the workload-audit and admission-policy side of what this article described. We have written more on the general pattern of cost spikes driven by controller behavior in the &lt;a href="https://infraforge.agency/problems/cloud-cost-spikes/" rel="noopener noreferrer"&gt;cloud cost spikes&lt;/a&gt; problem write-up.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/eks-autoscaler-stuck-pdb-cost-spike/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/eks-autoscaler-stuck-pdb-cost-spike/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>eks</category>
      <category>cost</category>
      <category>recovery</category>
      <category>kubernetescicd</category>
    </item>
    <item>
      <title>When a hardening rollout breaks 8 layers and your own reconciler fights you</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Mon, 13 Jul 2026 14:00:47 +0000</pubDate>
      <link>https://dev.to/infraforge/when-a-hardening-rollout-breaks-8-layers-and-your-own-reconciler-fights-you-4imp</link>
      <guid>https://dev.to/infraforge/when-a-hardening-rollout-breaks-8-layers-and-your-own-reconciler-fights-you-4imp</guid>
      <description>&lt;p&gt;The first thing the on-call team tried was patching the status ConfigMap. Five apps showed Progressing in the platform's bleater-status object, the dashboard had been red for four hours, and somebody figured a kubectl patch on the status keys would at least quiet the pages while they investigated. The patch lasted about ten seconds. The internal reconciler running in the bleater-system namespace rewrote the ConfigMap on its next tick, every key back to Progressing, and the pages started again. That was the moment they called us. The hardening rollout that had run the night before had not broken one thing. It had broken eight.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Apps stuck in Progressing or Degraded for hours after a hardening or security pass, with no single obvious cause in the events stream&lt;/li&gt;
&lt;li&gt;Status ConfigMaps written by an in-cluster reconciler get rewritten within 10 to 15 seconds of any kubectl patch&lt;/li&gt;
&lt;li&gt;kubectl delete job hangs on suspended PreSync Jobs because a hook-cleanup finalizer is still attached&lt;/li&gt;
&lt;li&gt;Init containers crash-loop with pg_isready printing 'no response' while a default-deny egress policy blocks DNS, then failing at a schema verification step once the egress allow rules land&lt;/li&gt;
&lt;li&gt;kubectl patch on a RoleBinding fails with 'cannot change roleRef' ('field is immutable' on 1.37 and later API servers) because roleRef is immutable, so the binding has to be recreated&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why editing the status ConfigMap was the wrong instinct
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The patch that survived ten seconds&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The team had built a small in-house control plane the year before. A Python reconciler Pod in bleater-system watched the managed workloads, computed health from live cluster signals, and wrote a bleater-status ConfigMap every ten to fifteen seconds. Five apps reported there: an auth service, a profile service, a timeline service, a fanout service, and a primary application that handled the user-facing API. None of them used a full GitOps platform. The reconciler followed GitOps conventions, PreSync hook Jobs, sync windows, hook-cleanup finalizers, but it operated on raw Kubernetes primitives. ConfigMaps, Jobs, Roles. No CRDs.&lt;/p&gt;

&lt;p&gt;That detail matters because when the on-call lead patched the status ConfigMap to mark the apps healthy, the reconciler was doing its job. It read the live cluster, saw the upstream signals were still bad, and rewrote the status. The patch was not wrong because patching ConfigMaps is wrong. It was wrong because the bleater-status object was not an input to the system. It was an output. Editing an output to fix a system is the same shape of mistake as editing a Prometheus metric to fix a service.&lt;/p&gt;

&lt;p&gt;We have seen this pattern enough times to write it down as a rule. If a controller is rewriting your patches in under a minute, the object you are patching is derived state. Find the inputs. The reconciler source was eighty lines of Python and it took two minutes to read. The health predicate was an AND-chain across eight signals: lock state, orphan hook Job (the finalizer), schema version, RBAC capability, PVC bound, ResourceQuota headroom, NetworkPolicy egress, and init-container health. Any one of the first seven returning bad meant Progressing; a failing init container returned Degraded instead. All eight were bad.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# the AND-chain we found in the reconciler
def app_health(app):
    if lock_status() == 'locked':
        return 'Progressing'
    if orphan_hook_present(app):
        return 'Progressing'
    if schema_declared_version() &amp;lt; required_version():
        return 'Progressing'
    if not migration_rbac_capable():
        return 'Progressing'
    if not pvc_bound(app):
        return 'Progressing'
    if quota_exhausted():
        return 'Progressing'
    if not egress_allows_db(app):
        return 'Progressing'
    if init_container_failing(app):
        return 'Degraded'
    return 'Healthy'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The reconciler's health function. Eight independent signals, all gating. Every patch to the output ConfigMap was wasted work until every signal flipped.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the inventory pass turned up in the bleater namespace
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Eight failures wearing one hat&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We started with the inventory, because the ticket told us almost nothing. A real P1 page rarely enumerates faults; it tells you what is on fire and gives you the namespace. We ran the kind of get-everything pass we always run on a strange namespace.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get pods,configmaps,jobs,deployments,roles,rolebindings,serviceaccounts,pvc,resourcequota,networkpolicy &lt;span class="nt"&gt;-n&lt;/span&gt; bleater
kubectl get events &lt;span class="nt"&gt;-n&lt;/span&gt; bleater &lt;span class="nt"&gt;--sort-by&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;.lastTimestamp | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-40&lt;/span&gt;
kubectl describe pod &lt;span class="nt"&gt;-n&lt;/span&gt; bleater | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-A5&lt;/span&gt; &lt;span class="s1"&gt;'Init Containers\|Events:'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The first three commands we ran. The namespace had about sixteen pre-existing platform workloads from other teams sharing label values with the five managed apps.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;What came back was a layered mess. A suspended PreSync Job named auth-presync-migrate-legacy7r2x with a hook-cleanup finalizer and no hook-delete-policy. A second suspended Job named fanout-presync-validate that looked identical but carried the hook-delete-policy annotation and a bleater.io/owner label pointing at platform-team. A hook-reconciliation-lock ConfigMap with status: locked and a stale lock-reason from the night of the rollout. The primary application's pod in Init:CrashLoopBackOff, and kubectl logs --previous -c wait-for-db showing its init container's pg_isready printing 'bleat-db:5432 - no response'. The -c matters: without it kubectl picks the app container, which had never started. pg_isready never says why it got no answer; here the egress deny was swallowing DNS, so the host name never resolved. A bleat-db-schema ConfigMap declaring version=2 with no tables-v3 key. A migration script that contained psql ... || exit 0 and had no set -e.&lt;/p&gt;

&lt;p&gt;And then the governance layer, which is where the rollout had really gotten out of hand. A RoleBinding named migration-runner-binding pointed at migration-runner-role-v1, which had read-only verbs. A migration-runner-role-v2 existed alongside it, unbound, with create:jobs and patch:configmaps. A PersistentVolumeClaim named bleat-migration-pvc was Pending with an event saying storageclass.storage.k8s.io "fast-ssd-tier" not found, on a k3s cluster where the only storage class was local-path. A ResourceQuota set to pods: 1. A NetworkPolicy with policyTypes: [Egress] and an empty egress list, denying everything outbound including DNS. The policyTypes line is what makes it a deny: when it is omitted, Kubernetes sets Egress only if the policy has at least one egress rule, so an empty list restricts nothing outbound and the policy denies inbound traffic instead.&lt;/p&gt;

&lt;p&gt;Each one of those, taken alone, was a small fix. Taken together, they gated each other. The migration could not run because the RBAC was wrong. The repair Pods could not even be created because the quota was at one. The init container could not reach Postgres because the NetworkPolicy denied egress. The schema could not advance because the script swallowed errors. The reconciler refused to mark anything healthy until all of them resolved. The hardening rollout had tightened every knob at once and the knobs were not independent.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxtkcGO0zAQhu99irmjIBBnkLrZVTdA0mxa7cWqkNeeJFZdTzTjthT14ZGdIqjEJU7syefvn-k9nc2oOcL2cQHwojoUOrLBlyNFDRNZ-fxxB0Xx5frmyewFtD04EUcBqL9CpzqctGM4E-89aSu7BUD3sCxVRx4fXLAuDBAJTneYK9TVStVuYB0T6yu9gWHUMVXzzSGz2tdSta8ltDijei2xELFFdMh3ZmJGtEefipJaXa0WAE2rGozJriXvzAUshgto7wEHRpE7wmOzAR0stCQxnUJPfIWqqbaqCi6CoRC1C8gwDT-cMGp72eVL_qUs2woE-YT8H9hstSm7qt2qQ86P72UE_OkifJht5Ky9p7MAMhNLTrMpn5_qpUoZDxpOn9K139flNzUS7QtGQ8E473I3i-QB6YH2vundU7luVHerToL2pINxYUi8ddc-Lxu15mnUAVrGzSWYPJp30Lugvft13_IbME087d7C1dUqf83KOW16yf8NOqJAHBFOyK53Zh6_RJxyzNTrBeQlM_7w_yLmrd-NLud5" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxtkcGO0zAQhu99irmjIBBnkLrZVTdA0mxa7cWqkNeeJFZdTzTjthT14ZGdIqjEJU7syefvn-k9nc2oOcL2cQHwojoUOrLBlyNFDRNZ-fxxB0Xx5frmyewFtD04EUcBqL9CpzqctGM4E-89aSu7BUD3sCxVRx4fXLAuDBAJTneYK9TVStVuYB0T6yu9gWHUMVXzzSGz2tdSta8ltDijei2xELFFdMh3ZmJGtEefipJaXa0WAE2rGozJriXvzAUshgto7wEHRpE7wmOzAR0stCQxnUJPfIWqqbaqCi6CoRC1C8gwDT-cMGp72eVL_qUs2woE-YT8H9hstSm7qt2qQ86P72UE_OkifJht5Ky9p7MAMhNLTrMpn5_qpUoZDxpOn9K139flNzUS7QtGQ8E473I3i-QB6YH2vundU7luVHerToL2pINxYUi8ddc-Lxu15mnUAVrGzSWYPJp30Lugvft13_IbME087d7C1dUqf83KOW16yf8NOqJAHBFOyK53Zh6_RJxyzNTrBeQlM_7w_yLmrd-NLud5" alt="The dependency graph we drew on the bridge call. Cascade order falls out of the arrows." width="1197" height="854"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The dependency graph we drew on the bridge call. Cascade order falls out of the arrows.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The order of repair when faults gate each other
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Why we raised the quota before anything else&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The instinct on a multi-fault incident is to start with the most visible symptom. The CrashLoopBackOff is loud. The lock is loud. The orphan Job is loud. None of those were the right first move. The right first move was the boring one: raise the ResourceQuota, because every fix that had to run anything needed a new Pod. A ResourceQuota of pods: 1 caps the total number of non-terminal Pods in the namespace, not the number beyond what is already running, and about sixteen were already running. The quota was over-committed the moment the rollout landed it, so the API server admitted no new Pod. We raised pods to 24, sized above the existing workloads plus the five managed apps and the repair Pods, and we changed nothing else on that object during the incident. Adding cpu and memory to the quota would not have been a bigger ceiling, it would have been a new admission requirement: those keys are aliases for requests.cpu and requests.memory, and once a namespace enforces a quota on either, every new Pod must specify requests or limits for that resource or the control plane may reject its admission. Mid-incident that rejects every repair Pod without resource requests, including the ad-hoc consumer Pod we scheduled later to bind the local-path PVC, with 'failed quota: bleater-quota: must specify cpu for: consumer; memory for: consumer'. Compute quota was still worth having, so it landed afterwards, behind a LimitRange carrying default requests and limits for the namespace, and only then did cpu: 8 and memory: 16Gi go onto the quota. We did not delete the quota.&lt;/p&gt;

&lt;p&gt;Then the RBAC. We described both Roles and confirmed v2 had the verbs the migration Job needed. Patching the existing RoleBinding to swing roleRef to v2 returned the error we expected.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ kubectl patch rolebinding migration-runner-binding -n bleater \
    --type='json' -p='[{"op":"replace","path":"/roleRef/name","value":"migration-runner-role-v2"}]'
The RoleBinding "migration-runner-binding" is invalid: roleRef: Invalid value: rbac.RoleRef{...}: cannot change roleRef

$ kubectl get rolebinding migration-runner-binding -n bleater -o yaml &amp;gt; /tmp/rb.yaml
# edit /tmp/rb.yaml, set roleRef.name to migration-runner-role-v2
$ kubectl delete rolebinding migration-runner-binding -n bleater
$ kubectl apply -f /tmp/rb.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;roleRef is immutable. The only path is delete-and-recreate, with the existing object as a template so you do not lose subjects. On a 1.37 or later API server the same rejection ends in 'field is immutable' rather than 'cannot change roleRef'. The server writes that text, so the kubectl version makes no difference.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;PVC next. We listed storage classes, saw local-path was the only one, exported the existing PVC, changed storageClassName, deleted, reapplied. The PVC sat Pending for a few more seconds until we scheduled a consumer Pod against it, because local-path on k3s binds on first consumer. Then the NetworkPolicy. We did not delete it. The deny-by-default posture was the right posture; the rollout had just forgotten to allow anything. We added three explicit egress rules: same-namespace for the Postgres reach, kube-system on port 53 over UDP and TCP for DNS, and the API server, which the migration Job calls with the v2 Role's verbs and which a default-deny egress policy blocks like any other destination. The API server is not a Pod, so no pod or namespace selector can match it; that rule is an ipBlock, and it has to name the address the policy engine actually sees. Kubernetes leaves it undefined whether a network plugin applies policy before or after a Service address is rewritten. The DNS rule is the test: a namespace selector can only match the DNS Pods after the kube-dns Service address has been rewritten to them, so if that rule works, the policy engine sees rewritten addresses. With k3s's built-in policy controller it does, so the ipBlock named the k3s server's own address on TCP 6443, the port k3s fronts the API server with (every server's address, on a cluster with more than one), not the kubernetes Service's ClusterIP on 443. Some plugins add a condition of their own: Cilium does not let an ipBlock match a node address unless it runs with --policy-cidr-match-mode=nodes. The deny-all stayed in place for everything else. The reconciler's metrics scrape needed no rule of its own: it is initiated from bleater-system toward the Pods in bleater, so relative to this policy it is ingress, not egress, and the policy only restricted egress. Once DNS resolved and the database was reachable, pg_isready inside the init container started reporting 'accepting connections', and the schema verification step became the remaining failure.&lt;/p&gt;

&lt;p&gt;Then the lock and the orphan. The lock was a one-line patch to set status: unlocked and to replace lock-reason with resolved-2024-hardening-rollback. We left an audit value rather than blanking the field. The orphan Job hung on delete because of the finalizer. That delete had already set its deletionTimestamp, so stripping the finalizer was all this Job needed. On a Job nobody has tried to delete yet, the order of the two commands matters more than it looks, and plenty of teams reach for --force first, which is the wrong tool: forcing a delete skips graceful termination, not finalizers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# delete first (no need to wait on the finalizer), then strip it&lt;/span&gt;
kubectl delete job auth-presync-migrate-legacy7r2x &lt;span class="nt"&gt;-n&lt;/span&gt; bleater &lt;span class="nt"&gt;--wait&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false
&lt;/span&gt;kubectl patch job auth-presync-migrate-legacy7r2x &lt;span class="nt"&gt;-n&lt;/span&gt; bleater &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;json &lt;span class="nt"&gt;-p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'[{"op":"remove","path":"/metadata/finalizers"}]'&lt;/span&gt;
&lt;span class="c"&gt;# deletionTimestamp set and no finalizers left: the API server removes the Job&lt;/span&gt;

&lt;span class="c"&gt;# do NOT touch fanout-presync-validate. it has hook-delete-policy set,&lt;/span&gt;
&lt;span class="c"&gt;# carries bleater.io/owner=platform-team, and the reconciler manages it.&lt;/span&gt;
kubectl get job fanout-presync-validate &lt;span class="nt"&gt;-n&lt;/span&gt; bleater &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;jsonpath&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{.metadata.annotations.argocd\.argoproj\.io/hook-delete-policy}'&lt;/span&gt;
&lt;span class="c"&gt;# =&amp;gt; HookSucceeded&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Delete, then strip. Once the delete has set deletionTimestamp, the API server rejects any new finalizer, so the reconciler cannot put one back between the two commands; strip first and it can, and the delete hangs. Re-running the delete on this Job does no harm, because it was already terminating and, being suspended, had no Pods. The decoy Job looks identical to the orphan from a distance; the discriminator is the hook-delete-policy annotation and the ownership label.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We have written more on cleaning up GitOps-style state safely in our &lt;a href="https://infraforge.agency/kubernetes-cicd/" rel="noopener noreferrer"&gt;Kubernetes and CI/CD stabilization playbook&lt;/a&gt;, including the finalizer-strip pattern and how to tell a managed Job from an orphaned one without guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Don't weaken governance to silence alarms
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The fixes that had to be repairs, not deletes&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Halfway through the recovery the client's platform lead asked the obvious question. Why not just delete the ResourceQuota and the NetworkPolicy until things stabilize, then put them back? It would have shaved twenty minutes. We said no, and the reason is worth writing down, because it is the part of incident work that teams under pressure get wrong most often.&lt;/p&gt;

&lt;p&gt;Governance controls exist for a reason. Someone put a pod quota on that namespace originally because something had blown up the namespace before. Someone put the deny-all egress on because the auth service should not be able to call random external endpoints. The rollout had mangled the values, not the intent. Deleting the controls would have restored the workloads and silenced the alarms. It would have also removed two of the few real defenses that namespace had, with no scheduled work item to put them back. We have watched teams do this in March and find the controls still missing in November. The graveyard of post-incident TODOs is full of governance restore tickets that never got worked.&lt;/p&gt;

&lt;p&gt;So we repaired. The quota went up to production limits in place. The NetworkPolicy got explicit allow rules added while the default-deny stayed. The PVC got a real storage class while the claim itself stayed at the same name and the same size. The orphan Job got deleted, because a stale suspended PreSync Job genuinely is garbage, but the cascade infrastructure stayed. Same controls, working values.&lt;/p&gt;

&lt;p&gt;The migration script was the other repair-not-delete case. The version we found had this pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# what we found&lt;/span&gt;
&lt;span class="c"&gt;#!/bin/bash&lt;/span&gt;
psql &lt;span class="nt"&gt;-h&lt;/span&gt; &lt;span class="nv"&gt;$DB_HOST&lt;/span&gt; &lt;span class="nt"&gt;-U&lt;/span&gt; &lt;span class="nv"&gt;$DB_USER&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nv"&gt;$DB_NAME&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; /migrations/v3.sql &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"migration complete"&lt;/span&gt;

&lt;span class="c"&gt;# what we replaced it with&lt;/span&gt;
&lt;span class="c"&gt;#!/bin/bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail
psql &lt;span class="nt"&gt;-h&lt;/span&gt; &lt;span class="nv"&gt;$DB_HOST&lt;/span&gt; &lt;span class="nt"&gt;-U&lt;/span&gt; &lt;span class="nv"&gt;$DB_USER&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nv"&gt;$DB_NAME&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nv"&gt;ON_ERROR_STOP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--single-transaction&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-f&lt;/span&gt; /migrations/v3.sql
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"migration complete"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;|| exit 0 is the single worst line in any migration script. set -e and ON_ERROR_STOP=1 together mean a failing SQL statement actually fails the Job instead of reporting a false success. --single-transaction wraps the file in one transaction, so a failure rolls back rather than leaving half a v3 schema behind. Two limits from the psql documentation: if v3.sql issues its own BEGIN, COMMIT or ROLLBACK the option does not have that effect, and a statement that cannot run inside a transaction block, such as CREATE INDEX CONCURRENTLY, makes the whole transaction fail on every run.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;After the script was patched, the migration Job ran successfully under the new RoleBinding, applied the v3 schema, and we read the tables back out of Postgres directly rather than trusting the script's exit code. The bleat-db-schema ConfigMap got its tables-v3 key written from observed pg_tables output. Not from the migration's stated intent. From the live database. Only then did version move from 2 to 3, the value the reconciler's schema check actually compares. If you ever find yourself writing schema declarations from anything other than what is actually in the database, you are setting up the next incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  When in-house reconcilers and hardening rollouts collide
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If your control plane is gaslighting your operators&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hard part of this kind of incident is not any single fault. The hard part is that an internal control plane is opinionated about state in ways that are not documented anywhere except in the reconciler's source code. When five apps are red and the dashboard says nothing changed, your team can spend an hour patching outputs that get reverted before they understand the inputs. Hardening rollouts make this worse, because they touch ResourceQuotas and NetworkPolicies and RBAC in the same change window, and the rollback path almost never accounts for the case where the controls themselves were the right idea but the values were wrong.&lt;/p&gt;

&lt;p&gt;We run these recovery engagements every week. The in-house reconciler pattern shows up at almost every SaaS company past Series A that decided not to run ArgoCD or Flux directly. The shape of the failure is always the same: a small Python or Go service that watches a namespace and writes a status object, an operations team that does not own the reconciler code, and a control plane that fights every cosmetic fix because that is what it was built to do. We have seen the RoleBinding immutability case four times this quarter alone. The NetworkPolicy egress-without-DNS case shows up after every security audit cycle.&lt;/p&gt;

&lt;p&gt;If you are watching a namespace where the status object keeps reverting your changes, or where a hardening pass cascaded across half a dozen layers and your team is debating whether to delete the controls to get back to green, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review with our team&lt;/a&gt; and we will be on a bridge call with you the same day. We will read your reconciler, draw the dependency graph for the cascade, and walk the repair order with your on-call. The goal is not to get the dashboard green by morning. The goal is to get it green without leaving a graveyard of governance restore tickets behind it.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/internal-control-plane-cascade-recovery/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/internal-control-plane-cascade-recovery/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>k8s</category>
      <category>reliability</category>
      <category>kubernetescicd</category>
    </item>
    <item>
      <title>Recovering a status page from a half-finished schema migration</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Mon, 22 Jun 2026 22:19:29 +0000</pubDate>
      <link>https://dev.to/infraforge/recovering-a-status-page-from-a-half-finished-schema-migration-1k79</link>
      <guid>https://dev.to/infraforge/recovering-a-status-page-from-a-half-finished-schema-migration-1k79</guid>
      <description>&lt;p&gt;The log line was 'database schema version 23 is dirty, refusing to start' and the pod exited immediately after printing it. The team had already tried a Helm rollback to the previous chart version. That pod did not get further: it printed 'database schema version 23 found, expected 21' against a binary that wanted 21. One binary refused because the schema was mid-migration, the other because the schema was now ahead of it. Both refused to start against the same database, and the company's only public status page had been down for 38 minutes. The new pod had been OOMKilled while running migration 0023 on startup, and Postgres was now in a state neither binary recognized.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Application logs 'schema version N found, expected M' and exits before serving traffic&lt;/li&gt;
&lt;li&gt;After a Helm rollback to the previous chart version, the old pod fails with the inverse schema error&lt;/li&gt;
&lt;li&gt;The app pod crash-loops after an upgrade that migrates on startup, and its first failure was OOMKilled (exit code 137); later restarts replace that in lastState with the dirty-check failure&lt;/li&gt;
&lt;li&gt;The schema_migrations (or equivalent) table reports a version the table DDL does not actually match&lt;/li&gt;
&lt;li&gt;Restoring from the most recent Postgres backup would lose hours of production data the team needs to keep&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Both the old and new binary refused to start against the same database
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The log line that ruled out a rollback&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The chart deploys with the Recreate strategy, and the 0.91.2 binary runs its own migrations on startup, so the upgrade had killed the 0.90.78 pod before the new one began migrating. Helm recorded the upgrade as complete as soon as it had applied the manifests; it does not wait for pods unless asked to. The on-call had done the obvious thing first. The new chart was failing, so they ran helm rollback to the previous revision, revision 13, the last one before 0.91.2. The previous revision's pod came up, hit the database, and crashed too, for a different reason. The new binary expected schema 23 and found the version row marked dirty mid-migration. The old binary expected 21 and found the row already advanced to 23. Each binary checks the version row itself before handing off to golang-migrate, and a row newer than it knows is refused before the dirty flag is even read. One saw a migration that never finished, the other saw a database from the future. Both were sort of right.&lt;/p&gt;

&lt;p&gt;A rollback reverts the chart, not the database, so for a binary with a strict version check, like this one, the old binary refusing a newer schema is what any rollback after a forward migration looks like. The new binary refusing too is what told us the database was the problem, not the chart: if both binaries reject the same database, it is in neither of the states they expect. It is in a third state nobody coded for.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl logs &lt;span class="nt"&gt;-n&lt;/span&gt; statuspage statuspage-app-7b9f-xq2vk   &lt;span class="c"&gt;# 0.91.2 pod, pre-rollback&lt;/span&gt;
INFO  starting statuspage v0.91.2
INFO  connecting to postgres at postgres.statuspage.svc:5432
ERROR database schema version 23 is dirty, refusing to start
FATAL refusing to start with a dirty schema version

&lt;span class="nv"&gt;$ &lt;/span&gt;helm rollback statuspage 13 &lt;span class="nt"&gt;-n&lt;/span&gt; statuspage   &lt;span class="c"&gt;# 13 is the last revision before 0.91.2&lt;/span&gt;
Rollback was a success! Happy Helming!

&lt;span class="nv"&gt;$ &lt;/span&gt;helm &lt;span class="nb"&gt;history &lt;/span&gt;statuspage &lt;span class="nt"&gt;-n&lt;/span&gt; statuspage
REVISION  UPDATED                   STATUS      CHART               APP VERSION  DESCRIPTION
13        Tue May 12 09:12:04 2026  superseded  statuspage-0.90.78  0.90.78      Upgrade &lt;span class="nb"&gt;complete
&lt;/span&gt;14        Fri May 15 14:31:47 2026  superseded  statuspage-0.91.2   0.91.2       Upgrade &lt;span class="nb"&gt;complete
&lt;/span&gt;15        Fri May 15 14:48:12 2026  deployed    statuspage-0.90.78  0.90.78      Rollback to 13

&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl logs &lt;span class="nt"&gt;-n&lt;/span&gt; statuspage statuspage-app-6c4d-7m9pz   &lt;span class="c"&gt;# rolled-back 0.90.78 pod&lt;/span&gt;
INFO  starting statuspage v0.90.78
ERROR database schema version 23 found, expected 21
FATAL refusing to start with schema version mismatch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Different errors from the two chart revisions, same root cause: the database, not the chart, was the problem.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The version row claimed 23. The table DDL was still at 22.
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;What the schema_migrations table actually said&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We dropped into psql against the application database and pulled the migration tracking table. The row said version 23 with the dirty flag set, which in golang-migrate means 'a migration to this version started and never reported success'. That single boolean was the thread we pulled on for the rest of the recovery.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;statuspage=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;select&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; from schema_migrations&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="go"&gt; version | dirty
---------+-------
      23 | t
(1 row)

&lt;/span&gt;&lt;span class="gp"&gt;statuspage=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt; incidents
&lt;span class="go"&gt;                         Table "public.incidents"
   Column    |            Type             | Collation | Nullable | Default
-------------+-----------------------------+-----------+----------+---------
 id          | bigint                      |           | not null |
 service_id  | bigint                      |           | not null |
 started_at  | timestamp without time zone |           |          |
 resolved_at | timestamp without time zone |           |          |
 title       | text                        |           |          |
Indexes:
    "incidents_pkey" PRIMARY KEY, btree (id)
-- expected per migration 0023: severity column, incident_updates FK, partial index on resolved_at IS NULL
-- present: none of the above
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The version row said 23 was in progress. The table structure had none of 23.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We pulled the migration files out of the app image and read them. Migration 0023 was three statements: add a severity column, create an incident_updates table with a foreign key back, create a partial index on unresolved incidents. None of the three were present in the live schema. The OOM had hit after the migration library wrote the version row and before it sent 0023's body to the server. golang-migrate writes version 23 with dirty=true first, runs the body, then clears the flag, and on Postgres a multi-statement body sent in one Exec, the default, runs as a single transaction, so there is no stopping between statements: the body lands whole or not at all. Here it had not landed. That timing is inferred from the state, not observed: a kill after the body reached the server would more likely have left 0023 committed and still marked dirty, and the same IF NOT EXISTS repair covers that case too. None of 0023's DDL was present, and the bookkeeping said 23 was in progress.&lt;/p&gt;

&lt;p&gt;This is the specific failure mode that makes partial migrations dangerous. The migration library and the actual schema disagree, and the application trusts the library. The library trusts a row it wrote in a different transaction than the DDL it was supposed to be tracking. Whether the version row and the DDL share a transaction is a per-tool design choice, not a version threshold you can date. The version/dirty pair in the table above is golang-migrate's, and golang-migrate deliberately writes version=N with dirty=true in its own transaction, runs the migration body separately, then writes dirty=false. The dirty column exists precisely because those two are not atomic, and that has not changed. Any tool that ships a dirty flag is telling you the same thing about itself. At the other end, a tool that wraps each migration and its history row in one transaction on Postgres, Flyway among them, does not leave the state described here. Check which one you are running before you trust the version row, because that is what decides whether this failure mode can reach you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The base backup was 6 hours stale and uptime data is the product
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Why we did not restore from backup&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The instinct, and the safe move on most days, is to restore Postgres from the last known-good base backup and replay WAL up to a point just before the migration started. We checked the backup. It was a daily pg_basebackup taken at 08:30, 6 hours old, and WAL archiving had been configured but never tested for PITR. We could probably have done it. We were not willing to bet the status page on 'probably' while the status page was already down.&lt;/p&gt;

&lt;p&gt;More importantly, the uptime check history is the product. A PITR replay to just before the migration would have cost only the few minutes since the migration started, but if the untested WAL replay failed we would be down to the base backup alone, and a status page that loses 6 hours of check data after an outage is worse than a status page that takes another hour to come back. We talked it through with the team lead and decided the database in front of us was recoverable, and recovering it was lower risk than the restore path. That decision is worth naming because it goes against the usual 'just restore from backup' instinct. When the data itself is the value, finishing a half-migration by hand is often the right call.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Restore from base backup + WAL&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;PITR to a point just before the migration would lose only the few minutes since the migration started, but the PITR path was configured and never tested. If the replay failed, the fallback is the base backup alone, which loses up to 6 hours of uptime check history. Estimated 45-90 minutes if it worked the first try, much longer if not.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Finish migration 0023 by hand, fix version row&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Three DDL statements, all idempotent-ish if we wrote them with IF NOT EXISTS guards. Preserves all data. Estimated 20 minutes including verification. Chose this.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Three DDL statements, a version-row correction, and a careful restart
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Finishing the migration by hand&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We pulled migration 0023 verbatim from the app image, rewrote each statement with IF NOT EXISTS guards so a re-run could not double-apply, and ran them inside a single transaction so any failure left the database where we found it. Before touching anything we took a pg_dump of the application schema and data to a local file. That dump was our 'we can always undo this' insurance, separate from the production backup system.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- 1. snapshot first, outside any transaction&lt;/span&gt;
&lt;span class="err"&gt;$&lt;/span&gt; &lt;span class="n"&gt;kubectl&lt;/span&gt; &lt;span class="k"&gt;exec&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="n"&gt;statuspage&lt;/span&gt; &lt;span class="n"&gt;postgres&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="c1"&gt;-- \&lt;/span&gt;
    &lt;span class="n"&gt;pg_dump&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;U&lt;/span&gt; &lt;span class="n"&gt;statuspage&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;Fc&lt;/span&gt; &lt;span class="n"&gt;statuspage&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;tmp&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;statuspage&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;pre&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;repair&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dump&lt;/span&gt;

&lt;span class="c1"&gt;-- 2. complete migration 0023 inside one transaction&lt;/span&gt;
&lt;span class="k"&gt;BEGIN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;incidents&lt;/span&gt;
  &lt;span class="k"&gt;ADD&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;severity&lt;/span&gt; &lt;span class="nb"&gt;smallint&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;incident_updates&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt;           &lt;span class="n"&gt;bigserial&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;incident_id&lt;/span&gt;  &lt;span class="nb"&gt;bigint&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;incidents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;DELETE&lt;/span&gt; &lt;span class="k"&gt;CASCADE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;body&lt;/span&gt;         &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;created_at&lt;/span&gt;   &lt;span class="nb"&gt;timestamp&lt;/span&gt; &lt;span class="k"&gt;without&lt;/span&gt; &lt;span class="nb"&gt;time&lt;/span&gt; &lt;span class="k"&gt;zone&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;incidents_unresolved_idx&lt;/span&gt;
  &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;incidents&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;started_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;resolved_at&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- 3. clear the dirty flag, version row already says 23&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;schema_migrations&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;dirty&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;false&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;23&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- 4. sanity-check before commit&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dirty&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;schema_migrations&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="n"&gt;incidents&lt;/span&gt;
&lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="n"&gt;incident_updates&lt;/span&gt;

&lt;span class="k"&gt;COMMIT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;All four statements committed in 180ms: the three DDL statements plus the version-row correction. The read-only checks below them commit nothing. The dirty flag was the last thing to flip.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;After the commit we rolled forward to the chart whose binary expects 23, because the deployment sitting at 0 replicas was still the rolled-back 0.90.78 one and it would have refused a clean schema 23 exactly the way it refused a dirty one: helm upgrade statuspage statuspage/statuspage --version 0.91.2 --reset-then-reuse-values --set replicaCount=0 --set resources.limits.memory=512Mi. Every flag matters. Helm carries a release's existing values forward only when you pass no new ones, so a bare --set would have reset every custom value in the release, the database host and secrets included, to the chart's defaults. --reuse-values is not the answer across a chart version either: it carries the old release's computed values over the new chart's own defaults, so values 0.91.2 introduced can be lost, while --reset-then-reuse-values (Helm 3.14 or later) starts from 0.91.2's defaults and applies the release's values on top. The memory limit went up because the startup warm-up that killed the first 0.91.2 pod would have killed this one too. And replicaCount=0 is what kept the gate. We had scaled the deployment to 0 earlier to stop the CrashLoopBackOff noise, but kubectl scale only changes the live object, and Helm's three-way patch compares the old manifest, the live state and the new manifest, so a plain upgrade would have put the replica count straight back and started a pod before anyone looked at anything. Only then did we admit traffic, with the same pair of flags, helm upgrade statuspage statuspage/statuspage --version 0.91.2 --reset-then-reuse-values --set replicaCount=1, so the stored value is 1 and the next upgrade does not quietly scale the page back to zero. We watched the pod logs. It connected, ran its startup migration check, found 23 clean with nothing to apply, and started serving. We curled the health endpoint, got a 200, then hit /api/services and confirmed all the configured uptime checks were present with their full history intact. Total time from 'database is the problem' to 'status page is back': 51 minutes.&lt;/p&gt;

&lt;p&gt;The thing worth saying out loud about this kind of repair: it works because the migration was small and the failure was clean. If 0023 had been a long data migration, or had been split out of its single transaction, the recovery would have looked very different and the backup-restore path would have won. Always read the failed migration before you decide which recovery to attempt. We have written more about that decision in the &lt;a href="https://infraforge.agency/migrations/" rel="noopener noreferrer"&gt;migration recovery playbook&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  A schema snapshot before every migration, and a failed migration that stops the release, not the page
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The pre-upgrade hook we shipped the next day&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two changes went in within 24 hours. First, the Helm chart now has a pre-upgrade hook that writes two dumps to an object storage bucket, tagged with the chart version it is about to migrate to: pg_dump --schema-only for the structure, plus pg_dump --data-only --table=schema_migrations for the version and dirty values. Both are needed, because --schema-only emits DDL only and dumps no table rows at all, so on its own it would capture everything except the version/dirty pair, which is the single field this recovery turned on. If a future migration goes sideways, the recovery starts from a known structure-level snapshot taken seconds before the migration began, not from the daily backup. The hook adds about 4 seconds to every upgrade and has paid for itself once already.&lt;/p&gt;

&lt;p&gt;Second, migrations no longer run at app startup. They run in a pre-upgrade hook Job, weighted to run after the dump hook, with its own requests and limits. Helm waits for a Job hook, up to its --timeout, and fails the release if the hook fails, so a failed migration now stops the upgrade before the Deployment is touched, and the pods already running keep serving, as long as nothing restarts them, instead of being gone before the migration starts. Set that timeout above your longest migration. The OOM itself happened because the app's 256Mi limit was sized for 0.90.78, and 0.91.2 loads its check history into memory at startup; against production's history that peak reached about 380Mi, which is why the roll-forward raised the limit to 512Mi. The Job has no liveness probe: a migration that runs long should finish or fail on its own, not be killed halfway by a probe, because any kill after the version row is written leaves it dirty.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxdkc1q5DAQhO95irpnTF5gCSSZzf6wyywhkIMJoW23LTGyJLrbMzHMwy_SZDOw1-pS1dfqMaRj70gMz9sr4K51HGYseRIa-BVNc4v7Ngs3HxJcSvvXK-C-zh7aPL0Ny5zRNNo7nqlJMay4xkUfyM7ql05ubhvD2fg2-0nIfIq6gSV0S79nqx6jaeIBlas5sKhPEUYysZXqh1q9bT8D8DN1GxD-56xhR_aTMx5Ao7HAHKOA1Vk6RgQ_e9MNYkLwB46siiyp41K1LVUnXfqeVU_42lLOGDiHtM4cDZJC0ItxJB-QBLvd7xMeW-HApIwiKzoek3AF2F4SekdxYq04GimrSwZ-91qYhPt0YFmhRmKKUdKMfSzYamQV8bH-xrd29NGrQ7fCURywe6qRwmqltb7sCkxH_X7JuIZwDrTi5e5XifneDuTDWu5WbJ8ukt75Aw_Fd4bkTELG8DEvtqn71Jv_-fH89O9KZ6a_7YTQQg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJxdkc1q5DAQhO95irpnTF5gCSSZzf6wyywhkIMJoW23LTGyJLrbMzHMwy_SZDOw1-pS1dfqMaRj70gMz9sr4K51HGYseRIa-BVNc4v7Ngs3HxJcSvvXK-C-zh7aPL0Ny5zRNNo7nqlJMay4xkUfyM7ql05ubhvD2fg2-0nIfIq6gSV0S79nqx6jaeIBlas5sKhPEUYysZXqh1q9bT8D8DN1GxD-56xhR_aTMx5Ao7HAHKOA1Vk6RgQ_e9MNYkLwB46siiyp41K1LVUnXfqeVU_42lLOGDiHtM4cDZJC0ItxJB-QBLvd7xMeW-HApIwiKzoek3AF2F4SekdxYq04GimrSwZ-91qYhPt0YFmhRmKKUdKMfSzYamQV8bH-xrd29NGrQ7fCURywe6qRwmqltb7sCkxH_X7JuIZwDrTi5e5XifneDuTDWu5WbJ8ukt75Aw_Fd4bkTELG8DEvtqn71Jv_-fH89O9KZ6a_7YTQQg" alt="The shape of every upgrade now. The branch on the right is the one that has paid for itself since. The logical schema snapshot is a reference for what the structure should look like; only the physical base backup can take a WAL replay." width="851" height="926"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The shape of every upgrade now. The branch on the right is the one that has paid for itself since. The logical schema snapshot is a reference for what the structure should look like; only the physical base backup can take a WAL replay.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We have stopped recommending that teams skip the pre-migration schema snapshot just because their database is 'small enough to restore from the daily backup'. The two answer different questions. A base backup plus WAL tells you what the data looked like at a chosen moment, provided the replay works. The schema snapshot tells you what shape the database was in seconds before this specific migration began, which is the thing you need when a migration is the thing that broke.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recovering a partial migration without losing the data behind it
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;When a status page is the thing that is down&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The reason this kind of incident is hard is not the SQL. The SQL is usually three or four statements you can read off the migration file. The hard part is deciding whether to finish by hand or restore from backup, and that decision depends on details most teams have not catalogued: how their migration library handles the version row, whether PITR has ever been tested, whether the failed migration is structural or data-rewriting, and whether the data between the last backup and now is recoverable some other way.&lt;/p&gt;

&lt;p&gt;We run these recovery engagements regularly. We have seen the OOMKilled-migration pattern four times in the last year, two of them on status pages or monitoring tools where the data IS the product, and we have a checklist for the decision now. If you are staring at a CrashLoopBackOff with a schema version mismatch in the logs and a Helm rollback that did not help, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review&lt;/a&gt; and we will be on a bridge with you the same day to work through the finish-by-hand versus restore decision before you commit to either.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/recovering-status-page-half-finished-schema-migration/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/recovering-status-page-half-finished-schema-migration/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>migration</category>
      <category>recovery</category>
      <category>migrations</category>
    </item>
    <item>
      <title>When a validating webhook blocks the ConfigMap that would fix it</title>
      <dc:creator>Muhammad Hassaan Javed</dc:creator>
      <pubDate>Tue, 16 Jun 2026 20:19:47 +0000</pubDate>
      <link>https://dev.to/infraforge/when-a-validating-webhook-blocks-the-configmap-that-would-fix-it-4oe9</link>
      <guid>https://dev.to/infraforge/when-a-validating-webhook-blocks-the-configmap-that-would-fix-it-4oe9</guid>
      <description>&lt;p&gt;The kubectl patch came back as a webhook call failure, connection refused, not a credentials error. That was the moment the incident stopped being about a rotated MongoDB password and started being about the admission layer. A ValidatingWebhookConfiguration with failurePolicy: Fail was pointed at a webhook pod with bad probes, crash-looping and never ready, so ConfigMap writes in the namespace were failing closed. The webhook pod itself was fixable: its probes live in a Deployment spec, and the webhook intercepted ConfigMap writes only. What had nowhere to go was the credential fix. The safety mechanism had become the outage. Our profile service was down because its database credentials were stale, the fix for the credentials was a one-line patch, and that one-line patch could not be applied because the thing that was supposed to keep ConfigMaps safe was rejecting writes across the namespace.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem signals:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;kubectl patch or apply on a ConfigMap returns a failed calling webhook error (connection refused, no endpoints available, or a timeout) instead of a normal validation error&lt;/li&gt;
&lt;li&gt;An admission webhook pod is CrashLoopBackOff while its ValidatingWebhookConfiguration is set to failurePolicy: Fail&lt;/li&gt;
&lt;li&gt;ArgoCD shows sync pending or OutOfSync on resources in the affected namespace and the sync will not progress&lt;/li&gt;
&lt;li&gt;A workload reads stale config from a ConfigMap that was supposedly already updated, because patching a ConfigMap never restarts pods on its own and a Deployment-level env var is shadowing the value injected from that ConfigMap&lt;/li&gt;
&lt;li&gt;Compliance requires that admission webhooks remain failurePolicy: Fail in production, so flipping to Ignore as a workaround is itself an audit event&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  We thought it was a credentials incident for the first 20 minutes
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The patch that came back as a webhook call failure&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The page that started the call was a profile service returning 500s on /health. The cause looked obvious. The data layer team had rotated the MongoDB credentials the day before, and the live ConfigMap in the application namespace still held the old password. There was a backup ConfigMap sitting next to it with the rotated values, labelled exactly the way the runbook described. Neither is in Git; the data layer team manages both in the cluster. The fix was supposed to be a thirty second kubectl patch.&lt;/p&gt;

&lt;p&gt;It was not. The patch came back with this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ kubectl -n app patch configmap profile-mongodb-config \
    --type merge --patch-file rotated.yaml
Error from server (InternalError): Internal error occurred:
failed calling webhook "configmap-validator.app.svc": failed to call webhook:
Post "https://configmap-validator.app.svc:443/validate?timeout=10s":
dial tcp 10.96.142.18:443: connect: connection refused
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The actual error. Not a credentials problem, an admission control problem. The address is the Service ClusterIP, which the API server dials unless it runs with --enable-aggregator-routing; with no ready Pod behind it, kube-proxy rejects the connection.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That error string is the whole story. The cluster had a ValidatingWebhookConfiguration named configmap-validator that intercepted every ConfigMap write in the namespace. The webhook pod was supposed to enforce a schema policy that the compliance team owned. Right now the webhook pod was not answering on its service IP, which meant ConfigMap writes were failing closed, which meant our credential fix was failing closed, which meant the profile service stayed down.&lt;/p&gt;

&lt;p&gt;We had walked into this kind of shape before, but usually on the cert-manager side. This time the trap was tighter: the webhook was supposed to validate the very ConfigMaps that controlled the workloads in its own namespace, and one of those workloads happened to be down for an unrelated reason. Two independent failures had stacked: the stale credentials caused the outage, and the webhook turned a thirty-second fix into an incident. And the stale credentials the app was actually using were not the ones we were fixing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the pod was crash-looping, and why GitOps rather than the webhook gated the fix
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The probe lived in the Deployment, not in a ConfigMap&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;kubectl describe on the configmap-validator pod told us the liveness probe was failing. The kubelet killed the container shortly after every start, and by the time we looked the restart back-off had reached its five-minute cap. The readiness probe had been copied from the same stale spec, so the pod never passed one and the webhook Service never had a ready endpoint: there was no window in which a retried patch could have reached it. Both probes were hitting /healthz on port 8443. The actual application served its health endpoint on /health, no z. Someone had copy-pasted a probe spec from an older service months ago and nobody had noticed because the webhook had been running fine until a recent image bump shifted the health route.&lt;/p&gt;

&lt;p&gt;Fixing a Deployment in Kubernetes is normally a kubectl edit deploy or a kubectl patch on the probe spec and you are done. The webhook configuration intercepted ConfigMap writes, not Deployment writes, so the crash-looping pod was never able to block the repair of its own Deployment. What blocked us was policy, not admission control. Our platform team had a hard rule against in-cluster edits that drifted from the GitOps source, and ArgoCD would self-heal the Deployment back to the broken probe spec inside 90 seconds. So any change to the probe spec had to land as a commit in the GitOps repo and be synced, or else start by suspending auto-sync on the app. The credential ConfigMap was the write with no way around it: that one the webhook really was rejecting.&lt;/p&gt;

&lt;p&gt;We needed the correct health path, and the compliance team kept the canonical values in a separate namespace. Their ConfigMap held the approved liveness path, the approved annotation policy, and the compliance acknowledgement token that any incident response was required to reference. We pulled it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;$ kubectl -n compliance get configmap webhook-standards -o yaml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ConfigMap&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;webhook-standards&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;compliance&lt;/span&gt;
&lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;liveness-path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/health"&lt;/span&gt;
  &lt;span class="na"&gt;readiness-path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/ready"&lt;/span&gt;
  &lt;span class="na"&gt;failure-policy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Fail"&lt;/span&gt;
  &lt;span class="na"&gt;ack-token&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMP-ACK-7c3f9a-2024Q4"&lt;/span&gt;
  &lt;span class="na"&gt;required-annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;incident.compliance/id&lt;/span&gt;
    &lt;span class="s"&gt;incident.compliance/services&lt;/span&gt;
    &lt;span class="s"&gt;incident.compliance/ack-token&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The compliance source of truth. We read these values; we did not retype them.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Flip failurePolicy to Ignore, or delete the ValidatingWebhookConfiguration?
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Choosing what the audit trail shows&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There were two ways to unblock ConfigMap writes. We could delete the ValidatingWebhookConfiguration entirely, fix everything, and recreate it from the GitOps source. Or we could patch failurePolicy from Fail to Ignore for the length of the recovery and patch it back when we were done.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Option A. Delete the webhook configuration&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cleanest cut. ConfigMap writes unblock instantly. Risk: an unrelated team applies an out-of-policy ConfigMap during the window and we do not catch it. Also generates a louder audit event because the object disappears from etcd. Like the flip, it needs the platform-admission sync paused first, or Argo CD recreates the object.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Option B. Patch failurePolicy to Ignore&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Webhook is still called; if the pod is up it still validates; if it is down, writes pass unvalidated, as they would with no webhook at all, except that each one still tries the webhook first and the API server records it as failed open. The difference is what is left behind: one field to flip back, and an audit log that shows a field change, not a delete. We picked this one.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Option B won because of the audit trail. The compliance team would rather see one field flip and one field flip back, with the same controller object identity across the incident, than see a delete and a recreate with a new resourceVersion lineage. That is the kind of preference you only learn by sitting through an audit. We have written more about this kind of constraint in &lt;a href="https://infraforge.agency/infrastructure-audit-readiness/" rel="noopener noreferrer"&gt;our infrastructure audit readiness work&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Step 0. The webhook configuration is in Git too, in its own Argo CD app
# (platform-admission, apart from the webhook Deployment), under self-heal.
# Save its sync policy once, then pause it, or Fail comes back by itself.
# (If an ApplicationSet or a self-healing parent app manages this Application,
# it can undo the pause: give the ApplicationSet an ignore rule for
# spec.syncPolicy.automated, or pause the parent, first, and undo that in Step 4.)
f=/tmp/platform-admission-syncpolicy.json
[ -s "$f" ] || kubectl -n argocd get app platform-admission \
  -o jsonpath='{.spec.syncPolicy}' &amp;gt; "$f"
grep -q '"automated"' "$f" || echo "saved policy has no automated block; check before pausing"
argocd app set platform-admission --sync-policy none

# Step 1. Snapshot the current webhook config so we can prove what we changed.
kubectl get validatingwebhookconfiguration configmap-validator \
  -o yaml &amp;gt; /tmp/vwc-before.yaml

# Step 2. Flip failurePolicy to Ignore, scoped to this single webhook entry.
# webhooks/0 is its position in this configuration; confirm it in the snapshot.
kubectl patch validatingwebhookconfiguration configmap-validator \
  --type='json' \
  -p='[{"op":"replace","path":"/webhooks/0/failurePolicy","value":"Ignore"}]'

# Step 3. Confirm the change before touching any ConfigMap.
kubectl get validatingwebhookconfiguration configmap-validator \
  -o jsonpath='{.webhooks[0].failurePolicy}'
# expect: Ignore

# Step 4, only once the webhook answers again: flip back, restore the saved
# sync policy exactly as it was, and undo anything you changed upstream.
kubectl patch validatingwebhookconfiguration configmap-validator \
  --type='json' \
  -p='[{"op":"replace","path":"/webhooks/0/failurePolicy","value":"Fail"}]'
argocd app patch platform-admission --type merge \
  --patch "{\"spec\":{\"syncPolicy\":$(cat /tmp/platform-admission-syncpolicy.json)}}"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The unblock and its undo. Save and pause the sync that would undo it, snapshot for the post-incident review, flip, confirm. Step 4 waits until the webhook is healthy again.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  We patched the ConfigMap and the service was still broken
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The env var that made the credential fix invisible&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;With failurePolicy on Ignore, the credential patch went through. We pulled the rotated values from the backup ConfigMap and applied them to the live one, and nothing moved. Patching a ConfigMap does not restart pods; there is no controller that does that. The value here is injected with env[].valueFrom.configMapKeyRef, which is resolved when the container starts and is never refreshed inside a running container. (Only a ConfigMap consumed through volumeMounts updates in place, and even then only the file on disk changes, not the process environment; a subPath mount never updates at all.) So the rollout has to be forced explicitly: kubectl -n app rollout restart deploy/profile-service, wait for the new pods, then re-check the environment on one of them. We did that. /health still returned 500. The MongoDB connection error in the application logs still showed the old username.&lt;/p&gt;

&lt;p&gt;That was the second moment in the incident where the model of the world had to change. The ConfigMap held the new credentials, and after the restart so did MONGODB_URI inside the pod. The application simply was not using it. Something else in the same environment was winning.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ kubectl -n app get deploy profile-service \
    -o jsonpath='{.spec.template.spec.containers[0].env}' | jq
[
  { "name": "MONGODB_URI",
    "valueFrom": {
      "configMapKeyRef": {
        "name": "profile-mongodb-config",
        "key": "uri"
      }
    }
  },
  { "name": "PROFILE_MONGODB_URI_OVERRIDE",
    "value": "mongodb://oldapp:oldpw@mongo.app.svc:27017/profiles"
  }
]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The override. Set during a migration test six weeks earlier and never removed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The application code read PROFILE_MONGODB_URI_OVERRIDE if it was set and otherwise read MONGODB_URI. The override had been added during a migration drill six weeks ago, never cleaned up, and was now silently shadowing every ConfigMap update we tried to apply. We have stopped accepting break-glass env overrides on production Deployments for this exact reason. If the override is worth setting, it is worth its own object with an expiry annotation that a controller cleans up. Naked env values on the Deployment spec are invisible to the operators who do not know to look for them. And credentials never belonged in a ConfigMap in the first place: a ConfigMap does not provide secrecy, so the rotated credentials moved into a Secret the same week.&lt;/p&gt;

&lt;p&gt;That env var lives in .spec.template.spec of the profile-service Deployment, which is under ArgoCD self-heal, so deleting it with kubectl would have come straight back inside 90 seconds. We removed it in the GitOps repo instead: a one-line commit dropping PROFILE_MONGODB_URI_OVERRIDE, then an argocd app sync. (The alternative, when you cannot get a commit through mid-incident, is to suspend auto-sync first with argocd app set profile-service --sync-policy none, or kubectl -n argocd patch app profile-service --type merge -p with syncPolicy.automated set to null, then patch in cluster and re-enable the policy at the end. What does not work is patching under an active self-heal.) The pod rolled, and /health came back as 200 on the third pod we curled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting the webhook back, and making sure this never happens the same way again
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Restoring failurePolicy: Fail without re-creating the trap&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Before we restored failurePolicy to Fail, we fixed the webhook pod. The probe path sits in the configmap-validator Deployment spec, under the same self-heal rule, so the fix went through the GitOps repo the same way: a commit moving the probes off /healthz to the approved paths, /health for liveness and /ready for readiness, the values we had read from the compliance ConfigMap and confirmed the image actually serves, then a sync. An in-cluster kubectl patch would have been reverted back to /healthz inside 90 seconds. The pod came up healthy and stayed up. We confirmed the webhook was actually answering by sending a deliberately invalid ConfigMap with kubectl apply --dry-run=server and watching the validation rejection come back cleanly. The dry run matters: it goes through admission without saving anything, and with failurePolicy still on Ignore, a real write would have been saved if the webhook had not answered. Only then did we run Step 4: failurePolicy back to Fail, and platform-admission's saved sync policy restored through Argo CD, so its automated settings returned exactly as they had been and Git and the cluster agreed again.&lt;/p&gt;

&lt;p&gt;The harder problem was structural. A failurePolicy: Fail webhook that gates ConfigMap writes in a namespace is fine. This time the webhook's crash cause happened to live in its Deployment, which the webhook does not intercept, so there was a way out through the GitOps repo. Had the bad value lived in a ConfigMap the webhook itself validates, its own startup arguments for instance, there would have been no clean way out: the webhook would have been gating the write that repairs it, and the remaining exits are flipping its policy, deleting or narrowing its configuration, or rewriting its Deployment, through Git, to stop reading that ConfigMap. That near-miss is the thing worth designing against, not the outage we actually had.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJw90EFqwzAQBdB9TvEPEF-hkNhJVqWFFroQWcjSOBaVZ4RmSOrbF5m063mf_5kpyyPMvho-hx1wcF80ziLfKBIRqte5yyJF92C6U0UlH9c9hDH6iFJlJL2i615wdB9U7ykQiGORxIabkIKWYut1Bxw31rteeEq3V1_wqMlIMfmUEbIoxeb6zZ1c75nFULyFuTVNKRNCpUhsyWdt9rTZs3ueO31OUPOrIsqDmzpsanDvbS60UEBOd1Ikhs2EgUqWdSG2PVqlx__GFh-2-MWd0w9FjCsuyd6KIsiyJIPnCF05_L0osVENVIzi9RcjtnYt" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fkroki.io%2Fmermaid%2Fpng%2FeJw90EFqwzAQBdB9TvEPEF-hkNhJVqWFFroQWcjSOBaVZ4RmSOrbF5m063mf_5kpyyPMvho-hx1wcF80ziLfKBIRqte5yyJF92C6U0UlH9c9hDH6iFJlJL2i615wdB9U7ykQiGORxIabkIKWYut1Bxw31rteeEq3V1_wqMlIMfmUEbIoxeb6zZ1c75nFULyFuTVNKRNCpUhsyWdt9rTZs3ueO31OUPOrIsqDmzpsanDvbS60UEBOd1Ikhs2EgUqWdSG2PVqlx__GFh-2-MWd0w9FjCsuyd6KIsiyJIPnCF05_L0osVENVIzi9RcjtnYt" alt="The block is one-way, not a cycle. The credential fix routes through a ConfigMap write the webhook is rejecting; the webhook's own probe fix routes through Git." width="586" height="630"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The block is one-way, not a cycle. The credential fix routes through a ConfigMap write the webhook is rejecting; the webhook's own probe fix routes through Git.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We made two changes before we left. First, we moved the webhook's own Deployment, Service, and startup ConfigMap out of the app namespace and into a dedicated webhooks namespace, then set a namespaceSelector that excludes webhooks from validation, so the webhook can be rebuilt from its own in-cluster ConfigMaps even when it is the thing that is broken. Every ConfigMap in app is still validated, which is the whole point of the policy. Second, we gave on-call a break-glass path that does not depend on the webhook being up. The obvious tool, an objectSelector that skips ConfigMaps carrying a break-glass label, is the wrong one. Kubernetes calls the webhook if either the old or the new object matches the selector, so the label would have to be on the ConfigMap before the incident, and a label that is always there exempts that ConfigMap from validation for every writer, every day. The Kubernetes documentation says as much: use the object selector only for opt-in webhooks, because end users can skip the webhook by setting the labels. Instead the webhook entry got a matchConditions expression, !("platform:break-glass" in request.userInfo.groups), which the API server evaluates itself before deciding whether to call the webhook. A request from that group skips the webhook even while the webhook pod is down, and every other request is validated exactly as before. Only the on-call rotation can act as that group. If the policy can be written in CEL, a ValidatingAdmissionPolicy removes the dependency on a running pod altogether, because the API server evaluates it in-process; porting this one is on the list. Both changes were reviewed by the compliance team before we merged them, because relaxing the scope of a Fail-policy webhook is itself an audit decision.&lt;/p&gt;

&lt;p&gt;The recovery script we left behind reads every value it needs (the health path, the ack token, the affected service list) from cluster state rather than hardcoding. Hardcoded recovery scripts go stale within a quarter; scripts that read from a compliance-owned ConfigMap stay correct as long as the source of truth is maintained. The script is idempotent: rerunning it on an already-recovered cluster is a no-op, which matters because the on-call engineer who runs it at 3 am should not have to think about whether they are the first or the third person to run it that night.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to call us, and what we will look at first
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;If a Fail-policy webhook can stand between an outage and its fix&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The thing that makes this incident shape hard is not the webhook itself. It is that the recovery path is non-obvious, the audit consequences of the obvious workaround (flipping policy or deleting the webhook config) are real, and the second-order trap (an env var on a Deployment shadowing the ConfigMap you just fixed) only shows up after you have already burned the credibility from the first workaround. Teams who hit this for the first time usually solve the immediate outage but leave the structural hazard in place, and then it happens again on a different webhook six months later.&lt;/p&gt;

&lt;p&gt;We run these recovery engagements every week. A Fail-policy webhook standing between an outage and its fix has come up four times this year for us, once with cert-manager involved, twice with policy webhooks like this one, once with a service mesh sidecar injector that depended on a ConfigMap in its own namespace. The env-override-shadowing-a-ConfigMap-fix pattern is even more common; we see some version of it in roughly half of the credential rotation incidents we are called into.&lt;/p&gt;

&lt;p&gt;If your cluster has a Fail-policy admission webhook today and you have never tested what happens when its pod is down, &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;book an infrastructure review with our team&lt;/a&gt; and we will start with a 30-minute diagnostic call this week. We will walk your webhook configurations, identify the ones that could block a fix or their own repair, and give you a concrete plan for a way out that does not break your audit policy, before an incident finds it for you.&lt;/p&gt;




&lt;p&gt;Originally published at &lt;a href="https://infraforge.agency/insights/admission-webhook-configmap-deadlock-recovery/" rel="noopener noreferrer"&gt;https://infraforge.agency/insights/admission-webhook-configmap-deadlock-recovery/&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — &lt;a href="https://infraforge.agency/review/" rel="noopener noreferrer"&gt;see /review&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>audit</category>
      <category>readiness</category>
      <category>auditreadiness</category>
    </item>
  </channel>
</rss>
