The kubectl patch came back as a webhook call failure, connection refused, not a credentials error. That was the moment the incident stopped being about a rotated MongoDB password and started being about the admission layer. A ValidatingWebhookConfiguration with failurePolicy: Fail was pointed at a webhook pod with bad probes, crash-looping and never ready, so ConfigMap writes in the namespace were failing closed. The webhook pod itself was fixable: its probes live in a Deployment spec, and the webhook intercepted ConfigMap writes only. What had nowhere to go was the credential fix. The safety mechanism had become the outage. Our profile service was down because its database credentials were stale, the fix for the credentials was a one-line patch, and that one-line patch could not be applied because the thing that was supposed to keep ConfigMaps safe was rejecting writes across the namespace.
Problem signals:
- kubectl patch or apply on a ConfigMap returns a failed calling webhook error (connection refused, no endpoints available, or a timeout) instead of a normal validation error
- An admission webhook pod is CrashLoopBackOff while its ValidatingWebhookConfiguration is set to failurePolicy: Fail
- ArgoCD shows sync pending or OutOfSync on resources in the affected namespace and the sync will not progress
- A workload reads stale config from a ConfigMap that was supposedly already updated, because patching a ConfigMap never restarts pods on its own and a Deployment-level env var is shadowing the value injected from that ConfigMap
- Compliance requires that admission webhooks remain failurePolicy: Fail in production, so flipping to Ignore as a workaround is itself an audit event
We thought it was a credentials incident for the first 20 minutes
The patch that came back as a webhook call failure
The page that started the call was a profile service returning 500s on /health. The cause looked obvious. The data layer team had rotated the MongoDB credentials the day before, and the live ConfigMap in the application namespace still held the old password. There was a backup ConfigMap sitting next to it with the rotated values, labelled exactly the way the runbook described. Neither is in Git; the data layer team manages both in the cluster. The fix was supposed to be a thirty second kubectl patch.
It was not. The patch came back with this:
$ kubectl -n app patch configmap profile-mongodb-config \
--type merge --patch-file rotated.yaml
Error from server (InternalError): Internal error occurred:
failed calling webhook "configmap-validator.app.svc": failed to call webhook:
Post "https://configmap-validator.app.svc:443/validate?timeout=10s":
dial tcp 10.96.142.18:443: connect: connection refused
The actual error. Not a credentials problem, an admission control problem. The address is the Service ClusterIP, which the API server dials unless it runs with --enable-aggregator-routing; with no ready Pod behind it, kube-proxy rejects the connection.
That error string is the whole story. The cluster had a ValidatingWebhookConfiguration named configmap-validator that intercepted every ConfigMap write in the namespace. The webhook pod was supposed to enforce a schema policy that the compliance team owned. Right now the webhook pod was not answering on its service IP, which meant ConfigMap writes were failing closed, which meant our credential fix was failing closed, which meant the profile service stayed down.
We had walked into this kind of shape before, but usually on the cert-manager side. This time the trap was tighter: the webhook was supposed to validate the very ConfigMaps that controlled the workloads in its own namespace, and one of those workloads happened to be down for an unrelated reason. Two independent failures had stacked: the stale credentials caused the outage, and the webhook turned a thirty-second fix into an incident. And the stale credentials the app was actually using were not the ones we were fixing.
Why the pod was crash-looping, and why GitOps rather than the webhook gated the fix
The probe lived in the Deployment, not in a ConfigMap
kubectl describe on the configmap-validator pod told us the liveness probe was failing. The kubelet killed the container shortly after every start, and by the time we looked the restart back-off had reached its five-minute cap. The readiness probe had been copied from the same stale spec, so the pod never passed one and the webhook Service never had a ready endpoint: there was no window in which a retried patch could have reached it. Both probes were hitting /healthz on port 8443. The actual application served its health endpoint on /health, no z. Someone had copy-pasted a probe spec from an older service months ago and nobody had noticed because the webhook had been running fine until a recent image bump shifted the health route.
Fixing a Deployment in Kubernetes is normally a kubectl edit deploy or a kubectl patch on the probe spec and you are done. The webhook configuration intercepted ConfigMap writes, not Deployment writes, so the crash-looping pod was never able to block the repair of its own Deployment. What blocked us was policy, not admission control. Our platform team had a hard rule against in-cluster edits that drifted from the GitOps source, and ArgoCD would self-heal the Deployment back to the broken probe spec inside 90 seconds. So any change to the probe spec had to land as a commit in the GitOps repo and be synced, or else start by suspending auto-sync on the app. The credential ConfigMap was the write with no way around it: that one the webhook really was rejecting.
We needed the correct health path, and the compliance team kept the canonical values in a separate namespace. Their ConfigMap held the approved liveness path, the approved annotation policy, and the compliance acknowledgement token that any incident response was required to reference. We pulled it:
$ kubectl -n compliance get configmap webhook-standards -o yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: webhook-standards
namespace: compliance
data:
liveness-path: "/health"
readiness-path: "/ready"
failure-policy: "Fail"
ack-token: "COMP-ACK-7c3f9a-2024Q4"
required-annotations: |
incident.compliance/id
incident.compliance/services
incident.compliance/ack-token
The compliance source of truth. We read these values; we did not retype them.
Flip failurePolicy to Ignore, or delete the ValidatingWebhookConfiguration?
Choosing what the audit trail shows
There were two ways to unblock ConfigMap writes. We could delete the ValidatingWebhookConfiguration entirely, fix everything, and recreate it from the GitOps source. Or we could patch failurePolicy from Fail to Ignore for the length of the recovery and patch it back when we were done.
| Step | What it does |
|---|---|
| Option A. Delete the webhook configuration | Cleanest cut. ConfigMap writes unblock instantly. Risk: an unrelated team applies an out-of-policy ConfigMap during the window and we do not catch it. Also generates a louder audit event because the object disappears from etcd. Like the flip, it needs the platform-admission sync paused first, or Argo CD recreates the object. |
| Option B. Patch failurePolicy to Ignore | Webhook is still called; if the pod is up it still validates; if it is down, writes pass unvalidated, as they would with no webhook at all, except that each one still tries the webhook first and the API server records it as failed open. The difference is what is left behind: one field to flip back, and an audit log that shows a field change, not a delete. We picked this one. |
Option B won because of the audit trail. The compliance team would rather see one field flip and one field flip back, with the same controller object identity across the incident, than see a delete and a recreate with a new resourceVersion lineage. That is the kind of preference you only learn by sitting through an audit. We have written more about this kind of constraint in our infrastructure audit readiness work.
# Step 0. The webhook configuration is in Git too, in its own Argo CD app
# (platform-admission, apart from the webhook Deployment), under self-heal.
# Save its sync policy once, then pause it, or Fail comes back by itself.
# (If an ApplicationSet or a self-healing parent app manages this Application,
# it can undo the pause: give the ApplicationSet an ignore rule for
# spec.syncPolicy.automated, or pause the parent, first, and undo that in Step 4.)
f=/tmp/platform-admission-syncpolicy.json
[ -s "$f" ] || kubectl -n argocd get app platform-admission \
-o jsonpath='{.spec.syncPolicy}' > "$f"
grep -q '"automated"' "$f" || echo "saved policy has no automated block; check before pausing"
argocd app set platform-admission --sync-policy none
# Step 1. Snapshot the current webhook config so we can prove what we changed.
kubectl get validatingwebhookconfiguration configmap-validator \
-o yaml > /tmp/vwc-before.yaml
# Step 2. Flip failurePolicy to Ignore, scoped to this single webhook entry.
# webhooks/0 is its position in this configuration; confirm it in the snapshot.
kubectl patch validatingwebhookconfiguration configmap-validator \
--type='json' \
-p='[{"op":"replace","path":"/webhooks/0/failurePolicy","value":"Ignore"}]'
# Step 3. Confirm the change before touching any ConfigMap.
kubectl get validatingwebhookconfiguration configmap-validator \
-o jsonpath='{.webhooks[0].failurePolicy}'
# expect: Ignore
# Step 4, only once the webhook answers again: flip back, restore the saved
# sync policy exactly as it was, and undo anything you changed upstream.
kubectl patch validatingwebhookconfiguration configmap-validator \
--type='json' \
-p='[{"op":"replace","path":"/webhooks/0/failurePolicy","value":"Fail"}]'
argocd app patch platform-admission --type merge \
--patch "{\"spec\":{\"syncPolicy\":$(cat /tmp/platform-admission-syncpolicy.json)}}"
The unblock and its undo. Save and pause the sync that would undo it, snapshot for the post-incident review, flip, confirm. Step 4 waits until the webhook is healthy again.
We patched the ConfigMap and the service was still broken
The env var that made the credential fix invisible
With failurePolicy on Ignore, the credential patch went through. We pulled the rotated values from the backup ConfigMap and applied them to the live one, and nothing moved. Patching a ConfigMap does not restart pods; there is no controller that does that. The value here is injected with env[].valueFrom.configMapKeyRef, which is resolved when the container starts and is never refreshed inside a running container. (Only a ConfigMap consumed through volumeMounts updates in place, and even then only the file on disk changes, not the process environment; a subPath mount never updates at all.) So the rollout has to be forced explicitly: kubectl -n app rollout restart deploy/profile-service, wait for the new pods, then re-check the environment on one of them. We did that. /health still returned 500. The MongoDB connection error in the application logs still showed the old username.
That was the second moment in the incident where the model of the world had to change. The ConfigMap held the new credentials, and after the restart so did MONGODB_URI inside the pod. The application simply was not using it. Something else in the same environment was winning.
$ kubectl -n app get deploy profile-service \
-o jsonpath='{.spec.template.spec.containers[0].env}' | jq
[
{ "name": "MONGODB_URI",
"valueFrom": {
"configMapKeyRef": {
"name": "profile-mongodb-config",
"key": "uri"
}
}
},
{ "name": "PROFILE_MONGODB_URI_OVERRIDE",
"value": "mongodb://oldapp:oldpw@mongo.app.svc:27017/profiles"
}
]
The override. Set during a migration test six weeks earlier and never removed.
The application code read PROFILE_MONGODB_URI_OVERRIDE if it was set and otherwise read MONGODB_URI. The override had been added during a migration drill six weeks ago, never cleaned up, and was now silently shadowing every ConfigMap update we tried to apply. We have stopped accepting break-glass env overrides on production Deployments for this exact reason. If the override is worth setting, it is worth its own object with an expiry annotation that a controller cleans up. Naked env values on the Deployment spec are invisible to the operators who do not know to look for them. And credentials never belonged in a ConfigMap in the first place: a ConfigMap does not provide secrecy, so the rotated credentials moved into a Secret the same week.
That env var lives in .spec.template.spec of the profile-service Deployment, which is under ArgoCD self-heal, so deleting it with kubectl would have come straight back inside 90 seconds. We removed it in the GitOps repo instead: a one-line commit dropping PROFILE_MONGODB_URI_OVERRIDE, then an argocd app sync. (The alternative, when you cannot get a commit through mid-incident, is to suspend auto-sync first with argocd app set profile-service --sync-policy none, or kubectl -n argocd patch app profile-service --type merge -p with syncPolicy.automated set to null, then patch in cluster and re-enable the policy at the end. What does not work is patching under an active self-heal.) The pod rolled, and /health came back as 200 on the third pod we curled.
Putting the webhook back, and making sure this never happens the same way again
Restoring failurePolicy: Fail without re-creating the trap
Before we restored failurePolicy to Fail, we fixed the webhook pod. The probe path sits in the configmap-validator Deployment spec, under the same self-heal rule, so the fix went through the GitOps repo the same way: a commit moving the probes off /healthz to the approved paths, /health for liveness and /ready for readiness, the values we had read from the compliance ConfigMap and confirmed the image actually serves, then a sync. An in-cluster kubectl patch would have been reverted back to /healthz inside 90 seconds. The pod came up healthy and stayed up. We confirmed the webhook was actually answering by sending a deliberately invalid ConfigMap with kubectl apply --dry-run=server and watching the validation rejection come back cleanly. The dry run matters: it goes through admission without saving anything, and with failurePolicy still on Ignore, a real write would have been saved if the webhook had not answered. Only then did we run Step 4: failurePolicy back to Fail, and platform-admission's saved sync policy restored through Argo CD, so its automated settings returned exactly as they had been and Git and the cluster agreed again.
The harder problem was structural. A failurePolicy: Fail webhook that gates ConfigMap writes in a namespace is fine. This time the webhook's crash cause happened to live in its Deployment, which the webhook does not intercept, so there was a way out through the GitOps repo. Had the bad value lived in a ConfigMap the webhook itself validates, its own startup arguments for instance, there would have been no clean way out: the webhook would have been gating the write that repairs it, and the remaining exits are flipping its policy, deleting or narrowing its configuration, or rewriting its Deployment, through Git, to stop reading that ConfigMap. That near-miss is the thing worth designing against, not the outage we actually had.
The block is one-way, not a cycle. The credential fix routes through a ConfigMap write the webhook is rejecting; the webhook's own probe fix routes through Git.
We made two changes before we left. First, we moved the webhook's own Deployment, Service, and startup ConfigMap out of the app namespace and into a dedicated webhooks namespace, then set a namespaceSelector that excludes webhooks from validation, so the webhook can be rebuilt from its own in-cluster ConfigMaps even when it is the thing that is broken. Every ConfigMap in app is still validated, which is the whole point of the policy. Second, we gave on-call a break-glass path that does not depend on the webhook being up. The obvious tool, an objectSelector that skips ConfigMaps carrying a break-glass label, is the wrong one. Kubernetes calls the webhook if either the old or the new object matches the selector, so the label would have to be on the ConfigMap before the incident, and a label that is always there exempts that ConfigMap from validation for every writer, every day. The Kubernetes documentation says as much: use the object selector only for opt-in webhooks, because end users can skip the webhook by setting the labels. Instead the webhook entry got a matchConditions expression, !("platform:break-glass" in request.userInfo.groups), which the API server evaluates itself before deciding whether to call the webhook. A request from that group skips the webhook even while the webhook pod is down, and every other request is validated exactly as before. Only the on-call rotation can act as that group. If the policy can be written in CEL, a ValidatingAdmissionPolicy removes the dependency on a running pod altogether, because the API server evaluates it in-process; porting this one is on the list. Both changes were reviewed by the compliance team before we merged them, because relaxing the scope of a Fail-policy webhook is itself an audit decision.
The recovery script we left behind reads every value it needs (the health path, the ack token, the affected service list) from cluster state rather than hardcoding. Hardcoded recovery scripts go stale within a quarter; scripts that read from a compliance-owned ConfigMap stay correct as long as the source of truth is maintained. The script is idempotent: rerunning it on an already-recovered cluster is a no-op, which matters because the on-call engineer who runs it at 3 am should not have to think about whether they are the first or the third person to run it that night.
When to call us, and what we will look at first
If a Fail-policy webhook can stand between an outage and its fix
The thing that makes this incident shape hard is not the webhook itself. It is that the recovery path is non-obvious, the audit consequences of the obvious workaround (flipping policy or deleting the webhook config) are real, and the second-order trap (an env var on a Deployment shadowing the ConfigMap you just fixed) only shows up after you have already burned the credibility from the first workaround. Teams who hit this for the first time usually solve the immediate outage but leave the structural hazard in place, and then it happens again on a different webhook six months later.
We run these recovery engagements every week. A Fail-policy webhook standing between an outage and its fix has come up four times this year for us, once with cert-manager involved, twice with policy webhooks like this one, once with a service mesh sidecar injector that depended on a ConfigMap in its own namespace. The env-override-shadowing-a-ConfigMap-fix pattern is even more common; we see some version of it in roughly half of the credential rotation incidents we are called into.
If your cluster has a Fail-policy admission webhook today and you have never tested what happens when its pod is down, book an infrastructure review with our team and we will start with a 30-minute diagnostic call this week. We will walk your webhook configurations, identify the ones that could block a fix or their own repair, and give you a concrete plan for a way out that does not break your audit policy, before an incident finds it for you.
Originally published at https://infraforge.agency/insights/admission-webhook-configmap-deadlock-recovery/.
If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — see /review.
Top comments (0)