The first thing the on-call team tried was patching the status ConfigMap. Five apps showed Progressing in the platform's bleater-status object, the dashboard had been red for four hours, and somebody figured a kubectl patch on the status keys would at least quiet the pages while they investigated. The patch lasted about ten seconds. The internal reconciler running in the bleater-system namespace rewrote the ConfigMap on its next tick, every key back to Progressing, and the pages started again. That was the moment they called us. The hardening rollout that had run the night before had not broken one thing. It had broken eight.
Problem signals:
- Apps stuck in Progressing or Degraded for hours after a hardening or security pass, with no single obvious cause in the events stream
- Status ConfigMaps written by an in-cluster reconciler get rewritten within 10 to 15 seconds of any kubectl patch
- kubectl delete job hangs on suspended PreSync Jobs because a hook-cleanup finalizer is still attached
- Init containers crash-loop with pg_isready printing 'no response' while a default-deny egress policy blocks DNS, then failing at a schema verification step once the egress allow rules land
- kubectl patch on a RoleBinding fails with 'cannot change roleRef' ('field is immutable' on 1.37 and later API servers) because roleRef is immutable, so the binding has to be recreated
Why editing the status ConfigMap was the wrong instinct
The patch that survived ten seconds
The team had built a small in-house control plane the year before. A Python reconciler Pod in bleater-system watched the managed workloads, computed health from live cluster signals, and wrote a bleater-status ConfigMap every ten to fifteen seconds. Five apps reported there: an auth service, a profile service, a timeline service, a fanout service, and a primary application that handled the user-facing API. None of them used a full GitOps platform. The reconciler followed GitOps conventions, PreSync hook Jobs, sync windows, hook-cleanup finalizers, but it operated on raw Kubernetes primitives. ConfigMaps, Jobs, Roles. No CRDs.
That detail matters because when the on-call lead patched the status ConfigMap to mark the apps healthy, the reconciler was doing its job. It read the live cluster, saw the upstream signals were still bad, and rewrote the status. The patch was not wrong because patching ConfigMaps is wrong. It was wrong because the bleater-status object was not an input to the system. It was an output. Editing an output to fix a system is the same shape of mistake as editing a Prometheus metric to fix a service.
We have seen this pattern enough times to write it down as a rule. If a controller is rewriting your patches in under a minute, the object you are patching is derived state. Find the inputs. The reconciler source was eighty lines of Python and it took two minutes to read. The health predicate was an AND-chain across eight signals: lock state, orphan hook Job (the finalizer), schema version, RBAC capability, PVC bound, ResourceQuota headroom, NetworkPolicy egress, and init-container health. Any one of the first seven returning bad meant Progressing; a failing init container returned Degraded instead. All eight were bad.
# the AND-chain we found in the reconciler
def app_health(app):
if lock_status() == 'locked':
return 'Progressing'
if orphan_hook_present(app):
return 'Progressing'
if schema_declared_version() < required_version():
return 'Progressing'
if not migration_rbac_capable():
return 'Progressing'
if not pvc_bound(app):
return 'Progressing'
if quota_exhausted():
return 'Progressing'
if not egress_allows_db(app):
return 'Progressing'
if init_container_failing(app):
return 'Degraded'
return 'Healthy'
The reconciler's health function. Eight independent signals, all gating. Every patch to the output ConfigMap was wasted work until every signal flipped.
What the inventory pass turned up in the bleater namespace
Eight failures wearing one hat
We started with the inventory, because the ticket told us almost nothing. A real P1 page rarely enumerates faults; it tells you what is on fire and gives you the namespace. We ran the kind of get-everything pass we always run on a strange namespace.
kubectl get pods,configmaps,jobs,deployments,roles,rolebindings,serviceaccounts,pvc,resourcequota,networkpolicy -n bleater
kubectl get events -n bleater --sort-by=.lastTimestamp | tail -40
kubectl describe pod -n bleater | grep -A5 'Init Containers\|Events:'
The first three commands we ran. The namespace had about sixteen pre-existing platform workloads from other teams sharing label values with the five managed apps.
What came back was a layered mess. A suspended PreSync Job named auth-presync-migrate-legacy7r2x with a hook-cleanup finalizer and no hook-delete-policy. A second suspended Job named fanout-presync-validate that looked identical but carried the hook-delete-policy annotation and a bleater.io/owner label pointing at platform-team. A hook-reconciliation-lock ConfigMap with status: locked and a stale lock-reason from the night of the rollout. The primary application's pod in Init:CrashLoopBackOff, and kubectl logs --previous -c wait-for-db showing its init container's pg_isready printing 'bleat-db:5432 - no response'. The -c matters: without it kubectl picks the app container, which had never started. pg_isready never says why it got no answer; here the egress deny was swallowing DNS, so the host name never resolved. A bleat-db-schema ConfigMap declaring version=2 with no tables-v3 key. A migration script that contained psql ... || exit 0 and had no set -e.
And then the governance layer, which is where the rollout had really gotten out of hand. A RoleBinding named migration-runner-binding pointed at migration-runner-role-v1, which had read-only verbs. A migration-runner-role-v2 existed alongside it, unbound, with create:jobs and patch:configmaps. A PersistentVolumeClaim named bleat-migration-pvc was Pending with an event saying storageclass.storage.k8s.io "fast-ssd-tier" not found, on a k3s cluster where the only storage class was local-path. A ResourceQuota set to pods: 1. A NetworkPolicy with policyTypes: [Egress] and an empty egress list, denying everything outbound including DNS. The policyTypes line is what makes it a deny: when it is omitted, Kubernetes sets Egress only if the policy has at least one egress rule, so an empty list restricts nothing outbound and the policy denies inbound traffic instead.
Each one of those, taken alone, was a small fix. Taken together, they gated each other. The migration could not run because the RBAC was wrong. The repair Pods could not even be created because the quota was at one. The init container could not reach Postgres because the NetworkPolicy denied egress. The schema could not advance because the script swallowed errors. The reconciler refused to mark anything healthy until all of them resolved. The hardening rollout had tightened every knob at once and the knobs were not independent.
The dependency graph we drew on the bridge call. Cascade order falls out of the arrows.
The order of repair when faults gate each other
Why we raised the quota before anything else
The instinct on a multi-fault incident is to start with the most visible symptom. The CrashLoopBackOff is loud. The lock is loud. The orphan Job is loud. None of those were the right first move. The right first move was the boring one: raise the ResourceQuota, because every fix that had to run anything needed a new Pod. A ResourceQuota of pods: 1 caps the total number of non-terminal Pods in the namespace, not the number beyond what is already running, and about sixteen were already running. The quota was over-committed the moment the rollout landed it, so the API server admitted no new Pod. We raised pods to 24, sized above the existing workloads plus the five managed apps and the repair Pods, and we changed nothing else on that object during the incident. Adding cpu and memory to the quota would not have been a bigger ceiling, it would have been a new admission requirement: those keys are aliases for requests.cpu and requests.memory, and once a namespace enforces a quota on either, every new Pod must specify requests or limits for that resource or the control plane may reject its admission. Mid-incident that rejects every repair Pod without resource requests, including the ad-hoc consumer Pod we scheduled later to bind the local-path PVC, with 'failed quota: bleater-quota: must specify cpu for: consumer; memory for: consumer'. Compute quota was still worth having, so it landed afterwards, behind a LimitRange carrying default requests and limits for the namespace, and only then did cpu: 8 and memory: 16Gi go onto the quota. We did not delete the quota.
Then the RBAC. We described both Roles and confirmed v2 had the verbs the migration Job needed. Patching the existing RoleBinding to swing roleRef to v2 returned the error we expected.
$ kubectl patch rolebinding migration-runner-binding -n bleater \
--type='json' -p='[{"op":"replace","path":"/roleRef/name","value":"migration-runner-role-v2"}]'
The RoleBinding "migration-runner-binding" is invalid: roleRef: Invalid value: rbac.RoleRef{...}: cannot change roleRef
$ kubectl get rolebinding migration-runner-binding -n bleater -o yaml > /tmp/rb.yaml
# edit /tmp/rb.yaml, set roleRef.name to migration-runner-role-v2
$ kubectl delete rolebinding migration-runner-binding -n bleater
$ kubectl apply -f /tmp/rb.yaml
roleRef is immutable. The only path is delete-and-recreate, with the existing object as a template so you do not lose subjects. On a 1.37 or later API server the same rejection ends in 'field is immutable' rather than 'cannot change roleRef'. The server writes that text, so the kubectl version makes no difference.
PVC next. We listed storage classes, saw local-path was the only one, exported the existing PVC, changed storageClassName, deleted, reapplied. The PVC sat Pending for a few more seconds until we scheduled a consumer Pod against it, because local-path on k3s binds on first consumer. Then the NetworkPolicy. We did not delete it. The deny-by-default posture was the right posture; the rollout had just forgotten to allow anything. We added three explicit egress rules: same-namespace for the Postgres reach, kube-system on port 53 over UDP and TCP for DNS, and the API server, which the migration Job calls with the v2 Role's verbs and which a default-deny egress policy blocks like any other destination. The API server is not a Pod, so no pod or namespace selector can match it; that rule is an ipBlock, and it has to name the address the policy engine actually sees. Kubernetes leaves it undefined whether a network plugin applies policy before or after a Service address is rewritten. The DNS rule is the test: a namespace selector can only match the DNS Pods after the kube-dns Service address has been rewritten to them, so if that rule works, the policy engine sees rewritten addresses. With k3s's built-in policy controller it does, so the ipBlock named the k3s server's own address on TCP 6443, the port k3s fronts the API server with (every server's address, on a cluster with more than one), not the kubernetes Service's ClusterIP on 443. Some plugins add a condition of their own: Cilium does not let an ipBlock match a node address unless it runs with --policy-cidr-match-mode=nodes. The deny-all stayed in place for everything else. The reconciler's metrics scrape needed no rule of its own: it is initiated from bleater-system toward the Pods in bleater, so relative to this policy it is ingress, not egress, and the policy only restricted egress. Once DNS resolved and the database was reachable, pg_isready inside the init container started reporting 'accepting connections', and the schema verification step became the remaining failure.
Then the lock and the orphan. The lock was a one-line patch to set status: unlocked and to replace lock-reason with resolved-2024-hardening-rollback. We left an audit value rather than blanking the field. The orphan Job hung on delete because of the finalizer. That delete had already set its deletionTimestamp, so stripping the finalizer was all this Job needed. On a Job nobody has tried to delete yet, the order of the two commands matters more than it looks, and plenty of teams reach for --force first, which is the wrong tool: forcing a delete skips graceful termination, not finalizers.
# delete first (no need to wait on the finalizer), then strip it
kubectl delete job auth-presync-migrate-legacy7r2x -n bleater --wait=false
kubectl patch job auth-presync-migrate-legacy7r2x -n bleater \
--type=json -p='[{"op":"remove","path":"/metadata/finalizers"}]'
# deletionTimestamp set and no finalizers left: the API server removes the Job
# do NOT touch fanout-presync-validate. it has hook-delete-policy set,
# carries bleater.io/owner=platform-team, and the reconciler manages it.
kubectl get job fanout-presync-validate -n bleater \
-o jsonpath='{.metadata.annotations.argocd\.argoproj\.io/hook-delete-policy}'
# => HookSucceeded
Delete, then strip. Once the delete has set deletionTimestamp, the API server rejects any new finalizer, so the reconciler cannot put one back between the two commands; strip first and it can, and the delete hangs. Re-running the delete on this Job does no harm, because it was already terminating and, being suspended, had no Pods. The decoy Job looks identical to the orphan from a distance; the discriminator is the hook-delete-policy annotation and the ownership label.
We have written more on cleaning up GitOps-style state safely in our Kubernetes and CI/CD stabilization playbook, including the finalizer-strip pattern and how to tell a managed Job from an orphaned one without guessing.
Don't weaken governance to silence alarms
The fixes that had to be repairs, not deletes
Halfway through the recovery the client's platform lead asked the obvious question. Why not just delete the ResourceQuota and the NetworkPolicy until things stabilize, then put them back? It would have shaved twenty minutes. We said no, and the reason is worth writing down, because it is the part of incident work that teams under pressure get wrong most often.
Governance controls exist for a reason. Someone put a pod quota on that namespace originally because something had blown up the namespace before. Someone put the deny-all egress on because the auth service should not be able to call random external endpoints. The rollout had mangled the values, not the intent. Deleting the controls would have restored the workloads and silenced the alarms. It would have also removed two of the few real defenses that namespace had, with no scheduled work item to put them back. We have watched teams do this in March and find the controls still missing in November. The graveyard of post-incident TODOs is full of governance restore tickets that never got worked.
So we repaired. The quota went up to production limits in place. The NetworkPolicy got explicit allow rules added while the default-deny stayed. The PVC got a real storage class while the claim itself stayed at the same name and the same size. The orphan Job got deleted, because a stale suspended PreSync Job genuinely is garbage, but the cascade infrastructure stayed. Same controls, working values.
The migration script was the other repair-not-delete case. The version we found had this pattern:
# what we found
#!/bin/bash
psql -h $DB_HOST -U $DB_USER -d $DB_NAME -f /migrations/v3.sql || exit 0
echo "migration complete"
# what we replaced it with
#!/bin/bash
set -euo pipefail
psql -h $DB_HOST -U $DB_USER -d $DB_NAME \
-v ON_ERROR_STOP=1 \
--single-transaction \
-f /migrations/v3.sql
echo "migration complete"
|| exit 0 is the single worst line in any migration script. set -e and ON_ERROR_STOP=1 together mean a failing SQL statement actually fails the Job instead of reporting a false success. --single-transaction wraps the file in one transaction, so a failure rolls back rather than leaving half a v3 schema behind. Two limits from the psql documentation: if v3.sql issues its own BEGIN, COMMIT or ROLLBACK the option does not have that effect, and a statement that cannot run inside a transaction block, such as CREATE INDEX CONCURRENTLY, makes the whole transaction fail on every run.
After the script was patched, the migration Job ran successfully under the new RoleBinding, applied the v3 schema, and we read the tables back out of Postgres directly rather than trusting the script's exit code. The bleat-db-schema ConfigMap got its tables-v3 key written from observed pg_tables output. Not from the migration's stated intent. From the live database. Only then did version move from 2 to 3, the value the reconciler's schema check actually compares. If you ever find yourself writing schema declarations from anything other than what is actually in the database, you are setting up the next incident.
When in-house reconcilers and hardening rollouts collide
If your control plane is gaslighting your operators
The hard part of this kind of incident is not any single fault. The hard part is that an internal control plane is opinionated about state in ways that are not documented anywhere except in the reconciler's source code. When five apps are red and the dashboard says nothing changed, your team can spend an hour patching outputs that get reverted before they understand the inputs. Hardening rollouts make this worse, because they touch ResourceQuotas and NetworkPolicies and RBAC in the same change window, and the rollback path almost never accounts for the case where the controls themselves were the right idea but the values were wrong.
We run these recovery engagements every week. The in-house reconciler pattern shows up at almost every SaaS company past Series A that decided not to run ArgoCD or Flux directly. The shape of the failure is always the same: a small Python or Go service that watches a namespace and writes a status object, an operations team that does not own the reconciler code, and a control plane that fights every cosmetic fix because that is what it was built to do. We have seen the RoleBinding immutability case four times this quarter alone. The NetworkPolicy egress-without-DNS case shows up after every security audit cycle.
If you are watching a namespace where the status object keeps reverting your changes, or where a hardening pass cascaded across half a dozen layers and your team is debating whether to delete the controls to get back to green, book an infrastructure review with our team and we will be on a bridge call with you the same day. We will read your reconciler, draw the dependency graph for the cascade, and walk the repair order with your on-call. The goal is not to get the dashboard green by morning. The goal is to get it green without leaving a graveyard of governance restore tickets behind it.
Originally published at https://infraforge.agency/insights/internal-control-plane-cascade-recovery/.
If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — see /review.
Top comments (0)