Your Terraform plan told you the security group changed. It has been telling you for eleven days. Nobody has reverted it, and nobody can tell you if reverting it is safe.
That is the actual state of infrastructure drift in most teams. The detection problem is genuinely solved — terraform plan, kubectl diff, cloud provider config rules, and GitOps reconcilers all give you a precise, timestamped diff of intended state versus actual state. We have excellent answers to "what changed?" We have almost no answer to "who changed it, why, and is it safe to put back?"
That gap is where outages are born.
The detection layer is genuinely good now
Let me be fair to the tooling. This part works.
-
terraform plan/terraform plan -refresh-onlygives you a diff between your state file and reality. It is fast, precise, and free. -
kubectl diffdoes the same for Kubernetes manifests against live cluster objects. - HCP Terraform drift detection runs those checks on a schedule and flags workspaces that have diverged, without a human running anything.
- Argo CD and Flux run continuous reconciliation loops — they detect drift every few minutes, not once a quarter.
- Cloud provider config rules (AWS Config, Azure Policy, GCP Security Command Center) catch out-of-band changes at the control plane.
The industry consensus is that drift detection is a solved problem. Microsoft's own SRE guidance puts it plainly: "Drift detection is a solved problem. terraform plan tells you what changed. But what happens next — who changed it, why, and whether it's safe to revert — that's still a manual investigation."
Read that sentence again. The tooling vendor is telling you the hard part was never detection.
Why remediation is the actual hard problem
A diff is a statement about difference. It says nothing about intent. Four things make the next step genuinely difficult:
1. Diffs can't distinguish a mistake from a deliberate hotfix.
An engineer paged at 3am widened an ingress rule to restore traffic. The change is real, it is out of band, and reverting it will cause a second outage. A pure reconciler does not know that — it knows the file says /32 and reality says 0.0.0.0/0, so it reverts.
2. Reverting is a mutation with a blast radius.
Drift remediation is not a read operation. Re-applying a Terraform plan against drifted state can destroy and recreate resources — databases, load balancers, volume attachments. The detection step is risk-free. The remediation step is one of the highest-risk operations you can run against production, and it usually gets executed by whoever happens to be on call.
3. State bloat and partial applies compound it.
A failed apply leaves half-applied changes. Re-running against that produces a plan full of noise — resources that look drifted but are actually mid-transition. Engineers learn to ignore the plan output, and once they ignore it, detection stops mattering.
4. Nobody owns the diff queue.
Detection tools surface findings into a dashboard. Dashboards need someone to triage them. Drift findings accumulate in exactly the same way security findings and Dependabot PRs do: a growing list nobody has time to work through. The change that caused the incident is sitting at position 47.
What teams actually do today
Look at the remediation strategies teams genuinely run, and the shape of the problem becomes obvious:
- Revert — re-apply the IaC source of truth. Correct, and highest risk. Breaks hotfixes.
- Align — update the IaC code to match reality. Preserves the fix, but turns accidental drift into permanent undocumented configuration.
- Ignore — accept and suppress the finding. Now your IaC is a lie.
Every one of these is a judgement call about intent, made by a human, under time pressure, with incomplete context. "Detect drift and notify" outsources the entire engineering problem to whoever reads the notification.
The research community knows this too. Recent work on IaC reconciliation (NSync and similar systems) focuses specifically on propagating out-of-band changes back into the IaC program — because the interesting problem was never finding the divergence. It is deciding what the divergence means.
What a useful remediation layer needs
If detection is cheap and remediation is expensive, the value is in the layer between them. That layer needs four properties:
Correlate the change to its cause. A drift finding should arrive attached to the event that produced it: which principal made the API call, at what time, in response to what alert. An ingress change made 90 seconds after a paging alert on 5xx rates is a hotfix. The same change made on a Tuesday afternoon with no correlated incident is something else entirely.
Verify before you assert. Any system claiming it remediated drift has to prove the end state remotely — not assume the API call returned 200. Check that the port is actually closed, the container is actually running, the response is actually 200. This is the same discipline as deployment verification, and it is the difference between "we reverted it" and "it is reverted."
Keep a memory of what worked. When the same drift signature recurs — and it does, because humans repeat their own hotfixes — the remediation that worked last time should be the first candidate, not a fresh manual investigation from zero.
Be honest about risk, and ask about intent. Some drift should be reverted automatically. Some should never be touched without a human. A system that cannot tell the difference will confidently cause the second incident. The objective is a small, high-confidence set of automatic remediations plus a well-contextualised queue for the rest.
The reframe
Stop measuring yourself on drift detection coverage. Almost everyone has that. Ask instead:
- When your reconciler finds drift, what is the median time from detection to either remediation or an explicit, recorded decision to keep it?
- Can you answer why a resource drifted, or only that it did?
- When something is reverted, did anything verify the end state remotely, or is the IaC apply exit code the only evidence?
If those questions are hard to answer, you have not shipped drift management. You have shipped drift notification, and the actual engineering work is still queued behind whoever is on call next.
Detection was the easy half. It got solved and we stopped paying attention to the other half.
KAIRO is an AI Infrastructure Engineer that closes the loop after detection — correlating infrastructure changes to the events that caused them, verifying remediations remotely instead of trusting an exit code, and reusing validated fixes for recurring failure signatures. See how KAIRO closes that loop →
Top comments (0)