The checkout pods came back up after a routine node rotation and started returning 502s on every request that touched the payments path. Same image tag, same application version, same deploy log as the week before. The pods were healthy. The CPU was flat. The payment client was timing out against a provider that answered every health check we threw at it.
We restarted pods. We rolled back a deploy that turned out to be unrelated. We checked the provider status page twice. The answer was an environment variable that only existed on the nodes we had just replaced.
The change nobody wrote down
Months before that morning, the payment provider had a slow afternoon. Requests were hanging on a client timeout longer than our load balancer was willing to wait, so every slow response turned into a 502 at the edge. A deploy through the pipeline meant review, a build, and a canary. An operator with console access ran one command against the running deployment instead, raising the client timeout, and watched the error rate fall.
It worked. The graph went green. Everyone moved on to the next page. The variable lived in the running environment and nowhere else.
That is how drift accumulates. It is rarely one dramatic edit. It is a kubectl set env here, a console slider there, an env var added to a single task definition because staging had the wrong value, a cron schedule someone patched directly on the box because the repo version had a typo in it. Each change is small, each is justified in the moment, and each is invisible to anyone who was not in the room.
The bill arrives later, and it arrives somewhere else.
Why the incident makes no sense
Drift failures are confusing because every signal you normally trust is silent. The code did not change, so the deploy log is clean. The image tag did not change, so the registry tells you nothing. Nobody touched the config repo, so the diff you would normally read is empty.
What changed is the ground under the code. A node rotation, a scale-up, an instance refresh, a region failover — anything that builds new instances from the repo's version of the truth instead of the running environment's version. The old instances still carry the hotfix. The new ones do not. For a while you run a mixed fleet, which is the worst version of the problem, because two pods behind the same service will answer the same request differently.
The symptom usually looks like a hard failure with a soft explanation. A timeout that comes back from the dead. A feature flag that turns itself off. A connection pool that quietly shrinks to its original size. If you catch yourself saying "it works on the old pods," you are looking at drift.
Compare the running config to the repo
The only comparison that finds drift is running state against committed state. Documentation describes intent at the time someone wrote it, and that someone may have left the company. The repo is the source. Everything else is hearsay.
So make the running environment describe itself, and diff that description against the repo:
# Dump the live environment and diff it against the committed one.
kubectl exec deploy/payments -- env | sort > /tmp/live.env
grep -E '^[A-Z_]+=' config/payments.env | sort > /tmp/repo.env
diff -u /tmp/repo.env /tmp/live.env
Run it on a schedule, not while you are on fire. Anything present in /tmp/live.env and missing from /tmp/repo.env is a finding, and so is any value that differs. The same shape works for database parameters, feature-flag defaults, and infrastructure a console manages.
Here is the file that diff compares against, and the value that leaked out of it:
# config/payments.yaml — what the service is supposed to run with.
payment:
clientTimeoutMs: 8000
retries: 3
circuitBreaker:
enabled: true
The console edit set clientTimeoutMs to 20000. The repo still says 8000. Every pod created from the repo gets 8000, and the provider is still slow.
The rule that stops it
If it is not in the repo, it does not exist. That is the whole rule. A setting that lives only inside a running process is not configuration. It is an undocumented mutation with an expiry date, and the expiry date is the next restart.
The escape hatch is real, so write it down before you need it. Emergencies happen, and sometimes changing a value by hand at 2 a.m. is the correct call. The hatch has a price: every manual change gets a commit opened against the repo in the same shift, linked in the incident channel, with the reason in the message. The incident does not close until that commit merges. No commit, no fix.
Then drift becomes visible in the one place people actually look. Reviewers see the change. The next operator inherits the reason, not just the value. And the next node rotation deploys an environment that matches the working one, because the working one was never a secret.
I write about production failures in Postgres, queues, and distributed systems.
Top comments (0)