DEV Community

Sergey Shinder
Sergey Shinder

Posted on

The incident ended and the spare never came back

At twenty past two on a Thursday, writes to our main Postgres cluster stopped for forty minutes. The cause was a routine node drain during monthly patching. It should have cost us nothing, and twelve days earlier it would have.

We run three nodes under Patroni, one primary and two replicas, with synchronous replication to at least one of them. On the second of the month an availability zone incident took a replica out. Patroni behaved exactly as intended, the cluster carried on, and the write up recorded seven seconds of impact and a successful automatic recovery. It was, honestly, a good day.

The replacement never joined. Its pod could not be scheduled while that zone had no capacity of the instance type we use, and by the time capacity returned nothing was retrying, so it sat Pending. Every alert we owned was satisfied. We watched the primary, replication lag on whatever was connected, connection counts and disk. Not one of them is capable of saying two out of three. The cluster was healthy, and it was healthy with nothing to spare, and those two states looked identical from where we were standing.

Twelve days later the patching automation drained the node holding one of the two survivors. With one replica left and a synchronous commit requirement it could no longer satisfy, the primary did the correct thing and stopped accepting writes.

Three changes came out of it. We alert on ready replicas against desired for every stateful workload, and on quorum margin rather than on whether the service is answering, so being one failure from an outage pages somebody during office hours. Drains and rollouts check that margin first, through a disruption budget that cluster never had. And the incident template now carries redundancy restored as a closing condition with a named owner, so an incident stays open while the system is still degraded, however well it is serving.

Redundancy is a stock, and you spend it. Nothing spends it more quietly than a recovery that worked, because everyone stops watching at the moment the customers stop noticing. Our monitoring could tell us when something was down, and had nothing whatever to say about how close we were.

– Sergey Shinder

Top comments (0)