DEV Community

Sergey Shinder
Sergey Shinder

Posted on

The plan said update in place and our database failed over at half past ten

In April a pull request intended to make our staging database bigger also made production's bigger, at half past ten on a Wednesday morning, with a failover in the middle of our order peak. Ninety four seconds of failed writes and about six hundred abandoned checkouts.

A refactor in March had moved the instance class out of the per environment variable files into a shared defaults file, because all three environments used the same size and the duplication looked untidy. The pull request changed that value, and its author believed, reasonably, that it was a staging change. The same pull request added a cost centre tag everywhere. So the production plan had forty three changes, forty two of them tag updates, and one line that read instance class from xlarge to 2xlarge, listed under update in place.

That line was accurate. Our database module sets apply_immediately to true, because somebody once waited a week for a maintenance window to apply a parameter, and on a Multi AZ instance an immediate class change means the standby is resized and then the database fails over to it. In Terraform's vocabulary, in place means the resource keeps its identity. It says nothing at all about whether your application notices the change being made.

The plan now goes through a check before anyone reviews it. A small job reads the plan as JSON and compares every changed attribute against a table we maintain of attributes that cause a restart, a failover or an empty cache: instance class and engine version on databases, node type on cache clusters, static parameters, instance type on anything that sits behind a single target. A hit labels the pull request disruptive, requires an approver from the owning team and an agreed apply window, and makes the production apply manual. apply_immediately is false in production and is set per change when someone actually wants it. Bulk changes such as tags go in their own pull requests, so a disruptive line cannot hide in a crowd of harmless ones. And sizing no longer lives in shared defaults.

A plan describes what will be different afterwards. It does not describe what happens to your traffic while the difference is being made, and the most expensive line in ours was the one that sounded least dramatic.

– Sergey Shinder

Top comments (0)