In May our product catalogue service started having a bad six minutes roughly an hour after each deploy. Database CPU went to its limit, page loads went from under a second to eight or nine, and then everything recovered by itself. Some days it happened again an hour later, a little smaller. We spent two days looking for an hourly cron job that did not exist.
The cause was a change we were proud of. After a deploy, new pods used to start with an empty cache, and the first few minutes were slow while it filled from real traffic. So we added a warm up step: before a pod took traffic, it loaded all hundred and eighty thousand product entries into the shared cache, each with a time to live of one hour. The cold start disappeared.
Every entry now had the same birthday. The warm up took about two minutes, so an hour later the entire catalogue expired within the same two minutes. Every request missed, every miss went to the database, and the database was sized for the few percent of traffic that misses on a normal day. The entries that got refilled during that storm were refilled together, which is why an echo came back an hour after that. The lazy cache we replaced had never done this, because traffic had spread its expiry times across the hour without anybody deciding to.
Four changes. Every time to live now has twenty percent random jitter, so entries loaded together expire apart. Misses for the same key are coalesced, so a thousand requests for one product produce one database query and nine hundred and ninety nine waits. Entries are served stale for up to five minutes past expiry while a single background refresh replaces them, which means an expiry is no longer something a user can feel. And cache fills go through their own small pool of database connections, so the worst a stampede can do is make the cache slow to refill, not take the database away from everything else.
Spreading work out over time is often something a system does by accident, and you only find out it was doing it when you tidy it up. Ours had been protecting the database with randomness we had never noticed and then carefully removed.
– Sergey Shinder
Top comments (0)