DEV Community

Sergey Shinder
Sergey Shinder

Posted on

A connector we switched off in May kept every change our database made

At four in the morning on a Saturday in September our main Postgres primary ran out of disk and stopped accepting writes. The data itself had barely grown. The disk was full of write ahead log, two point one terabytes of it, going back to the middle of May.

Postgres keeps write ahead log until everything that might need it has had it. Replicas need it, and so does anything holding a replication slot. A slot is a promise from the database to a consumer: I will keep every change from this point until you confirm you have read it. The database keeps that promise whether or not the consumer still exists.

In May we retired a change data capture connector that streamed orders into the old analytics platform. We stopped the connector, deleted its configuration and closed the ticket. Its slot stayed in the database, inactive, holding its place in the log from the last moment it had confirmed anything. From then on the database could not remove any log written after that moment. At eighteen gigabytes a day, it took four months to matter.

Our alerts covered disk usage, and the disk alert fired at eighty percent on a Thursday evening. The on call engineer saw a database that had been growing slowly for months, assumed the data had grown, and raised a ticket to expand the volume after the weekend. Nothing told him the growth was log rather than tables, or that it would never stop on its own.

We dropped the slot, watched the log fall away within minutes and brought writes back after forty seven minutes. Now max_slot_wal_keep_size is set on every instance, so a forgotten slot is eventually invalidated instead of filling the disk. We alert on any slot that is inactive for more than an hour and on the amount of log each slot is retaining, by name. Replication slots are created through our infrastructure code with an owning team, so removing a connector from the code removes its slot. And the disk panel splits usage into tables, indexes and log.

Decommissioning a consumer is two jobs, and the producer's half is easy to forget because it does not appear anywhere the consumer lived. Our database had kept its promise to a connector for four months after we stopped listening.

– Sergey Shinder

Top comments (0)