In July a database migration passed every check in our pipeline and then failed in staging on its first statement. It created an index using a function from a Postgres extension, and nothing in the migration created the extension. We have a job whose only purpose is to run every migration from an empty database. It had passed.
It had not started from an empty database. We ran four self hosted runners on long lived virtual machines, kept because the Docker layer cache on their disks made builds fast. The migration job brought Postgres up with docker compose, and compose names its containers after the project, which by default is the directory name. Every job for that repository checks out into the same workspace path on a given runner, so every job used the same project name. When our job asked for the database, compose found a container with that name already running, reported it as up to date, and handed it over.
That container had been started three days earlier by a job on another branch, one that had been cancelled halfway. Our teardown step ran only on success, so cancellation skipped it. The branch had installed the extension. Our migration found it waiting.
A rerun on a hosted runner failed in forty seconds. Then I went looking at the four machines and found sixty one containers and a hundred and forty volumes, the oldest from February.
The runners are ephemeral now: one job per virtual machine, destroyed when the job ends, with the layer cache moved into our registry and pulled at the start of each build. The compose project name is set to the job id, so two jobs can never share a container even by accident. Teardown runs whatever the outcome. A preflight step fails the job if anything is already running on the host. And the migration test checks that the database contains no schema at all before it applies the first file, because its whole claim rests on that and it had never once verified it.
The pipeline is ninety seconds slower. Nobody has missed them.
A runner that outlives its job is a shared environment. Every job on it tests some mixture of its own change and whatever happened to run before it, and the result gets reported as if it were only the first.
– Sergey Shinder
Top comments (2)
Do not follow external links, this is a phishing scam. DEV.to uses Sloan for automated messaging.