At ten past six on the last evening of March, two of our three ingress controller pods and half of the checkout API were terminated within the same minute. The events said it plainly: Preempted. Their replacements sat in Pending, because the room they had been using now belonged to forty pods from a reporting job.
The data team had moved their monthly reports into our cluster a week earlier. Their job carried a priorityClassName of high, and they had created that PriorityClass themselves with a value of a hundred thousand, copied from a tutorial about making sure batch jobs finish. None of our own workloads had a priority class at all, and a pod without one gets priority zero.
When a pod cannot be scheduled, the scheduler looks for a node where removing lower priority pods would make it fit, and removes them. It takes PodDisruptionBudgets into account, but only as a preference, and it will break them when nothing else works. At month end the cluster was busy, the autoscaler needed several minutes to add nodes, and the cheapest room available belonged to pods at priority zero. By the only measure the scheduler has, a monthly report mattered more than our front door.
New nodes arrived seven minutes later and the replacements started on them. Checkout lost about a third of its requests for those seven minutes. The report finished on time.
We now define every priority class ourselves, in the platform repository. Platform critical sits at the top, customer facing services below it, then a default class with globalDefault set to true so that nothing lands at zero by accident, and batch below the default with preemptionPolicy set to Never, which lets a job wait in the queue without evicting anybody. A ResourceQuota scoped by priority class means only the platform namespaces can create pods at the top two levels, and an admission policy rejects any PriorityClass created outside that repository. Preemption events go to an alert channel, listed by victim and by preemptor.
Priority in a cluster is not a label for how much a team cares about its work. It is a standing instruction to stop other people's pods, and ours had been issued by whoever arrived with the largest number.
– Sergey Shinder
Top comments (0)