DEV Community

Sergey Shinder
Sergey Shinder

Posted on

Our node pool rotation hit Docker Hub's limit and new nodes could not start their pods

Every month we replace all the nodes in our production cluster with fresh ones built on the latest image. It is routine and automated, and it usually takes about two hours. In February it stopped halfway. Fourteen new nodes had joined and their pods sat in ImagePullBackOff, with the message toomanyrequests from registry-1.docker.io.

Our own images live in our private registry, and I would have said we did not use Docker Hub at all. We did, in small pieces. The log shipping daemonset used an image from Docker Hub. So did an init container that waits for the database, a busybox image in two jobs and the metrics exporter for Redis. Each was a few megabytes and nobody had thought of them as dependencies.

Docker Hub limits anonymous pulls by source address. All our nodes reach the internet through two NAT gateways, so to Docker Hub the whole cluster was two addresses sharing one small allowance. On a quiet day that was plenty, because each node pulled those images once in its life. A rotation is the day every node is new at once. The CI runners, behind the same gateways, had spent the morning pulling their own Docker Hub images and had used most of the allowance before the rotation began.

The pods that could not start were the log shipper and the database wait container. Without the second, nothing on those nodes could start at all. We paused the rotation, kept the old nodes, and waited for the limit to reset.

The fix had three parts. Every third party image we run is now copied into our own registry by a scheduled job that records its digest, and manifests refer to our copy. Our registry also acts as a pull through cache for anything new, authenticated with a paid account, so even a missed image does not count against an anonymous limit. An admission policy rejects any pod whose image comes from a public registry, which found six more references on its first day, including two in charts we had installed and never read. CI runners pull through the same cache.

A dependency does not have to be large to stop you. The images we never counted were the ones whose availability we were borrowing from somebody else, and the day we needed all of them at once was the day we had chosen ourselves.

– Sergey Shinder

Top comments (1)

Collapse
 
amorizz profile image
Amorizz •

The pull-through cache / mirror angle is the part that usually gets skipped until the second outage. We hit the same Hub burst on a greenfield node pool and the symptom looked like a scheduler bug until docker pull by hand returned 429. Pinning a registry mirror in the daemon config plus pre-pulling the handful of base images in the AMI cut the blast radius more than bumping Hub plan did.