DEV Community

Cover image for Your Kubernetes Cluster Ran Out of IP Addresses Before It Ran Out of Nodes
Rakesh Tanwar
Rakesh Tanwar

Posted on

Your Kubernetes Cluster Ran Out of IP Addresses Before It Ran Out of Nodes

Most teams size Kubernetes clusters around compute.

They estimate CPU, memory, node count, and perhaps GPUs.

Then one day new Pods stop networking correctly even though the cluster has plenty of compute capacity.

The resource that disappeared was IP addresses.

I think IP capacity should be treated as a first-class Kubernetes sizing metric.

Every Pod needs network identity

The Kubernetes networking model normally gives each Pod an IP address.

Where those addresses come from depends on the CNI and infrastructure design.

In some cloud architectures, Pod addressing can consume subnet resources closely tied to the underlying virtual network.

That means adding larger nodes does not automatically solve the problem.

A cluster can have enough CPU to schedule hundreds of additional Pods while the available address pool can support only a fraction of them.

I calculate address capacity before production

I start with the CIDRs used by nodes, Pods, and Services.

They must be understood separately.

Then I ask how the selected CNI allocates Pod addresses.

Does it assign addresses directly from cloud subnets

Does it use an overlay

Does it reserve addresses in blocks

Does it keep warm addresses ready for faster Pod startup

Different answers produce very different capacity limits.

This is why VPC design and Kubernetes design cannot be separated. The AceCloud VPC comparison guide provides useful context around subnet design, address ranges, routing, and private connectivity.

Node density can surprise you

Suppose a team moves from many small nodes to fewer large nodes.

Compute efficiency might improve.

Network density might not.

The maximum number of Pods supported per node may be influenced by the networking implementation and available addresses.

I therefore model both resources.

A node is useful only if it has enough CPU, memory, and networking capacity to host the intended Pods.

Autoscaling can accelerate exhaustion

Node autoscaling makes IP planning even more important.

During a traffic spike, the system may create nodes and Pods rapidly.

If those nodes consume address capacity from an already crowded subnet, autoscaling can hit the network ceiling exactly when the application needs capacity most.

That is a nasty failure mode because the infrastructure appears to be scaling successfully while workloads remain unable to start correctly.

I include remaining IP capacity in cluster alerts for this reason.

I avoid using one subnet for everything

Where the infrastructure supports it, I prefer intentional address planning across environments and workload groups.

Production clusters should not inherit a tiny CIDR simply because it was convenient during initial setup.

I leave room for growth, rolling upgrades, temporary surge capacity, and node replacement.

A rolling node upgrade can temporarily require old and new workers to coexist.

If the address plan supports only normal steady state, maintenance itself can trigger exhaustion.

Dual stack can be strategic but not automatic

IPv6 and dual-stack Kubernetes can significantly change long-term address planning.

I do not view dual stack as an emergency fix for a poorly designed IPv4 network.

It introduces application, observability, security, and operational considerations of its own.

But for organizations building platforms expected to grow substantially, I believe IPv6 readiness belongs in the architecture conversation.

I monitor IP capacity beside CPU and memory

Most cluster dashboards make CPU and memory impossible to ignore.

I want network capacity to be similarly visible.

The exact metric depends on the CNI and cloud environment, but the principle is universal.

I want alerts before address utilization becomes critical.

I also test Pod creation during load exercises.

A cluster that can handle traffic only while no additional Pods need addresses is not genuinely ready for a spike.

Managed Kubernetes does not remove network planning

A managed control plane can eliminate considerable operational work.

It does not repeal network mathematics.

Teams still need suitable subnet sizes, CNI choices, node groups, and private network design.

When evaluating environments such as AceCloud Kubernetes, I would include projected Pod count and network isolation requirements alongside worker CPU and memory sizing.

My main lesson

Kubernetes capacity is multidimensional.

CPU is capacity.

Memory is capacity.

GPU devices are capacity.

Storage attachment is capacity.

IP addresses are capacity too.

If I design only for the first two, I can build a cluster with expensive idle nodes that still cannot create another usable Pod.

That is why I calculate IP headroom before deployment, monitor it during operation, and revisit it before major scaling events.

The best time to discover that a subnet is too small is during architecture review.

The worst time is when the autoscaler is trying to save production.

Top comments (0)