DEV Community

Dinesh Kumar Sarangapani
Dinesh Kumar Sarangapani

Posted on Originally published at dineshkumars.dev

Solving Kubernetes IPv4 Exhaustion at Scale with Secondary CGNAT CIDRs

In enterprise cloud environments, one of the most unexpected bottlenecks when scaling AI workloads is IPv4 address exhaustion.

When using cloud-native Kubernetes networking (such as the default AWS VPC CNI), every container pod receives its own private IP address directly from the host virtual network.

For standard web services, this design works well. Pods communicate at native wire speeds with zero encapsulation overhead.

However, when an enterprise AI platform scales up to support large-scale document parsing, batch vectorization, and multi-agent workflows, background job queues can easily demand 10,000+ concurrent pods.

In a large company, corporate network security teams manage strict IP address management (IPAM) policies. Routable enterprise IP ranges are scarce. Getting an entire /16 block (65,536 addresses) for a single Kubernetes cluster is rarely approved.

When your cluster exceeds its assigned subnet allocation, node provisioning freezes, new pods fail to schedule, and batch queues grind to a halt.

Here is the idea we used to solve this challenge, how it performed in production, and what to watch out for.


The Idea: Dual-CIDR Custom Networking with CGNAT

Instead of requesting scarce corporate routable IPs for every short-lived batch pod, we separated node infrastructure addressing from pod addressing.

The architecture uses a Hybrid Dual-CIDR Strategy:

  1. Primary Network (Corporate Routable Range): A small, standard corporate subnet is reserved strictly for worker node host interfaces, load balancers, and administrative bastions.
  2. Secondary Network (Non-Routable CGNAT Range): We attach a dedicated secondary CIDR block from the Carrier-Grade NAT (CGNAT) space defined in RFC 6598 (100.64.0.0/10) to the virtual network.

Because RFC 6598 addresses are designated for shared infrastructure, they do not conflict with standard private corporate networks (RFC 1918) or the public internet.

Using custom networking plugins, worker nodes bind their primary network interface to the corporate routable subnet, while secondary interfaces provision pod IPs exclusively from the vast CGNAT address pool.


How It Worked Well

  1. Massive Pod Scaling with Minimal Corporate IP Footprint: The cluster scaled to over 10,000 concurrent batch pods while consuming fewer than 150 corporate routable IP addresses for the underlying node fleet.
  2. Native Performance Without Overlay Penalties: Because the secondary CIDR is attached directly to the cloud provider's native VPC networking layer, packets travel without overlay encapsulation (such as VXLAN or Geneve tunnels), preserving native throughput and low latency.
  3. Transparent Egress via Automatic SNAT: When a pod needs to communicate with internal corporate databases or external cloud APIs, the node translates the pod's non-routable address to the node's legitimate corporate IP via Source Network Address Translation (SNAT). Upstream firewalls and databases see traffic from authorized node IPs without needing route updates.
  4. Isolated Pod Blast Radiuses: Because pods live in a non-routable address block, external networks cannot initiate inbound connections directly to individual pods, adding a natural layer of network perimeter defense.

What to Watch Out For

  1. Availability Zone Subnet Alignment: Pod network interfaces must reside in the same Availability Zone as the worker node hosting them. Ensure your secondary CIDR is carved into equal subnets across all active Availability Zones and that your node provisioner matches them correctly.
  2. Source NAT (SNAT) Exhaustion: When thousands of pods on a single node make high-frequency outbound connections to external APIs, they share the node's single primary IP. This can lead to port allocation exhaustion or connection tracking table saturation. Tune connection pooling, reuse keep-alive sockets, and avoid opening unpooled short-lived TCP handshakes.
  3. Non-VPC Native Peering Limitations: While cloud routing handles secondary CIDR translation within the same cloud network easily, traditional on-premises VPNs or direct hardware connections do not automatically route 100.64.0.0/10 traffic. Ensure all pod communication to legacy data centers goes through node-level SNAT or application load balancers.
  4. Third-Party CNI Compatibility: If your security architecture uses an in-cluster service mesh or container network policy engine, verify that your network policy agents support multi-interface secondary CIDR routing without dropping inter-pod packets.

Top comments (0)