If you’re using an external passthrough Network Load Balancer sitting in front of your Kuberentes cluster, facts are most of the times its not aware of your Service Pod’s container health which results in users receiving Connection Reset (RST) or invalid response.
This becomes increasingly more critical during peak hours of your traffic or especially during times when RPS is extremely high and you can’t afford to drop any network packet or cause connection reset as it can long-term hurt you business or revenue.
In our company, we’ve been noticing for a long time that some of our uptime monitoring checks starts failing randomly especially during node upgrades, rollouts, pod restarts or any upgrade/maintainance work going on. At first, we thought this could be a random network blip over the wire, but sooner it became clear that this was happnening specifically whenever a Kubernetes node goes upgrading or workloads migrate to existing/new nodes.
We started digging down and after a while it got clear to us that this was happening in our single-replica deployment meaning 1 proxy (NGNIX) pod on each node sitting behind 1 global external Network Load Balancer.
Let me get into straight
In a traditional NLB setup, the Load Balancer picks any random VM from the list (target-pool based) and forwards the TCP packet to that VM. Once the packet is forwarded to the VM, LB did its job and moves on to forwarding next packet. That’s all. So in short, the LB is not aware of the running Service Pod health, if Pod is Not Ready/Terminating LB will still end up forwarding traffic to that VM.
That’s what causing connection lost and reset.
Google Cloud offers another type of more modern Network Load Balancer and its called Regional service-based Network Load Balancer. So instead of maintaing a fixed target-pool of instances, it maintains NEGs (Network Endpoint Groups) which contains the VM IP inside. The interesting thing is the NEG list gets updated by a dedicated NEG-controller running inside GKE control plane which watches Kubernetes EndpointSlice specifically for the app defined in Service selector label for Load Balancer.
To enable regional external NLB, add this in your manifest and apply (immutable):
apiVersion: v1
kind: Service
metadata:
labels:
app: proxy
name: tls-proxy
spec:
externalTrafficPolicy: Local
loadBalancerClass: networking.gke.io/l4-regional-external # add this
Now let me walk you down through a scenario:
- Pod goes in Terminating state
- EndpointSlice gets updated marking that Pod Not Ready and Cilium reponds with 503 to mark node as unhealthy
- NEG controller watches EndpointSlice and updates the NEGs removing the VM IP of the terminating Pod
- LB send new connections to other nodes running health replica
You might will think we’re done here and no more connection reset right? The answer is no!
You just solved the problem for new connections routing to other nodes running healthy replica but what about existing connections with that VM with terminating Pod inside?
Now either could be your situation:
- Pod goes in Terminating state or shutting down during node upgrade, rollout, restart, etc. (we can fix)
- Pod dies immidiately, crashed, killed, server error (your luck!)
If a Pod is in terminating state and you want to gracefully handle all existing connections without dropping any packets you need to tweak and playaround your proxy settings and LB connection draining timeout.
In our case since we’re using NGNIX as our proxy layer for TLS and backend routing, we tweaked these settings and observed close to 99.99% uptime without connection loss.
- Added sleep 40s in preStop to let NGNIX stay alive and keep serving exisiting connections
- Set Connection Draining duration of regional LB to 30s to allow existing connections to finish before cutting off and removing from NEGs
- Added keepalive_time to 15s (or your choice) to gracefully close existing connections after 15s and let user/client open fresh TCP connection with other VMs running healthy replica
- During 40s NGNIX is still alive and serving so readiness probe keeps passing which means any packet entererd in VM until draining Cilium will forward it to NGNIX pod as EndpointSlice is still serving: true for terminating pod. Also we’ve externalTrafficPolicy: Local set which means no internal network hop allowed so the terminating NGNIX Pod stays serving and Cilium uses it as a fallback
- Make sure to increase terminationGracePeriodSeconds to sleep + N seconds so you allow enough time for shutdown before SIGKILL
- All existing connections gracefully closes within 40s sleep duration and no packet lost!
- New connections gets routed to other nodes running healthy replica of your LoadBalancer Service Pod
Once the new Pod is up and Ready, NEG controller will attach the same VM to NEGs again and new traffic starts coming.









Top comments (0)