The finance lead asked why AWS charged us $2,100 for NAT gateway data processing last month. Our normal was around $400. Nothing in the release calendar explained it: no new services, no traffic bump in our own metrics, no scaling events. The bill just quintupled. Then someone opened VPC Flow Logs and filtered on the NAT's ENI. Roughly 80% of the traffic it carried was to or from addresses that AWS's ip-ranges.json lists as S3 in us-east-1. We had an S3 gateway endpoint on the VPC. It was supposed to be handling that traffic on the AWS backbone for free.
Problem signals:
- NAT gateway data-processing charges (NatGateway-Bytes, with a region prefix outside us-east-1) jump to a new, higher daily level, or climb as pods land on an uncovered subnet, with no code or workload change
- VPC Flow Logs for the NAT gateway show heavy traffic to or from the S3 ranges that ip-ranges.json lists for your region (52.216.0.0/15 is one of them in us-east-1), traffic the gateway endpoint should be carrying
- aws ec2 describe-route-tables on a private subnet returns no route for the S3 managed prefix list (pl-xxx); check the main route table if the subnet has no explicit association
- NAT gateway BytesOutToDestination and BytesInFromDestination (uploads and downloads) stay high while VPC Flow Logs show S3 traffic still routing through NAT (the S3 gateway endpoint itself emits no CloudWatch metrics to check)
- The S3 gateway endpoint exists in the VPC, but the bill still shows $0.045/GB on service-to-service S3 traffic
$412 to $2,103 on a flat workload, all of it NatGateway-Bytes
The bill line item that should not have existed
The NatGateway-Bytes cost for us-east-1 climbed from $412 in March to $2,103 in April. Nothing else on the bill moved. Egress was flat. EC2 was flat. RDS was flat. The whole delta was NAT data processing at $0.045 per GB.
Cost Explorer's usage-type breakdown confirmed the shape. The entire spike was the NatGateway-Bytes usage type, in one region. At daily granularity it was not a ramp but a step: at the start of April the daily NatGateway-Bytes cost jumped to about five times its usual level and stayed there. Nobody looked at the daily view, because 'NAT costs a few hundred bucks' was our mental default.
So we mapped the destinations. AWS publishes the IP ranges for S3, DynamoDB and a couple of dozen other services in ip-ranges.json. We pulled an hour of the NAT's flow logs and kept only the records on its internet side, because a NAT gateway's flow logs record every byte twice, once between the pod and the NAT and once between the NAT and the internet. Then we joined the remote address of each record, the destination going out and the source coming back, against the S3 entries in those ranges, and found the answer. S3 was about 80% of the NAT-processed bytes, which is what the bill said too: everything above the old $412 baseline. That should have been impossible. We had a gateway endpoint for S3, and gateway endpoints are free. Their entire point is to keep S3 traffic off NAT.
Cross-region buckets, then the legacy global endpoint, then the truth
What we thought first, and why the first two theories died fast
The first theory was that a new batch job was writing to a bucket in a different region. Cross-region S3 traffic does not hit the same-region gateway endpoint; it goes out through NAT. Reasonable theory, wrong theory. We grepped Terraform for cross-region bucket references and found nothing. We checked CloudTrail for recent bucket creations. Nothing new. All our S3 traffic was to buckets in the same region as the workload.
The second theory was that an application was talking to S3 through the legacy global endpoint, https://s3.amazonaws.com, instead of the regional one. Also wrong, and it could not have been: we run in us-east-1, where the global endpoint resolves into the same S3 address ranges as the regional one, and the gateway endpoint covers both.
The third theory turned out to be right. The route table for one of our private subnets was not associated with the S3 gateway endpoint, so it had no S3 prefix-list route. Any pod scheduled onto a node in that subnet was reaching S3 through the NAT, at $0.045 per GB, for six weeks. Same workload the whole time. Different route table.
The route that has to exist, and the new route table that never got it
How a gateway endpoint quietly stops covering a subnet
A gateway endpoint is not attached to your subnets the way an interface endpoint is. There is no ENI. There is no private DNS record. The gateway endpoint lives in the VPC, and it covers a subnet only through that subnet's route table: you associate the route table with the endpoint, and AWS adds a route whose destination is the service's managed prefix list (pl-63a5400a for S3 in us-east-1) and whose target is the endpoint. You cannot add or edit that route by hand.
The whole 'your S3 traffic bypasses NAT' behavior depends entirely on that route existing. If the route is not there, S3 traffic falls through to the default route (0.0.0.0/0), which points at the NAT gateway. Unless something requires requests to arrive through the endpoint, there is no error, no warning, no CloudWatch alarm. The traffic just goes through NAT and gets billed per GB.
We checked all four of our private subnets' route tables for the S3 prefix-list route:
aws ec2 describe-route-tables \
--filters "Name=association.subnet-id,Values=subnet-0abc123" \
--query 'RouteTables[].{table: RouteTableId, s3: Routes[?DestinationPrefixListId==`pl-63a5400a`].GatewayId}' \
--output json
Repeat for each private subnet. An entry whose s3 list is empty means that subnet's table has no S3 endpoint route. A bare [] means no table is explicitly associated with that subnet, so it uses the VPC's main route table (check that one with --filters Name=vpc-id,Values= Name=association.main,Values=true), or the subnet ID, region or profile is wrong. pl-63a5400a is the us-east-1 S3 list; elsewhere, find yours with aws ec2 describe-managed-prefix-lists --region --filters Name=prefix-list-name,Values=com.amazonaws..s3.
Three subnets came back with the endpoint ID in their s3 list. The fourth came back with an empty one. Six weeks earlier, a network migration for another team's project had moved that subnet onto a new route table, built by a Terraform module that did not know the gateway endpoint existed. The endpoint's route-table associations were managed by a separate stack that still listed only the old tables, so nothing ever associated the endpoint with the new one, and the new table never got the S3 route.
Nothing failed. Nothing paged. The workload kept running. It just started paying NAT rates for every S3 GET and PUT the pods on those nodes made.
modify-vpc-endpoint, not create-route
The fix, and the command people reach for that does not apply here
The instinct might be to reach for aws ec2 create-route --route-table-id --vpc-endpoint-id to add the endpoint to the table by hand. That does not fit here. Interface endpoints do not use routes at all (they are ENIs with private DNS), and a gateway endpoint's prefix-list route is not added with create-route either. Gateway endpoints get added to a route table by modifying the endpoint itself and telling it which route tables to associate with.
The correct fix is one call:
aws ec2 modify-vpc-endpoint \
--vpc-endpoint-id vpce-0abc12345 \
--add-route-table-ids rtb-0def67890
AWS adds the prefix-list route to the table as part of the association. No second command. The switch drops open TCP connections to S3 from every subnet on that table and does not resume them, and S3 then sees those requests from private addresses instead of the NAT's public IP. A bucket or IAM policy that restricts access by aws:SourceIp can start refusing them, because that key is absent on requests through a VPC endpoint (use aws:VpcSourceIp or aws:SourceVpce instead), and the endpoint's own policy now applies to them. Run it when nothing critical is mid-transfer, or confirm your clients reconnect.
Verification took thirty seconds:
aws ec2 describe-route-tables --route-table-ids rtb-0def67890 \
--query 'RouteTables[].Routes[?GatewayId==`vpce-0abc12345`]'
Should now return the pl-xxx prefix-list route with the endpoint as GatewayId.
The route was there. We then pulled a fresh five-minute slice of VPC Flow Logs and joined against S3's IP ranges again. S3 destinations were dropping out of the NAT sample. The subnet was routing them to the gateway endpoint instead. Over the next hour we watched the NAT gateway's traffic in both directions, BytesOutToDestination plus BytesInFromDestination, fall by about 80% and level off at the old baseline.
The CLI call was the whole fix. The same day we added the new table to the endpoint stack's route-table list, because that stack's next apply would otherwise remove the association. Six weeks of overspend at about $56 a day, roughly $2,400 in avoidable NAT charges, closed in about four minutes of actual work.
Endpoint-to-route-table coupling in Terraform, plus an anomaly-detection NAT alarm
The two things we changed so this stops happening
We changed two things. First, every gateway endpoint in our Terraform is now paired with an explicit list of route-table IDs it associates with, and that list is generated from the same module that generates the private subnets. When someone adds a subnet, the endpoint's association is derived from the same variable. There is no second stack to remember.
Second, we set up a CloudWatch anomaly-detection alarm on each NAT gateway, on a metric math sum of BytesOutToDestination and BytesInFromDestination. Both terms matter: S3 uploads leave in the first, downloads come back in the second, and the NAT bills both. Not a fixed dollar threshold: the model learns each gateway's normal daily shape from up to two weeks of history, and we get paged when the sum stays above the band for two consecutive hours. Our jump was a step to about five times the usual level, so the alarm would have paged within hours of the migration instead of after six weeks of overspend. Act on that first page: the model retrains to absorb sudden changes, so a step nobody acts on becomes the band's new normal. After the fix, exclude the spike window from the model's training, so a repeat is measured against the normal level rather than the spike.
We considered AWS Cost Anomaly Detection. It can flag this shape of spike, but it can take up to a day to detect an anomaly after the usage, and with an absolute threshold its alerts wait for the anomaly's total cost impact, actual minus expected spend over the anomaly, to pass it: at about $56 a day, a $600 threshold can take eleven days to cross. Our own CloudWatch alarm fires in hours because NAT bytes are a real-time metric and the cost bill is not.
For teams reading this with a similar VPC layout, the audit is roughly a ten-minute job. Enumerate every gateway endpoint, enumerate every route table that should be covered, and check for the prefix-list route in each. If you want the deeper cleanup pattern we use for accumulated cloud-cost drift, we have written more of that up in the InfraForge services overview.
This kind of drift does not fail loudly. It just bills.
If your NAT bill just went sideways and nobody deployed anything
The specific shape of this problem is one people miss because it does not fail loudly. A gateway endpoint that stops covering a subnet produces no error as long as the subnet still routes to a NAT and nothing requires requests to arrive through the endpoint, such as an aws:SourceVpce or aws:SourceVpc condition in a bucket or IAM policy, or an access point restricted to the VPC. Requests keep succeeding through the NAT, and the bill grows. If you have not audited your route tables against your gateway endpoints in the last six months, there is a decent chance one of your subnets is quietly paying NAT rates for S3 or DynamoDB traffic right now.
We have seen this pattern three times this quarter. Two were the same shape as ours: a subnet moved onto a new route table that the endpoint was never associated with. One was a new subnet whose route table never got the association at all. Each was a five-figure annual overrun that took an afternoon to identify and minutes to fix.
If your NAT gateway bill jumped and nobody deployed anything, book an infrastructure review with our team and we will start with a 30-minute diagnostic call this week. We will measure how much of your NAT traffic goes to S3, DynamoDB and other AWS ranges, check your route tables against your gateway endpoints, and, if a subnet is bypassing its endpoint, tell you which one before the call ends.
Originally published at https://infraforge.agency/insights/nat-gateway-cost-spike-missing-vpc-endpoint-route/.
If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — see /review.
Top comments (0)