On a typical Kubernetes-on-AWS bill, 15–20% of the cost is network: cross-AZ data transfer, VPC peering, cross-region transfer and NAT Gateway bytes. On a $1M monthly AWS bill, that's $150–200k every month. It's serious money, and it's almost always hidden. It sits on a few line items that everybody sees and nobody owns, and every review ends the same way: "it's network, it's the platform's problem".
This post is about how we turned those lines into a list of named workloads with a cost next to each, and what we found once we could see it. The short version: within days of having per-app network data, we had more concrete savings opportunities than in months of looking at Cost Explorer.
Table of Contents
- The pain: you can see the bill, but not the cause
- The approach: count bytes where the pod still has a name
- What we found
- How FinOps and OBI solved it together
- Why network-level observability matters
- Takeaways
The pain: you can see the bill, but not the cause
-
Cost Explorer tells you what you pay for, never who caused it.
DataTransfer-Regional-Bytesis one number per account. It doesn't know about pods, deployments or teams. - Tag-based showback gets network cost wrong. Most FinOps reports assign cost by the EC2 instance's team tag. That works for compute. For network, it charges every byte leaving a node to whoever owns the node, including bytes sent by shared DaemonSets like log shippers and agents.
- VPC Flow Logs see IPs, not workloads. Pods churn, IPs get reused, and NAT and peering traffic is often SNATed to the node IP before Flow Logs see it. Joining Flow Logs to Kubernetes metadata after the fact is a project of its own, and keeping them on for a large cluster isn't cheap.
- So the conversation goes nowhere. "Your team's network cost went up." "We don't do anything network-heavy." Nobody can prove otherwise, so nothing changes.
The missing piece wasn't another dashboard of totals. It was attribution: this workload talks to that workload, across this billable path, costing this much.
The approach: count bytes where the pod still has a name
We used OBI (OpenTelemetry eBPF Instrumentation). A small eBPF program on each node counts bytes per network flow in the kernel, with no sidecars and no code changes. OBI then enriches each flow with Kubernetes metadata: namespace, owner workload, and pod labels like application and team.
Two placements made it work for cost:
- On the node NIC (egress only): every byte is counted once, on the sender. This is the source of truth for cross-AZ.
- On the pod's veth interface: the packet still carries the pod IP before SNAT, so NAT and peering traffic can still be attributed to the pod that sent it.
We gave OBI a CIDR map generated from the AWS API: every subnet with its AZ, every peered VPC (same region and cross region), the gateway endpoints, and 0.0.0.0/0 as "internet via NAT". An OpenTelemetry Collector turns that into one
metric with a cost class:
netcost_bytes_total{net_class="cross_az|nat|same_region_peering|cross_region",
net_team, net_app, net_workload,
net_remote_app, net_remote_workload, net_zone, ...}
Multiply by the AWS price per class and you get **cost per app, per team, and per "who talks to whom" pair*you already use.
The dashboard that mattered most was a set of plain tables, not charts: "Cross-AZ: who talks to whom", "alks to each peer VPC". Each row is a sender workload, a receiver workload, bytes, and estimated cost.
What we found
None of these were visible in Cost Explorer. All of them showed up in the first few days of data.
- Logs shipped uncompressed over VPC peering. The biggest peering sender wasn't an app. It was the logng every node's logs uncompressed to a search cluster in a peered VPC. Its cost was spread across everyteam's node bill. Fix: compress, drop noisy logs at the source, keep the log path local.
-
Pods that aren't topology-aware. With random endpoint choice across three AZs, about two-thirds of sossed a zone. Fix:
topologySpreadConstraintsplus zone-aware routing (trafficDistribution: PreferCloseor mesh locality load balancing). - Single-replica Redis. One app's biggest cross-AZ flow was to its own Redis: a single pod in one AZ, others. Fix: a read replica in every AZ and nearest-replica reads, which also removes a single point offailure.
- Unwanted cross-region transfer. Workloads calling another region when an in-region endpoint already d configs. Fix: use the in-region endpoint, and make cross-region traffic a reviewed decision.
- Ingress responses crossing zones. Responses (the big half of the traffic) went back to an ingress gateway in another AZ. Fix: locality-aware load balancing at the gateway.
How FinOps and OBI solved it together
FinOps on its own gives you a bill. OBI on its own gives you traffic metrics. The value came from using them in that order:
- FinOps answers "how much" and "where to look". The FinOps reports, built on the AWS bill (CUR), show each team's network spend and flag which team or account is growing. That's the starting point.
- OBI answers "who" and "why". Once FinOps points at a team, OBI digs in: every byte has a sender worklnd a billable path. And because AWS prices turn those bytes into cost, we can check the result adds upagainst the bill.
- Together, they turn a cost line into a ticket. Instead of "network cost went up", each finding names es, the cost, and the fix, and it goes to the team that can act on it.
- They also correct showback. When a team's bill includes traffic they didn't send, like shared log shipping on their nodes, the data shows it, and the cost moves to the owner who can fix it.
Why network-level observability matters
Most teams have good visibility into compute: CPU, memory, requests per second, cost per node pool. Network is usually a blind spot, even though it's often the fastest-growing part of the bill. Without network-level observability:
- Architecture decisions carry hidden cost. A single-replica cache, a service without zone awareness or a default cross-region endpoint all look fine in a design review. The cost only shows up on the bill, months later, with no owner attached.
- The cost grows silently. Every new service, replica or AZ adds traffic. Nothing alerts on it, and nobody sees it until the monthly bill.
- Showback is wrong. Tag-based attribution charges node owners for traffic they didn't send. Teams lose the real owners never find out.
- Troubleshooting is slower. The same data that explains cost also answers "which app suddenly started sending gigabytes to another region?" during an incident.
With network-level observability, network cost becomes like any other engineering metric: measured per workload, owned by a team, and reviewed when architecture changes. That's what lets a platform team push for zone-aware design, per-AZ
replicas and sensible log shipping with evidence instead of guidelines.
Takeaways
- Network is 15–20% of a typical AWS bill, and it's the part nobody owns. On a $1M bill, that's $150–200k without attribution.
- FinOps tells you which team's network spend to look at. eBPF (OBI) tells you which workload sends each byte, and why. You need both to turn the bill into action.
- The usual suspects are worth checking first: uncompressed logs over peering, pods that aren't topology-as, unwanted cross-region calls, and ingress responses crossing zones.
- Treat network as a first-class part of observability, alongside CPU and memory. It's where a lot of the hidden cost lives.
If your bill has a network line that nobody owns, you probably have the same findings waiting for you.
Top comments (0)