A hands-on build, a deliberate failure, and the cost, resilience and design trade-offs that never make it onto the architecture diagram.
Nobody designs a bad cloud network. They accumulate one reasonable decision at a time.
It starts with two VPCs and one peering connection, and it works beautifully. Then a third VPC appears for a new team. Then a fourth for a new environment. Six months later you own a web of routes that nobody can fully draw from memory, and a single missed entry can take down a service on a Friday afternoon.
I wanted to see exactly where that curve bends, so I built both designs in AWS: a peering topology, and a hub-and-spoke network around a Transit Gateway. I tested connectivity, broke it on purpose, and then did the part most tutorials skip: I looked hard at what each choice costs, how it fails, and when it is the wrong answer.
Here is what I learned, including the places where the "obvious" upgrade is not the right call.
Peering deserves more respect than it gets
Before the case against peering, the case for it, because it is a strong one.
A VPC peering connection is a direct, private link between two VPCs. AWS builds it on the existing VPC infrastructure rather than on a gateway or separate appliance, and its own documentation states there is no single point of failure and no bandwidth bottleneck. There is also no charge to create one, and data that stays within an Availability Zone is free.
For a web tier talking to a database tier, peering is close to perfect. I would still choose it there.
So why leave it? Because of one design property that quietly becomes your biggest problem.
The trap: peering is non-transitive
If VPC A is peered with B, and B is peered with C, A still cannot reach C. Connectivity never passes through a middleman. The only way to connect everything is a full mesh, where every VPC peers with every other VPC:
connections = N ร (N โ 1) รท 2
| VPCs | Peering connections | Route entries to maintain* |
|---|---|---|
| 3 | 3 | 6 |
| 10 | 45 | 90 |
| 50 | 1,225 | 2,450 |
| 100 | 4,950 | 9,900 |
*At least one route per side of every connection, before accounting for multiple subnets and route tables.
Growth here is quadratic, not linear, which is why a network that felt effortless at five VPCs feels unmanageable at twenty. And the connection count is not even the real cost. The real cost is the human one: every new VPC means touching every existing VPC, and every manual touch is a chance to get a route wrong. AWS also caps the number of active peering connections per VPC, so a mesh eventually hits a hard wall as well as a human one.
The fix: put a router in the middle
AWS Transit Gateway replaces the mesh with hub-and-spoke. Each VPC attaches once to a central, regional router, and the router handles traffic between them.
The operational shift is the whole point. Adding the 11th VPC stops meaning "update ten other VPCs" and starts meaning "attach it to the hub and point its routes there."
What I built
A three-VPC lab around one Transit Gateway:
- VPC 1, the bastion: a public subnet with an Internet Gateway, the only way in.
- VPC 2 and VPC 3, the spokes: fully private, with no direct internet access.
All three VPCs attached to the gateway, and I pointed each VPC's route tables at the hub. To verify it, I SSH'd into the bastion and pinged the private instances in both spokes. Traffic crossed the hub with 0% packet loss.
That is the classic secure-access pattern: one hardened entry point, private workloads behind it, a central router in between. But getting it working was the easy part. The better lesson came from breaking it.
The experiment that taught me the most
I attached VPC 3 to the Transit Gateway but deliberately left the return route to VPC 1 out of its route table. The result: an immediate timeout.
The attachment existed. The gateway was healthy. But packets had no path back, so the connection failed. Two lessons came out of that single missing line:
- Attachment is not connectivity. Being attached to the hub means nothing until routes on both sides say where traffic goes. When a connection times out, check the return path first. It is the most common culprit, and the one people check last.
- Routing is your security boundary. You can attach dozens of VPCs to one Transit Gateway and still keep dev and prod unable to reach each other, purely through which routes exist. In production you do this with separate Transit Gateway route tables, controlling which attachments associate with and propagate into each one. Isolation becomes a deliberate design decision, not a happy accident of topology.
Debugging tip worth remembering
A silent timeout on a Transit Gateway network is usually a routing gap, not a firewall. Walk the path in order: source subnet route table, Transit Gateway route table, destination subnet route table, and finally the return path. Then check security groups and network ACLs.
The cost question nobody answers honestly
"Transit Gateway is better" is easy to say. "Transit Gateway is worth it for you" requires numbers. Using current AWS list pricing for US East (N. Virginia):
| VPC Peering | Transit Gateway | |
|---|---|---|
| Fixed cost | None | $0.05 per attachment-hour (about $36.50/month each) |
| Data cost | Free within an AZ; $0.01/GB in each direction across AZs | $0.02/GB processed |
| 3 VPCs, fixed cost | $0 | About $110/month |
| 10 VPCs, fixed cost | $0 | About $365/month |
| 10 TB through the hub | Depends on AZ placement | About $200 in processing |
Prices vary by Region and change over time. Confirm on the AWS pricing pages before relying on any figure.
Three things stand out.
- The fixed fee is small next to the cost of getting it wrong. A few hundred dollars a month is less than a single day of an engineer untangling a routing outage. Judge it against operational risk, not against zero.
- The data charge is the one that scales. If two VPCs exchange terabytes constantly, such as database replication or analytics pipelines, that per-GB processing fee adds up. Those pairs are candidates for a direct peering link alongside the hub.
- Shared connectivity changes the math. One VPN or Direct Connect attachment on the gateway serves every attached VPC, instead of building that link repeatedly.
The pragmatic answer is often a hybrid: Transit Gateway as the backbone, with selective peering for the few heavy, latency-sensitive pairs.
Availability and performance: what really changes
Peering has a reputation for resilience, and it is earned: no gateway means no central component to fail. Moving to a hub does introduce a shared dependency, so it deserves a clear-eyed look.
- Capacity is not the concern. A VPC attachment supports up to 100 Gbps per Availability Zone and up to 7.5 million packets per second by default.
- Attach in every AZ you use. A Transit Gateway only routes to Availability Zones where the VPC has an attachment subnet. Put one in each AZ your workloads run in, or traffic to the missing zone has nowhere to go.
- Mind the MTU. A Transit Gateway supports an MTU of 8500 bytes between VPCs. When migrating from peering, a size mismatch can drop jumbo packets on asymmetric paths, so update both VPCs together and test large transfers.
- Centralization concentrates risk. One wrong route in a shared table can affect many VPCs at once. Manage the gateway as code, review every change, and use blackhole routes to block traffic deliberately.
The hidden factors that decide real-world success
- IP planning comes first. Neither approach can route between overlapping CIDR ranges. A clean, non-overlapping address plan is the cheapest decision you will ever make, and the most painful to fix later.
- Multi-account is where it shines. A Transit Gateway can be shared across accounts with AWS Resource Access Manager, so each team attaches its own VPC to a central network owned by the platform team.
- Central inspection becomes possible. Hubs let you steer traffic through a shared firewall layer instead of duplicating security tooling in every VPC.
- Visibility improves. One routing hub means one place to look, log and audit, instead of dozens of point-to-point links.
Which should you choose?
| Stay with VPC Peering | Move to Transit Gateway | |
|---|---|---|
| VPC count | 2 to 4, and stable | 5+, or growing |
| Traffic pattern | Few pairs, heavy data volume | Many VPCs, mixed traffic |
| Team shape | One team, one account | Multiple teams or accounts |
| Connectivity | VPC to VPC only | Shared VPN or Direct Connect |
| Isolation needs | Simple | Segmented environments (dev, prod, shared) |
| Priority | Lowest cost | Lowest operational risk |
A simple rule: if you can already see your fourth, fifth, or tenth VPC coming, design the hub now. Migrating a live mesh means changing routes on running production traffic. Starting with a hub means adding an attachment.
If you are migrating, do it in this order
- Audit CIDRs and routes. Find overlaps and document every existing peering path before touching anything.
- Build the gateway alongside the mesh. Attach the VPCs, create the route tables, and leave live traffic untouched.
- Cut over one pair at a time. Update both sides together, validate with pings, flow logs and a large transfer, then move on.
- Retire the peering links last. Remove old routes and connections only after the new path has proven itself.
The bottom line
Peering solves today's problem. Transit Gateway solves the one you will have when your environment doubles.
Network design is cheap to get right early and expensive to fix later. The mesh does not announce itself. It simply grows until the day it hurts. If you are running more VPCs than you can draw from memory, treat that as your signal.
I would like to hear from you: are you running a peering mesh today, or have you already moved to a hub-and-spoke design? What finally pushed you to change?

Top comments (0)