
Photo by Daniel Hatcher on Unsplash
I'd already picked the GPU. I was checking a motherboard's specs for the second slot I'd need down the line, and the it said "x16/x16." A few YouTube videos, a dozen Google searches, and one straight answer from the AI god later, it turned out that only holds when one slot is populated — put a card in the second slot and both drop to x8. Nothing about the GPU I'd chosen was wrong. The board I was about to buy would have quietly halved it.
I've spent years picking AWS instance types without ever thinking about what's inside one. g5.2xlarge just works — AWS already matched the GPU, vCPUs, RAM, and network throughput for you. Building a box yourself means every one of those ratios is now your problem, and the one that bites first isn't the one people warn you about.
The setup
Goal: run 13B-70B quantized models locally, on one GPU today, with a real path to a second GPU later if I need the pooled VRAM for something bigger. (I didn't fix a hard dollar budget going in — the number that mattered more was price-per-GB of VRAM, below.)
Everyone tells you to start with VRAM. Fair — VRAM is the hard ceiling: if the model plus its KV cache doesn't fit, it either doesn't load or spills to system RAM and gets 10-50x slower. So you pick a GPU on VRAM alone.
I landed on a used RTX 3090: 24GB GDDR6X, 936 GB/s bandwidth, $700-900 on the secondary market. The alternative was a 4090 or 5090 for more bandwidth and newer memory, but the 5090's price has come apart from its spec sheet — $1,999 MSRP, but street prices as of this fall are running $3,800-11,565 depending on retailer, driven by GDDR7 supply constraints. Every guide still quoting $1,999 is stale. The 3090 isn't touching that shortage, so its price has stayed put. Two 3090s pooled (48GB, ~$1,600 total) would even beat a single 4090 on VRAM for the same money — which is what put a second GPU on my radar at all.
What actually happened
With the GPU picked, the next line item was the motherboard. I wasn't buying a second GPU yet, but I wanted the option without having to rebuild the whole machine later. That's when the "x16/x16" label turned out to mean "x16/x16, if you only use one slot" — populate both, and you're at x8/x8. For a single card that's irrelevant; a GPU at x16 doesn't come close to saturating a consumer CPU's lane budget on its own. But it meant the board I'd almost bought on price alone would have made the expansion I was planning for pointless.
Why this happens
A GPU spec sheet sells you VRAM and bandwidth. It doesn't sell you the thing that decides whether you can use that GPU alongside anything else: how many PCIe lanes your CPU actually has to hand out.
Consumer CPUs (AMD Ryzen, Intel Core) typically expose 20-24 usable PCIe lanes total — enough for one GPU at full x16, barely enough left for an NVMe drive. Workstation and server CPUs (Threadripper, Xeon, EPYC) expose 64-128+ lanes, which is the actual reason those chips exist for this use case — not raw clock speed.
This is the same distinction as EBS-optimized vs. non-optimized instances on AWS: the compute was never the bottleneck, the path to storage was. Here, the GPU was never the bottleneck — the path from CPU to GPU is.
This isn't just a homelab-scale problem. a16z hit the same wall building an 8x RTX 4090 server: at 4 GPUs per board, full x16 lanes each, the only CPU that could supply enough lanes was a dual-EPYC server board (128+ lanes). Their fix at that scale was custom PCIe boards wired directly to the motherboard, because extender cables were silently downgrading connections to PCIe 3.0. Same constraint, same fix — just bigger numbers.
Memory bandwidth compounds this. LLM decoding is memory-bandwidth-bound, not compute-bound — each generated token requires reading the entire model's weights once. A GPU starved of PCIe lanes to feed it (in a multi-GPU setup) or backed by slow system RAM (in a CPU-offload setup) hits this same wall from a different direction.
What to do instead
Option 1 — single GPU, consumer platform. If you're not planning multi-GPU, this doesn't matter: one GPU at x16 on a consumer board is fine. Spend the saved money on VRAM instead of lanes.
Option 2 — multi-GPU, plan the lanes first. Pick the CPU/motherboard combo for lane count before picking GPUs. Check the motherboard's manual for the electrical (not just physical) PCIe configuration — "x16/x16" printed on the slot doesn't mean both run at x16 simultaneously.
Option 3 — read bandwidth alongside VRAM, always. When comparing GPUs, put memory bandwidth (GB/s) next to VRAM (GB) in the same table. Two cards with identical VRAM can differ 2x in tokens/sec purely on bandwidth.
Here's what I landed on — single RTX 3090 now, motherboard and PSU sized so a second one is a drop-in later instead of a rebuild:
- GPU — used RTX 3090, 24GB, 936 GB/s. Best price-per-GB on the market right now; the 4090/5090 premium isn't worth it while 5090 pricing is this unstable.
- CPU — any current-gen mid-range Ryzen or Intel Core. One GPU at x16 uses less than half a consumer CPU's lane budget. Overspending here buys nothing a single card can use.
- RAM — 64GB DDR5, dual-channel, non-ECC. Matches VRAM headroom for OS and any CPU-offloaded layers. ECC only matters for multi-day training runs, not inference.
- Storage — 2TB NVMe Gen4. 70B model files run 40-140GB each; this leaves room for a few quantizations without constantly re-downloading.
- Motherboard — ATX board with two full-length x16 slots, confirmed x8/x8 electrical when both are populated, and enough slot spacing for a second triple-slot card. This is the one part that's expensive to get wrong later — swapping it means rebuilding the machine. Mid-to-high consumer chipsets (X670/X870, Z790/Z890) wire this correctly; budget B-series boards often don't.
- PSU — 850W 80+ Gold. Covers the one 3090 comfortably now. Sized against the formula (GPU draw + ~400W baseline + 20% margin) so going to two 3090s later means (2 × 350W) + 400W + margin — around 1,200W — which is a PSU swap, not a surprise.
- Cooling — ATX case, 3+ fans, front-to-back airflow. Overkill for one GPU is wasted money; this only becomes a real decision at multi-GPU or sustained 24/7 load.
x8/x8 instead of true dual x16 would cost me something if I were splitting one huge model across two cards with constant inter-GPU traffic. That's not my use case — I want pooled VRAM for a bigger model, or two models running in parallel, neither of which hammers that link continuously. For that, x8/x8 is right-sized, not a compromise I'm quietly eating.
Takeaway
Before you spec a GPU, spec the path to it — the CPU's PCIe lane count decides what you're allowed to build around that GPU, not the other way around.

Top comments (0)