<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: rishi764</title>
    <description>The latest articles on DEV Community by rishi764 (@sindabad764).</description>
    <link>https://dev.to/sindabad764</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3926224%2F664029ac-a8eb-4c81-95b7-ddb511d46205.png</url>
      <title>DEV Community: rishi764</title>
      <link>https://dev.to/sindabad764</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sindabad764"/>
    <language>en</language>
    <item>
      <title>PCIe lanes, not GPU VRAM, is the spec that kills your homelab GPU plan</title>
      <dc:creator>rishi764</dc:creator>
      <pubDate>Sun, 04 Oct 2026 05:07:28 +0000</pubDate>
      <link>https://dev.to/sindabad764/pcie-lanes-not-gpu-vram-is-the-spec-that-kills-your-homelab-gpu-plan-3h9f</link>
      <guid>https://dev.to/sindabad764/pcie-lanes-not-gpu-vram-is-the-spec-that-kills-your-homelab-gpu-plan-3h9f</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1gbr1d9suqs8967drmp.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1gbr1d9suqs8967drmp.jpg" alt=" " width="800" height="534"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Photo by &lt;a href="https://unsplash.com/@dphatcher" rel="noopener noreferrer"&gt;Daniel Hatcher&lt;/a&gt; on &lt;a href="https://unsplash.com/photos/lighted-black-and-gray-graphics-card-zPHftoPajis" rel="noopener noreferrer"&gt;Unsplash&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I'd already picked the GPU. I was checking a motherboard's specs for the second slot I'd need down the line, and the it said "x16/x16." A few YouTube videos, a dozen Google searches, and one straight answer from the AI god later, it turned out that only holds when one slot is populated — put a card in the second slot and both drop to x8. Nothing about the GPU I'd chosen was wrong. The board I was about to buy would have quietly halved it.&lt;/p&gt;

&lt;p&gt;I've spent years picking AWS instance types without ever thinking about what's inside one. &lt;code&gt;g5.2xlarge&lt;/code&gt; just works — AWS already matched the GPU, vCPUs, RAM, and network throughput for you. Building a box yourself means every one of those ratios is now your problem, and the one that bites first isn't the one people warn you about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Goal: run 13B-70B quantized models locally, on one GPU today, with a real path to a second GPU later if I need the pooled VRAM for something bigger. (I didn't fix a hard dollar budget going in — the number that mattered more was price-per-GB of VRAM, below.)&lt;/p&gt;

&lt;p&gt;Everyone tells you to start with VRAM. Fair — VRAM is the hard ceiling: if the model plus its KV cache doesn't fit, it either doesn't load or spills to system RAM and gets 10-50x slower. So you pick a GPU on VRAM alone.&lt;/p&gt;

&lt;p&gt;I landed on a used RTX 3090: 24GB GDDR6X, 936 GB/s bandwidth, $700-900 on the secondary market. The alternative was a 4090 or 5090 for more bandwidth and newer memory, but the 5090's price has come apart from its spec sheet — $1,999 MSRP, but street prices as of this fall are running $3,800-11,565 depending on retailer, driven by GDDR7 supply constraints. Every guide still quoting $1,999 is stale. The 3090 isn't touching that shortage, so its price has stayed put. Two 3090s pooled (48GB, ~$1,600 total) would even beat a single 4090 on VRAM for the same money — which is what put a second GPU on my radar at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually happened
&lt;/h2&gt;

&lt;p&gt;With the GPU picked, the next line item was the motherboard. I wasn't buying a second GPU yet, but I wanted the option without having to rebuild the whole machine later. That's when the "x16/x16" label turned out to mean "x16/x16, if you only use one slot" — populate both, and you're at x8/x8. For a single card that's irrelevant; a GPU at x16 doesn't come close to saturating a consumer CPU's lane budget on its own. But it meant the board I'd almost bought on price alone would have made the expansion I was planning for pointless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this happens
&lt;/h2&gt;

&lt;p&gt;A GPU spec sheet sells you VRAM and bandwidth. It doesn't sell you the thing that decides whether you can &lt;em&gt;use&lt;/em&gt; that GPU alongside anything else: how many PCIe lanes your CPU actually has to hand out.&lt;/p&gt;

&lt;p&gt;Consumer CPUs (AMD Ryzen, Intel Core) typically expose 20-24 usable PCIe lanes total — enough for one GPU at full x16, barely enough left for an NVMe drive. Workstation and server CPUs (Threadripper, Xeon, EPYC) expose 64-128+ lanes, which is the actual reason those chips exist for this use case — not raw clock speed.&lt;/p&gt;

&lt;p&gt;This is the same distinction as EBS-optimized vs. non-optimized instances on AWS: the compute was never the bottleneck, the path to storage was. Here, the GPU was never the bottleneck — the path from CPU to GPU is.&lt;/p&gt;

&lt;p&gt;This isn't just a homelab-scale problem. a16z hit the same wall building an 8x RTX 4090 server: at 4 GPUs per board, full x16 lanes each, the only CPU that could supply enough lanes was a dual-EPYC server board (128+ lanes). Their fix at that scale was custom PCIe boards wired directly to the motherboard, because extender cables were silently downgrading connections to PCIe 3.0. Same constraint, same fix — just bigger numbers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcm6kvzf7c499f2x4di9v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcm6kvzf7c499f2x4di9v.png" alt=" " width="800" height="538"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Memory bandwidth compounds this. LLM decoding is memory-bandwidth-bound, not compute-bound — each generated token requires reading the entire model's weights once. A GPU starved of PCIe lanes to feed it (in a multi-GPU setup) or backed by slow system RAM (in a CPU-offload setup) hits this same wall from a different direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do instead
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Option 1 — single GPU, consumer platform.&lt;/strong&gt; If you're not planning multi-GPU, this doesn't matter: one GPU at x16 on a consumer board is fine. Spend the saved money on VRAM instead of lanes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Option 2 — multi-GPU, plan the lanes first.&lt;/strong&gt; Pick the CPU/motherboard combo for lane count &lt;em&gt;before&lt;/em&gt; picking GPUs. Check the motherboard's manual for the electrical (not just physical) PCIe configuration — "x16/x16" printed on the slot doesn't mean both run at x16 simultaneously.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Option 3 — read bandwidth alongside VRAM, always.&lt;/strong&gt; When comparing GPUs, put memory bandwidth (GB/s) next to VRAM (GB) in the same table. Two cards with identical VRAM can differ 2x in tokens/sec purely on bandwidth.&lt;/p&gt;

&lt;p&gt;Here's what I landed on — single RTX 3090 now, motherboard and PSU sized so a second one is a drop-in later instead of a rebuild:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPU — used RTX 3090, 24GB, 936 GB/s.&lt;/strong&gt; Best price-per-GB on the market right now; the 4090/5090 premium isn't worth it while 5090 pricing is this unstable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CPU — any current-gen mid-range Ryzen or Intel Core.&lt;/strong&gt; One GPU at x16 uses less than half a consumer CPU's lane budget. Overspending here buys nothing a single card can use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAM — 64GB DDR5, dual-channel, non-ECC.&lt;/strong&gt; Matches VRAM headroom for OS and any CPU-offloaded layers. ECC only matters for multi-day training runs, not inference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage — 2TB NVMe Gen4.&lt;/strong&gt; 70B model files run 40-140GB each; this leaves room for a few quantizations without constantly re-downloading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Motherboard — ATX board with two full-length x16 slots, confirmed x8/x8 electrical when both are populated, and enough slot spacing for a second triple-slot card.&lt;/strong&gt; This is the one part that's expensive to get wrong later — swapping it means rebuilding the machine. Mid-to-high consumer chipsets (X670/X870, Z790/Z890) wire this correctly; budget B-series boards often don't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PSU — 850W 80+ Gold.&lt;/strong&gt; Covers the one 3090 comfortably now. Sized against the formula (GPU draw + ~400W baseline + 20% margin) so going to two 3090s later means (2 × 350W) + 400W + margin — around 1,200W — which is a PSU swap, not a surprise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cooling — ATX case, 3+ fans, front-to-back airflow.&lt;/strong&gt; Overkill for one GPU is wasted money; this only becomes a real decision at multi-GPU or sustained 24/7 load.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;x8/x8 instead of true dual x16 would cost me something if I were splitting one huge model across two cards with constant inter-GPU traffic. That's not my use case — I want pooled VRAM for a bigger model, or two models running in parallel, neither of which hammers that link continuously. For that, x8/x8 is right-sized, not a compromise I'm quietly eating.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;Before you spec a GPU, spec the path to it — the CPU's PCIe lane count decides what you're allowed to build around that GPU, not the other way around.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>homelab</category>
      <category>gpu</category>
      <category>hardware</category>
    </item>
  </channel>
</rss>
