Your 7B Model Doesn't Need an H100: A Practical GPU Sizing Guide for LLM Inference
Researched October 2026. Sizing rules from public docs and community practice, not my own benchmarks.
The most expensive mistake in LLM inference isn't picking the wrong cloud — it's renting too much GPU. I keep seeing people spin up H100s to serve 7B chatbots. Here's a practical guide to right-sizing your GPU.
The One Formula That Matters
Weights in FP16 take roughly 2 bytes per parameter:
| Model size | FP16 weights | Minimum sensible GPU |
|---|---|---|
| 7–8B | ~14–16 GB | RTX 4090 / L4 (24 GB) |
| 13B | ~26 GB | A100 40GB |
| 34B | ~68 GB | 2× A100 40GB or 1× H100 80GB |
| 70B | ~140 GB | 2× A100/H100 80GB |
Then add ~20–30% headroom for KV cache (grows with context length and batch size). Long-context workloads can double your memory needs — that's the part most guides skip.
Quantization Changes Everything
FP8 / AWQ / GPTQ roughly halve the memory requirement:
- A 70B model in FP8 (~70 GB) can fit on a single H100 80GB
- A 13B AWQ model runs comfortably on a 24 GB card
If your accuracy budget allows it, quantization is the single biggest cost lever you have — bigger than switching cloud providers.
Throughput vs. Latency: Pick Your Bottleneck
- Chatbot (low concurrency): you need enough VRAM, that's it. A 4090 serving a 7B model at 60+ tok/s is overkill for 5 concurrent users — but cheap overkill.
- Batch/API workloads: memory bandwidth is king. This is where A100/H100 earn their premium.
Don't pay for bandwidth you won't use.
A Decision Checklist
- What's the largest model you'll serve in the next 6 months?
- FP16 or can you quantize?
- Max context length × max batch size → KV cache estimate
- Latency SLA? (real-time chat vs. background jobs)
- Spot-tolerant? (see my previous post on spot pricing)
What This Costs in Practice
Combining sizing with cloud prices (researched Oct 2026): a 7B model on a Vast.ai spot 4090 (~$0.35/hr) vs. an H100 on-demand (~$2.50+/hr) is a 7x cost difference for a workload that might not even notice.
I'm building a free tool that asks about your use case and ranks GPU options: https://qisuanai.com/advisor
What are you serving, and on what? Curious where these rules break down in practice.
Top comments (0)