DEV Community

Cover image for What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You
AI Tech News
AI Tech News

Posted on

What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You

TL;DR

Self-hosting an open-weight LLM is rarely expensive because the GPU is expensive. It is expensive because most people run the GPU at single-digit utilization. The headline rental price of an accelerator is fixed per hour, so your true cost per million tokens is set almost entirely by how many tokens you push through that hour. Continuous batching, a realistic input/output token mix, and honest utilization numbers move the price by more than 10x. This post gives you a back-of-the-envelope model you can plug your own numbers into, plus the three server settings that actually change the result.

Why is the per-hour GPU price the wrong number to quote?

When someone says "an H100 costs about 2 to 3 dollars per hour," they have quoted an input to the problem, not the answer. What you sell or consume is tokens, not GPU-seconds. The conversion factor between them is throughput, and throughput is not a property of the GPU alone. It is a property of the GPU plus the model plus the request pattern plus your server configuration.

The single formula that matters:

cost_per_1M_tokens = (gpu_hourly_cost / tokens_per_hour) * 1_000_000
tokens_per_hour    = throughput_tokens_per_sec * 3600 * utilization
Enter fullscreen mode Exit fullscreen mode

Two deployments renting the identical GPU at the identical hourly price can differ by an order of magnitude in cost_per_1M_tokens, purely because one achieves 800 tokens per second at 70 percent utilization and the other achieves 90 tokens per second at 15 percent utilization. The hardware invoice is the same. The unit economics are not.

How much does continuous batching really change the math?

A naive server processes one request, waits for it to finish generating, then starts the next. The GPU spends most of its time idle between the compute-heavy prefill and the memory-bound decode steps. Continuous batching, popularized by the vLLM project, instead keeps a running set of in-flight sequences and injects new requests into the batch as soon as slots free up, token by token. The vLLM docs describe this as iteration-level scheduling, and it is the main reason a well-tuned open server reaches throughput that feels implausible if you have only ever benchmarked batch size 1.

Here is the intuition in a tiny simulator. It is not a GPU model, it just shows why aggregate throughput climbs with concurrency until something saturates:

def tokens_per_hour(single_stream_tps, max_concurrency, scaling_efficiency):
    # scaling_efficiency < 1 captures memory-bandwidth and KV-cache limits
    effective = single_stream_tps * max_concurrency * scaling_efficiency
    return effective * 3600

for c in (1, 8, 32, 64):
    tph = tokens_per_hour(42, c, scaling_efficiency=0.55)
    cost = (2.5 / tph) * 1_000_000   # $2.5/hr GPU, illustrative
    print(f"concurrency={c:>3}  tok/hr={tph/1e6:5.2f}M  $/1M={cost:6.2f}")
Enter fullscreen mode Exit fullscreen mode

The numbers above are illustrative, not measured, but the shape is real: moving from one in-flight request to a few dozen can turn a 20-dollar price per million tokens into something close to 1 dollar. The cover chart shows the same relationship. The curve flattens once you hit a bottleneck, which is usually KV-cache memory rather than raw compute.

What caps the batch size before compute does?

The quiet ceiling is the KV cache. Every active sequence stores keys and values for all previous tokens, and that memory grows with context length and batch size. When the cache fills, the scheduler either queues new requests or preempts running ones, and your effective concurrency stops climbing.

A rough KV-cache size estimate for a transformer:

kv_bytes_per_token = 2 * num_layers * num_kv_heads * head_dim * bytes_per_elem
# factor of 2 is for keys AND values
Enter fullscreen mode Exit fullscreen mode

Two levers shrink this and therefore raise the concurrency ceiling:

  1. Grouped-query attention (fewer KV heads than query heads) directly reduces num_kv_heads, which is why many recent open models adopt it.
  2. KV-cache quantization (storing the cache in FP8 instead of FP16) halves bytes_per_elem. vLLM exposes this through kv_cache_dtype.

If your p99 latency is fine but throughput plateaus early, you are almost certainly cache-bound, not compute-bound. Measure gpu_cache_usage in the server metrics before you reach for a bigger GPU.

Does quantizing the weights lower cost, or just fit a bigger model?

Weight quantization (for example 4-bit GGUF for llama.cpp, or AWQ and GPTQ for GPU serving) does two separate things, and people conflate them.

  • It reduces the memory the weights occupy, which frees room for a larger KV cache and therefore higher concurrency. This lowers cost per token.
  • It reduces memory-bandwidth pressure during the decode phase, which is bandwidth-bound. This can raise single-stream throughput.

What it does not reliably do is cut cost on a GPU you were already underutilizing. If you serve one user at a time, a 4-bit model on an idle GPU still produces an embarrassing price per million tokens, because the denominator (tokens per hour) is still tiny. Quantization pays off when you combine it with batching so the freed memory becomes extra concurrency.

GGUF on CPU is a related but different trade. llama.cpp can serve quantized models with no GPU at all, which is attractive for low-traffic internal tools where a GPU would sit idle. The crossover point is traffic volume: below some requests-per-minute threshold, a CPU box you already own beats a rented GPU you barely use. Above it, the GPU wins decisively because its tokens-per-hour ceiling is so much higher.

How do I estimate my own number in ten minutes?

You need four inputs, three of which you can measure and one you must assume:

gpu_hourly = 2.5          # your actual rental or amortized price
measured_tps = 650        # steady-state tokens/sec under realistic load
utilization = 0.60        # fraction of the hour the GPU is actually busy
overhead = 1.15           # networking, retries, idle head start, replicas

tph = measured_tps * 3600 * utilization
cost_per_1M = (gpu_hourly / tph) * 1_000_000 * overhead
print(round(cost_per_1M, 2), "USD per 1M tokens")
Enter fullscreen mode Exit fullscreen mode

The one input people fake is measured_tps. Do not copy a vendor benchmark run at batch 256 with 128-token outputs if your workload is batch 4 with 1,000-token outputs. Output-heavy workloads spend far more time in the bandwidth-bound decode phase, so their tokens-per-hour is lower and their cost is higher. Benchmark with your own input/output length distribution or the number is fiction.

Which three settings move the result the most?

After tuning many open-model deployments, the settings with the largest effect on cost per token are consistently these:

  1. max_num_seqs (or the equivalent max concurrency). Set it too low and you leave throughput on the table. Set it too high and you thrash the KV cache and trigger preemption. Sweep it against your real traffic.
  2. gpu_memory_utilization. vLLM pre-allocates the KV cache from this fraction. Nudging it from 0.80 to 0.90 can meaningfully raise the concurrency ceiling, as long as you keep headroom for activation spikes.
  3. kv_cache_dtype set to fp8. This is often the single cheapest way to raise concurrency on memory-bound models, with a quality impact that is small for many workloads but must be validated per model.

Everything else (speculative decoding, chunked prefill, tensor parallel degree) matters, but these three are where the first large wins live.

FAQ

Is self-hosting cheaper than a hosted API?
It depends entirely on utilization. A hosted API amortizes one GPU across thousands of tenants, so it runs at high utilization you would struggle to match. Self-hosting wins on cost only when your sustained traffic keeps your own GPU busy, or when data residency and latency, not price, are the reason.

Why is my cost per token so much higher than the blog posts I read?
Almost always because those posts report peak throughput at large batch sizes with short outputs, while your workload has long outputs and modest concurrency. Re-run the benchmark with your own token distribution.

Does a bigger GPU always lower cost per token?
No. A bigger GPU only helps if you can fill it. If you are already at low utilization, a larger accelerator raises your hourly cost while the tokens-per-hour barely moves, making the unit cost worse.

Where do continuous batching and quantization overlap?
Quantization frees memory, continuous batching turns that freed memory into extra concurrent sequences. Used together they compound. Used alone on an idle GPU, neither fixes the underlying utilization problem.

How do I know if I am compute-bound or memory-bound?
Watch KV-cache usage and GPU compute utilization together. If cache usage hits the ceiling while compute sits below 100 percent, you are memory-bound and should attack the cache (GQA models, FP8 cache, shorter contexts) before buying compute.

What is a realistic utilization target?
For bursty interactive traffic, sustained 50 to 70 percent is a reasonable and honest target. Claiming 95 percent usually means you are either batching offline jobs or quoting a benchmark, not a production SLA.

Further reading

All throughput and cost figures in this post are labeled illustrative and are meant to show the shape of the relationships. Replace them with your own measured values before making a budget decision.

Top comments (0)