TL;DR
The KV cache is often the real memory ceiling in LLM serving, not the weights. Quantizing it to FP8 or INT8 roughly halves KV memory, which lets you serve longer contexts or more concurrent requests on the same GPU. The catch: the quality impact is real but small when measured correctly, and it is easy to measure it wrong. A single kv_cache_dtype flag, or a scaling factor baked into a model's config.json, can quietly change your outputs while every dashboard still shows green. This post covers where the savings come from, what actually degrades, the silent-config trap, and a fair measurement protocol.
Why is the KV cache the thing you should quantize?
When a transformer generates tokens autoregressively, it stores the key and value tensors for every past token so it does not recompute them. That store is the KV cache, and it grows linearly with sequence length and batch size.
A rough per-token KV size:
bytes_per_token = 2 (K and V)
* num_layers
* num_kv_heads
* head_dim
* dtype_bytes
For a mid-size model with 32 layers, 8 KV heads (grouped-query attention), and head_dim 128, in FP16 that is:
2 * 32 * 8 * 128 * 2 = 131,072 bytes (128 KB per token)
At 8K context that is 1 GB per sequence, before you batch anything. Weights are a fixed cost you pay once; the KV cache is a per-request cost that scales with your traffic. That is why it is the highest-leverage place to spend a quantization budget. Drop the dtype from 2 bytes to 1 byte and you halve that 128 KB to 64 KB per token, which directly buys you longer contexts or a bigger batch.
FP8 vs INT8 for the KV cache: what is the real tradeoff?
Both formats store each cached value in 8 bits, so the memory savings are effectively identical, about 2x versus FP16/BF16. The difference is in how they represent numbers and what that does to accuracy.
FP8 (typically e4m3 or e5m2) keeps an exponent field, so it covers a wide dynamic range with graceful precision falloff. The vLLM docs note that FP8 KV cache is supported on CUDA 11.8 and later and on ROCm, and that e4m3 is the common default on NVIDIA hardware. Because attention keys and values can have occasional large-magnitude outliers, the exponent bits help: an outlier does not blow out the whole tensor's scale.
INT8 uses a uniform grid with a per-tensor or per-channel scale. When the distribution is well behaved, INT8 can actually be more precise than FP8 inside its range because it does not spend bits on an exponent. The risk is outliers: one large value forces a coarse scale on everything else. INT8 KV usually wants calibration to pick good scales.
Practical reading:
- FP8 is the lower-friction default. On recent NVIDIA and AMD GPUs it often needs no calibration and tends to be robust to outliers.
- INT8 can match or beat FP8 quality with good per-channel scales, and it is useful where FP8 hardware paths are missing, but it asks more of you up front.
How much quality do you actually lose?
Less than people fear, if you measure the right thing. KV quantization only touches the cache, not the weights or the activations on the compute path, so the error is confined to attention's read of past tokens. In practice:
- Short-context, factual tasks: differences are usually within benchmark noise.
- Long-context retrieval and exact-copy tasks: this is where FP8/INT8 KV shows up first, because tiny per-token errors accumulate across thousands of cached positions.
- Reasoning chains: a small per-step perturbation can flip a final answer even when average token accuracy barely moves, so pass@1 on reasoning sets is more sensitive than perplexity.
The headline number ("FP8 KV costs X percent") is meaningless without naming the task, the context length, and whether you measured a distribution or a single greedy run.
The silent config.json trap
Here is the failure that burns teams. You benchmark model A and model B with what looks like the identical serving command, same flags, same prompts, and you get different accuracy. The flags were identical. The conditions were not.
Two things commonly hide inside a model directory rather than on your command line:
A checkpoint-level KV cache dtype. Some published checkpoints ship a quantization config in
config.json(for example aquantization_configor a KV-specific field) that sets the cache to FP8 even when your launch command says nothing about it. Your serving engine reads it and honors it. Your benchmark harness never sees it.Baked-in KV scaling factors. FP8 KV can use static per-layer scales stored in the checkpoint. If one checkpoint has calibrated scales and another falls back to default scales, their attention precision differs even though both report "fp8 KV" in the logs. vLLM exposes
--calculate-kv-scalesto compute scales at runtime for thee4m3path precisely so you are not silently depending on whatever shipped in the file.
The lesson that keeps recurring: identical flags are not identical conditions. The authoritative record of what ran is the resolved config inside the model folder plus the engine's startup log, not the command you typed.
A one-line guard before trusting any KV benchmark:
# Dump the resolved KV settings the engine actually used
python - <<'PY'
import json, glob
for f in glob.glob("models/**/config.json", recursive=True):
c = json.load(open(f))
qc = c.get("quantization_config", {})
print(f, "| kv:", c.get("kv_cache_dtype"), "| qc_kv:", qc.get("kv_cache_dtype"))
PY
Then confirm against the server's own report:
vllm serve <model> --kv-cache-dtype fp8_e4m3 --calculate-kv-scales 2>&1 \
| grep -iE "kv.cache|kv_scale|quantiz"
If the two disagree, the file wins, and your benchmark was comparing two different things.
How do you measure KV quantization fairly?
Treat it as an A/B where the only thing allowed to change is the KV dtype. Everything else gets pinned.
# Fair KV-quant comparison: pin everything except kv_cache_dtype
from vllm import LLM, SamplingParams
COMMON = dict(
model="your/model",
max_model_len=8192,
gpu_memory_utilization=0.90,
seed=1234, # pin the seed
enforce_eager=True, # avoid graph-capture variance while measuring
)
sp = SamplingParams(temperature=0.0, max_tokens=512) # greedy, fixed budget
baseline = LLM(**COMMON, kv_cache_dtype="auto") # FP16/BF16 KV
fp8 = LLM(**COMMON, kv_cache_dtype="fp8_e4m3", calculate_kv_scales=True)
Rules that make the result trustworthy:
- Pin the seed and the decoding params. Compare greedy to greedy, or run N samples per prompt and compare distributions. Never compare one greedy run to one sampled run.
-
Give reasoning models room. A
max_tokensthat is too small truncates the chain and manufactures a fake quality gap that has nothing to do with the KV cache. - Disable prefix caching, or warm both runs identically. A shared prefix cache can make the second configuration look faster or different for reasons unrelated to quantization.
- Measure where it bites. Include a long-context retrieval task and a reasoning set, not just short QA. If your eval cannot tell FP16 from FP8, your eval is too easy, not your quantization too good.
- Report two standard errors, not a point. KV-quant effects are small. A 0.4 point difference with an SE of 0.6 is noise, not a regression.
- Separate memory from quality. Show the KV memory you saved and the accuracy you paid in the same table, because the whole point is the tradeoff.
A minimal results table makes the decision obvious:
| Config | KV bytes/token | Max batch at 8K | Long-ctx retrieval | Reasoning pass@1 |
|---|---|---|---|---|
| FP16 KV | 128 KB | 1x | 94.1 +/- 0.5 | 71.2 +/- 0.6 |
| FP8 e4m3 KV | 64 KB | ~2x | 93.8 +/- 0.5 | 70.9 +/- 0.6 |
(Numbers above are illustrative placeholders; fill them from your own run.)
If the accuracy columns overlap within two standard errors and the batch column doubles, FP8 KV is a free win for that workload. If the long-context column drops outside the error bars, you have found the ceiling for that task and should keep FP16 KV there.
FAQ
Does KV cache quantization quantize the model weights too?
No. It only changes how cached keys and values are stored. Weights and the active compute path are untouched unless you separately apply weight quantization. That is why the quality impact is usually smaller than full weight quantization.
FP8 or INT8, which should I pick first?
Start with FP8 e4m3 on recent NVIDIA or AMD GPUs. It is robust to attention outliers and often needs no calibration. Reach for INT8 when FP8 hardware paths are unavailable or when you have invested in good per-channel calibration.
Why did two checkpoints give different accuracy with the same command?
Almost always a setting inside the model folder. A config.json KV dtype or baked-in KV scales override your intent silently. Dump the resolved config and the engine startup log before trusting the comparison.
What does --calculate-kv-scales do?
For the FP8 e4m3 path it computes the KV scaling factors at runtime instead of relying on whatever scales shipped in the checkpoint, which keeps your comparison honest and avoids depending on an unknown baked-in value.
Will KV quantization make serving faster, or just smaller?
Primarily smaller, which indirectly raises throughput: a smaller KV cache means a larger batch or longer context fits, and higher batch occupancy is where serving throughput actually comes from. Do not expect a large per-token latency drop by itself.
How do I know my eval is sensitive enough to detect a regression?
Include a task you know FP16 passes and a weak baseline fails. If FP8 KV and FP16 KV score identically on everything, verify the eval can discriminate at all before concluding there is no loss.
Further reading
- vLLM documentation: FP8 KV cache and
kv_cache_dtypeoptions (docs.vllm.ai) - vLLM documentation:
--calculate-kv-scalesand quantization support matrix (docs.vllm.ai)
Top comments (0)