DEV Community

Cover image for Five Ways Your LLM Serving Benchmark Is Lying to You (and How to Catch Each One)
AI Tech News
AI Tech News

Posted on

Five Ways Your LLM Serving Benchmark Is Lying to You (and How to Catch Each One)

TL;DR

In my benchmarking work on vLLM serving stacks, every one of the following produced a confident, wrong conclusion before anyone noticed:

  1. Prefix caching inflated an A/B test unequally. The harness reused the same prompt on every call. Cache hit rate was 67% on one arm and 45% on the other, and paired throughput came out 2.0 to 2.8x higher than a sweep run on the identical server.
  2. The harness counted SSE chunks, not tokens. Under speculative decoding one chunk carries several tokens, so measured output was cut nearly in half (698 vs 362). On a 1:1 chat workload that error is about 20% of total throughput: enough to report a +49% win as a loss.
  3. Identical CLI flags, different runtime. A config.json inside one model folder declared an 8-bit KV cache scheme, and vLLM silently enabled FP8 KV. KV capacity was 1.775x larger on that arm (940,736 vs 529,856 tokens).
  4. Five knobs changed at once. Result: -16% to -22% on long-context reads, with no way to tell which knob did it.
  5. A "stability-first" auto-selection gate picked the slow config. It was the only candidate with run-to-run spread under 3% because it was 20% slower everywhere.

Bonus: Restart=always turned a configuration bug into a silent crash loop that systemctl is-active reported as healthy.

Each section below gives the symptom, the mechanism, and a code-level fix.


Why does prefix caching inflate vLLM benchmarks, and why unequally?

vLLM enables Automatic Prefix Caching by default in recent V1 releases (docs). If two requests share a prefix, the second one skips prefill for the cached blocks. That is great in production and poison in a synthetic benchmark.

Here is what happened. A team I worked with ran a paired A/B of two serving configs and wanted one shot per configuration, so they used --reps 1 --warm 0. The harness derived the prompt seed from the batch index b. With one rep and no warmup, b was always 0. Every call sent the exact same prompt.

Measured prefix cache hit rate:

Prefix cache hit rate per arm

The clue was a ratio that made no sense: on the same server with the same settings, paired-comparison throughput was 2.0 to 2.8x the sweep throughput. At concurrency 32 on a long-context workload, the paired run reported 61,347 tok/s against 22,159 tok/s in the sweep.

The part most people miss is that the contamination is not symmetric. Cache capacity is whatever KV memory is left over, so the arm with more free KV blocks keeps cached prefixes alive longer and gets a larger bonus. You stop measuring "which config is faster" and start measuring "which config retains cache better". If the two arms have different memory headroom, that alone biases the result.

Fix: seed every call independently and log the hit rate

import random

LARGE_PRIME = 1_000_003

def make_prompt(seed: int, batch: int, idx: int, n_words: int) -> str:
    # Separate axes for run seed, batch and request index, so
    # no two calls in any phase or concurrency level share a prefix.
    rng = random.Random(seed * LARGE_PRIME + batch * 10_000 + idx)
    # Put the randomness at the START of the prompt; a shared system
    # prompt in front would still produce prefix hits.
    nonce = f"[req {rng.getrandbits(64):016x}] "
    words = [rng.choice(VOCAB) for _ in range(n_words)]
    return nonce + " ".join(words)
Enter fullscreen mode Exit fullscreen mode

Then prove it. Scrape the server's Prometheus endpoint before and after each run and record the delta as part of the result artifact:

import re, requests

def prefix_cache_counters(base_url: str) -> dict:
    text = requests.get(f"{base_url}/metrics", timeout=5).text
    out = {}
    for name in ("vllm:prefix_cache_queries_total", "vllm:prefix_cache_hits_total"):
        m = re.search(rf"^{re.escape(name)}(?:{{[^}}]*}})?\s+([0-9.e+]+)$", text, re.M)
        out[name] = float(m.group(1)) if m else float("nan")
    return out
# hit_rate = d(hits) / d(queries) over the run window
Enter fullscreen mode Exit fullscreen mode

(Metric names vary across vLLM versions; check your /metrics output.) If you want a cache-free number, start the server with --no-enable-prefix-caching, but still log the counters: "we disabled it" is a claim, the counter is evidence.

One more lesson from that week: before you discount a competitor's published numbers because "they use prefix caching and we don't", print your own hit rate. In this case the team had made exactly that argument while running at 67%.


Why does my tokens-per-second drop when I enable speculative decoding?

Sometimes it doesn't drop. Your counter does.

A common streaming harness pattern:

async for chunk in stream:          # OpenAI-compatible SSE stream
    if chunk.choices[0].delta.content:
        ntok += 1                   # BUG: counts chunks, not tokens
Enter fullscreen mode Exit fullscreen mode

With normal decoding, one SSE event roughly equals one token, so this works by accident. Speculative decoding (EAGLE, MTP, n-gram drafting; see the vLLM speculative decoding docs) accepts several tokens per step, and the server emits them in a single chunk. The harness undercounts only the accelerated arm.

In the case I saw, measured output dropped from 698 to 362 tok/s for the speculative arm while total throughput moved only 6%. Those two numbers cannot both be true: you cannot halve output and keep total flat. That contradiction was the entrance to the bug.

Why didn't it show up sooner? The test was long-context with a 32:1 input:output ratio, so output tokens were a tiny share of the total and the error on total was about 1%. On a 1:1 chat-shaped workload, the same bug costs about 20% of total throughput. The real speculative result there was +49%; a 20% haircut on one arm is big enough to flip the headline.

Long-context throughput: the cache artifact

Fix: count tokens from the source of truth

payload = {
    "model": MODEL,
    "prompt": prompt,
    "max_tokens": OUT_LEN,
    "ignore_eos": True,             # vLLM extension: fixes output length
    "stream": True,
    "stream_options": {"include_usage": True},
}
# The final SSE event carries usage.completion_tokens. Use it.
# Cross-check: with ignore_eos, completion_tokens must equal OUT_LEN.
Enter fullscreen mode Exit fullscreen mode

Because ignore_eos=True plus max_tokens=N pins the output length, you can often recompute historical results without rerunning: true_total = in_tok_per_s * (in_len + out_len) / in_len, provided input token counts were exact. Apply the correction to every arm with the same formula. Correcting only the arm you suspect is a new contamination.

General rule: any counter that answers "how much" by counting "how many times" (chunks, callbacks, events) breaks under acceleration techniques, because those techniques work precisely by packing more into each event. Batching, speculative decoding and stream coalescing all do this, and each one breaks a different counter in a different direction. Log at least two quantities that must agree with each other.


Can identical vLLM command-line flags produce different runtime configs?

Yes, and command-line diffs will never show it.

Two quantized checkpoints of the same base model were served with scripts identical except for model path, port and GPU index. The comparison was presented as a clean weights-only A/B. Then someone diffed the startup logs:

Arm KV cache dtype (from log) KV cache capacity (tokens)
A kv_cache_dtype=auto (16-bit) 529,856
B Using standard fp8 KV cache 940,736 (1.775x)

The cause: arm B's config.json carried a quantization block with a kv_cache_scheme of 8 bits. vLLM reads quantization metadata from the checkpoint and enabled FP8 KV on its own (quantized KV cache docs). Arm A had no such entry. Same flags, different conditions.

KV cache capacity with identical CLI flags

The same report had a second confound: one row was "weights B + default settings", another was "weights A + tuned settings", and the conclusion said "tuning lost". Two axes moved at once; that sentence was never proven.

Fix: diff the realized config, not the intended one

# Run after both servers are up. Compare what the engine actually did.
for f in serverA.log serverB.log; do
  echo "== $f"
  grep -Ei "kv_cache_dtype|fp8 KV|GPU KV cache size|Maximum concurrency|attention backend|quantization=|tensor_parallel_size|max_num_seqs|enable_prefix_caching" "$f"
done

# And check the checkpoint itself before you start:
jq '.quantization_config | {quant_method, kv_cache_scheme}' modelB/config.json
Enter fullscreen mode Exit fullscreen mode

If you want to force parity, pass --kv-cache-dtype explicitly on both arms rather than relying on auto. And label every row of a results table with both (weights, config), never just one.

A useful habit: when a difference appears, check whether it is a number you already know. That 1.775x matched an earlier experiment where increasing KV capacity by about 78% did not raise the concurrency ceiling at all. So "B wins because it has more KV" was ruled out in minutes.


Why should you tune vLLM settings one at a time?

An "apply the recommended profile" change modified five settings together: FP8 KV cache, --gpu-memory-utilization 0.95, lower --max-model-len, higher --max-num-seqs, and async scheduling. Result: -16% to -22% across every long-context point. Nobody could say which knob caused it, and each server restart cost about 10 minutes, so bisecting was expensive.

Serving knobs trade against each other (memory you gain, compute you may lose), so their effects have mixed signs and do not add up neatly.

Two things made the failure useful anyway:

  • The failure named the bottleneck. Growing KV capacity by 78% (523,264 to 933,120 tokens) left the concurrency-64 throughput ceiling unchanged. Therefore the bottleneck was compute, not memory. A success would never have told us that.
  • The same setting flips sign by workload. The identical profile was -22% on prefill-dominated long-context reads and +6.8% on a balanced read/write workload.

Same tuning, opposite sign by workload

Fix: one axis per run, and ask the code before restarting

# Baseline, then vary exactly one flag per run.
BASE="--max-model-len 32768 --max-num-seqs 128 --gpu-memory-utilization 0.90"
for VARIANT in "" "--kv-cache-dtype fp8" "--max-num-seqs 256" "--async-scheduling"; do
  vllm serve "$MODEL" $BASE $VARIANT --port 8000 > "log_${VARIANT// /_}.txt" 2>&1 &
  wait_for_ready && run_bench --shape longctx && run_bench --shape chat
  kill %1; wait
done
Enter fullscreen mode Exit fullscreen mode

Flag names drift across releases, so check vllm serve --help for your version. For expensive restarts (for example, choosing between attention backends) try to answer feasibility in-process before launching a server: in one case, calling the backend's support-check function directly classified four candidates in about a minute, after two full restarts had been wasted trying them one by one.

Always benchmark at least two workload shapes (prefill-heavy and decode-heavy). A single-shape result is half a result.


Why do "stability-first" auto-tuning gates pick the slowest config?

An auto-selection script chose among three batch-limit candidates with this rule:

Among candidates whose run-to-run spread is under 3% of peak, pick the one with the highest peak.

Only one candidate passed (spread 0.1%). The other two had spreads around 20%. The winner was stable because it was 20% slower at every point. Throttling a system removes variance along with throughput. With no performance floor, "do nothing" is always the most stable option.

The resulting combined profile lost to the best single change on four of five measured points: -13.0%, +2.0%, -3.1%, -3.3%, -13.9%.

Auto-selected combo vs best single change

The same gate had a second hole: "adopt if it wins on both shapes". Actual margins were +0.0% to +0.5%, all inside the noise, and they were adopted because the sign was positive.

This is the same mistake at a smaller scale that I have seen in model evaluation: a probe at 89.5% vs a baseline at 89.1% was marked "pass" because A > B. On 276 questions the standard error was about 1.9 points, the bootstrap 95% interval for the difference was [-4.3, +5.1] points, and McNemar z was 0.00. Not a win.

Fix: gates need a floor and an interval, written in code

import numpy as np

def bootstrap_ci(a, b, n=10_000, seed=0):
    """Paired bootstrap of mean(a - b); a, b are per-run (or per-item) values."""
    rng = np.random.default_rng(seed)
    d = np.asarray(a) - np.asarray(b)
    idx = rng.integers(0, len(d), size=(n, len(d)))
    means = d[idx].mean(axis=1)
    return np.percentile(means, [2.5, 97.5])

def accept(candidate_runs, baseline_runs, baseline_peak):
    peak = max(candidate_runs)
    spread = (max(candidate_runs) - min(candidate_runs)) / peak
    lo, _ = bootstrap_ci(candidate_runs, baseline_runs)
    return (
        spread < 0.03
        and peak >= 0.95 * baseline_peak   # performance floor
        and lo > 0                          # interval, not sign
    )
Enter fullscreen mode Exit fullscreen mode

Two more rules: always measure the auto-selected combination next to the best single change (combos can be worse than their parts), and do not assume more knobs means more speed. Of the five knobs tested that day, one helped; the rest were neutral or harmful.


Can systemd Restart=always hide a broken benchmark worker?

Yes. Restart=always is medicine for transient failures. A configuration error is permanent, and the restart policy disguises it as a retry.

In one fleet, an automatic rebalancer assigned zero work slots to three workers. The generated unit line ended up with an empty argument, which the shell dropped entirely. The worker script ran with set -u, hit $4: unbound variable, and exited immediately. systemd plus a watchdog revived it every 15 seconds, forever. Load average went to 67 with every GPU at 0%, and systemctl is-active said active.

A related trap: a watchdog that "restores the service when the queue is idle" cannot tell idle from "I deliberately stopped things to run a benchmark". In one case it revived a 50 GB service onto the GPU under test, and the control server got roughly one tenth of its normal KV capacity (89,536 vs 919,232 tokens). Left alone, that would have produced a headline claiming speculative decoding was 10x faster.

Fix

[Unit]
StartLimitIntervalSec=300
StartLimitBurst=5          # after 5 failures in 5 min, stay failed and alert

[Service]
Restart=on-failure
RestartSec=15
ExecStartPre=/usr/local/bin/validate-args.sh %i   # fail loudly on empty args
Enter fullscreen mode Exit fullscreen mode

Validate generated argument lists at generation time (raise on empty values), disable watchdogs and auto-restorers before a measurement window, and record nvidia-smi memory per GPU in the benchmark artifact so a stray tenant shows up in the data.


Pre-flight checklist for LLM serving benchmarks

  • [ ] Every request has a unique prompt prefix; seeds use separate axes for run, batch and request index.
  • [ ] Prefix cache hit rate is scraped from /metrics and stored with the results.
  • [ ] Output tokens come from usage.completion_tokens, not chunk counts; ignore_eos plus max_tokens pins length.
  • [ ] At least two quantities in each run must agree (output tok/s, total tok/s, completion_tokens x requests / wall time).
  • [ ] Startup logs of both arms are diffed for KV dtype, KV capacity, attention backend, parallelism and max sequences.
  • [ ] Checkpoint config.json is inspected for quantization_config and KV cache schemes.
  • [ ] Results table rows state both weights and config.
  • [ ] One knob per run; at least one prefill-heavy and one decode-heavy shape.
  • [ ] Acceptance gates use a confidence interval lower bound above zero, plus a performance floor relative to baseline.
  • [ ] Auto-selected combos are benchmarked against the best single change.
  • [ ] Watchdogs off during measurement; per-GPU memory snapshot recorded; restart policies have burst limits.

FAQ

Should I disable prefix caching for every vLLM benchmark?

Not necessarily. If production traffic shares system prompts, a cache-on benchmark is more realistic. What matters is that both arms see the same cache conditions and that you report the measured hit rate. Unique prompts with caching left on, plus a logged hit rate near zero, is a clean way to measure raw engine speed.

How do I count tokens correctly with streaming responses?

Request stream_options: {"include_usage": true} and read usage.completion_tokens from the final event. If you need per-token timing, tokenize the accumulated text with the model's tokenizer instead of counting events, and validate against the usage field.

How can I tell if vLLM enabled FP8 KV cache on its own?

Grep the startup log for the KV cache dtype and the reported KV cache size in tokens, and inspect quantization_config in the checkpoint's config.json. Passing --kv-cache-dtype explicitly on every arm removes the ambiguity.

How many repetitions do I need before declaring a winner?

Enough that the confidence interval of the difference excludes zero. Start with at least five runs per arm, compute a paired bootstrap interval, and treat any gap smaller than the observed run-to-run spread as a tie.

Is a lower-variance config always better for production?

No. Low variance is easy to buy by throttling throughput. Require a performance floor (for example, at least 95% of the best peak) before stability becomes a tiebreaker.

Why did my tuning help chat traffic but hurt long documents?

Prefill-dominated and decode-dominated workloads hit different bottlenecks. In the case above the same profile measured -22% on long-context reads and +6.8% on a balanced workload. Benchmark both shapes and report them separately.

Top comments (1)

Collapse
 
theagentloop profile image
The Agent Loop •

This is the most useful list of benchmark traps I've read, and the systemctl is-active bonus is the one I'd have believed without evidence.

There's a sixth trap missing from the list, and it sits outside all five of yours. Every one of your examples is a bias inside a run that completed. The one that worries me is the runs that never completed at all, because they're invisible by construction.

If your test rig loses the host to an OOM kill, drops the connection, or gets its process reaped, that run produces no result. And no result is very easy to drop from the denominator. The A/B still comes out, still looks clean, still gets a p-value. You just measured the subset of runs that survived being measured. Every completion-rate and throughput number you report inherits that bias, and the bias isn't random: it favours whichever arm is more likely to kill the host, which is usually the one doing more work per request.

We keep our own logs because of how invisible this is. Ten separate agent-host deaths on our box from Chromium workloads, and the first eight produced no error message at all. No traceback, nothing in the log, the process simply gone. Same class as your crash loop: a health check reporting fine because the thing that was supposed to say otherwise had already stopped. If absence of an error were treated as evidence of success, we would have logged eight clean runs.

The cheap guard, if you want one: make "no result" a counted outcome rather than a filtered one. Log an explicit failure row when the expected event doesn't arrive, then assert on the count of those rows before you compare anything. It's the difference between a rig that reports what happened and one that reports what survived.

Also worth saying since you did: the stability-first gate picking the slow config is the same failure as the crash loop wearing a lab coat. Both are "the metric that was supposed to be a proxy started measuring the proxy."