I rented an A100 for under an hour to answer one question.
Does a single vLLM flag really change what a token costs?
The flag was max_num_seqs: how many requests the server works on at once. I set it to 1, which is a deliberately bad setting, and then to 8. Same GPU, same model (Qwen2.5-0.5B), same prompts, eight requests in flight the whole time.
The obvious way to test this is: run the old setting, run the new one, compare. I don't trust that anymore. I once watched a server I hadn't touched "get" 27% more expensive between two runs. GPUs warm up, other work comes and goes, and a single before/after quietly bakes all of that into your answer.
So I interleaved it. Six runs, alternating the order:
baseline → change → baseline
change → baseline → change
If the machine drifts during the session, the drift lands on both sides instead of making one look better.
Here's what came back, in dollars per million output tokens at $1.39/hr:
max_num_seqs=1 $0.748 $0.746 $0.745
max_num_seqs=8 $0.237 $0.229 $0.237
Six runs, two tight clusters, no overlap. About 68% cheaper per token, and it held every single time.
Two honest caveats:
- A baseline of 1 is unrealistically bad. Nobody should run production like that. The point was to prove the method, not to promise you 68%.
- It's a 0.5B model. Bigger models, longer prompts and real traffic will give you a different number. That's exactly why you measure your own.
What surprised me most wasn't the drop. It was how calm the numbers became once the runs were interleaved. Same server, same day, no drama.
I turned this into a small open-source tool so I'd stop doing it by hand. It measures cost per million tokens with confidence intervals, and it only says CHEAPER when a change actually beats the noise:
pipx install throttle-pro
All six raw runs are public, if you want to check my math.
Top comments (0)