I rent GPUs by the hour.
For a long time I couldn't tell you what a single token cost me. I knew the hourly rate, and I knew the API providers' price sheets. The number in between, my cost per million tokens on my server, I was guessing.
Turns out almost everyone self-hosting is guessing too.
The math itself is one line:
$ per million tokens = ($/hr ÷ 3600) ÷ tokens per second × 1,000,000
The hard part is the "tokens per second". It moves with everything: batch size, a vLLM flag, a quantization, the engine version, the GPU, even what else the machine is doing.
So I started measuring it properly. A few things surprised me.
The same server doesn't give the same number twice. I ran one check four times on an Ollama box I hadn't touched. It came back $8.77, $8.86, $11.26, $10.02 per million tokens. Had I changed a setting between runs three and four, I'd have happily told you it saved 11%.
One number is not a measurement. You need repeat runs, a confidence interval, and a noise floor. If a change doesn't beat the noise, the honest answer is "no winner".
When a change is real, it's obvious. On an A100 with vLLM, moving max_num_seqs from 1 to 8 took output cost from $0.746 to $0.234 per million tokens, holding across six interleaved runs. That baseline was deliberately bad. The point isn't the 68%; it's that you can finally see it.
Caching quietly lies. Send the same prompts twice and a prefix cache makes run two look cheaper than it is. Every request needs to be cold unless you're measuring the cache on purpose.
I got tired of doing this by hand, so I built a small tool that does it: measure, change one thing, measure again, and tell me CHEAPER, MORE EXPENSIVE or NO WINNER. It refuses to call a winner it can't prove, which turned out to be the whole point.
It's open source and runs on your machine:
pipx install throttle-pro
throttle check --url http://localhost:8000 --model <your-model> --gpu-hourly-rate <your $/hr>
Run it three or four times before you change anything. The first surprise is usually how much your own server wobbles.
If you're self-hosting and you've made a change you're not sure paid off, I'd genuinely like to hear about it.
Top comments (1)
The confidence-interval point is the real operational lesson. It is easy to optimize a self-hosted stack against one flattering run and then mistake noise for progress. I’d keep throughput, latency, utilization, and quality in the same benchmark record, because the cheapest token is not useful if it arrives too slowly or degrades the task enough to create rework.