Say a model charges $2 per million input tokens and $10 per million output tokens. Now compare two requests:
| Input tokens | Output tokens | Cost | |
|---|---|---|---|
| Short question, long answer | 1,000 | 2,000 | $0.022 |
| Long question, short answer | 3,000 | 100 | $0.007 |
The "small" prompt costs about three times more. Output tokens are the expensive half, and reply length is the number I kept forgetting to account for.
I wanted to see that math before sending anything, so I built PromptCost: a free calculator that shows token count and estimated cost for 24 models (GPT-5, Claude, Gemini, Grok, DeepSeek and others) as you type.
👉 https://promtcost.pages.dev
How the counting works
Tokenization runs in the browser. Where a model uses o200k_base, the count is exact for plain text (it leaves out the small per-message overhead Chat Completions adds). Some providers have no offline tokenizer, so for those I use o200k_base as a proxy and the page labels the result as an estimate. Treat those numbers as estimates, not invoices.
Prices are tagged too: either from the provider's own documentation, or from a third-party tracker. Prices change often, so check the provider's page before you set a budget.
Your prompt text stays in the browser. It's a client-side app on Cloudflare Pages, so there's no backend.
Your turn
It's new, so rough edges are likely. Which request type has surprised you most on an LLM bill: long outputs, big system prompts, something else? And if a price looks wrong or a model is missing, tell me in the comments and I'll fix it.
I used AI tools while building this project and writing this post.
Top comments (0)