DEV Community

Bilawal Builds
Bilawal Builds

Posted on

Why a short prompt can cost more than a long one

Say a model charges $2 per million input tokens and $10 per million output tokens. Now compare two requests:

Input tokens Output tokens Cost
Short question, long answer 1,000 2,000 $0.022
Long question, short answer 3,000 100 $0.007

The "small" prompt costs about three times more. Output tokens are the expensive half, and reply length is the number I kept forgetting to account for.

I wanted to see that math before sending anything, so I built PromptCost: a free calculator that shows token count and estimated cost for 24 models (GPT-5, Claude, Gemini, Grok, DeepSeek and others) as you type.

👉 https://promtcost.pages.dev

How the counting works

Tokenization runs in the browser. Where a model uses o200k_base, the count is exact for plain text (it leaves out the small per-message overhead Chat Completions adds). Some providers have no offline tokenizer, so for those I use o200k_base as a proxy and the page labels the result as an estimate. Treat those numbers as estimates, not invoices.

Prices are tagged too: either from the provider's own documentation, or from a third-party tracker. Prices change often, so check the provider's page before you set a budget.

Your prompt text stays in the browser. It's a client-side app on Cloudflare Pages, so there's no backend.

Your turn

It's new, so rough edges are likely. Which request type has surprised you most on an LLM bill: long outputs, big system prompts, something else? And if a price looks wrong or a model is missing, tell me in the comments and I'll fix it.

I used AI tools while building this project and writing this post.

Top comments (0)