You cannot control what you do not measure — and GPT APIs bill you for something you never see.
Every request is priced, rate-limited, and truncated by token count, not by words, characters, or number of calls. If you only learn the real number from the API response, you already paid for it. Counting GPT tokens before you send gives you three things: a cost you can quote, a context window you know you fit inside, and a prompt you can trim while trimming is still cheap.
Here is how to do it in Python and JavaScript, which encoding to use, and what your request actually contains.
Tokens are not words
A token is a subword unit from a byte-pair encoding (BPE) tokenizer. Common English words usually become one token. Less common ones get split — tokenization becomes token + ization, two tokens for one word.
For rough planning, English prose runs about 4 characters per token, or roughly 0.75 words per token. That average breaks the moment the input stops looking like English prose:
- Identifiers, URLs, file paths — a UUID is several tokens; every path segment splits separately
- Code — indentation, punctuation, and operators each cost tokens
-
Numbers — digits are often split into groups, so
2026may not be one token - CJK text — Chinese characters commonly cost 1–2 tokens each, so character heuristics undercount badly
- Emoji and accents — expand into multiple tokens once UTF-8 encoded
- Pretty-printed JSON — indentation across a big object adds up fast; collapse it before the call
You also cannot estimate in the other direction. A character or word count gives you a ceiling, not the billed number.
Which encoding your model uses
Counts differ between tokenizer generations. Using the wrong one gives you a plausible number that is off by 10–20% on English text and much more on anything else.
| Encoding | Vocab (approx.) | Used by |
|---|---|---|
o200k_base |
~200,000 | gpt-4o, gpt-4o-mini, gpt-4.1 family, gpt-5 family, o1, o3, o3-mini, o4-mini |
cl100k_base |
~100,000 | gpt-4, gpt-4-turbo, gpt-3.5-turbo, text-embedding-3-small/large |
p50k_base |
~50,000 | text-davinci-002/003, code-davinci-002 |
r50k_base (gpt2) |
~50,000 | gpt-2, davinci, curie, babbage, ada |
You rarely need to memorize it — tiktoken resolves the encoding from the model name by prefix matching, so anything starting with gpt-5-, gpt-4o-, gpt-4.1-, o1-, o3-, or o4-mini- resolves to o200k_base.
import tiktoken
for model in ("gpt-4o", "gpt-4o-mini", "gpt-4.1", "gpt-5-mini", "o3-mini", "gpt-4", "gpt-3.5-turbo"):
print(model, tiktoken.encoding_for_model(model).name)
Count GPT tokens in Python
pip install tiktoken
import tiktoken
enc = tiktoken.encoding_for_model("gpt-4o-mini") # resolves to o200k_base
text = open("prompt.txt", encoding="utf-8").read()
print(len(enc.encode(text)), "tokens")
It runs offline — no API key, no network request after the encoding file is cached.
When a count surprises you, look at how the string actually splits:
ids = enc.encode("Unbelievable")
print(ids)
print([enc.decode_single_token_bytes(t) for t in ids])
# [59026, 17536] -> [b'Un', b'believable']
Counting a whole chat request
Each message carries structural overhead beyond its content. This approximation follows OpenAI's cookbook: roughly 3 tokens per message, plus 1 if you set a name, plus 3 priming the assistant reply.
import tiktoken
def count_chat_tokens(messages, model="gpt-4o-mini"):
try:
enc = tiktoken.encoding_for_model(model)
except KeyError:
enc = tiktoken.get_encoding("o200k_base") # fallback for new model names
tokens_per_message, tokens_per_name, total = 3, 1, 0
for message in messages:
total += tokens_per_message
for key, value in message.items():
if isinstance(value, str):
total += len(enc.encode(value))
if key == "name":
total += tokens_per_name
total += 3 # assistant reply primer
return total
Treat it as a planning figure. The authoritative count is always the usage object in the response — log it and reconcile against your estimate.
Count GPT tokens in JavaScript
Counting where the prompt is built lets you warn users before they send something enormous, without shipping their text anywhere.
npm install js-tiktoken
import { encodingForModel } from "js-tiktoken";
const enc = await encodingForModel("gpt-4o");
console.log(enc.encode("Count me before I cost you money.").length, "tokens");
gpt-tokenizer is the other common option and ships per-model entry points, which helps when bundle size matters:
import { encode } from "gpt-tokenizer";
const n = encode(text).length;
The tokenizer is just a data table and a merge algorithm, so none of this needs a network call.
What your request actually contains
The user message is usually the smallest part of the bill. All of this counts against the same context window and the same input price:
- System prompt — resent every request, so length here is a permanent tax
- Conversation history — grows quadratically with turns if replayed in full
- Tool/function schemas — twenty tools with detailed descriptions can occupy thousands of tokens even when none are called
- Retrieved context — every pasted chunk
- Reasoning tokens — on reasoning models these occupy context and bill as output
-
The reply — output usually costs several times more per token than input, and
max_tokensis a cap, not a budget
Budgeting the context window
The window is shared, not per-message:
free_tokens = context_window
- system + history + tools + retrieved_chunks
- max_tokens_you_allow_for_the_reply
If that lands near zero you get either a rejected request or silently dropped messages — your "memory" vanishes mid-conversation. Keep 10–15% slack so unusual inputs do not tip you over. Better: size retrieval from the count. After counting the fixed parts, whatever remains is exactly how many tokens of context you can afford.
Turning tokens into money
cost = (input_tokens / 1_000_000) * input_price
+ (output_tokens / 1_000_000) * output_price
Standard-tier OpenAI list prices, October 2026 (always re-check the official pricing page):
| Model | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| gpt-5-nano | $0.05 | $0.005 | $0.40 |
| gpt-5-mini | $0.25 | $0.025 | $2.00 |
| gpt-4o-mini | $0.15 | $0.075 | $0.60 |
| gpt-4o | $2.50 | $1.25 | $10.00 |
| o3 | $2.00 | $0.50 | $8.00 |
Concrete case: 10,000 support answers per day, 600-token prompt, 300-token reply.
- gpt-4o-mini: 6M input × $0.15 = $0.90/day + 3M output × $0.60 = $1.80/day → $2.70/day (~$81/month)
- gpt-4o: 6M × $2.50 = $15.00/day + 3M × $10.00 = $30.00/day → $45.00/day (~$1,350/month)
Same prompt, same answer quality for a routing task, 16× difference. This is why counting comes before model selection — you cannot compare prices without knowing volume.
Six traps that inflate prompts
- Replaying the whole history — summarize old turns or use a rolling window
- Minifying JSON by hand — whitespace is real money at scale; strip it, drop unused field names
- Registering every tool on every request — send only this turn's schemas
- Trusting one "average tokens per request" number — track mean, p90, p99; the p99 breaks your budget
- Trusting character heuristics for non-English text — 4 chars/token is an English figure
-
Setting
max_tokensto the max — an uncapped reply is the classic surprise bill
Checklist before every production request
- Count GPT tokens with the encoding of the model you actually serve
- Include system prompt, history, tools, and retrieved chunks — not just the user message
- Leave headroom: prompt + expected reply < context window
- Log response
usageand compare against your estimate weekly - Hard-cap both input length and
max_tokens - Re-check prices quarterly and after any model-family switch
This post is the written-up version of how I actually budget prompts before shipping anything that calls an LLM. The full guide with extra sections is on my site: How to Count GPT Tokens Before You Hit the API.
If you work in the browser, two tools from the same toolkit are useful alongside this: a JSON Formatter & Validator for collapsing whitespace out of payloads, and a Word & Character Counter for the fast ceiling estimate. Both run locally — nothing uploaded.
How do you budget tokens in your pipeline — pre-flight counting, response usage logging, or something else?
Top comments (0)