DEV Community

Cover image for How to Count GPT Tokens Before You Hit the API
Zhang Enquan
Zhang Enquan

Posted on Edited on Originally published at aisubtools.xyz

How to Count GPT Tokens Before You Hit the API

You cannot control what you do not measure — and GPT APIs bill you for something you never see.

Every request is priced, rate-limited, and truncated by token count, not by words, characters, or number of calls. If you only learn the real number from the API response, you already paid for it. Counting GPT tokens before you send gives you three things: a cost you can quote, a context window you know you fit inside, and a prompt you can trim while trimming is still cheap.

Here is how to do it in Python and JavaScript, which encoding to use, and what your request actually contains.

Tokens are not words

A token is a subword unit from a byte-pair encoding (BPE) tokenizer. Common English words usually become one token. Less common ones get split — tokenization becomes token + ization, two tokens for one word.

For rough planning, English prose runs about 4 characters per token, or roughly 0.75 words per token. That average breaks the moment the input stops looking like English prose:

  • Identifiers, URLs, file paths — a UUID is several tokens; every path segment splits separately
  • Code — indentation, punctuation, and operators each cost tokens
  • Numbers — digits are often split into groups, so 2026 may not be one token
  • CJK text — Chinese characters commonly cost 1–2 tokens each, so character heuristics undercount badly
  • Emoji and accents — expand into multiple tokens once UTF-8 encoded
  • Pretty-printed JSON — indentation across a big object adds up fast; collapse it before the call

You also cannot estimate in the other direction. A character or word count gives you a ceiling, not the billed number.

Which encoding your model uses

Counts differ between tokenizer generations. Using the wrong one gives you a plausible number that is off by 10–20% on English text and much more on anything else.

Encoding Vocab (approx.) Used by
o200k_base ~200,000 gpt-4o, gpt-4o-mini, gpt-4.1 family, gpt-5 family, o1, o3, o3-mini, o4-mini
cl100k_base ~100,000 gpt-4, gpt-4-turbo, gpt-3.5-turbo, text-embedding-3-small/large
p50k_base ~50,000 text-davinci-002/003, code-davinci-002
r50k_base (gpt2) ~50,000 gpt-2, davinci, curie, babbage, ada

You rarely need to memorize it — tiktoken resolves the encoding from the model name by prefix matching, so anything starting with gpt-5-, gpt-4o-, gpt-4.1-, o1-, o3-, or o4-mini- resolves to o200k_base.

import tiktoken

for model in ("gpt-4o", "gpt-4o-mini", "gpt-4.1", "gpt-5-mini", "o3-mini", "gpt-4", "gpt-3.5-turbo"):
    print(model, tiktoken.encoding_for_model(model).name)
Enter fullscreen mode Exit fullscreen mode

Count GPT tokens in Python

pip install tiktoken
Enter fullscreen mode Exit fullscreen mode
import tiktoken

enc = tiktoken.encoding_for_model("gpt-4o-mini")   # resolves to o200k_base
text = open("prompt.txt", encoding="utf-8").read()
print(len(enc.encode(text)), "tokens")
Enter fullscreen mode Exit fullscreen mode

It runs offline — no API key, no network request after the encoding file is cached.

When a count surprises you, look at how the string actually splits:

ids = enc.encode("Unbelievable")
print(ids)
print([enc.decode_single_token_bytes(t) for t in ids])
# [59026, 17536]  ->  [b'Un', b'believable']
Enter fullscreen mode Exit fullscreen mode

Counting a whole chat request

Each message carries structural overhead beyond its content. This approximation follows OpenAI's cookbook: roughly 3 tokens per message, plus 1 if you set a name, plus 3 priming the assistant reply.

import tiktoken

def count_chat_tokens(messages, model="gpt-4o-mini"):
    try:
        enc = tiktoken.encoding_for_model(model)
    except KeyError:
        enc = tiktoken.get_encoding("o200k_base")   # fallback for new model names

    tokens_per_message, tokens_per_name, total = 3, 1, 0
    for message in messages:
        total += tokens_per_message
        for key, value in message.items():
            if isinstance(value, str):
                total += len(enc.encode(value))
            if key == "name":
                total += tokens_per_name
    total += 3  # assistant reply primer
    return total
Enter fullscreen mode Exit fullscreen mode

Treat it as a planning figure. The authoritative count is always the usage object in the response — log it and reconcile against your estimate.

Count GPT tokens in JavaScript

Counting where the prompt is built lets you warn users before they send something enormous, without shipping their text anywhere.

npm install js-tiktoken
Enter fullscreen mode Exit fullscreen mode
import { encodingForModel } from "js-tiktoken";

const enc = await encodingForModel("gpt-4o");
console.log(enc.encode("Count me before I cost you money.").length, "tokens");
Enter fullscreen mode Exit fullscreen mode

gpt-tokenizer is the other common option and ships per-model entry points, which helps when bundle size matters:

import { encode } from "gpt-tokenizer";
const n = encode(text).length;
Enter fullscreen mode Exit fullscreen mode

The tokenizer is just a data table and a merge algorithm, so none of this needs a network call.

What your request actually contains

The user message is usually the smallest part of the bill. All of this counts against the same context window and the same input price:

  • System prompt — resent every request, so length here is a permanent tax
  • Conversation history — grows quadratically with turns if replayed in full
  • Tool/function schemas — twenty tools with detailed descriptions can occupy thousands of tokens even when none are called
  • Retrieved context — every pasted chunk
  • Reasoning tokens — on reasoning models these occupy context and bill as output
  • The reply — output usually costs several times more per token than input, and max_tokens is a cap, not a budget

Budgeting the context window

The window is shared, not per-message:

free_tokens = context_window
            - system + history + tools + retrieved_chunks
            - max_tokens_you_allow_for_the_reply
Enter fullscreen mode Exit fullscreen mode

If that lands near zero you get either a rejected request or silently dropped messages — your "memory" vanishes mid-conversation. Keep 10–15% slack so unusual inputs do not tip you over. Better: size retrieval from the count. After counting the fixed parts, whatever remains is exactly how many tokens of context you can afford.

Turning tokens into money

cost = (input_tokens / 1_000_000) * input_price
     + (output_tokens / 1_000_000) * output_price
Enter fullscreen mode Exit fullscreen mode

Standard-tier OpenAI list prices, October 2026 (always re-check the official pricing page):

Model Input / 1M Cached input / 1M Output / 1M
gpt-5-nano $0.05 $0.005 $0.40
gpt-5-mini $0.25 $0.025 $2.00
gpt-4o-mini $0.15 $0.075 $0.60
gpt-4o $2.50 $1.25 $10.00
o3 $2.00 $0.50 $8.00

Concrete case: 10,000 support answers per day, 600-token prompt, 300-token reply.

  • gpt-4o-mini: 6M input × $0.15 = $0.90/day + 3M output × $0.60 = $1.80/day → $2.70/day (~$81/month)
  • gpt-4o: 6M × $2.50 = $15.00/day + 3M × $10.00 = $30.00/day → $45.00/day (~$1,350/month)

Same prompt, same answer quality for a routing task, 16× difference. This is why counting comes before model selection — you cannot compare prices without knowing volume.

Six traps that inflate prompts

  1. Replaying the whole history — summarize old turns or use a rolling window
  2. Minifying JSON by hand — whitespace is real money at scale; strip it, drop unused field names
  3. Registering every tool on every request — send only this turn's schemas
  4. Trusting one "average tokens per request" number — track mean, p90, p99; the p99 breaks your budget
  5. Trusting character heuristics for non-English text — 4 chars/token is an English figure
  6. Setting max_tokens to the max — an uncapped reply is the classic surprise bill

Checklist before every production request

  • Count GPT tokens with the encoding of the model you actually serve
  • Include system prompt, history, tools, and retrieved chunks — not just the user message
  • Leave headroom: prompt + expected reply < context window
  • Log response usage and compare against your estimate weekly
  • Hard-cap both input length and max_tokens
  • Re-check prices quarterly and after any model-family switch

This post is the written-up version of how I actually budget prompts before shipping anything that calls an LLM. The full guide with extra sections is on my site: How to Count GPT Tokens Before You Hit the API.

If you work in the browser, two tools from the same toolkit are useful alongside this: a JSON Formatter & Validator for collapsing whitespace out of payloads, and a Word & Character Counter for the fast ceiling estimate. Both run locally — nothing uploaded.

How do you budget tokens in your pipeline — pre-flight counting, response usage logging, or something else?

Top comments (0)