LLMs do not read letters or whole words. They read tokens: subword chunks of text that are converted into numbers. Understanding tokens matters because they decide two things at once: how much the model can remember, and how much your AI app costs to run.
In plain terms: Think of a taxi meter that counts in small chunks of text, not in kilometres. Every chunk you send, and every chunk the model writes back, adds to the fare.
Diagram: A tokenizer splits text into subword chunks and maps each to a number. Common words are one token, rare or compound words split into pieces, and numbers and punctuation often take several. See the animated version.
As a rule of thumb, one token is about three quarters of an English word, so 100 tokens is roughly 75 words. Common words are a single token. Rare or compound words are split into pieces, and numbers, symbols and code often take several tokens each. Different models use different tokenizers, so the exact counts vary.
The context window
The context window is the model's short-term memory limit: the total number of tokens it can consider at once. A model with a 128,000-token window can hold roughly the equivalent of a few hundred pages of text in a single call. Everything counts towards that limit: the system instructions, the chat history, any documents you paste in, and the reply itself. Go past it and the oldest content is dropped, so the model appears to forget earlier instructions.
Diagram: The context window is how much the model can hold at once. Instructions, chat history, documents and the reply all count. When it overflows, the oldest text is dropped and effectively forgotten. See the animated version.
Token economics
Providers bill you for input tokens (what you send) and output tokens (what the model generates). Output tokens usually cost several times more, often three to four times, because producing text takes more GPU work than reading it. Prices differ by provider and change often, so always check the current price list.
Diagram: Providers bill separately for input and output tokens, and output tokens usually cost several times more. In this illustrative call, output is a fifth of the tokens but about half of the cost. See the animated version.
Optimization
Small habits add up quickly at scale:
- Trim the history. Do not send the whole chat every time. Keep the recent turns, and summarize or drop old ones.
- Remove redundant whitespace and noise from injected text and code.
- Use prompt caching for static system instructions, so repeated long prompts cost less.
- Cap the answer length with a maximum-tokens setting.
- Route easy tasks to smaller models and keep the flagship model for hard problems.
Diagram: Trimming chat history, removing whitespace, caching the static prompt, capping answer length and routing easy tasks to smaller models each take a bite out of the bill. The percentages here are illustrative. See the animated version.
Here is the same idea in Python: counting tokens with a real tokenizer, and estimating a call's cost. The prices are made-up units, only to show the shape of the calculation.
import tiktoken
enc = tiktoken.get_encoding("cl100k_base") # a common tokenizer; others split text differently
for text in ["Network latency is rising", "microservices"]:
tokens = enc.encode(text)
print(len(tokens), "tokens:", [enc.decode([t]) for t in tokens])
# Cost of one call. The prices below are made-up units: output costs 4x as much as input.
PRICE_IN, PRICE_OUT = 1.0, 4.0 # per 1,000 tokens (illustrative)
def call_cost(tokens_in: int, tokens_out: int) -> float:
return tokens_in / 1000 * PRICE_IN + tokens_out / 1000 * PRICE_OUT
print("2,000 in + 500 out:", call_cost(2000, 500)) # 2.0 for input + 2.0 for output
print("trim the history to 1,200 in:", call_cost(1200, 500))
Coming Up Next
Day 11: Embeddings and vector representations: how machines map meaning.
Tokenization #GenerativeAI #CostOptimization #SoftwareArchitecture #CloudEconomics
Originally published at https://sureshpallapothu.in/blog/day-10-tokens-and-cost, where this post includes animated diagrams.
Top comments (0)