I'm a solo developer in Korea, and for a long time I assumed my Korean prompts were "about 2-3x more expensive" than English. Everyone says non-English text eats tokens. But I had never actually measured it.
So I did. I took one ordinary prompt, translated it into 27 languages, and ran every version through OpenAI's tokenizers. Some results matched what I expected. Several did not.
The setup
The prompt is a typical customer-support task, 34 tokens in English:
Please summarize the customer email below in three bullet points and suggest a polite reply. The customer says the order arrived two days late and one item was missing from the box.
Each translation was tokenized with two encodings using js-tiktoken:
- o200k_base, the tokenizer OpenAI has used since GPT-4o
- cl100k_base, the older GPT-4 / GPT-3.5 tokenizer, for comparison
import { getEncoding } from "js-tiktoken";
const enc = getEncoding("o200k_base");
const tokens = (text) => enc.encode(text).length;
tokens(english); // 34
tokens(korean); // 49
tokens(czech); // 68
The results
| Language | Tokens (o200k) | vs English | Old GPT-4 tokenizer (cl100k) |
|---|---|---|---|
| English | 34 | 1.00× | 1.00× |
| Chinese (Simplified) | 35 | 1.03× | 1.53× |
| Indonesian | 39 | 1.15× | 1.38× |
| Spanish | 40 | 1.18× | 1.29× |
| Portuguese | 41 | 1.21× | 1.41× |
| Persian | 42 | 1.24× | 2.79× |
| German | 43 | 1.26× | 1.50× |
| Arabic | 43 | 1.26× | 3.03× |
| French | 44 | 1.29× | 1.44× |
| Dutch | 44 | 1.29× | 1.74× |
| Russian | 45 | 1.32× | 2.15× |
| Swedish | 45 | 1.32× | 1.50× |
| Chinese (Traditional) | 46 | 1.35× | 2.12× |
| Vietnamese | 46 | 1.35× | 2.32× |
| Italian | 47 | 1.38× | 1.59× |
| Korean | 49 | 1.44× | 2.50× |
| Turkish | 50 | 1.47× | 2.06× |
| Hindi | 51 | 1.50× | 4.59× |
| Filipino | 52 | 1.53× | 1.76× |
| Hebrew | 53 | 1.56× | 3.76× |
| Urdu | 54 | 1.59× | 4.24× |
| Bengali | 57 | 1.68× | 6.09× |
| Thai | 59 | 1.74× | 3.71× |
| Japanese | 61 | 1.79× | 2.21× |
| Ukrainian | 64 | 1.88× | 3.15× |
| Polish | 64 | 1.88× | 2.12× |
| Czech | 68 | 2.00× | 2.59× |
What surprised me
1. The tokenizer upgrade was huge for Indic, Arabic and Thai scripts.
With the old GPT-4 tokenizer, Bengali needed 6.1x the tokens of English and Hindi 4.6x. With o200k they are down to 1.7x and 1.5x. Arabic went from 3.0x to 1.26x. If you last measured this in 2023, your numbers are badly out of date.
2. Simplified Chinese is basically free.
35 tokens vs 34 for English. Traditional Chinese, for the same sentence, costs 1.35x. Same language family, very different price, most likely because of how much of each was in the training data.
3. The new "expensive" languages use Latin and Cyrillic letters.
The top of the table is not Japanese or Thai. It's Czech (2.0x), Polish (1.88x) and Ukrainian (1.88x). Heavily inflected languages with diacritics split into many small pieces. Russian and Ukrainian share an alphabet, yet Russian costs 1.32x and Ukrainian 1.88x.
4. Korean is cheaper than I thought: 1.44x, not 2-3x.
The "2-3x" number I had in my head was true for the old tokenizer (2.5x). It's stale folk wisdom now.
What it means in money
Say you send this prompt 1 million times on a model priced at $2 per 1M input tokens:
| Language | Input tokens | Input cost |
|---|---|---|
| English | 34M | $68 |
| Korean | 49M | $98 |
| Japanese | 61M | $122 |
| Czech | 68M | $136 |
And that's only the input. If the model also answers in the same language, the output (usually 4-5x more expensive per token) carries the same multiplier.
Practical takeaways
- Write system prompts and fixed instructions in English, and keep only the user's content in their language. The instructions are sent with every request; the multiplier compounds.
- For internal steps (classification, extraction, tool calls), ask for output in English or JSON keys in English, and translate only what the end user sees.
- Cache the long, fixed part of your prompt if your provider supports prompt caching; cached input is often 90%+ cheaper.
- Measure your own text. One prompt is a small sample, and technical text, names and code behave differently.
Caveats
- This is one prompt. A single sentence can swing a ratio by 0.1-0.2x. Treat the ranking as a rough map, not a precise price list.
- Translations were machine-assisted and spot-checked. If you're a native speaker and a translation reads unnaturally, tell me and I'll re-run it.
- Claude and Gemini use different tokenizers, which aren't available to run locally in the browser, so these numbers are only exact for OpenAI models.
The tool I built out of this
I turned this into a small free tool: TokenSave. You paste text and it shows:
- exact GPT token counts (o200k runs in your browser; nothing is uploaded) and estimates for Claude and Gemini
- a language overhead badge ("1.4x tokens vs English") calibrated with the numbers above
- a cost planner: input tokens, expected output tokens and requests per month
- a chat (JSON) mode that counts per-message formatting tokens
The UI is available in 27 languages, since the people who pay this tax are mostly not reading in English. There are also side calculators for AI video and image generation prices.
I'd love to hear:
- What ratio do you see for your language on real production prompts?
- Has anyone measured the same thing for Claude's or Gemini's tokenizer via their token-count APIs?
The 27 prompts I used (click to expand)

Top comments (0)