DEV Community

Cover image for Non-English Prompts Cost More Tokens, Context and Money in Most LLMs
Khasky
Khasky

Posted on

Non-English Prompts Cost More Tokens, Context and Money in Most LLMs

Claude Opus 5 costs $5 per million input tokens whether a prompt arrives in English or in Russian. What changes between the two languages is how many tokens the same paragraph becomes, and the bill is that count multiplied by the price.

A recent test on Habr measured the gap across 11 tokenizers. Its numbers line up with research going back to 2023: on most popular models, a language other than English costs more to send, fills the context window sooner and reaches rate limits with less text.

How the test worked

A token is the chunk of text a model reads and bills by, usually a piece of a word. The tokenizer is the component that cuts text into tokens with a vocabulary fixed before the model is trained.

The author, Habr user donseo, wrote 4 parallel text pairs in Russian and English: a technical article, a support chat, a business notice and Python code with comments. The Russian side holds about 2,500 characters. Open tokenizers ran locally through tiktoken and Hugging Face. For closed models, the script sent each text with a short fixed prompt and the prompt alone, then subtracted the two token counts the API reported. The code and the corpus are public on GitHub under MIT.

Russian against English, same text

Model                      RU/EN tokens   Russian vs English
YandexGPT 5 Lite           0.91x          -9%
GigaChat 3                 0.96x          -4%
Grok 4.7                   1.18x          +18%
GPT-5.x / GPT-6 (o200k)    1.19x          +19%
Gemini 3.7 Flash           1.20x          +20%
GLM-5.3                    1.23x          +23%
DeepSeek V3/V4             1.38x          +38%
Qwen 3                     1.52x          +52%
Kimi K3                    1.86x          +86%
GPT-4 (cl100k)             2.09x          +109%
Claude Opus 5              2.96x          +196%
Enter fullscreen mode Exit fullscreen mode

Claude Opus 5 sits far from the rest at 2.96x. Its technical article took 327 tokens in English and 872 in Russian. Even its English is expensive: 3.25 characters per token, against 5.07 for the OpenAI o200k tokenizer behind GPT-5.x. In Russian it falls to 1.03 characters per token, close to one token per letter.

The Python sample showed the smallest gap on Claude Opus 5, 424 tokens in Russian against 231 in English. Prose is where the premium bites hardest.

At the same per-token price, Russian text on Claude Opus 5 costs about 3.9x what it costs on GPT-5.5, from the tokenizer alone. That 3.9x compares two models on the same Russian text. Reposts of the table often present it as Russian against English, which overstates the language gap and hides the model gap.

Why the tokenizer favors English

A tokenizer builds its vocabulary from a training corpus, and frequent sequences in that corpus get their own tokens. Petrov et al. found that "tokenizers are heavily influenced by the biases of the corpus source". Common English words become one token each. A Cyrillic word breaks into fragments, and in GPT-4's cl100k_base tokenizer some Cyrillic letters take two tokens on their own.

None of this depends on how well the model understands the language. The same paper puts it plainly: "the unequal treatment of languages arises at the tokenization stage, well before the language model sees any data at all."

Anthropic's pricing page says the same thing in vendor terms. One token is "approximately 4 characters or 0.75 words in English", and "the exact count varies by language and content type". It also notes that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text", a change that touches every language.

What extra tokens cost

Price, context and rate limits in an LLM API are all counted in tokens.

  • Price. Anthropic bills per million tokens, $5 input and $25 output on Claude Opus 5. A prompt that becomes 2.96x the tokens costs 2.96x as much to send.
  • Answers. Output is tokenized the same way, and on Claude Opus 5 an output token costs 5x an input token. A long answer in Russian carries the premium at the higher rate.
  • Context window. The window is a token budget. At 2.96x, it holds about a third as much Russian text as English, so long documents and chat histories run out of room sooner.
  • Rate limits. Anthropic measures them in input and output tokens per minute. A team working in Russian reaches the cap with less work done. 💸

Not a Russian problem

Petrov et al. measured parallel text in many languages for NeurIPS 2023. The tokenizer behind ChatGPT and GPT-4 used about 1.6x the tokens of English for Italian, 2.6x for Bulgarian and 3x for Arabic. For some languages the gap reached 15x. Their Russian figure on cl100k_base was 2.49x, on a different corpus from the Habr test. The authors' conclusion about money is direct: "the tokenization premiums discussed in Section 4 directly map to cost premiums."

Newer tokenizers have narrowed the gap for some vendors. OpenAI's o200k came in at 1.19x for Russian in the Habr test. The two studies used different texts, so this shows a direction rather than a precise trend.

Price is not the only cost. Ahia et al. at EMNLP 2023 found that speakers of many supported languages "are overcharged while obtaining poorer results". Anthropic publishes scores relative to English for Claude Sonnet 4.5: Spanish at 98.2%, Chinese at 96.9%, Hindi at 96.7%, Swahili at 91.1% and Yoruba at 79.7%. Major languages stay close to English, smaller ones fall behind. Russian is not in that table. Wendler et al. add a hint about why English stays the default: inside the models they studied, "the abstract 'concept space' lies closer to English than to other languages".

Where the rule breaks

YandexGPT 5 Lite and GigaChat 3 use fewer tokens for Russian than for English. They are the two models in the test built for Russian, and their tokenizers encode Russian at least as compactly as English. That is the same mechanism from the other side: the language a tokenizer saw most is the language it encodes cheaply.

The test has known limits. Its author lists them: the corpus is small, and the coefficients may move by about 5% on other texts. Closed-model counts came through third-party API gateways. The GigaChat figure uses the open GigaChat 3 tokenizer, which may differ from the production model. Claude's tokenizer is not public, so its number rests on API usage counts.

What to do with it

  • Write system prompts, instructions and tool schemas in English when the model tokenizes your language poorly. The Habr author recommends exactly this.
  • Compare models by the cost of a typical request on your own text, not by the price per million tokens.
  • Count before choosing. Anthropic's token counting endpoint is free, and tiktoken covers OpenAI's tokenizers.
  • For workloads that must stay in Russian, the Russian-built models start with a tokenizer advantage. 🧭

A price per token is only comparable between languages when the token counts are.

For anyone paying per token, the language of the prompt is a pricing decision, and on Claude Opus 5 a Russian prompt buys about a third of the text the same money buys in English.

References


Follow me for more on AI and Software Development:

khasky — LinkedIn / GitHub / Patreon / Bluesky / Mastodon / Medium / Devto

khaskydev — X / Threads / Instagram / Pinterest / Tumblr / Facebook / VK

Top comments (0)