DEV Community

jidonglab
jidonglab

Posted on

tiktoken Undercounts Claude Tokens: 37 of 418 Requests Blew the Limit

My nightly digest job died at 2:14 a.m. with a 400 and the message prompt is too long. My packer had measured the prompt at 181,874 tokens, comfortably under the 200K window. Claude said it was 207,431.

That's a 14% gap, and it came from one lazy decision I made months earlier: counting Claude tokens with tiktoken. It's fast and local, and it's already in every Python environment I own. On my code-heavy inputs it was also wrong in the same direction every single time.

This post covers the failure, the audit of 600 chunks, and the fix that took me from 37 failures to zero.

TL;DR

  • tiktoken is OpenAI's tokenizer. Using cl100k_base to count Claude tokens gives you a different vocabulary's answer, not an approximation of Claude's.
  • On my corpus it undercounted by a median of 14%: about 9% on prose, 16% on Python, 27% on JSON and lockfiles, and 38% on the worst single chunk.
  • The error only bites when you pack close to the limit. 61 of my 418 requests filled the budget, and 37 of those 61 failed.
  • It skews cost estimates too. My dashboard said $41.20 for the month. The invoice said $46.90.
  • Fix: do a cheap tiktoken pass with a multiplier, then verify the fully assembled request with Anthropic's free messages.count_tokens endpoint and trim until it fits.

What was I building?

I have a Python job that runs every night across my side-project repos. For each repo it collects the day's diffs, the touched files, and any related docs, then packs them into one Claude request that writes a digest: what changed, what looks risky, what I forgot to finish.

Packing was greedy. Sort the files by relevance and add them until the budget is full:

import tiktoken

enc = tiktoken.get_encoding("cl100k_base")
BUDGET = 180_000  # leave ~20K for system prompt + slack

def pack(files):
    chosen, used = [], 0
    for f in files:
        n = len(enc.encode(f.text))
        if used + n > BUDGET:
            break
        chosen.append(f)
        used += n
    return chosen, used
Enter fullscreen mode Exit fullscreen mode

I'd picked cl100k_base because a blog post once said it was "close enough" for Claude. I never checked. For two months it looked fine, because most nights the repos were small and the packer never got near the budget.

Then I pointed it at a monorepo with a fat package-lock.json and a folder of JSON test fixtures.

Why does tiktoken undercount Claude tokens?

tiktoken undercounts Claude tokens because it implements OpenAI's BPE vocabularies, and Claude uses a different tokenizer that splits the same text into more pieces. Two tokenizers trained on different data with different merge rules will disagree, and the disagreement depends on the content. Anthropic doesn't ship a local tokenizer for current Claude models, so there's no offline library that gives you the exact number.

I wanted the real size of the gap on my data instead of a vibe, so I sampled 600 chunks from the job's inputs. For each one I counted with cl100k_base and with Anthropic's count_tokens endpoint, then took the ratio of Claude's count to tiktoken's.

Content type Chunks Median ratio (Claude / tiktoken) Worst chunk
Markdown docs, commit messages 180 1.09 1.15
Python source 210 1.16 1.24
TypeScript source 110 1.17 1.26
JSON, lockfiles, fixtures 100 1.27 1.38
All 600 1.14 1.38

The ratio is never below 1.0 on my corpus. tiktoken never once overcounted. That's the worst kind of error, because a symmetric error would at least cancel out sometimes.

Structured, whitespace-heavy, and symbol-dense text had the biggest gap. A minified JSON fixture full of UUIDs hit 1.38. Plain English prose stayed under 1.10.

One caveat matters more than the table: this ratio isn't a constant. It's specific to my content and to the model I counted against. Tokenizers change between model generations, so a ratio you calibrated last quarter can drift when you upgrade. Measure it on your own data, against the model you actually call.

Why did only some requests fail?

Only requests that packed close to the context limit were at risk, because a 14% undercount on a half-empty prompt never reaches 200K. Out of 418 requests over 19 nights, 357 never filled the budget and none of them failed. The other 61 filled it, and 37 of those blew up.

The 24 that survived were mostly docs-heavy repos. At a 1.09 ratio, 180K tiktoken tokens comes out to about 196K real tokens. That's tight but it fits.

Code and JSON pushed the ratio to 1.16 or higher, which puts the same 180K over 208K. The monorepo failed every night it ran.

Then there was the overhead I hadn't counted at all. My system prompt was counted with tiktoken too, so it had the same error. My three tool definitions (a file lookup, a git log reader, and a "flag this" tool) added 1,180 tokens that never went through any counter. Tool schemas are part of the input. I'd just never thought of them as text.

What did it do to my cost estimates?

My cost estimate was low by about the same 14%, because I priced every request off the tiktoken count. My little dashboard multiplied tiktoken counts by the per-token input price and reported $41.20 for the month. The real bill for that job was $46.90.

$5.70 isn't going to bankrupt anyone. But the dashboard was the number I used to decide whether to add more repos to the job, and it was quietly flattering me. If you're quoting AI features to a client off tiktoken math, you're under-quoting every one of them.

How do you count Claude tokens accurately?

Use Anthropic's messages.count_tokens endpoint with the exact model, system prompt, messages, and tools you're about to send. It's free to call (rate-limited separately from message creation) and returns the input token count for the whole assembled request, tool definitions included. That last part is the one I needed.

I didn't want a network round trip for every file during packing, though. So the fix has two passes: a fast pessimistic estimate to fill the prompt, then one real count on the final request, trimming if needed.

import anthropic
import tiktoken

client = anthropic.Anthropic()
enc = tiktoken.get_encoding("cl100k_base")

MODEL = "claude-sonnet-5-5"
TARGET = 194_000   # real tokens; headroom under the 200K window
FAST_RATIO = 1.2   # pessimistic tiktoken multiplier for the first pass

def real_count(system, messages, tools):
    resp = client.messages.count_tokens(
        model=MODEL, system=system, messages=messages, tools=tools,
    )
    return resp.input_tokens

def pack(files, system, tools):
    # Pass 1: cheap local estimate, deliberately inflated
    chosen, est = [], 0
    for f in files:
        n = int(len(enc.encode(f.text)) * FAST_RATIO)
        if est + n > TARGET:
            break
        chosen.append(f)
        est += n

    # Pass 2: ask Claude's own tokenizer, drop the least relevant file until it fits
    while chosen:
        messages = build_messages(chosen)
        n = real_count(system, messages, tools)
        if n <= TARGET:
            return messages, n
        chosen.pop()
    raise ValueError("system prompt + tools alone exceed TARGET")
Enter fullscreen mode Exit fullscreen mode

A few details that took me an iteration to get right:

  • Count the assembled request, not the pieces. Message wrappers, file headers, and tool schemas all add tokens. Summing per-file counts misses them.
  • Keep headroom. Anthropic's docs describe the count as an estimate that can differ slightly from what you're billed. I target 194K, not 199,999.
  • Pass the real model ID. The count belongs to the tokenizer of the model you send to. Counting against one model and sending to another brings back the original bug.
  • Trim by relevance. chosen is sorted most-relevant first, so pop() drops the least useful file.

Did the fix actually work?

Yes: zero overflow failures across the next 21 nights and 463 requests. The second pass needed one count_tokens call on 79% of the requests that filled the budget and two on the rest. Nothing needed a third.

The costs are small:

  • Latency: the count call added a median of 170 ms to each request. The digest takes a minute to generate, so nobody will ever notice.
  • Packing efficiency: budget-filling prompts now average 191K real tokens against the 194K target. Before the fix they averaged "whatever tiktoken thought, plus a coin flip."
  • Cost tracking: the dashboard now logs input_tokens from the API response's usage field instead of my estimate. This month's figure matched the invoice to the cent.

That last change deserves its own line. You never need to estimate tokens after the fact. Every response tells you exactly what you were billed for, so log the actual number.

What would I tell past me?

Don't use a tokenizer from a different model family as a ruler, even if it's "close enough." It's close enough until you pack to the limit, and a packer exists to pack to the limit.

If you just need a rough guess for a UI ("about 40K tokens"), tiktoken times a calibrated ratio is fine. If a request fails, or a bill comes out wrong, when the number is off, ask the API.

So, can you count Claude tokens with tiktoken?

You can get a rough estimate, but you can't get an accurate count. tiktoken implements OpenAI's tokenizers, and Claude splits text differently. On my corpus cl100k_base undercounted Claude by a median of 14%, and by up to 38% on JSON. That broke 37 of the 61 requests I packed near the 200K window and understated my bill by $5.70 a month. For anything that has to fit a limit or match an invoice, call Anthropic's free messages.count_tokens endpoint on the fully assembled request (system prompt, messages, and tools) for the exact model you're calling, keep a few thousand tokens of headroom, and log real usage from the response's usage field.


Written by the developer behind Preterview, an interview prep platform.

Top comments (0)