You don't need an API key to find out your AI feature is quietly expensive.
You don't need a production account.
You don't need to wait for the first bill.
You need:
- Google Colab
- Python
pip install tiktoken- About 10 minutes
The last article in this series, "Give Your Chatbot a Memory in Google Colab Before Your Next AI Interview," was about managing a conversation once it gets long. (Link it to that post's live Dev.to URL when you paste this in — it isn't in this repo.) This one rewinds further — before memory, before RAG, before agents — to the mechanism everything else in an LLM system is built on top of: tokens, and what they actually cost you. If you can't answer "how many tokens is this, and what does that mean in dollars," the rest of the stack doesn't matter yet.
Let's measure it.
What Are We Building?
Every LLM call is billed and bounded by tokens, not words or characters. We'll build three small tools:
- A real tokenizer that counts tokens the way a production model actually would
- A cost calculator that turns a token count into a daily and yearly dollar figure
- A tiny simulation of what "temperature" actually does to next-token sampling
No fine print, no mocked numbers — every count below came from running the code, not estimating it.
Step 1: Count Tokens for Real
Skip guessing "about 4 characters per token." Use the same tokenizer family production APIs use.
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
text_a = "The meeting is scheduled for Thursday at 3pm. Please bring your notes from the last session and any open questions you want to discuss with the team."
text_b = "The rate-limiting factor in transformer inference is memory bandwidth during KV-cache reads, not FLOPs. Prefill is compute-bound; autoregressive decode is memory-bandwidth-bound."
text_c = "Here is the function signature: async def fetch_user(user_id: UUID, db: AsyncSession) -> Optional[UserModel]:. It should return None if the user does not exist, not raise an exception."
for name, t in [("A - plain English", text_a), ("B - technical jargon", text_b), ("C - mixed code", text_c)]:
toks = enc.encode(t)
print(f"{name}: {len(toks)} tokens, {len(t.split())} words -> {len(toks)/len(t.split()):.2f} tokens/word")
Output, run just now:
A - plain English: 31 tokens, 27 words -> 1.15 tokens/word
B - technical jargon: 37 tokens, 21 words -> 1.76 tokens/word
C - mixed code: 43 tokens, 27 words -> 1.59 tokens/word
Same rough word count across all three (21–27 words), but Text B costs 53% more tokens per word than Text A. Compound terms like "rate-limiting" get chopped into multiple sub-word pieces; UUIDs and type annotations in Text C do the same. Plain English is the cheapest register you can write in — technical prompts are a variable tax, not a fixed multiplier.
Step 2: Turn Tokens Into Dollars
A token count is abstract until it's attached to a query volume and a price sheet.
PRICING = {
"claude-sonnet": {"input": 3.00, "output": 15.00}, # $ per 1M tokens
"claude-haiku": {"input": 0.80, "output": 4.00},
}
def estimate_cost(text, queries_per_day, model, days=365):
tokens = len(enc.encode(text))
daily_tokens = tokens * queries_per_day
daily_cost = daily_tokens * (PRICING[model]["input"] / 1_000_000)
return tokens, daily_tokens, daily_cost, daily_cost * days
system_prompt = (
"You are a support assistant for Acme Cloud. Always answer in a friendly, "
"professional tone. Never reveal internal pricing formulas. If the user asks "
"about billing, refer them to the billing dashboard at app.acme.com/billing. "
"If the user reports an outage, check the status page before responding, and "
"always include the current incident ID if one is open. Keep responses under "
"150 words unless the user explicitly asks for more detail."
)
for model in ["claude-sonnet", "claude-haiku"]:
t, daily_tok, daily_cost, yearly_cost = estimate_cost(system_prompt, 1000, model)
print(f"{model}: {t} tokens/call, {daily_tok:,} tokens/day, ${daily_cost:.2f}/day, ${yearly_cost:,.0f}/year")
Real output:
claude-sonnet: 87 tokens/call, 87,000 tokens/day, $0.26/day, $95/year
claude-haiku: 87 tokens/call, 87,000 tokens/day, $0.07/day, $25/year
One system prompt, sent unchanged on every one of 1,000 daily queries. $95/year on Sonnet doesn't sound alarming — until this is one of six prompts in your app, or volume is 100,000/day instead of 1,000. The formula doesn't change; the multiplier does.
Step 3: Cut the Bill Without Cutting the Feature
trimmed_prompt = (
"You are Acme Cloud's support assistant. Friendly, professional tone. "
"Never reveal pricing formulas. Billing questions -> app.acme.com/billing. "
"Outage reports -> check status page, include open incident ID. Max 150 words."
)
t2, _, daily_cost2, yearly_cost2 = estimate_cost(trimmed_prompt, 1000, "claude-sonnet")
print(f"Trimmed: {t2} tokens/call ({(1 - t2/87)*100:.0f}% shorter), ${daily_cost2:.2f}/day, ${yearly_cost2:,.0f}/year")
Trimmed: 47 tokens/call (46% shorter), $0.14/day, $51/year
Denser prose cut the prompt nearly in half and the yearly cost from $95 to $51 — same rules enforced, zero functionality lost.
Step 4: What Temperature Actually Does
Temperature doesn't make a model smarter or dumber — it reshapes how sharply it commits to its top guess. Simulate it with a toy 5-word distribution over "The weather today is ___":
import math, random
random.seed(7)
candidates = ["sunny", "cloudy", "cold", "unpredictable", "magnificent"]
logits = [2.2, 2.0, 1.5, 1.0, 0.3]
def softmax_with_temperature(logits, temperature):
scaled = [l / max(temperature, 1e-6) for l in logits]
m = max(scaled)
exps = [math.exp(s - m) for s in scaled]
total = sum(exps)
return [e / total for e in exps]
def sample(candidates, probs):
r = random.random()
cum = 0.0
for c, p in zip(candidates, probs):
cum += p
if r <= cum:
return c
return candidates[-1]
for temp in [0.2, 1.0]:
probs = softmax_with_temperature(logits, temp)
picks = [sample(candidates, probs) for _ in range(8)]
print(f"temp={temp}: probs={dict(zip(candidates, [round(p,2) for p in probs]))}")
print(f" 8 samples: {picks}")
temp=0.2: probs={'sunny': 0.71, 'cloudy': 0.26, 'cold': 0.02, 'unpredictable': 0.0, 'magnificent': 0.0}
8 samples: ['sunny', 'sunny', 'sunny', 'sunny', 'sunny', 'sunny', 'sunny', 'sunny']
temp=1.0: probs={'sunny': 0.36, 'cloudy': 0.3, 'cold': 0.18, 'unpredictable': 0.11, 'magnificent': 0.05}
8 samples: ['sunny', 'cloudy', 'sunny', 'sunny', 'cloudy', 'cold', 'sunny', 'sunny']
Same logits both times. At temp=0.2 the distribution collapses onto "sunny" — 8 for 8. At temp=1.0 the same model "opinion" produces four different words across 8 draws. Nothing about the model's knowledge changed — only how willing it is to pick outside its top guess.
Make It Reusable
def tokenize_and_cost(text, queries_per_day, model="claude-sonnet"):
tokens, daily_tokens, daily_cost, yearly_cost = estimate_cost(text, queries_per_day, model)
return {
"tokens_per_call": tokens,
"daily_cost_usd": round(daily_cost, 2),
"yearly_cost_usd": round(yearly_cost, 0),
}
One function, three inputs (text, volume, model), a real dollar figure out. Swap in your own system prompt before your next design review.
Colab Experiments to Try
- Paste your own longest system prompt into Step 1 — is its tokens/word ratio closer to plain English or technical jargon?
- Run
estimate_costat 10,000 queries/day. Watch the yearly figure move from "line item" to "budget conversation." - Add
gpt-4o-mini($0.15input) toPRICINGand compare its yearly cost againstclaude-haiku. - Change
logitsso all 5 candidates are nearly equal, then re-run both temperatures — does low temperature still look deterministic?
Interview Questions Hidden Inside This Notebook
Why does a technical or code-heavy prompt cost more than plain English at the same word count?
Sub-word tokenization splits compound technical terms and punctuation-heavy syntax (UUIDs, type hints) into more pieces than common English words. Word count and token count are only loosely correlated — measure tokens directly, never estimate from word count.
What lever actually reduces LLM cost on a fixed feature set?
Prompt density. The Step 3 rewrite cut cost 46% by removing redundant phrasing, not functionality. Token count is a writing-quality problem before it's an infrastructure problem.
Does temperature affect what the model "knows"?
No. It reshapes sampling over an unchanged probability distribution — it doesn't touch the model's underlying weights. Low temperature for tasks needing the same answer every time; higher temperature only where variety is the actual goal.
The Calculator Is the Easy Part
tokenize_and_cost() is maybe six lines. A real production cost model also has to account for output tokens (billed 3-5x the input rate, harder to bound), prompt caching (steep discounts on repeated system prompts this notebook doesn't model), and retries or tool-call round-trips that resend and re-bill context.
Interviewers ask about token cost because a candidate who can name a real number, on the spot, has clearly built something — not just read about it.
Want to Go Deeper?
This notebook covers the mechanics. The full session covers the "how does an LLM generate a response?" 5-beat answer end to end, including the hallucination root cause this notebook didn't touch:
https://confidentprep.com/courses/ai-ml-for-interview/1-how-llms-actually-work/
There's also a companion piece, "How Does an LLM Actually Generate a Response?" — The Interview Question Everyone Gets 80% Right, with 8 interview questions pulled straight from this same session. (Link it to that post's Dev.to URL once both are published.)
Measure your own prompt before your next standup. A real number beats a guess — in a cost review and in an interview.
Top comments (0)