Last month I opened my OpenAI usage dashboard and found something embarrassing: I was sending the same ~400-token system prompt on every single call — even when the user just typed "yes" or "thanks." That's not a model problem. That's a prompt problem, and it's an easy one to fix once you can see it.
Here's the thing nobody tells you when you start building with LLMs: you pay for every token you send, not just the ones that matter. A bloated system prompt, a few paragraphs of "context" copy-pasted into every call, or a chat history you never trim will quietly double or triple your bill while your actual answers stay the same length.
Step 1: measure before you guess
You can't fix what you don't measure. Here's a 10-line script using tiktoken (OpenAI's own tokenizer) to see exactly how many tokens a prompt costs before you send it:
import tiktoken
def count_tokens(text, model="gpt-4o"):
enc = tiktoken.encoding_for_model(model)
return len(enc.encode(text))
system_prompt = """You are a helpful assistant. Always be polite, thorough,
and explain your reasoning step by step. Never make assumptions about
what the user wants. If you are unsure, ask a clarifying question first.
Always format your answers with markdown headers and bullet points so
they are easy to read. Remember to be concise but also comprehensive."""
print(count_tokens(system_prompt)) # 73 tokens, sent on EVERY call
Run that against your own system prompt right now. If it's over 100 tokens and you're making thousands of calls a day, that's thousands of tokens you're paying for that add zero value to a simple "what's 2+2" query.
Step 2: the checklist that actually moved the needle
After going through my own agent scripts, I wrote down every place tokens were leaking and turned it into a one-page checklist. The short version:
- Measure first — count tokens before you optimize anything (see above).
- Trim the system prompt — most system prompts repeat instructions the model already follows by default ("be helpful," "be polite"). Cut anything that isn't a hard constraint.
- Don't resend static context — if the same reference doc goes into every call, that's a candidate for fine-tuning, embeddings, or just... not sending it every time.
- Cap chat history — truncate or summarize old turns instead of replaying the whole conversation every time.
- Short-circuit cheap questions — route "yes/no/thanks" style turns to a smaller/cheaper model instead of your main one.
- Watch your output tokens too — a verbose "explain step by step" instruction costs you on the way out, not just in.
- Batch when you can — one call with 5 questions often costs less than 5 separate calls with repeated boilerplate.
- Re-measure after every change — token diets are measured in percent, not vibes.
A real before/after
Taking the system prompt above and trimming it to only the hard constraints:
system_prompt_trimmed = "Be concise. Ask a clarifying question only if the request is ambiguous."
print(count_tokens(system_prompt_trimmed)) # 14 tokens vs 73 — an 80% cut, on every single call
That's the kind of fix that costs you 5 minutes and compounds forever, because it applies to every future call, not just the one you're debugging right now.
I turned the full checklist (all 8 points, with notes on when each one applies) into a free one-page PDF/checklist you can run through your own agent or chatbot code: Prompt Token Diet Checklist (free, no signup wall).
What's the biggest token-bloat you've found in your own prompts — a stale system prompt, an unbounded chat history, or something else?
Top comments (0)