I got fooled by a benchmark that looked incredible.
Same agent workflow. Same model. Same code path.
Second run was dramatically faster.
My first reaction was the same dumb little hit of engineer dopamine most of us get: nice, we optimized something.
We didn’t.
The prompt prefix matched, the cache kicked in, and my “model improvement” was mostly reuse.
That sounds obvious after the fact. But a lot of LLM benchmark posts still miss it. People run the same agent twice, latency drops, and suddenly GPT-4o or Claude or Gemini is supposedly faster, cheaper, or more stable.
Usually the benchmark just got warm.
If you build agents in n8n, Make, Zapier, OpenClaw, or your own Python workers, this gets even easier to miss because so much of the request stays identical between calls.
The benchmark looked great right up until I checked usage
Agent workloads are basically built to trigger prompt caching.
A typical request includes:
- a large system prompt
- tool schemas
- examples
- conversation history
- planner / critic prompts
- structured output schemas
The user’s latest message is often the smallest part of the request.
The expensive part is the scaffolding.
So when the second run gets much faster, that doesn’t automatically mean the model improved. It often means the provider reused the big stable prefix.
For production, that’s good.
For benchmarking, it can absolutely lie to you.
What to log if you want numbers that mean anything
If you’re not logging cache-related usage fields, you’re guessing.
Here are the ones that matter:
| Provider | What to look for |
|---|---|
| OpenAI Prompt Caching | usage.prompt_tokens_details.cached_tokens |
| Anthropic Prompt Caching | cache usage fields in the response usage object |
| Google Gemini Implicit Caching |
usage.total_cached_tokens or Vertex AI cachedContentTokenCount
|
| OpenRouter | sticky routing behavior and shared session_id
|
OpenAI is the easiest example.
Prompt caching is enabled by default on supported models. It kicks in for prompts over 1,024 tokens, and the response tells you how many prompt tokens were served from cache.
Example response shape:
{
"usage": {
"prompt_tokens": 2006,
"completion_tokens": 300,
"prompt_tokens_details": {
"cached_tokens": 1920,
"audio_tokens": 0
}
}
}
If cached_tokens goes from 0 on run one to 1920 on run two, you did not discover a magical optimization.
You discovered that the prefix matched.
That’s the story.
Why the second run also looked cheaper
Because caching changes both latency and cost.
OpenAI has been explicit about this: cached input is cheaper than uncached input on supported models.
Google says implicit cache hits on Vertex AI can get major discounts too.
So when someone says:
Gemini 2.5 Pro got way cheaper in our pipeline
my first question is not:
what changed in the model?
It’s:
did you log
usage.total_cached_tokens?
If the answer is no, the pricing story is incomplete.
Quick Python example: log the cache fields or don’t trust the benchmark
Here’s a minimal OpenAI-compatible example.
from openai import OpenAI
import time
import json
client = OpenAI()
messages = [
{"role": "system", "content": "You are a careful code review assistant."},
{"role": "user", "content": "Summarize this repository architecture and list the top risks."}
]
def run_once():
start = time.time()
resp = client.chat.completions.create(
model="gpt-4o",
messages=messages,
temperature=0
)
elapsed = time.time() - start
usage = resp.usage
cached = 0
if hasattr(usage, "prompt_tokens_details") and usage.prompt_tokens_details:
cached = getattr(usage.prompt_tokens_details, "cached_tokens", 0)
print(json.dumps({
"latency_s": round(elapsed, 3),
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
"cached_tokens": cached
}, indent=2))
run_once()
run_once()
If the second call is faster and cached_tokens jumps, that result is warm-cache behavior.
Not raw model speed.
If you use curl, same rule
You don’t need a full benchmark harness to catch this.
curl https://api.openai.com/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "system", "content": "You are a careful code review assistant."},
{"role": "user", "content": "Summarize this repository architecture and list the top risks."}
]
}' | jq '.usage'
Run it twice.
Check whether cached token fields appear or increase.
That one step will save you from a lot of fake benchmark confidence.
Anthropic and Gemini can fool you just as easily
This is not an OpenAI-only problem.
Anthropic
Anthropic supports prompt caching with cache_control. Cache lifetime is usually around 5 minutes.
from anthropic import Anthropic
client = Anthropic()
resp = client.messages.create(
model="claude-opus-4-1",
max_tokens=1024,
cache_control={"type": "ephemeral"},
system="You are an AI assistant...",
messages=[
{"role": "user", "content": "Analyze the major themes in Pride and Prejudice."}
]
)
That means follow-up requests can look much faster if the cacheable prefix is unchanged.
There’s also an annoying timing edge case: if your streamed response takes several minutes, you can accidentally fall out of the cache window and think the benchmark is random.
It isn’t random. Your timing is just sloppy.
Gemini
Gemini 2.5+ models also have implicit caching behavior.
So yes, the second call may be faster.
No, that does not mean Gemini “learned” your workflow.
If you’re on Vertex AI, inspect cached token fields. If you’re not logging them, you’re missing the main reason the benchmark changed.
OpenRouter can make the benchmark look cleaner than it really is
This one matters if you benchmark through OpenRouter.
OpenRouter uses sticky routing after a cached request so later requests for the same model and conversation can keep hitting the same provider-side cache.
That’s useful in production.
It’s bad for sloppy benchmarking because now the router is helping your benchmark stay warm.
Example payload:
{
"model": "openai/gpt-4o",
"messages": [
{"role": "user", "content": "Summarize this codebase"}
],
"session_id": "agent-run-42"
}
If later calls improve, you need to ask whether the model got better or whether OpenRouter kept sending you back to the same warm provider.
Those are very different explanations.
n8n can fake the win before the request even reaches the model
This is the part people really hate hearing.
Sometimes the model didn’t get faster.
Sometimes your workflow got shorter.
In n8n, nodes like Remove Duplicates can reduce how many items actually reach the LLM across runs.
So the benchmark story becomes:
- n8n filtered duplicate items
- fewer requests hit GPT-4o or Claude
- later runs looked faster and cheaper
- everyone credited the model
That’s not model improvement.
That’s upstream dedupe.
The same thing can happen in Make, Zapier, or custom queues if retries, filters, or batching behavior change between runs.
My benchmark rule now: separate cold from warm or don’t publish it
This is the simplest fix.
If your benchmark doesn’t separate cold and warm runs, it’s unfinished.
Here’s the checklist I use now.
1. Log cache metrics on every request
At minimum:
- OpenAI:
usage.prompt_tokens_details.cached_tokens - Gemini:
usage.total_cached_tokensorcachedContentTokenCount - Anthropic: cache usage fields in the response usage object
- OpenRouter: routing/session behavior
2. Hash the fully rendered prompt
Do not just version the user message.
Hash the final request body after templating, including:
- system prompt
- developer prompt
- tool definitions
- tool order
- examples
- conversation history
- response schema
If any of those change, cache behavior changes too.
3. Benchmark cold and warm separately
Run one set with unique prefixes.
Run another set with stable prefixes.
Label them honestly:
- cold-start latency
- warm-cache latency
Do not average them into one vanity number.
4. Control the timing window
Cache TTL matters.
If your benchmark drifts across cache expiration windows, your numbers will drift too.
5. Verify the workflow didn’t reduce actual model calls
Before celebrating a cost or latency win, confirm:
- no dedupe node removed items
- no retry logic changed request count
- no batching layer altered call volume
- no router stickiness hid the real path
A practical benchmark harness pattern
This is the shape I’d recommend for agent teams.
import hashlib
import json
import time
from openai import OpenAI
client = OpenAI()
def prompt_fingerprint(payload: dict) -> str:
raw = json.dumps(payload, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(raw.encode()).hexdigest()
def run(payload: dict):
fp = prompt_fingerprint(payload)
start = time.time()
resp = client.chat.completions.create(**payload)
elapsed = time.time() - start
usage = resp.usage
cached = 0
if getattr(usage, "prompt_tokens_details", None):
cached = getattr(usage.prompt_tokens_details, "cached_tokens", 0)
return {
"fingerprint": fp,
"latency_s": round(elapsed, 3),
"prompt_tokens": usage.prompt_tokens,
"completion_tokens": usage.completion_tokens,
"cached_tokens": cached,
}
payload = {
"model": "gpt-4o",
"messages": [
{"role": "system", "content": "You are a strict API design reviewer."},
{"role": "user", "content": "Review this OpenAPI spec and list breaking changes."}
],
"temperature": 0
}
print(run(payload))
print(run(payload))
This won’t solve every benchmarking problem.
But it will stop you from calling a warm cache hit a model breakthrough.
The annoying truth: prompt caching is good, fake benchmark stories are not
Prompt caching is not cheating.
It’s one of the best features for real agent workloads.
If you run long-lived automations, coding agents, support bots, or multi-step workflows, you absolutely want big stable prefixes to be reused.
That’s a real production win.
The mistake is claiming you measured model quality, raw speed, or provider pricing when what you really measured was cache friendliness.
Those are different things.
Once you notice this, a lot of benchmark posts start looking suspiciously magical.
- the second run is always faster
- the giant agent gets “smarter” after turn three
- Claude suddenly becomes cheaper in a multi-turn evaluator
- Gemini “settles in” after a few calls
Maybe.
But usually the boring answer wins.
The prefix matched.
The router stayed sticky.
The workflow dropped duplicates.
Why this matters even more for teams building agents
If you’re building serious automations, you need predictable behavior, not benchmark theater.
That’s one reason I like what Standard Compute is doing.
Standard Compute gives you unlimited AI compute for a flat monthly price through an OpenAI-compatible API. So if you’re running agents in n8n, Make, Zapier, OpenClaw, or custom workers, you can keep the same SDK patterns without living in per-token panic.
That doesn’t remove the need for honest benchmarking.
It just means you can optimize for actual workflow performance instead of obsessing over every token spike from repeated agent calls.
For teams running lots of long-context, multi-step automations, that tradeoff makes a lot of sense.
The takeaway
If you benchmark repeated LLM calls without checking cache fields, your numbers are probably lying.
Log the usage fields.
Separate cold from warm.
Hash the full prompt.
Check the workflow layer.
Then publish the numbers.
Anything else is just benchmark fan fiction.
Top comments (0)