I learned this the annoying way: I inspected a support-ticket automation that was doing cleanup, classification, extraction, summarization, and escalation decisions with the same expensive model.
It worked.
That was the problem.
When a workflow works, nobody asks whether it’s wasting money and context on glorified janitor tasks.
And most automations absolutely are.
If you’re building with n8n, Make, Zapier, LangChain, OpenClaw, or plain API calls, the highest-leverage optimization usually isn’t a better prompt.
It’s this:
- use cheap, fast models for boring steps
- use expensive reasoning models only for ambiguous or high-risk steps
- shrink context before it reaches your strongest model
That pattern saves money, reduces context bloat, and usually makes the workflow easier to reason about.
The real waste: treating every step like frontier reasoning
A lot of teams say they’re building agents with GPT-5, Claude Opus, or Gemini.
Then you inspect the actual pipeline and half the calls are doing this:
- classify this ticket
- extract these fields
- rewrite this into valid JSON
- tag this document
- summarize this thread
- check whether required keys exist
That is not the same job as:
reason through an ambiguous escalation with conflicting evidence
Yet people keep paying frontier-model prices for both.
That makes no sense.
Anthropic’s pricing makes the argument pretty clearly:
- Claude Haiku 4.5: $1 input / $5 output per 1M tokens
- Claude Opus 4.6: $5 input / $25 output per 1M tokens
That’s a 5x spread.
So if your workflow starts with extraction, tagging, cleanup, or first-pass summarization, that’s where I’d optimize first.
Not the final reasoning step.
Context bloat is usually self-inflicted
The second problem is even more common: people keep sending giant prompts to the wrong model.
Typical anti-pattern:
- take a messy email thread
- add CRM notes
- add internal comments
- add a giant system prompt
- send all of it to the best model
- repeat for every step
So classification sees the entire novel.
Extraction sees the entire novel.
Summarization sees the entire novel again.
That’s backwards.
The cheap model should do the trimming.
Use a fast model to:
- normalize text
- dedupe repeated content
- extract fields
- compress long threads
- convert unstructured junk into structured JSON
Then pass the smaller result to Claude Opus 4.6 or GPT-5 only if the task is actually hard.
That’s how you reduce context window pressure in practice.
Not with heroic prompt surgery. With better pipeline design.
The vendors are basically telling us to do this
Google is pretty blunt here.
The Gemini API pricing page offers Batch API pricing at a 50% cost reduction for supported models.
That’s basically Google saying:
stop using premium synchronous calls for bulk background work
Example pricing called out in the original research:
- Gemini 3.8 Flash standard: $0.75 input / $3.75 output per 1M tokens
- Gemini 3.8 Flash Batch: $0.375 input / $1.875 output per 1M tokens
If you’re summarizing 10,000 support conversations overnight, that should be batch work.
If you’re evaluating a single churn-risk enterprise escalation, that’s where premium reasoning belongs.
Anthropic sends the same signal with prompt caching.
If your automation repeats long system instructions or reusable context, Claude Opus 4.6 prompt cache hits/refreshes at $0.50 per 1M tokens can change the economics a lot.
Same idea:
- don’t recompute what you can cache
- don’t reason deeply about what you can preprocess cheaply
Routing is now an engineering primitive
This gets more interesting once routing becomes part of the API layer.
OpenRouter’s provider routing lets you optimize for:
- price
- throughput
- max latency
- fallbacks
That means routing is no longer just “I like Claude for writing.”
It becomes operational.
Cheap first. Fast second. Strongest only when needed.
A support workflow can prefer low-cost providers for first-pass classification, then escalate if:
- latency spikes
- output fails validation
- confidence is low
- the case matches a risky category
That’s much better than blind loyalty to one model vendor.
Example request shape:
{
"model": "openai/gpt-5.4",
"messages": [
{
"role": "user",
"content": "Classify this ticket"
}
],
"provider": {
"sort": "price",
"allow_fallbacks": true,
"preferred_max_latency": 2
}
}
That little provider block matters more than most people think.
It means your workflow can optimize for cost, speed, and reliability at the same time.
The pattern I keep coming back to: cheap model -> validate -> escalate
If I had to recommend one pattern to most teams, it would be this:
- cheap model does extraction or classification
- validate the output
- escalate only failures or ambiguous cases
That’s it.
This is the cleanest way to stop paying premium-model prices for routine work.
Example: support ticket triage
First pass:
- classify urgency
- extract account ID
- detect billing / outage / legal / enterprise keywords
- produce structured JSON
If the JSON validates and confidence is high, continue.
If validation fails or the case is risky, escalate to Claude Opus 4.6 or GPT-5.
Example schema:
{
"urgency": "low | medium | high | critical",
"category": "billing | bug | feature_request | cancellation | legal | other",
"account_id": "string | null",
"needs_human": true,
"confidence": 0.91,
"reason": "Customer mentions production outage and enterprise SLA"
}
Example: document extraction pipeline
Cheap model:
- parse invoice, PDF, email, or form
- extract fields to JSON
- normalize dates, totals, IDs
Validation step:
- required keys exist
- number formats are valid
- dates parse correctly
- totals aren’t contradictory
Expensive model:
- only repair malformed outputs
- only reason over ambiguous or conflicting data
That’s better engineering, not just cheaper inference.
Practical orchestration examples
n8n pattern
n8n is good at this because you can split the workflow into explicit stages.
A simple shape:
- Webhook node receives support ticket
- Cheap LLM call extracts structure
- Code node validates schema
- IF node checks confidence / category / validation result
- Strong model handles exceptions
- Ticket gets routed to human or auto-resolved
Pseudo-flow:
Webhook
-> LLM (cheap model)
-> Code (validate JSON)
-> IF (valid && confidence > 0.85 && category != legal)
-> Auto-route
-> Else -> LLM (Claude Opus 4.6 / GPT-5)
LangChain example
If you’re using LangChain, swapping models by step is straightforward.
from langchain.chat_models import init_chat_model
classifier = init_chat_model("anthropic:claude-3-5-haiku-latest")
reasoner = init_chat_model("openai:gpt-5")
raw_ticket = "Customer says their production workspace is down and legal will review the SLA."
classification_prompt = f"""
Extract JSON with fields:
- urgency
- category
- needs_human
- confidence
- reason
Ticket:
{raw_ticket}
"""
first_pass = classifier.invoke(classification_prompt)
print(first_pass)
Then validate before escalating:
import json
def should_escalate(payload: dict) -> bool:
if payload.get("confidence", 0) < 0.85:
return True
if payload.get("category") in {"legal", "billing"}:
return True
if payload.get("urgency") == "critical":
return True
return False
payload = json.loads(first_pass.content)
if should_escalate(payload):
final = reasoner.invoke(
f"Review this support case and decide next action:\n\n{raw_ticket}\n\nFirst pass:\n{json.dumps(payload)}"
)
print(final)
Batch the boring stuff
If the user does not need the result right now, batch it.
Examples:
- nightly summarization of support threads
- CRM note cleanup
- bulk document compression
- asynchronous tagging jobs
- historical conversation labeling
That’s exactly the kind of workload where lower-cost batch processing wins.
A simple stack I’d actually use
If I were building this today, I’d reach for something like:
- Claude Haiku 4.5 for extraction, tagging, classification, and structured cleanup
- Gemini Flash Batch for bulk asynchronous summarization
- Claude Opus 4.6 or GPT-5 for exception handling and edge-case reasoning
- OpenRouter for routing, fallback, and price-aware provider selection
- n8n or LangChain for orchestration
That gives you a clean separation:
- cheap model for discipline
- expensive model for judgment
That’s the split most automations need.
Yes, this adds complexity
It does.
Using one strong model everywhere is simpler.
Routing adds:
- more prompts
- more failure modes
- more testing
- more evaluation work between steps
For small workflows, one strong model is often fine.
If your automation is:
- low volume
- rarely triggered
- expensive to debug
- not context heavy
then simplicity may be the better tradeoff.
But once repeated boring steps show up thousands of times a day, routing starts paying for itself very quickly.
One warning: cheap isn’t always cheap
Raw token pricing can lie to you.
Tokenizer differences, retries, malformed outputs, and escalation rates all matter.
The original post also pointed out that Anthropic notes newer tokenizers can produce about 30% more tokens for the same text in some cases.
So don’t stop at pricing tables.
Measure end-to-end workflow economics:
- how often first pass succeeds
- how often you escalate
- how much context you removed before reasoning
- how many retries happen
- how much latency each stage adds
- how much human review you avoided
That’s the real unit of cost.
Routing cheat sheet
| Option | What it’s actually good for |
|---|---|
| Claude Haiku 4.5 | Cheap extraction, tagging, classification, and structured cleanup at $1 input / $5 output per 1M tokens |
| Claude Opus 4.6 | Hard reasoning, ambiguous validation, and exception handling at $5 input / $25 output per 1M tokens |
| Claude Opus 4.6 prompt caching | Repeated system prompts and reusable context, with cache hits/refreshes at $0.50 per 1M tokens |
| Gemini 3.8 Flash standard | Fast synchronous summarization or lightweight user-facing steps at $0.75 input / $3.75 output |
| Gemini 3.8 Flash Batch | Background summarization and bulk processing with 50% lower cost at $0.375 input / $1.875 output |
| OpenRouter routing | Price sorting, latency controls, throughput-aware routing, and fallbacks across providers |
What I’d optimize first
If you want the least painful, highest-impact change, start with one of these.
1. Support triage
Send first-pass classification and tagging to a low-cost model.
Escalate only:
- billing disputes
- legal threats
- enterprise accounts
- critical incidents
- low-confidence outputs
2. Document extraction
Use a cheap model to turn PDFs, emails, and forms into JSON.
Validate the schema.
Send only broken outputs to a stronger model.
3. Background summarization
Move async summaries, note cleanup, and bulk compression onto a batch-friendly model.
Keep synchronous decision-making on your strongest model.
4. Context cleanup before reasoning
This one is underrated.
Before the expensive model sees anything, run a cheap pass that:
- removes duplicates
- extracts facts
- compresses long threads
- strips irrelevant metadata
That alone can dramatically reduce context usage and improve downstream consistency.
My blunt rule
Use the cheapest model that can reliably produce valid structure.
Use the strongest model only where mistakes are actually expensive.
Batch anything the user does not need immediately.
Cache repeated prompts whenever the vendor supports it.
And measure escalations, not just token price.
That last one matters most.
Because the real question isn’t:
which model is best?
It’s:
which step deserves intelligence, and which step just needs discipline?
Once you look at automations that way, sending everything to one expensive model starts to look less like sophistication and more like laziness.
One more thing: if you’re tired of building this kind of routing while also babysitting per-token costs, that’s basically the problem Standard Compute is trying to remove. It gives you an OpenAI-compatible API with flat monthly pricing, so you can run automations and agents without constantly watching token spend. Same SDK shape, less pricing anxiety.
That part is especially appealing if you’re running n8n, Make, Zapier, OpenClaw, or custom agent workflows all day.
Quick test plan
If you want to try this this week, do this:
# 1) pick one workflow with repeated boring steps
# 2) identify the expensive model calls
# 3) split first-pass extraction/classification from final reasoning
# 4) add schema validation
# 5) log escalation rate
# 6) compare cost + latency + quality after 3-7 days
Metrics to log:
first_pass_success_rate
schema_validation_failure_rate
escalation_rate
avg_tokens_before_cleanup
avg_tokens_after_cleanup
p50_latency
p95_latency
human_review_rate
That gives you a real answer fast.
And in most cases, the answer is the same one I learned the annoying way:
your best model should not be doing your boring work.
Top comments (0)