DEV Community

Cover image for TIL the best automations reduce context window by using cheap models for boring steps and saving Claude Opus for the hard part
Lars Winstand
Lars Winstand

Posted on Originally published at standardcompute.com

TIL the best automations reduce context window by using cheap models for boring steps and saving Claude Opus for the hard part

I learned this the annoying way: I inspected a support-ticket automation that was doing cleanup, classification, extraction, summarization, and escalation decisions with the same expensive model.

It worked.

That was the problem.

When a workflow works, nobody asks whether it’s wasting money and context on glorified janitor tasks.

And most automations absolutely are.

If you’re building with n8n, Make, Zapier, LangChain, OpenClaw, or plain API calls, the highest-leverage optimization usually isn’t a better prompt.

It’s this:

  • use cheap, fast models for boring steps
  • use expensive reasoning models only for ambiguous or high-risk steps
  • shrink context before it reaches your strongest model

That pattern saves money, reduces context bloat, and usually makes the workflow easier to reason about.

The real waste: treating every step like frontier reasoning

A lot of teams say they’re building agents with GPT-5, Claude Opus, or Gemini.

Then you inspect the actual pipeline and half the calls are doing this:

  • classify this ticket
  • extract these fields
  • rewrite this into valid JSON
  • tag this document
  • summarize this thread
  • check whether required keys exist

That is not the same job as:

reason through an ambiguous escalation with conflicting evidence

Yet people keep paying frontier-model prices for both.

That makes no sense.

Anthropic’s pricing makes the argument pretty clearly:

  • Claude Haiku 4.5: $1 input / $5 output per 1M tokens
  • Claude Opus 4.6: $5 input / $25 output per 1M tokens

That’s a 5x spread.

So if your workflow starts with extraction, tagging, cleanup, or first-pass summarization, that’s where I’d optimize first.

Not the final reasoning step.

Context bloat is usually self-inflicted

The second problem is even more common: people keep sending giant prompts to the wrong model.

Typical anti-pattern:

  1. take a messy email thread
  2. add CRM notes
  3. add internal comments
  4. add a giant system prompt
  5. send all of it to the best model
  6. repeat for every step

So classification sees the entire novel.
Extraction sees the entire novel.
Summarization sees the entire novel again.

That’s backwards.

The cheap model should do the trimming.

Use a fast model to:

  • normalize text
  • dedupe repeated content
  • extract fields
  • compress long threads
  • convert unstructured junk into structured JSON

Then pass the smaller result to Claude Opus 4.6 or GPT-5 only if the task is actually hard.

That’s how you reduce context window pressure in practice.

Not with heroic prompt surgery. With better pipeline design.

The vendors are basically telling us to do this

Google is pretty blunt here.

The Gemini API pricing page offers Batch API pricing at a 50% cost reduction for supported models.

That’s basically Google saying:

stop using premium synchronous calls for bulk background work

Example pricing called out in the original research:

  • Gemini 3.8 Flash standard: $0.75 input / $3.75 output per 1M tokens
  • Gemini 3.8 Flash Batch: $0.375 input / $1.875 output per 1M tokens

If you’re summarizing 10,000 support conversations overnight, that should be batch work.
If you’re evaluating a single churn-risk enterprise escalation, that’s where premium reasoning belongs.

Anthropic sends the same signal with prompt caching.

If your automation repeats long system instructions or reusable context, Claude Opus 4.6 prompt cache hits/refreshes at $0.50 per 1M tokens can change the economics a lot.

Same idea:

  • don’t recompute what you can cache
  • don’t reason deeply about what you can preprocess cheaply

Routing is now an engineering primitive

This gets more interesting once routing becomes part of the API layer.

OpenRouter’s provider routing lets you optimize for:

  • price
  • throughput
  • max latency
  • fallbacks

That means routing is no longer just “I like Claude for writing.”

It becomes operational.

Cheap first. Fast second. Strongest only when needed.

A support workflow can prefer low-cost providers for first-pass classification, then escalate if:

  • latency spikes
  • output fails validation
  • confidence is low
  • the case matches a risky category

That’s much better than blind loyalty to one model vendor.

Example request shape:

{
  "model": "openai/gpt-5.4",
  "messages": [
    {
      "role": "user",
      "content": "Classify this ticket"
    }
  ],
  "provider": {
    "sort": "price",
    "allow_fallbacks": true,
    "preferred_max_latency": 2
  }
}
Enter fullscreen mode Exit fullscreen mode

That little provider block matters more than most people think.

It means your workflow can optimize for cost, speed, and reliability at the same time.

The pattern I keep coming back to: cheap model -> validate -> escalate

If I had to recommend one pattern to most teams, it would be this:

  1. cheap model does extraction or classification
  2. validate the output
  3. escalate only failures or ambiguous cases

That’s it.

This is the cleanest way to stop paying premium-model prices for routine work.

Example: support ticket triage

First pass:

  • classify urgency
  • extract account ID
  • detect billing / outage / legal / enterprise keywords
  • produce structured JSON

If the JSON validates and confidence is high, continue.
If validation fails or the case is risky, escalate to Claude Opus 4.6 or GPT-5.

Example schema:

{
  "urgency": "low | medium | high | critical",
  "category": "billing | bug | feature_request | cancellation | legal | other",
  "account_id": "string | null",
  "needs_human": true,
  "confidence": 0.91,
  "reason": "Customer mentions production outage and enterprise SLA"
}
Enter fullscreen mode Exit fullscreen mode

Example: document extraction pipeline

Cheap model:

  • parse invoice, PDF, email, or form
  • extract fields to JSON
  • normalize dates, totals, IDs

Validation step:

  • required keys exist
  • number formats are valid
  • dates parse correctly
  • totals aren’t contradictory

Expensive model:

  • only repair malformed outputs
  • only reason over ambiguous or conflicting data

That’s better engineering, not just cheaper inference.

Practical orchestration examples

n8n pattern

n8n is good at this because you can split the workflow into explicit stages.

A simple shape:

  1. Webhook node receives support ticket
  2. Cheap LLM call extracts structure
  3. Code node validates schema
  4. IF node checks confidence / category / validation result
  5. Strong model handles exceptions
  6. Ticket gets routed to human or auto-resolved

Pseudo-flow:

Webhook
  -> LLM (cheap model)
  -> Code (validate JSON)
  -> IF (valid && confidence > 0.85 && category != legal)
      -> Auto-route
      -> Else -> LLM (Claude Opus 4.6 / GPT-5)
Enter fullscreen mode Exit fullscreen mode

LangChain example

If you’re using LangChain, swapping models by step is straightforward.

from langchain.chat_models import init_chat_model

classifier = init_chat_model("anthropic:claude-3-5-haiku-latest")
reasoner = init_chat_model("openai:gpt-5")

raw_ticket = "Customer says their production workspace is down and legal will review the SLA."

classification_prompt = f"""
Extract JSON with fields:
- urgency
- category
- needs_human
- confidence
- reason

Ticket:
{raw_ticket}
"""

first_pass = classifier.invoke(classification_prompt)
print(first_pass)
Enter fullscreen mode Exit fullscreen mode

Then validate before escalating:

import json

def should_escalate(payload: dict) -> bool:
    if payload.get("confidence", 0) < 0.85:
        return True
    if payload.get("category") in {"legal", "billing"}:
        return True
    if payload.get("urgency") == "critical":
        return True
    return False

payload = json.loads(first_pass.content)

if should_escalate(payload):
    final = reasoner.invoke(
        f"Review this support case and decide next action:\n\n{raw_ticket}\n\nFirst pass:\n{json.dumps(payload)}"
    )
    print(final)
Enter fullscreen mode Exit fullscreen mode

Batch the boring stuff

If the user does not need the result right now, batch it.

Examples:

  • nightly summarization of support threads
  • CRM note cleanup
  • bulk document compression
  • asynchronous tagging jobs
  • historical conversation labeling

That’s exactly the kind of workload where lower-cost batch processing wins.

A simple stack I’d actually use

If I were building this today, I’d reach for something like:

  • Claude Haiku 4.5 for extraction, tagging, classification, and structured cleanup
  • Gemini Flash Batch for bulk asynchronous summarization
  • Claude Opus 4.6 or GPT-5 for exception handling and edge-case reasoning
  • OpenRouter for routing, fallback, and price-aware provider selection
  • n8n or LangChain for orchestration

That gives you a clean separation:

  • cheap model for discipline
  • expensive model for judgment

That’s the split most automations need.

Yes, this adds complexity

It does.

Using one strong model everywhere is simpler.

Routing adds:

  • more prompts
  • more failure modes
  • more testing
  • more evaluation work between steps

For small workflows, one strong model is often fine.

If your automation is:

  • low volume
  • rarely triggered
  • expensive to debug
  • not context heavy

then simplicity may be the better tradeoff.

But once repeated boring steps show up thousands of times a day, routing starts paying for itself very quickly.

One warning: cheap isn’t always cheap

Raw token pricing can lie to you.

Tokenizer differences, retries, malformed outputs, and escalation rates all matter.

The original post also pointed out that Anthropic notes newer tokenizers can produce about 30% more tokens for the same text in some cases.

So don’t stop at pricing tables.

Measure end-to-end workflow economics:

  • how often first pass succeeds
  • how often you escalate
  • how much context you removed before reasoning
  • how many retries happen
  • how much latency each stage adds
  • how much human review you avoided

That’s the real unit of cost.

Routing cheat sheet

Option What it’s actually good for
Claude Haiku 4.5 Cheap extraction, tagging, classification, and structured cleanup at $1 input / $5 output per 1M tokens
Claude Opus 4.6 Hard reasoning, ambiguous validation, and exception handling at $5 input / $25 output per 1M tokens
Claude Opus 4.6 prompt caching Repeated system prompts and reusable context, with cache hits/refreshes at $0.50 per 1M tokens
Gemini 3.8 Flash standard Fast synchronous summarization or lightweight user-facing steps at $0.75 input / $3.75 output
Gemini 3.8 Flash Batch Background summarization and bulk processing with 50% lower cost at $0.375 input / $1.875 output
OpenRouter routing Price sorting, latency controls, throughput-aware routing, and fallbacks across providers

What I’d optimize first

If you want the least painful, highest-impact change, start with one of these.

1. Support triage

Send first-pass classification and tagging to a low-cost model.
Escalate only:

  • billing disputes
  • legal threats
  • enterprise accounts
  • critical incidents
  • low-confidence outputs

2. Document extraction

Use a cheap model to turn PDFs, emails, and forms into JSON.
Validate the schema.
Send only broken outputs to a stronger model.

3. Background summarization

Move async summaries, note cleanup, and bulk compression onto a batch-friendly model.
Keep synchronous decision-making on your strongest model.

4. Context cleanup before reasoning

This one is underrated.

Before the expensive model sees anything, run a cheap pass that:

  • removes duplicates
  • extracts facts
  • compresses long threads
  • strips irrelevant metadata

That alone can dramatically reduce context usage and improve downstream consistency.

My blunt rule

Use the cheapest model that can reliably produce valid structure.

Use the strongest model only where mistakes are actually expensive.

Batch anything the user does not need immediately.

Cache repeated prompts whenever the vendor supports it.

And measure escalations, not just token price.

That last one matters most.

Because the real question isn’t:

which model is best?

It’s:

which step deserves intelligence, and which step just needs discipline?

Once you look at automations that way, sending everything to one expensive model starts to look less like sophistication and more like laziness.

One more thing: if you’re tired of building this kind of routing while also babysitting per-token costs, that’s basically the problem Standard Compute is trying to remove. It gives you an OpenAI-compatible API with flat monthly pricing, so you can run automations and agents without constantly watching token spend. Same SDK shape, less pricing anxiety.

That part is especially appealing if you’re running n8n, Make, Zapier, OpenClaw, or custom agent workflows all day.

Quick test plan

If you want to try this this week, do this:

# 1) pick one workflow with repeated boring steps
# 2) identify the expensive model calls
# 3) split first-pass extraction/classification from final reasoning
# 4) add schema validation
# 5) log escalation rate
# 6) compare cost + latency + quality after 3-7 days
Enter fullscreen mode Exit fullscreen mode

Metrics to log:

first_pass_success_rate
schema_validation_failure_rate
escalation_rate
avg_tokens_before_cleanup
avg_tokens_after_cleanup
p50_latency
p95_latency
human_review_rate
Enter fullscreen mode Exit fullscreen mode

That gives you a real answer fast.

And in most cases, the answer is the same one I learned the annoying way:

your best model should not be doing your boring work.

Top comments (0)