DEV Community

Cover image for I thought my cheap AI workflow was fine until I counted 3,000 tiny calls
Lars Winstand
Lars Winstand

Posted on Originally published at standardcompute.com

I thought my cheap AI workflow was fine until I counted 3,000 tiny calls

I didn’t get burned by one giant GPT-5 prompt.

I got burned by a workflow that looked cheap.

You know the kind:

  • fetch a batch of records
  • classify each one
  • extract a field or two
  • write a short summary
  • move on

Lead enrichment. Support triage. CRM cleanup. Scraped page normalization.

Nothing fancy. Nothing that looks like it should trigger budget panic.

And that’s exactly why it’s dangerous.

The trap: every step is cheap, but the workflow isn’t

The mistake is simple: people inspect one AI call at a time.

They say:

  • “It’s just a classifier.”
  • “It’s only extracting pain points.”
  • “The summary is two sentences.”

All true.

But nobody multiplies.

Here’s a totally normal n8n flow:

Fetch 1,000 leads
-> Loop Over Items (Batch Size: 1)
-> OpenAI node: classify company
-> OpenAI node: extract pain points
-> OpenAI node: draft outreach angle
Enter fullscreen mode Exit fullscreen mode

That is already:

1,000 records * 3 AI steps = 3,000 model calls
Enter fullscreen mode Exit fullscreen mode

And that’s before:

  • retries
  • validation passes
  • fallback prompts
  • confidence scoring
  • second-model QA

This is not an edge case. This is a standard automation pattern.

If you build with n8n, Zapier, Make, or custom worker queues, you’ve probably done this already.

Tiny prompts can be the expensive part

Not just in dollars.

Also in:

  • request throughput
  • rate limits
  • queueing behavior
  • retry storms
  • operational complexity

A lot of teams focus on token count because that’s what pricing pages train you to do.

But provider limits are usually not just about tokens.

OpenAI separates RPM and TPM for a reason. You can be nowhere near your token-per-minute cap and still hit request-per-minute limits. Their docs explicitly call out the idea that if your RPM is 20, then 20 requests of only 100 tokens each can still max you out.

That changes the architecture discussion.

If your workload is “analyze one giant contract,” token cost is the problem.

If your workload is “touch 8,000 CRM rows and make 3 tiny decisions on each,” request multiplication is usually the real problem.

Why this fools people in testing

Because staging lies.

A test run on 20 records looks cheap.

Then someone points the workflow at:

  • the whole HubSpot table
  • the full Zendesk queue
  • a scraped Apify dataset
  • a backlog of support tickets

And suddenly the cost shape changes.

Not because prompts got bigger.

Because you turned on a machine that makes tiny calls thousands of times.

Quick way to estimate whether your workflow is actually cheap

I now do this before I trust any AI automation:

items_per_day=1000
ai_steps_per_item=3
retry_rate=0.1
validation_calls_per_item=1

base_calls=$((items_per_day * ai_steps_per_item))
validation_calls=$((items_per_day * validation_calls_per_item))
retry_calls=$(python3 - <<'PY'
items=1000
steps=3
retry_rate=0.1
print(int(items * steps * retry_rate))
PY
)

echo "Base calls: $base_calls"
echo "Validation calls: $validation_calls"
echo "Retry calls: $retry_calls"
Enter fullscreen mode Exit fullscreen mode

Even rough math is enough.

If the answer is “we’re making 4,000 to 10,000 model requests a day,” you do not have a tiny workflow.

You have a high-frequency AI system.

The official fixes are real, but they solve specific problems

I’m not anti-OpenAI Batch API or anti-Anthropic prompt caching.

Both are good.

But they are not universal fixes for “my automation explodes into thousands of micro-calls.”

1) OpenAI Batch API: great if you can wait

OpenAI’s Batch API is legitimately useful.

It offers a 50% discount versus synchronous calls, and OpenAI explicitly positions it for jobs like large-scale classification and embeddings.

That maps well to:

  • nightly enrichment
  • backfills
  • offline cleanup jobs
  • large async processing queues

Example request shape:

{"custom_id":"request-1","method":"POST","url":"/v1/chat/completions","body":{"model":"gpt-3.5-turbo-0125","messages":[{"role":"system","content":"You are a helpful assistant."},{"role":"user","content":"Hello world!"}],"max_tokens":1000}}
Enter fullscreen mode Exit fullscreen mode

But there are tradeoffs:

  • asynchronous only
  • can take up to 24 hours
  • no streaming
  • not a drop-in answer for real-time routing

If your support workflow needs to classify a ticket now, Batch is not your answer.

2) Anthropic prompt caching: amazing when the prefix repeats

Anthropic prompt caching is also very real.

When you have a large repeated prompt prefix, it can cut both latency and cost dramatically.

That’s excellent for:

  • long system prompts
  • many-shot examples
  • repeated document context
  • “chat with a book” style workloads

Example:

import anthropic

client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    cache_control={"type": "ephemeral"},
    system="You are an AI assistant tasked with analyzing literary works.",
    messages=[
        {"role": "user", "content": "Analyze the major themes in Pride and Prejudice."}
    ],
)
Enter fullscreen mode Exit fullscreen mode

But caching helps when requests share stable context.

It does not magically fix workloads where every row is different:

  • every Salesforce lead is unique
  • every Zendesk ticket is unique
  • every scraped page is unique

You can’t cache uniqueness.

The real question: what cost problem do you actually have?

Most teams skip this and jump straight to model comparisons.

That’s backwards.

First figure out the shape of your workload.

Option Best fit
OpenAI synchronous API Real-time flows where latency matters, but you still deal with per-token billing and RPM/TPM limits
OpenAI Batch API Large asynchronous jobs where lower cost matters more than immediate completion
Anthropic prompt caching Repeated prompt prefixes, long shared context, and workloads that benefit from cache reuse
Per-item multi-step automation in n8n, Make, or Zapier Easy to build, easy to underestimate, and very likely to multiply request count fast

If your workload is a few giant prompts, token pricing is the main issue.

If your workload is thousands of tiny calls, the issue is usually a mix of:

  • request count
  • retries
  • fan-out per item
  • provider rate limits
  • automation platform task limits
  • engineers getting conservative because every extra step feels billable

That last one matters more than people admit.

Per-token pricing changes how people design automations

This is the hidden tax.

Teams start making worse technical decisions because they’re trying not to trigger more model calls.

I’ve seen teams do all of these:

  1. Remove QA steps because “it’s probably fine.”
  2. Skip classifying every inbound support ticket, so humans still sort manually.
  3. Avoid enriching lower-priority leads because the long tail feels too expensive.
  4. Collapse separate steps into one messy prompt because another call feels too costly.

That is not clean engineering.

That is workflow design shaped by billing anxiety.

And the model bill is not the only meter.

You may also be managing:

  • Zapier task limits
  • n8n execution volume
  • OpenAI RPM and TPM
  • Anthropic spend caps
  • fallback routing logic across providers

At some point this stops being a prompt engineering problem and becomes systems design.

What a lot of automations really are: tiny AI committees

This is the mental model that finally made it click for me.

A lot of record-level automations do not make one AI decision.

They make a committee.

For one item, you might do:

  • classify intent
  • extract fields
  • normalize to JSON
  • score confidence
  • draft a response
  • run safety review

Every call is defensible.

Together, they behave like a swarm.

That’s why I’m increasingly opinionated about this:

For high-volume operational workflows, pricing model matters almost as much as model quality.

Not because GPT-5, Claude Opus, Grok, Qwen, or Llama are bad.

Because once your team stops fearing each micro-call, you build better automations:

  • cleaner step separation
  • better validation
  • more second-pass QA
  • longer-running agents
  • fewer hacks to save pennies

What I’d do in practice

If I’m building a real system, I’d break the problem down like this.

Use batching when:

  • the job is asynchronous
  • latency does not matter
  • the workload is large and predictable
  • you want the cheapest possible processing path

Use caching when:

  • the prompt prefix is large and repeated
  • many requests share the same context
  • latency and cost both benefit from reuse

Rethink pricing architecture when:

  • each item triggers multiple small calls
  • workflows run continuously
  • retries and validation are common
  • the team is hesitating to add useful AI steps because of cost anxiety

That third category is where a lot of agent and automation teams actually live.

A practical workflow audit you can run this week

If you have an n8n, Make, Zapier, or custom agent workflow, map it like this:

workflow_audit:
  items_per_day: 1000
  ai_steps_per_item: 3
  average_retries_per_100_calls: 12
  validation_calls_per_item: 1
  fallback_model_enabled: true
  real_time_steps:
    - classify_ticket
    - route_priority
  async_steps:
    - nightly_summary
    - enrichment_backfill
  repeated_prompt_prefixes:
    - support_policy_context
    - extraction_schema
Enter fullscreen mode Exit fullscreen mode

Then answer these questions honestly:

  • How many items arrive per day?
  • How many AI nodes run per item?
  • How many retries happen in production?
  • How many validation or fallback calls happen after failures?
  • Which steps truly need real-time latency?
  • Which steps could be batched?
  • Which prompts actually share enough prefix to benefit from caching?

That exercise usually reveals one of two stories.

Story A: your problem is giant repeated context

If that’s true, prompt caching can help a lot.

Story B: your problem is death by 10,000 tiny calls

If that’s true, you need to stop evaluating cost one prompt at a time.

You need to think in workflow volume.

Where Standard Compute fits

This is exactly why products like Standard Compute are interesting for agent and automation workloads.

If you’re running lots of small calls across n8n, Make, Zapier, OpenClaw, or custom workers, flat-rate unlimited compute changes the design space.

Instead of asking:

  • “Can we afford one more validation step?”
  • “Should we skip QA on lower-value records?”
  • “Can this agent run all day?”

You can build the workflow you actually want.

Standard Compute is a drop-in OpenAI-compatible API, so you can usually swap it into existing SDKs or HTTP clients without rebuilding your stack. Under the hood it routes across models like GPT-5.4, Claude Opus 4.6, and Grok 4.20, with batching, prompt optimization, and throttling designed for exactly this kind of high-frequency automation work.

That matters if your bottleneck is not one giant prompt.

It matters if your bottleneck is thousands of tiny decisions.

My takeaway

Before you optimize prompts, count calls.

Before you compare GPT-5 vs Claude, count calls.

Before you celebrate a cheap staging run, count calls.

Most teams think their cost problem is “big prompts are expensive.”

A surprising number actually have a different problem:

small prompts, repeated constantly, inside workflows that looked harmless.

That was my mistake.

If you’re building AI automations, don’t price a single request.

Price the swarm.

Top comments (0)