DEV Community

Haku
Haku

Posted on

Your agent bill has an arbitrage in it: a three-tier audit

Your agent bill has an arbitrage in it: a three-tier audit

In August 2026, the FT reported that 56% of the tokens flowing through Vercel's AI Gateway ran on open-weight models — for 14% of the spend. Read the fine print first, because it's the whole point: that's gateway-specific, not universal. It's an observed ratio, not a list price. But it asks the right question: how much of your agent bill is paying flagship prices for work that doesn't need a flagship?

October made the question urgent. Flagship pricing is in freefall, and every price move has a trapdoor in it. Here's how to price your own workloads across three tiers, with this month's numbers as the worked example.

The price war in one table

Prices in USD per 1M tokens. Every row is sourced and dated — quote none of these without the fine print.

Tier Model In / 1M Out / 1M Source Fine print
A GPT-6 Astra $10 $50 OpenAI Sep 2026. Predecessor flagship.
B GPT-6.1 Sol $2 $10 OpenAI Sep 29, 2026. 2x repricing cliff above 272K tokens per request.
B Gemini 4 Argon $2 $10 DeepMind Sep 30, 2026. Promo rate — jumps to $4/$20 when the promo ends (a 50% cliff). 1M-token output limit.
C Open weights — — FT on Vercel AI Gateway Aug 2026. Not a list price. Observed ratio ≈ 0.13x frontier (~7.8x cheaper). Gateway-specific.
C Gemma 4 (26B/31B server, E2B/E4B edge) — — Google Oct 3, 2026. $0 license — you pay hosting.
⚠ Anthropic enterprise discounts — — The Information Oct 2, 2026. ~15% volume discounts being unwound at the usage cap. Expect mid-cycle bill shock on capped contracts.

Tier A is what you're paying now. Tier B is the cheapest equivalent flagship. Tier C is open weights, priced by the gateway-observed ratio for estimation — or by your actual hosting invoice once you self-host. Note that pricing is moving weekly in October 2026: this is a snapshot, not a feed.

The three-tier audit

The audit has three steps. It takes an afternoon.

Step 1: classify your workloads. Not your models — your workloads:

  • T1 — bulk / batchable: embeddings, classification, summarization drafts, eval harnesses. Nobody is watching in real time.
  • T2 — interactive: user-facing chat, tool-calling agents, anything where a human waits on the answer.
  • T3 — regulated: workloads pinned to contracted or approved models by policy, customer agreement, or compliance scope. These don't move. The audit tells you what fraction of spend is immovable, which is itself useful.

Step 2: price each workload under all three tiers. Here's the script (fictional demo volumes):

PRICES = {
    "astra": {"in": 10.0, "out": 50.0},  # OpenAI, Sep 2026
    "sol":   {"in": 2.0,  "out": 10.0},  # OpenAI, Sep 29 2026
    "argon": {"in": 2.0,  "out": 10.0},  # DeepMind promo, Sep 30 2026
}
OPEN_WEIGHT_RATIO = 0.13   # FT, Aug 2026: gateway-observed, NOT a list price
SOL_CLIFF_TOKENS = 272_000 # OpenAI, Sep 29: 2x repricing above this per request

def price(in_m, out_m, model, max_request_tokens=100_000):
    p = PRICES[model]
    cliff = 2.0 if (model == "sol" and max_request_tokens > SOL_CLIFF_TOKENS) else 1.0
    return (in_m * p["in"] + out_m * p["out"]) * cliff

def audit(in_m, out_m, **kw):
    a = price(in_m, out_m, "astra", **kw)   # Tier A: current flagship
    b = price(in_m, out_m, "sol", **kw)     # Tier B: cheap flagship
    c = a * OPEN_WEIGHT_RATIO               # Tier C: open weights (ratio est.)
    return {"tier_a": a, "tier_b": b, "tier_c": round(c, 2)}

print(audit(10, 5))
# {'tier_a': 350.0, 'tier_b': 70.0, 'tier_c': 45.5}
Enter fullscreen mode Exit fullscreen mode

Step 3: read the worked example. A workload doing 10M input / 5M output tokens a month on Astra costs $350/mo. The same workload on Sol: $70/mo. On open weights, estimated by the gateway ratio: ~$45/mo. The Tier A → Tier B move is the easy 80% of the savings and requires no infrastructure. The Tier B → Tier C move is smaller in absolute dollars and costs you a migration — which is why you run the numbers before you touch anything.

One trap the script makes visible: pass max_request_tokens=300_000 for Sol and Tier B doubles. Long-context agent loops hit the 272K cliff quietly. Price your actual request shapes, not your monthly totals.

The migration runbook (condensed)

If Tier C wins for a T1 workload, migrate in four stages: shadow → 10% → 50% → 100%.

  • Shadow: route a copy of production traffic to the new model, compare outputs, spend nothing extra on user-facing quality. Promote only if the quality gate passes on your evals, not the vendor's benchmarks.
  • 10% / 50%: canary on real traffic with per-stage hold criteria — quality delta within your threshold, p95 latency within budget, cost per successful task down.
  • 100%: full cutover with the old model as the instant fallback.

Rollback criteria (write these down before you start): quality delta beyond your threshold for two consecutive eval windows, p95 latency breach, or cost per successful task rising instead of falling. Any one of these fires and traffic goes back — no meeting required.

The enforcement half is a one-page routing policy, evaluated per request:

# evaluated per request; fail closed
- if: workload == "T1" and quality_gate == "pass"
  route: tier_c
- elif: workload == "T2"
  route: tier_b
- else: route: tier_a   # T3, or gate failure: fail closed to flagship
Enter fullscreen mode Exit fullscreen mode

A T1 request whose quality gate fails doesn't get retried on a cheaper model — it falls back to the flagship. Fail closed, always.

What breaks

The honest section, because this is where audits usually lie:

  • Quality deltas are per-task, not per-model. "Open weights are fine for summarization" is true on average and false for your specific summarization task often enough to matter. The shadow stage exists because you cannot borrow someone else's eval.
  • Promo cliffs are real money. Budget Argon at $4/$20, not $2/$10 — the promo ends and your bill jumps 50% overnight. Same discipline for Sol: any request shape near 272K tokens gets priced with the cliff, not without it.
  • $0 license is not $0 cost. Gemma 4 costs nothing to download and real money to host: GPUs, inference infra, ops time. The 0.13x ratio is a gateway observation, not your invoice. Get a hosting quote before you claim the savings.
  • Discount unwinds bite mid-cycle. If you're on an Anthropic enterprise contract near its usage cap, that ~15% discount (The Information, Oct 2) may already be gone from your next invoice. Check before you compare against it.

What I built

I packaged this into OpenArb — the audit kit: the sourced price table above (kept as a dated snapshot), the audit script and CSV template, the four-stage migration runbook, the routing-policy template, and an arbitrage-report template for the writeup you'd hand your team.

It ships with a free interactive self-assessment demo — plug in your token volumes, get your three-tier number with every stat sourced and dated.

The full kit is $49 one-time: https://vittoriali.gumroad.com/l/openarb

Two companion pieces if you're building the whole stack: GateKeep handles the runtime side — cheap calibrated models gating expensive frontier calls (https://vittoriali.gumroad.com/l/gatekeep) — and DotSpend puts hard spend caps on workloads with a 402 cutoff when a budget breaches (https://vittoriali.gumroad.com/l/dotspend). Audit first, route second, cap third.

— Haku

Pricing is a snapshot of October 2026 and moves weekly — re-verify every row before a migration decision. All stats sourced with outlet and date above. Fictional demo data in code.

Top comments (0)