GLM-5.3 Flash 50% Discount Ends Tomorrow — The Complete September 2026 AI API Price War Map
September 8, 2026 — Six major pricing events in 10 days. GPT-6 Astra at $10/$50, Gemini 3.8 Flash at $0.75, GLM-5.3 Flash's 50% promo ends tomorrow, and DeepSeek's peak pricing is now the new normal. Here's the full battlefield and what you should route where.
September has been the most concentrated pricing event of the year. Between September 1 and September 3, three frontier labs compressed token costs within a 72-hour window. The Silicon Data LLM Token Expenditure Index — a benchmark tracking the blended cost per million tokens across the industry — dipped below $1 for the first time on September 1.
But the biggest deadline for budget-conscious developers is tomorrow, September 9: GLM-5.3 Flash's 50% introductory discount ends. Let me break down what changed, what's staying, and where your API budget stretches furthest.
🚨 Deadline: GLM-5.3 Flash 50% Off Ends September 9
GLM-5.3 Flash (the former "Ox Alpha") has been the biggest story of the last two weeks. A model that emerged anonymously on OpenRouter, hit #1 in weekly usage with 11.6 trillion tokens, and turned out to be Z.ai's MIT-licensed 320B MoE (18B active) beast.
Current pricing (through Sep 9):
| Metric | Price (50% off) | Normal price after Sep 9 |
|---|---|---|
| Input | $0.075/M | $0.15/M |
| Output | $0.25/M | $0.50/M |
| Cached input | $0.015/M | $0.03/M |
Yes — right now it's $0.075 per million input tokens. That's 667x cheaper than GPT-6 Astra's $50/M output, 640x cheaper than Claude Fable 5.1's $50/M output, and about 2x cheaper than DeepSeek V4 Flash's off-peak $0.22/M input.
What happens after Sep 9: The price doubles to $0.15/$0.50. But even at full price, it's still the cheapest model at its intelligence level — comparable to Claude Opus 4.8 on agent benchmarks while costing 33x less on input and 50x less on output.
Where to use it: Agent workflows, coding assistants, high-volume customer support, document extraction — anything with long context and repeated tool calls. The 1M context window and hybrid sparse+linear attention make it insanely efficient for long-running agent sessions.
📊 The Complete September 2026 Pricing Table
I've updated the master price comparison table. All prices per million tokens, USD:
| Model | Input | Cached Input | Output | 1M+1M Cost | Tier |
|---|---|---|---|---|---|
| GLM-5.3-Flash (50% off) | $0.075 | $0.015 | $0.25 | $0.325 | Budget 🔥 |
| GLM-5.3-Flash (standard) | $0.15 | $0.03 | $0.50 | $0.65 | Budget |
| DeepSeek V4 Flash (off-peak) | $0.22 | $0.007 | $0.66 | $0.88 | Budget |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $1.40 | Budget |
| DeepSeek V4 Flash (peak) | $0.44 | $0.014 | $1.32 | $1.76 | Budget |
| Gemini 3.8 Flash (promo) | $0.75 | — | $3.75 | $4.50 | Budget |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 | $5.25 | Mid |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | $6.00 | Mid |
| Gemini 3.8 Flash (standard) | $1.50 | — | $7.50 | $9.00 | Mid |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | $12.00 | Mid |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | $14.00 | Mid |
| DeepSeek V4 Pro (peak) | $1.32 | $0.044 | $3.96 | $5.28 | Mid |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | $30.00 | Premium |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | $60.00 | Frontier |
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | $60.00 | Frontier |
The 143x gap we reported in June is now 68x, and it closed from the bottom, not the top. DeepSeek raised prices, while Western labs mostly held or cut.
🔄 What Changed in the Last 10 Days
September 1 — Anthropic Fable 5.1 + Sonnet 5 Price Lock
Two announcements from Anthropic on the same day:
- Fable 5.1 launched at $10/$50 — same base rates as Fable 5, but cache reads dropped 75% from $1.00 to $0.25 per million tokens. For agentic workloads with heavy context reuse, this means 25-45% total savings.
- Sonnet 5's $2/$10 made permanent — the scheduled September 1 price hike to $3/$15 was cancelled. This is huge: Sonnet 5 stays at the mid-tier sweet spot permanently.
September 2 — Google Gemini 3.8 Flash
Google's newest Flash model at $0.75/$3.75 — valid through December 31, then doubles to $1.50/$7.50. A pure volume play: grab market share before the rate resets.
September 3 — OpenAI GPT-6 Astra
OpenAI's new flagship at $10/$50, matching Fable 5.1's base price. Cached input at $1.00/M. Rolling out over the following days.
September 4 — Meta Muse Spark Contributor Tier
Meta launched a controversial pricing model: $0.10/$0.20 per million tokens — but only if you share your agent prompts and outputs as training data. Standard tier is $1.25/$4.25. A 92-95% discount for data-sharing.
💰 What This Means for Your API Budget
The budget tier is more competitive than ever. GLM-5.3 Flash at $0.075/$0.25 (promo) or $0.15/$0.50 (standard) sets a new floor for quality models. Even at standard pricing, it's 2-3x cheaper than DeepSeek V4 Flash off-peak despite comparable agent capabilities.
The "Flash" model category is now the default. The 36kr article put it well: "过去带有Flash或Mini后缀的模型往往被视为全量旗舰模型的裁剪阉割版...然而如今,Flash模型已然冲上牌桌中央,成为巨头抢占生态入口、争夺API吞吐量的核心主力。"
Translation: Flash models are no longer the "lite version." They're the main event. 85%+ of enterprise workloads don't need frontier intelligence — they need low latency, high throughput, and a bill that doesn't make your CFO cry.
The routing strategy is clearer than ever:
- Simple tasks, high volume → GLM-5.3 Flash (cheapest quality model, grab the promo while it lasts)
- General production → DeepSeek V4 Flash (off-peak $0.22/$0.66, stable pricing)
- Long context, agent workflows → GLM-5.3 Flash (1M context, hybrid attention, MIT weights)
- Mid-tier need → Sonnet 5 / Gemini 3.8 Flash (both at $2/$10 range)
- Complex reasoning → DeepSeek V4 Pro (off-peak $0.66/$1.98, still 12.6x cheaper than Fable 5)
- Frontier → Use sparingly. Astra and Fable 5.1 at $10/$50 are for when quality literally cannot be compromised.
🚀 How to Access These Models
All the models mentioned above are available through TunanAPI — a single OpenAI-compatible endpoint that routes to Chinese AI models at significantly lower rates than Western providers.
from openai import OpenAI
client = OpenAI(
base_url="https://api.tunanapi.com/v1",
api_key="your-key"
)
# GLM-5.3 Flash — grab the promo pricing before Sep 9
response = client.chat.completions.create(
model="glm-5.3-flash",
messages=[{"role": "user", "content": "Build a Python script that monitors API pricing changes"}]
)
Current TunanAPI pricing highlights:
| Model | Input | Output | Best For |
|-------|-------|--------|----------|
| GLM-5.3-Flash | From $0.15/M | From $0.50/M | Agent workflows, coding |
| DeepSeek V4 Flash | $0.20/M | $0.40/M | General production |
| DeepSeek V4 Pro (off-peak) | $0.66/M | $1.98/M | Complex reasoning |
| Qwen3.8-Max | $2.00/M | $6.00/M | General purpose, 1M context |
| MiniMax M3 | $1.20/M | $4.80/M | Coding & reasoning |
Bottom line: GLM-5.3 Flash's 50% promo ends tomorrow. If you've been testing it, lock in the rate now. The pricing landscape has shifted more in the last 10 days than in the previous 3 months — and the clear winner is the developer who builds a smart routing layer, not the one who picks a single provider.
This article is part of a weekly series tracking AI API pricing. All prices sourced from vendor pricing pages as of September 8, 2026. Pricing may vary by region, context length, and usage tier.
Built with TunanAPI — access the best Chinese AI models through a single OpenAI-compatible API. tunanapi.com
Top comments (1)
Great article! Having quick access to free APIs for testing and prototyping really speeds up development. API aggregator platforms are worth checking out for discovering services by category. Keep it up! 🚀