AI API Digest – 2026-08-20
Most cost‑impacting change
- Qwen: Qwen3.6 27B – Prompt price rose from $0.30 to $0.60 /1M tokens (+$0.30); completion price rose from $2.00 to $3.60 /1M tokens (+$1.60). Who should care: Teams using this model for long‑form generation will see per‑million costs jump by ~$1.90; consider alternatives or adjust usage budgets.
Added models
- Z.ai: GLM Latest – Prompt $1.40 /1M, completion $4.40 /1M, 1M‑token context. Who should care: Users needing very long context (1M tokens) at moderate pricing.
Removed models
- AI21: Jamba Large 1.7 – Prompt $2.00 /1M, completion $8.00 /1M, 256k context. Who should care: Migrate workloads; this model is no longer available.
- Mancer: Weaver (alpha) – Prompt $0.50 /1M, completion $0.75 /1M, 8k context. Who should care: Alpha users must switch to another provider.
Price changes (per‑million tokens)
- DeepSeek V4 Pro 0813 – Prompt ↓$0.13 (1.32→1.19), completion ↓$0.40 (3.96→3.56).
- DeepSeek V4 Flash Latest – Prompt ↓$0.01 (0.076→0.065), completion ↓$0.01 (0.153→0.14).
- DeepSeek V4 Pro 0423 – Prompt ↑$0.28 (1.32→1.60), completion ↓$0.76 (3.96→3.20).
- DeepSeek V4 Flash 0423 – Prompt ↑$0.01 (0.083→0.089), completion ↑$0.01 (0.165→0.177).
- Qwen3.5‑35B‑A3B – Prompt ↑$0.03 (0.225→0.250), completion ↓$0.55 (1.80→1.25).
- MiniMax M2.5 – Prompt ↑$0.01 (0.220→0.225), completion unchanged (0.90).
- DeepSeek V3 0324 – Prompt ↓$0.02 (0.270→0.250), completion ↓$0.12 (1.12→1.00).
Cheapest models today (prompt / completion per 1M tokens)
- inclusionAI: Ling‑2.6‑flash – $0.01 / $0.03
- IBM: Granite 4.0 Micro – $0.02 / $0.11
- Mistral: Mistral Nemo – $0.02 / $0.03
Total models tracked: 414.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)