Token Ledger Digest – 2026-09-10
Most cost‑impacting change
- Qwen: Qwen3.5-122B-A10B – Completion price fell from $2.40 / 1M to $2.08 / 1M (‑$0.32/M); prompt price fell from $0.29 / M to $0.26 / M (‑$0.03/M). Who should care: Teams running large‑scale Qwen inference; lowers cost per token for long‑form generation.
Added models
- DeepSeek V4.1 Flash – Prompt $0.15 / M, Completion $0.60 / M.
- Mistral Small 4 (batch) – Prompt $0.075 / M, Completion $0.30 / M.
- Ministral 3 8B 2512 (batch) – Prompt $0.075 / M, Completion $0.075 / M.
- Mistral Large 3 2512 (batch) – Prompt $0.25 / M, Completion $0.75 / M.
- Mistral Medium 3.1 (batch) – Prompt $0.20 / M, Completion $1.00 / M.
- Codestral 2508 (batch) – Prompt $0.15 / M, Completion $0.45 / M. Who should care: Developers seeking new Mistral or DeepSeek options; batch variants offer lower latency‑cost tradeoffs.
Removed model
- Nous: Hermes 4 70B – Prompt $0.13 / M, Completion $0.40 / M (no longer available). Who should care: Anyone relying on this model must migrate to alternatives.
Other price changes
- Z.ai: GLM 5.3 Flash – Prompt ↑$0.075/M (0.075→0.15), Completion ↑$0.25/M (0.25→0.50).
- Z.ai: GLM Latest – Prompt ↓$0.106/M (1.113→1.008), Completion ↓$0.088/M (3.498→3.410).
- MiniMax: MiniMax M2.5 – Prompt ↑$0.03/M (0.27→0.30), Completion ↑$0.12/M (1.08→1.20).
- DeepSeek: DeepSeek V3 – Prompt ↓$0.063/M (0.32→0.257), Completion ↑$0.139/M (0.890→1.029). Who should care: Users of these models should adjust cost estimates; GLM 5.3 Flash becomes notably more expensive for completion‑heavy workloads.
Cheapest models today (per‑million tokens)
- IBM Granite 4.0 Micro – Prompt $0.017/M, Completion $0.112/M
- Mistral Nemo – Prompt $0.019/M, Completion $0.030/M
- inclusionAI Ling 3.0 Flash – Prompt $0.021/M, Completion $0.063/M
Total models tracked: 436.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)