The Token Ledger Digest – 2026-08-30
The most cost-impacting change today is Z.ai GLM 5.1 completion price dropping from $3.96/1M to $3.04/1M (prompt fell from $1.26/1M to $0.97/1M). Developers running long‑generation workloads should see lower bills.
- DeepSeek V4 Flash Latest – completion rose $0.10/1M → $0.16/1M (prompt steady at $0.03/1M). Affects completion‑heavy apps.
- DeepSeek V4 Flash 0731 – prompt $0.045/1M → $0.065/1M; completion $0.09/1M → $0.18/1M. Impacts both prompt and completion usage.
- Thinking Machines: Inkling – prompt $0.95/1M → $1.00/1M (completion unchanged at $4.05/1M). Minor prompt cost rise for latency‑sensitive calls.
- DeepSeek V4 Pro 0423 – prompt $0.61/1M → $0.42/1M; completion $1.22/1M → $0.83/1M. Significant savings for balanced workloads.
- DeepSeek V4 Flash 0423 – prompt $0.083/1M → $0.080/1M; completion $0.166/1M → $0.159/1M. Small reductions for flash‑mode users.
- Arcee AI: Trinity Large Thinking – prompt $0.22/1M → $0.25/1M; completion $0.85/1M → $0.80/1M. Slight prompt increase, completion decrease.
- Mistral: Devstral 2 2512 – prompt $0.44/1M → $0.40/1M; completion $2.20/1M → $2.00/1M. Moderate savings across the board.
- Qwen: Qwen3 Next 80B A3B Instruct – prompt $0.10/1M → $0.09/1M (completion steady at $1.10/1M). Minor prompt discount.
Cheapest models today (per‑million tokens):
- IBM Granite 4.0 Micro – prompt $0.017/1M, completion $0.112/1M
- Mistral Nemo – prompt $0.019/1M, completion $0.030/1M
- Ling‑3.0‑flash – prompt $0.021/1M, completion $0.063/1M
No models were added or removed. Total models tracked: 396.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)