The Token Ledger Digest – 2026-09-02
Most cost‑impacting change: DeepSeek: DeepSeek V4 Pro 0813 – prompt price rose from $0.66 to $1.12/1M tokens (+$0.46) and completion from $1.98 to $3.35/1M (+$1.37). Developers using this model for high‑volume generation will see inference costs jump ~92%.
Added models
- Anthropic: Claude Fable 5.1 – prompt $10.00/1M, completion $50.00/1M. Teams needing long‑context (1M tokens) with strong reasoning should evaluate.
- Anthropic: Claude Fable 5.1 (batch) – prompt $5.00/1M, completion $25.00/1M. Batch users can halve costs vs. synchronous version.
- Inception: Mercury 2.5 Preview – prompt $0.04/1M, completion $0.15/1M. Ultra‑cheap for short‑form tasks; attractive for prototyping.
- Z.ai: GLM Flash Latest – prompt $0.075/1M, completion $0.25/1M. Good balance of cost and 1.3M‑token context for retrieval‑augmented workflows.
Removed models
- Claude Opus 5 (Fast) – prompt $10.00/1M, completion $50.00/1M. Users must migrate to Claude Fable 5.1 or alternatives.
- Anthropic: Claude Opus 4.8 (Fast) – same pricing as Opus 5 Fast; same migration impact.
Price changes
- Z.ai: GLM Latest – prompt $1.15/1M (down $0.02), completion $3.50/1M (down $0.46). Cost‑sensitive apps benefit slightly.
- Qwen: Qwen3.8 2.4T A95B (batch) – prompt $2.00/1M (down $0.50), completion $6.00/1M (down $0.25). Batch users see ~12% lower spend.
- DeepSeek: DeepSeek V4 Pro 0813 – prompt $1.12/1M (+$0.46), completion $3.35/1M (+$1.37). Notable cost increase; reconsider usage.
- Z.ai: GLM 5.2 – prompt $0.97/1M (down $0.22), completion $3.04/1M (down $0.70). Moderate savings for existing users.
- NVIDIA: Nemotron 3 Ultra – prompt $0.63/1M (+$0.13), completion $3.13/1M (+$0.93). Higher cost for those needing large model capacity.
- DeepSeek: DeepSeek V4 Pro 0423 – prompt $1.04/1M (down $0.56), completion $2.08/1M (down $1.12). Significant reduction; attractive for cost‑conscious deployments.
- DeepSeek: DeepSeek V4 Flash 0423 – prompt $0.089/1M (+$0.008), completion $0.177/1M (+$0.015). Minimal impact.
- DeepSeek: DeepSeek V3.1 – prompt $0.25/1M (down $0.30), completion $0.95/1M (down $0.70). Large savings for chat‑oriented workloads.
Cheapest models today (per‑million)
- IBM: Granite 4.0 Micro – prompt $0.017/1M, completion $0.112/1M
- Mistral: Mistral Nemo – prompt $0.019/1M, completion $0.030/1M
- Ling-3.0-flash – prompt $0.021/1M, completion $0.063/1M
Total models tracked: 421.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)