The Token Ledger – 2026-07-20
Price Changes
- Z.ai: GLM 5.2 – Prompt price rose from $0.27/M to $1.12/M (+$0.85/M); completion rose from $0.85/M to $3.52/M (+$2.67/M). Affects: developers using GLM 5.2 for high‑volume workloads.
- NVIDIA: Nemotron 3 Super – Prompt price fell from $0.21/M to $0.085/M (−$0.125/M); completion fell from $0.455/M to $0.40/M (−$0.055/M). Affects: cost‑sensitive users of Nemotron 3 Super.
- Z.ai: GLM 5 – Prompt unchanged at $0.95/M; completion dropped from $3.15/M to $2.55/M (−$0.60/M). Affects: users of GLM 5 completion‑heavy tasks.
- Qwen: Qwen3 235B A22B Thinking 2507 – Prompt increased from $0.15/M to $0.30/M (+$0.15/M); completion jumped from $1.50/M to $3.00/M (+$1.50/M). Affects: users of the large thinking variant.
- Qwen: Qwen3 14B – Prompt decreased from $0.23/M to $0.12/M (−$0.11/M); completion decreased from $0.91/M to $0.24/M (−$0.67/M). Affects: users seeking cheaper 14B inference.
Model Removals (Free Tier)
Six free models were removed from the catalog:
- Qwen: Qwen3 Next 80B A3B Instruct (free)
- Qwen: Qwen3 Coder 480B A35B (free)
- Venice: Uncensored (free)
- Meta: Llama 3.3 70B Instruct (free)
- Meta: Llama 3.2 3B Instruct (free)
- Nous: Hermes 3 405B Instruct (free) All had $0 prompt and completion pricing. Affects: anyone relying on these free offerings; consider switching to paid alternatives or other free models.
Cheapest Models Today (per‑million token cost)
- inclusionAI: Ling-2.6-flash – Prompt $0.01/M, Completion $0.03/M
- IBM: Granite 4.0 Micro – Prompt $0.017/M, Completion $0.112/M
- Mistral: Mistral Nemo – Prompt $0.019/M, Completion $0.03/M
Total models tracked: 338.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)