The Token Ledger Digest – 2026-08-02
Price Drop – Z.ai: GLM 5.2
- Prompt: fell from $0.76 /1M to $0.28 /1M (‑63%).
- Completion: fell from $2.39 /1M to $0.89 /1M (‑63%).
- Who should care: Teams running large‑scale inference or fine‑tuning on GLM 5.2 see per‑token costs cut by roughly two‑thirds.
Price Drop – DeepSeek: DeepSeek V4 Flash 0731
- Prompt: dropped from $0.14 /1M to $0.09 /1M (‑36%).
- Completion: dropped from $0.28 /1M to $0.18 /1M (‑36%).
- Who should care: Users of the 0731 checkpoint benefit from lower latency‑cost trade‑offs for flash‑style workloads.
Price Increase – NVIDIA: Nemotron 3 Ultra
- Prompt: rose from $0.50 /1M to $0.60 /1M (+20%).
- Completion: rose from $2.20 /1M to $3.60 /1M (+64%).
- Who should care: Cost‑sensitive applications relying on Nemotron 3 Ultra for completion‑heavy tasks will see noticeably higher bills.
Added Model – DeepSeek V4 Flash Latest
- Prompt: $0.09 /1M
- Completion: $0.18 /1M
- Context: 1,048,576 tokens
- Who should care: Developers needing ultra‑long context with flash‑level pricing can now access this newest variant.
Total models tracked: 337.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)