Token Ledger Digest – 2026-07-22
Most cost‑impacting change
- Google Gemini Flash Latest – completion price fell from 9.0 $/1M to 7.5 $/1M (‑1.5 $/1M). Prompt price unchanged at 1.5 $/1M. Who should care: Teams running high‑volume completion workloads (chatbots, code generation) see immediate savings.
Other price changes
- Z.ai: GLM 5.2 – prompt ↓ from 0.937 $/1M to 0.794 $/1M (‑0.143 $/1M); completion ↓ from 2.944 $/1M to 2.495 $/1M (‑0.449 $/1M). Who should care: Cost‑sensitive inference pipelines using this model benefit from ~0.59 $/1M total reduction.
- DeepSeek: DeepSeek V4 Flash – prompt ↑ from 0.094 $/1M to 0.098 $/1M (+0.004 $/1M); completion ↑ from 0.188 $/1M to 0.196 $/1M (+0.008 $/1M). Who should care: Minor uptick; watch for budget drift in latency‑critical apps.
- Qwen: Qwen3 Next 80B A3B Instruct – prompt ↑ from 0.098 $/1M to 0.100 $/1M (+0.003 $/1M); completion ↑ from 0.78 $/1M to 1.10 $/1M (+0.32 $/1M). Who should care: Large‑model users see a notable rise in completion cost; consider alternatives for long‑form tasks.
Added models
- Poolside: Laguna S 2.1 – prompt 0.10 $/1M, completion 0.20 $/1M, 1M‑token context. Who should care: Developers needing long context at low cost.
- Poolside: Laguna S 2.1 (free) – zero‑price tier, 256K context. Who should care: Prototyping or education with no spend.
- Google: Gemini 3.6 Flash – prompt 1.5 $/1M, completion 7.5 $/1M, 1M‑token context. Who should care: Users seeking Gemini‑class performance with updated pricing.
- Google: Gemini 3.5 Flash‑Lite – prompt 0.30 $/1M, completion 2.5 $/1M, 1M‑token context. Who should care: Budget‑conscious flash workloads.
Total models tracked: 342.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)