Token Ledger Digest – 2026-08-14
Most cost‑impacting change: Google Gemini Flash Latest (~google/gemini-flash-latest) slashed completion pricing from $7.50 to $1.875 per 1M tokens (‑$5.625) and prompt pricing from $1.50 to $0.375 per 1M tokens (‑$1.125). Developers using latency‑sensitive flash workloads now see a ~6.75× cost reduction.
Price changes
- DeepSeek V4 Flash 0731 – Prompt ↑$0.08→$0.14/1M (+$0.06); Completion ↑$0.18→$0.28/1M (+$0.10). Cost‑sensitive batch jobs may need re‑budgeting.
- Gemini 3.6 Flash – Prompt ↓$1.50→$0.75/1M (‑$0.75); Completion ↓$7.50→$3.75/1M (‑$3.75). Teams running medium‑length flash calls benefit from ~40% lower spend.
- Gemini 3.6 Flash (batch) – Prompt ↓$0.75→$0.375/1M (‑$0.375); Completion ↓$3.75→$1.875/1M (‑$1.875). Batch pipelines see ~50% cost cut.
- Z.ai GLM 5.2 – Prompt ↑$0.50→$0.63/1M (+$0.13); Completion ↓$3.15→$1.98/1M (‑$1.17). Mixed workloads: slight prompt rise, notable completion saving.
- Qwen3 VL 30B A3B Instruct – Prompt ↓$0.15→$0.13/1M (‑$0.02); Completion ↓$0.60→$0.52/1M (‑$0.08). Vision‑language users gain modest savings.
- Qwen3 Next 80B A3B Instruct – Prompt ↑$0.09→$0.10/1M (+$0.01); Completion unchanged at $1.10/1M. Negligible impact; monitor for future shifts.
Added models
- Gemini 3.7 Flash – Prompt $0.375/1M, Completion $1.875/1M. Developers needing 1M‑token context with strong reasoning.
- Gemini 3.7 Flash (batch) – Prompt $0.1875/1M, Completion $0.9375/1M. Batch‑oriented users seeking half‑price flash.
No removals
No models were removed today.
Cheapest models today (per 1M tokens)
- inclusionAI: Ling-2.6-flash – Prompt $0.01, Completion $0.03
- IBM: Granite 4.0 Micro – Prompt $0.017, Completion $0.112
- Mistral: Mistral Nemo – Prompt $0.019, Completion $0.03
Total models tracked: 411.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)