DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

Token Ledger Digest – 2026-07-22

Token Ledger Digest – 2026-07-22

Most cost‑impacting change

  • Google Gemini Flash Latest – completion price fell from 9.0 $/1M to 7.5 $/1M (‑1.5 $/1M). Prompt price unchanged at 1.5 $/1M. Who should care: Teams running high‑volume completion workloads (chatbots, code generation) see immediate savings.

Other price changes

  • Z.ai: GLM 5.2 – prompt ↓ from 0.937 $/1M to 0.794 $/1M (‑0.143 $/1M); completion ↓ from 2.944 $/1M to 2.495 $/1M (‑0.449 $/1M). Who should care: Cost‑sensitive inference pipelines using this model benefit from ~0.59 $/1M total reduction.
  • DeepSeek: DeepSeek V4 Flash – prompt ↑ from 0.094 $/1M to 0.098 $/1M (+0.004 $/1M); completion ↑ from 0.188 $/1M to 0.196 $/1M (+0.008 $/1M). Who should care: Minor uptick; watch for budget drift in latency‑critical apps.
  • Qwen: Qwen3 Next 80B A3B Instruct – prompt ↑ from 0.098 $/1M to 0.100 $/1M (+0.003 $/1M); completion ↑ from 0.78 $/1M to 1.10 $/1M (+0.32 $/1M). Who should care: Large‑model users see a notable rise in completion cost; consider alternatives for long‑form tasks.

Added models

  • Poolside: Laguna S 2.1 – prompt 0.10 $/1M, completion 0.20 $/1M, 1M‑token context. Who should care: Developers needing long context at low cost.
  • Poolside: Laguna S 2.1 (free) – zero‑price tier, 256K context. Who should care: Prototyping or education with no spend.
  • Google: Gemini 3.6 Flash – prompt 1.5 $/1M, completion 7.5 $/1M, 1M‑token context. Who should care: Users seeking Gemini‑class performance with updated pricing.
  • Google: Gemini 3.5 Flash‑Lite – prompt 0.30 $/1M, completion 2.5 $/1M, 1M‑token context. Who should care: Budget‑conscious flash workloads.

Total models tracked: 342.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)