DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

Token Ledger Digest – 2026-08-01

Token Ledger Digest – 2026-08-01

Most impactful change

  • Z.ai: GLM 5.2 – Prompt price fell from $1.232 / 1M to $0.760 / 1M; completion price fell from $3.872 / 1M to $2.389 / 1M. Who should care: Cost‑sensitive applications using this model see ~48% lower prompt and ~38% lower completion costs.

Added

  • Thinking Machines: Inkling Small – New model with 524 k context, prompt $0.50 / 1M, completion $1.20 / 1M. Who should care: Teams needing long‑context generation at low cost.

Removed

  • 30 batch models were deleted, including Google Gemini 3.6 Flash batch, Anthropic Claude Sonnet 5 batch, and OpenAI GPT‑5.5 batch. Their prompt prices ranged from $0.15 / 1M to $2.50 / 1M and completion prices from $1.25 / 1M to $25.00 / 1M. Who should care: Users relying on batch pricing must migrate to alternative endpoints or non‑batch versions.

Other price changes

  • NVIDIA: Nemotron 3 Ultra – Prompt $0.60 → $0.50 / 1M; completion $3.60 → $2.20 / 1M.
  • MoonshotAI: Kimi K2.6 – Prompt $0.95 → $0.60 / 1M; completion $4.00 → $3.41 / 1M.
  • Qwen: Qwen3 Coder 30B A3B Instruct – Prompt unchanged $0.07 / 1M; completion rose $0.27 → $0.28 / 1M.
  • Mistral: Mistral Small 3.2 24B – Prompt $0.10 → $0.075 / 1M; completion $0.30 → $0.20 / 1M.

Cheapest models today (for reference)

  1. inclusionAI: Ling‑2.6‑flash – $0.01 / 1M prompt, $0.03 / 1M completion
  2. IBM: Granite 4.0 Micro – $0.017 / 1M prompt, $0.112 / 1M completion
  3. Mistral: Mistral Nemo – $0.019 / 1M prompt, $0.03 / 1M completion

Total models tracked: 336.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)