DEV Community

4663437Mehdi
4663437Mehdi

Posted on Originally published at 4663437mehdi.github.io

The Token Ledger Digest – 2026-08-30

The Token Ledger Digest – 2026-08-30

The most cost-impacting change today is Z.ai GLM 5.1 completion price dropping from $3.96/1M to $3.04/1M (prompt fell from $1.26/1M to $0.97/1M). Developers running long‑generation workloads should see lower bills.

  • DeepSeek V4 Flash Latest – completion rose $0.10/1M → $0.16/1M (prompt steady at $0.03/1M). Affects completion‑heavy apps.
  • DeepSeek V4 Flash 0731 – prompt $0.045/1M → $0.065/1M; completion $0.09/1M → $0.18/1M. Impacts both prompt and completion usage.
  • Thinking Machines: Inkling – prompt $0.95/1M → $1.00/1M (completion unchanged at $4.05/1M). Minor prompt cost rise for latency‑sensitive calls.
  • DeepSeek V4 Pro 0423 – prompt $0.61/1M → $0.42/1M; completion $1.22/1M → $0.83/1M. Significant savings for balanced workloads.
  • DeepSeek V4 Flash 0423 – prompt $0.083/1M → $0.080/1M; completion $0.166/1M → $0.159/1M. Small reductions for flash‑mode users.
  • Arcee AI: Trinity Large Thinking – prompt $0.22/1M → $0.25/1M; completion $0.85/1M → $0.80/1M. Slight prompt increase, completion decrease.
  • Mistral: Devstral 2 2512 – prompt $0.44/1M → $0.40/1M; completion $2.20/1M → $2.00/1M. Moderate savings across the board.
  • Qwen: Qwen3 Next 80B A3B Instruct – prompt $0.10/1M → $0.09/1M (completion steady at $1.10/1M). Minor prompt discount.

Cheapest models today (per‑million tokens):

  1. IBM Granite 4.0 Micro – prompt $0.017/1M, completion $0.112/1M
  2. Mistral Nemo – prompt $0.019/1M, completion $0.030/1M
  3. Ling‑3.0‑flash – prompt $0.021/1M, completion $0.063/1M

No models were added or removed. Total models tracked: 396.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)