DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

Token Ledger Digest – 2026-07-24

Token Ledger Digest – 2026-07-24

Most cost-impacting change: Qwen: Qwen3.6 27B – prompt price fell from $0.60 to $0.29 per 1M tokens, completion from $3.60 to $2.40 per 1M tokens.

Added

  • Ling-3.0-flash (free) – prompt $0.00 / 1M, completion $0.00 / 1M. Who should care: Teams needing zero‑cost experimentation or prototyping.

Price Changes

  • Tencent: Hy3 – prompt $0.14 → $0.13 / 1M (−$0.01); completion $0.58 → $0.53 / 1M (−$0.05). Who should care: Users of Hy3 for latency‑sensitive tasks seeing modest savings.
  • Z.ai: GLM 5.2 – prompt $0.80 → $0.75 / 1M (−$0.05); completion $2.52 → $2.37 / 1M (−$0.15). Who should care: Developers running large‑scale GLM 5.2 workloads.
  • MoonshotAI: Kimi K2.7 Code – prompt $0.82 → $0.71 / 1M (−$0.11); completion $3.75 → $3.49 / 1M (−$0.26). Who should care: Code‑generation pipelines benefiting from lower token cost.
  • Qwen: Qwen3.6 27B – prompt $0.60 → $0.29 / 1M (−$0.31); completion $3.60 → $2.40 / 1M (−$1.20). Who should care: Cost‑conscious teams using this model for bulk inference.
  • Google: Gemma 4 31B – prompt $0.12 → $0.10 / 1M (−$0.02); completion unchanged $0.35 / 1M. Who should care: Light‑usage scenarios where prompt savings matter.
  • MiniMax: MiniMax M2.7 – prompt $0.25 → $0.30 / 1M (+$0.05); completion $1.00 → $1.20 / 1M (+$0.20). Who should care: Budget‑watchers noting a cost rise for this variant.
  • Qwen: Qwen3.5-27B – prompt $0.26 → $0.20 / 1M (−$0.06); completion $2.60 → $1.56 / 1M (−$1.04). Who should care: Users seeing significant completion‑price relief.
  • Z.ai: GLM 5 – prompt unchanged $0.95 / 1M; completion $3.15 → $2.55 / 1M (−$0.60). Who should care: Completion‑heavy applications saving on GLM 5.
  • MiniMax: MiniMax M2 – prompt $0.30 → $0.26 / 1M (−$0.05); completion $1.20 → $1.02 / 1M (−$0.18). Who should care: Moderate savings across both prompt and completion.
  • OpenAI: gpt-oss-20b – prompt $0.03 → $0.03 / 1M (−$0.001); completion $0.13 → $0.14 / 1M (+$0.01). Who should care: Near‑neutral impact; slight completion cost increase.
  • Qwen: Qwen3 Coder 480B A35B – prompt $0.30 → $0.38 / 1M (+$0.08); completion $1.00 → $1.55 / 1M (+$0.55). Who should care: Cost increase for large‑code‑model users.
  • Google: Gemma 3 27B – prompt $0.10 → $0.08 / 1M (−$0.02); completion $0.30 → $0.45 / 1M (+$0.15). Who should care: Trade‑off: cheaper prompts, pricier completions.

Cheapest Models Today (per 1M tokens)

  1. inclusionAI: Ling-2.6-flash – prompt $0.01, completion $0.03
  2. IBM: Granite 4.0 Micro – prompt $0.02, completion $0.11
  3. Mistral: Mistral Nemo – prompt $0.02, completion $0.03

Total models tracked: 343.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)