Token Ledger Digest – 2026-07-24
Most cost-impacting change: Qwen: Qwen3.6 27B – prompt price fell from $0.60 to $0.29 per 1M tokens, completion from $3.60 to $2.40 per 1M tokens.
Added
- Ling-3.0-flash (free) – prompt $0.00 / 1M, completion $0.00 / 1M. Who should care: Teams needing zero‑cost experimentation or prototyping.
Price Changes
- Tencent: Hy3 – prompt $0.14 → $0.13 / 1M (−$0.01); completion $0.58 → $0.53 / 1M (−$0.05). Who should care: Users of Hy3 for latency‑sensitive tasks seeing modest savings.
- Z.ai: GLM 5.2 – prompt $0.80 → $0.75 / 1M (−$0.05); completion $2.52 → $2.37 / 1M (−$0.15). Who should care: Developers running large‑scale GLM 5.2 workloads.
- MoonshotAI: Kimi K2.7 Code – prompt $0.82 → $0.71 / 1M (−$0.11); completion $3.75 → $3.49 / 1M (−$0.26). Who should care: Code‑generation pipelines benefiting from lower token cost.
- Qwen: Qwen3.6 27B – prompt $0.60 → $0.29 / 1M (−$0.31); completion $3.60 → $2.40 / 1M (−$1.20). Who should care: Cost‑conscious teams using this model for bulk inference.
- Google: Gemma 4 31B – prompt $0.12 → $0.10 / 1M (−$0.02); completion unchanged $0.35 / 1M. Who should care: Light‑usage scenarios where prompt savings matter.
- MiniMax: MiniMax M2.7 – prompt $0.25 → $0.30 / 1M (+$0.05); completion $1.00 → $1.20 / 1M (+$0.20). Who should care: Budget‑watchers noting a cost rise for this variant.
- Qwen: Qwen3.5-27B – prompt $0.26 → $0.20 / 1M (−$0.06); completion $2.60 → $1.56 / 1M (−$1.04). Who should care: Users seeing significant completion‑price relief.
- Z.ai: GLM 5 – prompt unchanged $0.95 / 1M; completion $3.15 → $2.55 / 1M (−$0.60). Who should care: Completion‑heavy applications saving on GLM 5.
- MiniMax: MiniMax M2 – prompt $0.30 → $0.26 / 1M (−$0.05); completion $1.20 → $1.02 / 1M (−$0.18). Who should care: Moderate savings across both prompt and completion.
- OpenAI: gpt-oss-20b – prompt $0.03 → $0.03 / 1M (−$0.001); completion $0.13 → $0.14 / 1M (+$0.01). Who should care: Near‑neutral impact; slight completion cost increase.
- Qwen: Qwen3 Coder 480B A35B – prompt $0.30 → $0.38 / 1M (+$0.08); completion $1.00 → $1.55 / 1M (+$0.55). Who should care: Cost increase for large‑code‑model users.
- Google: Gemma 3 27B – prompt $0.10 → $0.08 / 1M (−$0.02); completion $0.30 → $0.45 / 1M (+$0.15). Who should care: Trade‑off: cheaper prompts, pricier completions.
Cheapest Models Today (per 1M tokens)
- inclusionAI: Ling-2.6-flash – prompt $0.01, completion $0.03
- IBM: Granite 4.0 Micro – prompt $0.02, completion $0.11
- Mistral: Mistral Nemo – prompt $0.02, completion $0.03
Total models tracked: 343.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)