The Token Ledger – 2026-07-26
Most cost‑impacting change: Qwen’s Qwen3.6 27B completion price fell from $2.40 to $2.00 per 1M tokens (‑$0.40/M), a 16.7% reduction.
Price changes
-
Z.ai: GLM 5.2
- Prompt: $0.749 → $0.685 /M
- Completion: $2.354 → $2.152 /M
- Who should care: Teams running latency‑critical inference where both prompt and completion costs matter.
-
MoonshotAI: Kimi K2.7 Code
- Prompt: $0.78 → $0.75 /M
- Completion: unchanged at $3.50 /M
- Who should care: Code‑generation workflows sensitive to prompt pricing.
-
Qwen: Qwen3.6 27B (lead change)
- Prompt: $0.289 → $0.300 /M (+$0.011/M)
- Completion: $2.400 → $2.000 /M (‑$0.400/M)
- Who should care: Applications generating long completions (e.g., chat, summarization) that benefit from lower output cost.
-
DeepSeek: DeepSeek V4 Flash
- Prompt: $0.094 → $0.14 /M (+$0.046/M)
- Completion: $0.188 → $0.28 /M (+$0.092/M)
- Who should care: Cost‑conscious users of flash‑mode models; higher per‑token cost may shift usage to alternatives.
-
OpenAI: gpt‑oss‑20b
- Prompt: unchanged at $0.030 /M
- Completion: $0.130 → $0.140 /M (+$0.010/M)
- Who should care: Light‑weight completion tasks where a small cost increase may affect budgeting.
-
Qwen: Qwen3 30B A3B Instruct 2507
- Prompt: $0.100 → $0.048 /M (‑$0.052/M)
- Completion: $0.300 → $0.193 /M (‑$0.107/M)
- Who should care: Instruction‑following pipelines seeing notable savings on both input and output.
-
Qwen: Qwen3 30B A3B
- Prompt: $0.130 → $0.120 /M (‑$0.010/M)
- Completion: $0.520 → $0.500 /M (‑$0.020/M)
- Who should care: General‑purpose deployments where modest per‑token reductions accumulate over high volume.
No models were added or removed today.
Three cheapest models (per‑token pricing):
- inclusionAI: Ling‑2.6‑flash – Prompt $0.00001, Completion $0.00003 /M
- IBM: Granite 4.0 Micro – Prompt $0.000017, Completion $0.000112 /M
- Mistral: Mistral Nemo – Prompt $0.000019, Completion $0.00003 /M
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)