The Token Ledger – 2026-08-06
Most cost‑impacting change: Qwen: Qwen3.6 27B saw a major price increase. Prompt rose from $0.289 to $0.60 per 1M tokens (+$0.311); completion rose from $2.40 to $3.60 per 1M tokens (+$1.20). Total cost per 1M tokens is now $4.20, up $1.511. Teams running large‑scale inference or fine‑tuning on this model should reassess budget allocations.
Added models
- Meta: Muse Spark 1.2 – prompt $1.25 / 1M, completion $4.25 / 1M. Ideal for developers needing long‑context (1M tokens) generative tasks.
- inclusionAI: Ling-3.0-flash – prompt $0.075 / 1M, completion $0.22 / 1M. Suitable for low‑latency, cost‑sensitive applications with moderate context (128K).
Price changes
- MoonshotAI: Kimi K2.7 Code – prompt down $0.03 to $0.70 / 1M; completion unchanged at $3.50 / 1M. Minor savings for code‑generation workloads.
- DeepSeek: DeepSeek V4 Flash 0423 – prompt down $0.0518 to $0.0882 / 1M; completion down $0.1036 to $0.1764 / 1M. Total reduction $0.1554 / 1M; beneficial for high‑volume flash inference.
- Z.ai: GLM 5.1 – prompt down $0.014 to $0.952 / 1M; completion down $0.044 to $2.992 / 1M. Small overall cut ($0.058 / 1M) for balanced prompt/completion use.
- MiniMax: MiniMax M2.5 – prompt up $0.07 to $0.22 / 1M; completion steady at $0.90 / 1M. Slight cost rise for users of this model.
- Qwen: Qwen3 235B A22B Instruct 2507 – prompt down $0.0595 to $0.09 / 1M; completion down $0.048 to $0.55 / 1M. Total saving $0.1075 / 1M, relevant for large‑model deployments.
Cheapest models today (for reference): inclusionAI: Ling-2.6-flash ($0.01 / 1M prompt, $0.03 / 1M completion), IBM: Granite 4.0 Micro ($0.017 / 1M prompt, $0.112 / 1M completion), Mistral: Mistral Nemo ($0.019 / 1M prompt, $0.03 / 1M completion).
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)