The Token Ledger Digest – 2026-08-04
Z.ai: GLM 5.2 – Completion price fell from $3.74 to $2.42 per 1M tokens (‑$1.32/M). Prompt price dropped from $1.19 to $0.76 per 1M tokens (‑$0.43/M). Who should care: Teams running cost‑sensitive, long‑completion workloads on Z.ai.
Qwen: Qwen3.5‑122B‑A10B – Prompt ↓ $0.40→$0.26/M (‑$0.14/M). Completion ↓ $3.20→$2.08/M (‑$1.12/M). Who should care: Users of large‑scale Qwen models seeking lower generation cost.
Qwen: Qwen2.5 VL 72B Instruct – Prompt ↓ $0.80→$0.25/M (‑$0.55/M). Completion ↓ $1.00→$0.75/M (‑$0.25/M). Who should care: Vision‑language apps needing cheaper prompt processing.
MoonshotAI: Kimi K2.6 – Prompt ↓ $0.60→$0.589/M (‑$0.011/M). Completion ↓ $3.41→$2.48/M (‑$0.93/M). Who should care: Developers balancing prompt and completion costs.
Qwen: Qwen3.6 27B – Prompt ↓ $0.30→$0.289/M (‑$0.011/M). Completion ↑ $2.00→$2.40/M (+$0.40/M). Who should care: Slightly cheaper prompts but higher completion expense.
Meta: Llama 3.3 70B Instruct – Prompt ↓ $0.13→$0.10/M (‑$0.03/M). Completion ↓ $0.40→$0.32/M (‑$0.08/M). Who should care: General‑purpose Llama users seeing modest savings.
Qwen: Qwen3 Next 80B A3B Instruct – Prompt ↓ $0.10→$0.09/M (‑$0.01/M). Completion unchanged at $1.10/M. Who should care: Minor prompt‑cost reduction.
Qwen: Qwen3 Coder 30B A3B Instruct – Prompt unchanged $0.07/M. Completion ↓ $0.28→$0.27/M (‑$0.01/M). Who should care: Tiny completion‑cost trim.
Qwen: Qwen3 VL 30B A3B Instruct – Prompt ↑ $0.13→$0.15/M (+$0.02/M). Completion ↑ $0.52→$0.60/M (+$0.0. Who should care: Slight costlier.
*M0.08/M). *Who should care: Small increase for vision‑language tasks.
MythoMax 13B – Prompt ↑ $0.06→$0.08/M (+$0.02/M). Completion ↑ $0.06→$0.11/M (+$0.05/M). Who should care: Minor cost rise for this model.
Added: Qwen: Qwen3.8 Max – Prompt $2.00/M, Completion $6.00/M, 1M‑token context. Who should care: New high‑capacity option for enterprises needing massive context windows.
Originally published at The Token Ledger. Subscribe for the daily digest.
Top comments (0)