DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

The Token Ledger – 2026-08-06

The Token Ledger – 2026-08-06

Most cost‑impacting change: Qwen: Qwen3.6 27B saw a major price increase. Prompt rose from $0.289 to $0.60 per 1M tokens (+$0.311); completion rose from $2.40 to $3.60 per 1M tokens (+$1.20). Total cost per 1M tokens is now $4.20, up $1.511. Teams running large‑scale inference or fine‑tuning on this model should reassess budget allocations.

Added models

  • Meta: Muse Spark 1.2 – prompt $1.25 / 1M, completion $4.25 / 1M. Ideal for developers needing long‑context (1M tokens) generative tasks.
  • inclusionAI: Ling-3.0-flash – prompt $0.075 / 1M, completion $0.22 / 1M. Suitable for low‑latency, cost‑sensitive applications with moderate context (128K).

Price changes

  • MoonshotAI: Kimi K2.7 Code – prompt down $0.03 to $0.70 / 1M; completion unchanged at $3.50 / 1M. Minor savings for code‑generation workloads.
  • DeepSeek: DeepSeek V4 Flash 0423 – prompt down $0.0518 to $0.0882 / 1M; completion down $0.1036 to $0.1764 / 1M. Total reduction $0.1554 / 1M; beneficial for high‑volume flash inference.
  • Z.ai: GLM 5.1 – prompt down $0.014 to $0.952 / 1M; completion down $0.044 to $2.992 / 1M. Small overall cut ($0.058 / 1M) for balanced prompt/completion use.
  • MiniMax: MiniMax M2.5 – prompt up $0.07 to $0.22 / 1M; completion steady at $0.90 / 1M. Slight cost rise for users of this model.
  • Qwen: Qwen3 235B A22B Instruct 2507 – prompt down $0.0595 to $0.09 / 1M; completion down $0.048 to $0.55 / 1M. Total saving $0.1075 / 1M, relevant for large‑model deployments.

Cheapest models today (for reference): inclusionAI: Ling-2.6-flash ($0.01 / 1M prompt, $0.03 / 1M completion), IBM: Granite 4.0 Micro ($0.017 / 1M prompt, $0.112 / 1M completion), Mistral: Mistral Nemo ($0.019 / 1M prompt, $0.03 / 1M completion).


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)