DEV Community

4663437Mehdi
4663437Mehdi

Posted on • Originally published at 4663437mehdi.github.io

Token Ledger Digest – 2026-08-10

Token Ledger Digest – 2026-08-10

The most cost‑impacting change today is a price increase for Z.ai GLM 5.2 completion tokens (+$2.20 per 1M tokens).

Price changes

  • DeepSeek: DeepSeek V4 Flash 0731

    • Prompt: $0.09 → $0.08 / 1M (−$0.01)
    • Completion: unchanged $0.18 / 1M
    • Who should care: Teams running prompt‑heavy workloads see a slight saving.
  • Z.ai: GLM 5.2

    • Prompt: $0.07 → $0.76 / 1M (+$0.69)
    • Completion: $0.22 → $2.42 / 1M (+$2.20)
    • Who should care: Any application using this model for completion‑heavy tasks (chat, code generation) will see costs rise sharply.
  • MoonshotAI: Kimi K2.6

    • Prompt: $0.58 → $0.95 / 1M (+$0.37)
    • Completion: $2.44 → $4.00 / 1M (+$1.56)
    • Who should care: Users generating long outputs will notice a notable cost increase.
  • Z.ai: GLM 5.1

    • Prompt: $0.95 → $1.40 / 1M (+$0.45)
    • Completion: $2.99 → $4.40 / 1M (+$1.41)
    • Who should care: Both prompt and completion costs rise; affects all usage patterns.
  • Google: Gemma 4 26B A4B

    • Prompt: $0.07 → $0.12 / 1M (+$0.05)
    • Completion: $0.34 → $0.40 / 1M (+$0.06)
    • Who should care: Minor uptick; relevant for budget‑sensitive deployments.
  • NVIDIA: Nemotron 3 Super

    • Prompt: $0.30 → $0.09 / 1M (−$0.22)
    • Completion: $0.90 → $0.40 / 1M (−$0.50)
    • Who should care: Significant reduction benefits both prompt‑ and completion‑heavy jobs.
  • Qwen: Qwen3 14B

    • Prompt: $0.23 → $0.12 / 1M (−$0.11)
    • Completion: $0.91 → $0.24 / 1M (−$0.67)
    • Who should care: Large savings, especially for completion‑intensive workloads.
  • Meta: Llama 4 Maverick

    • Prompt: unchanged $0.20 / 1M
    • Completion: $0.80 → $0.70 / 1M (−$0.10)
    • Who should care: Small completion cost cut.
  • MythoMax 13B

    • Prompt: $0.08 → $0.06 / 1M (−$0.02)
    • Completion: $0.11 → $0.06 / 1M (−$0.05)
    • Who should care: Minor reductions across the board.

Cheapest models today (per‑token prices)

  1. inclusionAI: Ling‑2.6‑flash – prompt $0.01 / 1M, completion $0.03 / 1M
  2. IBM: Granite 4.0 Micro – prompt $0.02 / 1M, completion $0.11 / 1M
  3. Mistral: Mistral Nemo – prompt $0.02 / 1M, completion $0.03 / 1M

Total models tracked: 400.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)