DEV Community

4663437Mehdi
4663437Mehdi

Posted on Originally published at 4663437mehdi.github.io

The Token Ledger Digest – 2026-08-28

The Token Ledger Digest – 2026-08-28

Lead: Google Gemini 3.7 Flash and Gemini Flash Latest doubled both prompt and completion rates.

  • Google: Gemini 3.7 Flash – prompt rose from $0.375/M to $0.75/M (+$0.375/M); completion from $1.875/M to $3.75/M (+$1.875/M).

    Who should care: Teams using Gemini Flash for high‑volume inference; cost per million tokens jumps ~2.25×.

  • Google Gemini Flash Latest – identical change (prompt $0.375→$0.75/M, completion $1.875→$3.75/M).

    Who should care: Same as above; affects any workload pinned to the “latest” tag.

Other notable price shifts

  • DeepSeek V4 Pro 0813 – prompt fell $1.122/M → $0.66/M (−$0.462/M); completion $3.366/M → $1.98/M (−$1.386/M).

    Who should care: Cost‑sensitive users of DeepSeek’s larger V4 model see ~45% lower prompt and ~41% lower completion cost.

  • NVIDIA Nemotron 3 Ultra – prompt $0.60/M → $0.50/M (−$0.10/M); completion $3.60/M → $2.20/M (−$1.40/M).

    Who should care: Users of Nemotron Ultra benefit from a 61% drop in completion cost.

Additions & removals

  • 8 new models added, highlighted by Tencent Hy4 preview (prompt $0.834/M, completion $2.501/M) and several Mistral batch variants ranging from $0.15/M prompt to $0.75/M completion.
  • 37 models removed, chiefly the entire OpenAI GPT‑5.6 Luna/Terra/Sol families and GPT‑5.5/5.4 series, eliminating high‑priced options (e.g., GPT‑5.5 Pro at $15/M prompt, $90/M completion).

Cheapest models today (per‑million): IBM Granite 4.0 Micro ($0.017 prompt, $0.112 completion), Mistral Nemo ($0.019 prompt, $0.03 completion), Ling‑3.0‑flash ($0.021 prompt, $0.063 completion).

Total models tracked: 388.


Originally published at The Token Ledger. Subscribe for the daily digest.

Top comments (0)