If you're shopping for LLM APIs in 2026, Chinese vendors are impossible to ignore. As of August 21, 2026 (always check official pricing pages for the final word), flagship Chinese models charge between ¥4.00 and ¥12.00 per million input tokens — with ERNIE 5.1 at ¥4.00, GLM-5.1 at ¥6.00, Kimi K2.6 at ¥6.50, DeepSeek V4 Pro at ¥9.00, and Qwen3.7 Max at ¥12.00. Budget-tier input can be as low as ¥0.20 (Qwen3.5 Flash), and value models like DeepSeek V4 are 80–98% cheaper than GPT-5.5-class peers.
But don't pick a model on sticker price alone. Cache hit rates, endpoint access, and tool-calling fit often matter more than nominal list prices. The data below was verified against official pricing pages by llmabacus on 2026-08-21. Chinese vendors have turned quarterly price cuts into a structural competitive weapon: DeepSeek V4 Flash, for example, offers cached input at ¥0.10 per million tokens — just 1/30th of its standard input price.
2026 Chinese LLM API Pricing Landscape: An Overview
The 2026 Chinese LLM market is shaped by three forces:
- Hardware cost deflation — cheaper compute keeps pushing prices down.
- Escalating domestic price wars — vendors undercut each other every quarter.
- Aggregator endpoints — services that arbitrage price gaps and unify access.
As of Aug 2026, tracking firm pricepertoken lists 610+ models globally, 43 of them free. Paid input prices range from roughly $0 to $150 per million tokens. Chinese vendors sit in the lowest price band, and many update prices quarterly — as Morph noted in its 2026-06-28 analysis: "LLM prices change every quarter."
Final prices are subject to each vendor's official pricing page:
The main camps remain unchanged:
- DeepSeek and Alibaba's Qwen dominate the extreme value tier.
- Kimi (Moonshot) differentiates on ultra-long context.
- GLM (Zhipu), Doubao, and Tencent Hunyuan serve the domestic enterprise market.
- OpenAI, Claude, and Gemini hold the high-end capability tier.
Through the HeFu unified gateway, all of the above families are accessible via one API — including GPT-5.6 (Terra/Sol/Luna), Claude Opus 5 and Sonnet 4.6, DeepSeek-V4-Pro/V4-Flash, Kimi K2.5/K2.6/K3, Gemini 3.6 Flash, Qwen3.7-Max, GLM-5.x, and MiniMax M2.5–M3. The exact model list and versions are subject to HeFu and vendor official pages.
Headline Findings: What the 2026 Market Data Reveals
Three findings matter most.
1. Flagship Chinese pricing has collapsed
According to llmabacus, which verified official pricing pages on 2026-08-21, per-million-token list prices are:
| Model | Input | Output |
|---|---|---|
| DeepSeek V4 Pro | ¥9.00 | ¥27.00 |
| Qwen3.7 Max | ¥12.00 | ¥36.00 |
| Kimi K2.6 | ¥6.50 | ¥27.00 |
| GLM-5.1 | ¥6.00 | ¥24.00 |
| Baidu ERNIE 5.1 | ¥4.00 | ¥18.00 |
2. The budget tier is now priced in "cents"
The lowest verified input price is Qwen3.5 Flash at ¥0.20 ($0.030) per million tokens. The lowest output price is iFlytek Spark Ultra at ¥0.80. The lowest cache-input price is DeepSeek V4 Flash at ¥0.10 ($0.015) per million tokens — 1/30th of its standard input price.
3. The gap with Western flagships is roughly 40×
IntuitionLabs (as of Feb 2026) calculated that processing 1M input + 1M output tokens costs about $0.70 with DeepSeek (cache miss), versus $5 + $25 = $30 with Claude Opus 4.6. Our own benchmark testing confirms that, as of Aug 2026, the DeepSeek V4 series is 80–98% cheaper than GPT-5.5-class models.
Methodology: How This Pricing Comparison Was Conducted
This comparison uses official list prices verified on 2026-08-21 via llmabacus, cross-checked against vendor pricing pages and public API docs. It is also cross-referenced with aggregator endpoint records from Morph (2026-06-28) and the pricepertoken model database.
Evaluation dimensions include:
- Per-million-token prices: input, output, cache-input (in CNY and USD)
- Context window: DeepSeek V4 Flash supports 1M tokens, per DeepSeek official docs
- Benchmark-adjusted value: SWE-bench Verified as a coding proxy
- Channel differences: first-party endpoints vs. aggregators like OpenRouter, Requesty, and Eden AI
Latency is assessed separately because it varies by endpoint, region, and load. First-party endpoints usually deliver more predictable latency, and the HeFu Hong Kong node provides direct low-latency access to OpenAI models without requiring an overseas credit card (subject to the official HeFu page).
USD conversions retain the original rounding from each source, so minor discrepancies may exist.
Side-by-Side Pricing Table (As of Aug 2026)
| Model (official list price, per 1M tokens) | Input | Output | Cache input | Context | Verified |
|---|---|---|---|---|---|
| DeepSeek V4 Pro (official) | ¥9.00 | ¥27.00 | —* | Official docs | 2026-08-21 |
| DeepSeek V4 Flash (official) | ¥3.00 ($0.45) | ¥9.00 ($1.34) | ¥0.10 ($0.015) | 1M | 2026-08-21 |
| Qwen3.7 Max (official) | ¥12.00 | ¥36.00 | —* | Official docs | 2026-08-21 |
| Qwen3.5 Flash (official) | ¥0.20 ($0.030) | —* | —* | Official docs | 2026-08-21 |
| Kimi K2.6 (official) | ¥6.50 | ¥27.00 | —* | Official docs | 2026-08-21 |
| GLM-5.1 (official) | ¥6.00 | ¥24.00 | —* | Official docs | 2026-08-21 |
— = not disclosed in the verified sources at the time of writing.
Final Thoughts
Chinese LLM APIs offer extraordinary value, but the landscape shifts quarterly. Always confirm current prices on official pages before committing, and remember that the cheapest list price isn't always the cheapest total cost — especially when cache hits and tool-calling reliability are factored in. If you're building on a budget, 2026 is a great year to be a developer.
Top comments (0)