DEV Community

liekeai
liekeai

Posted on Originally published at lieke-ai.com

China's AI Models Just Out-Called the US for 19 Straight Weeks: What It Means for Your Stack

China's AI Models Out-Called the US for 19 Straight Weeks: Enterprise LLM Guide (Sep 2026) | Lieke Tech

  • # China's AI Models Just Out-Called the US for 19 Straight Weeks: What It Means for Your Stack

56.7T vs 16.5T tokens per week, four of the global top five, and Tencent's Hy4 taking #1. Token volume is money voting — here's how enterprises capture the "high-intelligence, low-cost" dividend.
✅ 19 weeks at #1
✅ Model Studio from $0.15/M input
🚀 1M free tokens per model

Claim Free Credits →
View ECS Plans

📌 Key Takeaways

The numbers: per OpenRouter's latest weekly data (Aug 31 – Sep 6), global LLM usage hit 115 trillion tokens. Chinese models processed 56.7T (+2.83% WoW) versus 16.5T for US models (-3.1%) — China has now led the world for 19 consecutive weeks.

  • Four of the top five are Chinese: Tencent Hunyuan Hy4 preview took #1 with 14.7T tokens (+379% WoW); Zhipu GLM-5.3 Flash is #3; DeepSeek-V4-Flash holds #4 and #5.

  • Why it matters: token volume reflects real usage intensity, stickiness, and commercial value far better than user counts. The eastward shift means developers worldwide are moving production workloads to high-value, low-cost Chinese models.

  • How to plug in: one API key on Alibaba Cloud Model Studio gives access to Qwen, DeepSeek, MiniMax and more, with OpenAI/Anthropic-compatible endpoints. New accounts get free credits; Qwen3.8-Flash starts at $0.15/M input tokens.

  • Team plans: Model Studio Token Plan team seats bundle credits for AI coding/agent tools (Qwen Code, Claude Code, Qoder, OpenClaw), with nightly 50%-off discounts on selected DeepSeek models.

📊 The Data: Not a One-Week Fluke

OpenRouter is one of the world's largest LLM API aggregators, and its weekly usage is widely watched as the bellwether of real AI consumption. Last week's key figures:

Metric This week WoW
Global total usage 115T tokens +1.77%
Chinese models 56.7T tokens +2.83%
US models 16.5T tokens -3.1%
China : US ratio ~3.4 : 1 19th straight week ahead
Chinese models in global top 5 4 of 5 Hy4 / GLM / DeepSeek ×2

Global model usage — Top 6

# Model Vendor · Country Weekly tokens WoW
1 Hunyuan Hy4 preview Tencent · China 14.7T +379%
2 GPT-5.6 Luna OpenAI · US 12.9T +66%
3 GLM-5.3 Flash Zhipu · China 12.4T +101%
4 DeepSeek-V4-Flash (official) DeepSeek · China 12.4T
5 DeepSeek-V4-Flash (preview) DeepSeek · China 5.19T
6 MiniMax M3 MiniMax · China 5.02T +95%

Xiaomi MiMo-V2.5 and Gemini 3.7 Flash dropped out of the list this week. Source: OpenRouter, via public reporting (Sep 7, 2026).

How to read this chart: a token is the smallest unit a model processes. Compared with "user counts", token volume reveals actual usage intensity, stickiness, and commercial value — free-trial users don't generate steady tokens; only customers who wired AI into production (coding, support, content, agents) do. The sustained eastward shift is developers voting with their actual cloud bills.

🚀 What Is Hy4, the Week's Breakout?

The week's biggest story is Tencent Hunyuan's Hy4 preview: released and open-sourced on August 28, it took the #1 spot on OpenRouter within a week at 14.7T tokens (+379% WoW). Key specs:

Dimension Specification
Architecture MoE, 770B total / 49B activated parameters
Context 1M tokens (~1.5M Chinese characters)
License Apache 2.0 open weights (HuggingFace / GitHub / ModelScope)
Focus Agents, coding, productivity; generates playable game prototypes from one prompt
Blind eval 163 experts, 203 engineering tasks, avg 2.99/4.00 — ahead of GLM-5.3 and Kimi K3
Lightweight build Sep 1: weights compressed from 1.5TB to ~214GB via in-house Sherry ternary quantization, lowering self-hosting barriers

GLM-5.3 Flash, DeepSeek-V4-Flash and MiniMax M3 likewise crowded the leaderboard — all following the same playbook: open weights, rock-bottom API prices, and agent/coding-specific optimization. China's rise is a cluster breakthrough, not a single hit.

💰 The Enterprise Opportunity: Cheap Enough to Use Freely

Behind the usage surge is a fact that favors every business: top-tier model APIs are now cheap enough to use liberally. The question is no longer "can we afford AI" but "which channel is the most stable, cost-effective, and compliant".

For teams in China and globally, Alibaba Cloud Model Studio (Bailian) is the simplest unified entry point — one platform, one API key, multiple leading models:

Capability Details
Model aggregation Qwen3.8 family, DeepSeek-V4, MiniMax and more — text, image, video, voice; switch on demand
API compatibility OpenAI- and Anthropic-compatible endpoints — migrate existing code with near-zero changes
Tool ecosystem Works directly with Qwen Code, Claude Code, Qoder, OpenClaw and other leading AI coding/agent tools
New-user credits Free tokens per model (90-day validity)
Cost controls "Stop when free quota runs out" switch; alerts at 20% remaining and at exhaustion — no surprise bills
Enterprise-ready Multi-tenant isolation, no queueing at peak; no training on your conversation data; contracts, invoices, SLA

Model Studio pricing reference (international, USD)

Model / item Input Output Cached input
Qwen3.8-Flash $0.15 /M tokens $0.47 $0.016
Context / limits 262K native (1M via YaRN); 5,000 RPM / 5M TPM
Free credits Per-model free quota for new accounts; check console for eligibility

Figures from Alibaba Cloud Model Studio official pages (Sep 2026). The model price war is intense and adjustments are frequent — always confirm live pricing on the official campaign page.

🧭 Recommendations for Three Types of Teams

① Traditional businesses not yet using LLMs: start with free credits on one high-frequency scenario (support Q&A, knowledge base, document extraction). At $0.15/M input tokens, a 10-person internal agent typically costs just a few dollars a month in API calls; a budget ECS instance (from ~$14/year) hosts the business layer — no GPU needed.

② Teams already on closed overseas models: open Model Studio and run a shadow comparison — same prompts against Qwen/DeepSeek versus your current model. Many teams find Chinese models dramatically more cost-effective for coding and Chinese-language tasks; route by job (flagship for hard reasoning, Flash for bulk work).

③ Product/SaaS teams going global: use Alibaba Cloud International's Model Studio (USD billing, new-user credits) to serve worldwide customers, or self-host open-weight Hy4/DeepSeek/Qwen on Alibaba Cloud GPU instances with data staying inside your own VPC.

⚠️ Selection Pitfalls to Avoid

  • Don't look at unit price alone. In agent workloads, output tokens and cache misses drive the bill. Route long system prompts and knowledge-base prefixes through context caching (cached input as low as $0.016/M) — higher hit rates mean dramatically lower costs.

  • Marketplace token cards ≠ official channels. Third-party top-ups are fine for individuals; for production use the official cloud channel — free credits, invoices, SLA, 5,000 RPM quotas, and an auditable data chain.

  • Enable the quota stop-switch. You get alerts at 20% and at exhaustion; unverified accounts halt automatically while verified ones default to pay-as-you-go — set the switch to avoid surprise charges.

  • Free credits have a 90-day clock from activation — plan your pilot early.

  • Confirm live pricing. Prices change fast in this market; figures here are from public sources as of September 2026.

❓ FAQ

Does out-calling the US mean Chinese models lead on all technology?

Token volume reflects real-world usage scale and value, not universal technical supremacy. Chinese models clearly lead on open weights, price, agent/coding optimization, and Chinese-language scenarios — developers vote with their budgets. Frontier closed models still excel at some hard-reasoning tasks. The pragmatic enterprise approach is multi-model routing by task.

Is Model Studio limited to Alibaba's own models?

No. Model Studio aggregates Qwen, DeepSeek, MiniMax and other leading models across text, image, video, and voice — one API key, unified metering, switch on demand, with OpenAI/Anthropic-compatible endpoints.

Do individuals or small teams need the team Token Plan?

Individuals and small teams should start with pay-as-you-go plus free credits — costs are minimal. Team plans suit organizations needing seat management, usage analytics, and budget caps, with admins assigning/recycling seats and monitoring per-member usage.

Self-host open weights or call the API?

For most teams, the API wins — no GPUs to buy, no ops, pay-per-use, new models instantly available. Self-hosting open weights (Hy4, DeepSeek, Qwen) on Alibaba Cloud GPU instances makes sense only with hard data-residency requirements or volumes large enough to amortize GPU costs.

Start on Model Studio — Capture the Chinese AI Model Dividend

Free credits for new users · OpenAI/Anthropic-compatible · One key for Qwen, DeepSeek, MiniMax

Claim Free Credits →
View ECS Plans

© 2026 Lieke Tech · About · Privacy · Contact

Sources: OpenRouter weekly data via public reporting (Sep 2026), Tencent Hunyuan official materials, Alibaba Cloud Model Studio official docs; live pricing on official campaign pages prevails

Top comments (0)