China's AI Models Out-Called the US for 19 Straight Weeks: Enterprise LLM Guide (Sep 2026) | Lieke Tech
- # China's AI Models Just Out-Called the US for 19 Straight Weeks: What It Means for Your Stack
56.7T vs 16.5T tokens per week, four of the global top five, and Tencent's Hy4 taking #1. Token volume is money voting — here's how enterprises capture the "high-intelligence, low-cost" dividend.
✅ 19 weeks at #1
✅ Model Studio from $0.15/M input
🚀 1M free tokens per model
Claim Free Credits →
View ECS Plans
📌 Key Takeaways
The numbers: per OpenRouter's latest weekly data (Aug 31 – Sep 6), global LLM usage hit 115 trillion tokens. Chinese models processed 56.7T (+2.83% WoW) versus 16.5T for US models (-3.1%) — China has now led the world for 19 consecutive weeks.
Four of the top five are Chinese: Tencent Hunyuan Hy4 preview took #1 with 14.7T tokens (+379% WoW); Zhipu GLM-5.3 Flash is #3; DeepSeek-V4-Flash holds #4 and #5.
Why it matters: token volume reflects real usage intensity, stickiness, and commercial value far better than user counts. The eastward shift means developers worldwide are moving production workloads to high-value, low-cost Chinese models.
How to plug in: one API key on Alibaba Cloud Model Studio gives access to Qwen, DeepSeek, MiniMax and more, with OpenAI/Anthropic-compatible endpoints. New accounts get free credits; Qwen3.8-Flash starts at $0.15/M input tokens.
Team plans: Model Studio Token Plan team seats bundle credits for AI coding/agent tools (Qwen Code, Claude Code, Qoder, OpenClaw), with nightly 50%-off discounts on selected DeepSeek models.
📊 The Data: Not a One-Week Fluke
OpenRouter is one of the world's largest LLM API aggregators, and its weekly usage is widely watched as the bellwether of real AI consumption. Last week's key figures:
| Metric | This week | WoW |
|---|---|---|
| Global total usage | 115T tokens | +1.77% |
| Chinese models | 56.7T tokens | +2.83% |
| US models | 16.5T tokens | -3.1% |
| China : US ratio | ~3.4 : 1 | 19th straight week ahead |
| Chinese models in global top 5 | 4 of 5 | Hy4 / GLM / DeepSeek ×2 |
Global model usage — Top 6
| # | Model | Vendor · Country | Weekly tokens | WoW |
|---|---|---|---|---|
| 1 | Hunyuan Hy4 preview | Tencent · China | 14.7T | +379% |
| 2 | GPT-5.6 Luna | OpenAI · US | 12.9T | +66% |
| 3 | GLM-5.3 Flash | Zhipu · China | 12.4T | +101% |
| 4 | DeepSeek-V4-Flash (official) | DeepSeek · China | 12.4T | — |
| 5 | DeepSeek-V4-Flash (preview) | DeepSeek · China | 5.19T | — |
| 6 | MiniMax M3 | MiniMax · China | 5.02T | +95% |
Xiaomi MiMo-V2.5 and Gemini 3.7 Flash dropped out of the list this week. Source: OpenRouter, via public reporting (Sep 7, 2026).
How to read this chart: a token is the smallest unit a model processes. Compared with "user counts", token volume reveals actual usage intensity, stickiness, and commercial value — free-trial users don't generate steady tokens; only customers who wired AI into production (coding, support, content, agents) do. The sustained eastward shift is developers voting with their actual cloud bills.
🚀 What Is Hy4, the Week's Breakout?
The week's biggest story is Tencent Hunyuan's Hy4 preview: released and open-sourced on August 28, it took the #1 spot on OpenRouter within a week at 14.7T tokens (+379% WoW). Key specs:
| Dimension | Specification |
|---|---|
| Architecture | MoE, 770B total / 49B activated parameters |
| Context | 1M tokens (~1.5M Chinese characters) |
| License | Apache 2.0 open weights (HuggingFace / GitHub / ModelScope) |
| Focus | Agents, coding, productivity; generates playable game prototypes from one prompt |
| Blind eval | 163 experts, 203 engineering tasks, avg 2.99/4.00 — ahead of GLM-5.3 and Kimi K3 |
| Lightweight build | Sep 1: weights compressed from 1.5TB to ~214GB via in-house Sherry ternary quantization, lowering self-hosting barriers |
GLM-5.3 Flash, DeepSeek-V4-Flash and MiniMax M3 likewise crowded the leaderboard — all following the same playbook: open weights, rock-bottom API prices, and agent/coding-specific optimization. China's rise is a cluster breakthrough, not a single hit.
💰 The Enterprise Opportunity: Cheap Enough to Use Freely
Behind the usage surge is a fact that favors every business: top-tier model APIs are now cheap enough to use liberally. The question is no longer "can we afford AI" but "which channel is the most stable, cost-effective, and compliant".
For teams in China and globally, Alibaba Cloud Model Studio (Bailian) is the simplest unified entry point — one platform, one API key, multiple leading models:
| Capability | Details |
|---|---|
| Model aggregation | Qwen3.8 family, DeepSeek-V4, MiniMax and more — text, image, video, voice; switch on demand |
| API compatibility | OpenAI- and Anthropic-compatible endpoints — migrate existing code with near-zero changes |
| Tool ecosystem | Works directly with Qwen Code, Claude Code, Qoder, OpenClaw and other leading AI coding/agent tools |
| New-user credits | Free tokens per model (90-day validity) |
| Cost controls | "Stop when free quota runs out" switch; alerts at 20% remaining and at exhaustion — no surprise bills |
| Enterprise-ready | Multi-tenant isolation, no queueing at peak; no training on your conversation data; contracts, invoices, SLA |
Model Studio pricing reference (international, USD)
| Model / item | Input | Output | Cached input |
|---|---|---|---|
| Qwen3.8-Flash | $0.15 /M tokens | $0.47 | $0.016 |
| Context / limits | 262K native (1M via YaRN); 5,000 RPM / 5M TPM | ||
| Free credits | Per-model free quota for new accounts; check console for eligibility |
Figures from Alibaba Cloud Model Studio official pages (Sep 2026). The model price war is intense and adjustments are frequent — always confirm live pricing on the official campaign page.
🧭 Recommendations for Three Types of Teams
① Traditional businesses not yet using LLMs: start with free credits on one high-frequency scenario (support Q&A, knowledge base, document extraction). At $0.15/M input tokens, a 10-person internal agent typically costs just a few dollars a month in API calls; a budget ECS instance (from ~$14/year) hosts the business layer — no GPU needed.
② Teams already on closed overseas models: open Model Studio and run a shadow comparison — same prompts against Qwen/DeepSeek versus your current model. Many teams find Chinese models dramatically more cost-effective for coding and Chinese-language tasks; route by job (flagship for hard reasoning, Flash for bulk work).
③ Product/SaaS teams going global: use Alibaba Cloud International's Model Studio (USD billing, new-user credits) to serve worldwide customers, or self-host open-weight Hy4/DeepSeek/Qwen on Alibaba Cloud GPU instances with data staying inside your own VPC.
⚠️ Selection Pitfalls to Avoid
Don't look at unit price alone. In agent workloads, output tokens and cache misses drive the bill. Route long system prompts and knowledge-base prefixes through context caching (cached input as low as $0.016/M) — higher hit rates mean dramatically lower costs.
Marketplace token cards ≠ official channels. Third-party top-ups are fine for individuals; for production use the official cloud channel — free credits, invoices, SLA, 5,000 RPM quotas, and an auditable data chain.
Enable the quota stop-switch. You get alerts at 20% and at exhaustion; unverified accounts halt automatically while verified ones default to pay-as-you-go — set the switch to avoid surprise charges.
Free credits have a 90-day clock from activation — plan your pilot early.
Confirm live pricing. Prices change fast in this market; figures here are from public sources as of September 2026.
❓ FAQ
Does out-calling the US mean Chinese models lead on all technology?
Token volume reflects real-world usage scale and value, not universal technical supremacy. Chinese models clearly lead on open weights, price, agent/coding optimization, and Chinese-language scenarios — developers vote with their budgets. Frontier closed models still excel at some hard-reasoning tasks. The pragmatic enterprise approach is multi-model routing by task.
Is Model Studio limited to Alibaba's own models?
No. Model Studio aggregates Qwen, DeepSeek, MiniMax and other leading models across text, image, video, and voice — one API key, unified metering, switch on demand, with OpenAI/Anthropic-compatible endpoints.
Do individuals or small teams need the team Token Plan?
Individuals and small teams should start with pay-as-you-go plus free credits — costs are minimal. Team plans suit organizations needing seat management, usage analytics, and budget caps, with admins assigning/recycling seats and monitoring per-member usage.
Self-host open weights or call the API?
For most teams, the API wins — no GPUs to buy, no ops, pay-per-use, new models instantly available. Self-hosting open weights (Hy4, DeepSeek, Qwen) on Alibaba Cloud GPU instances makes sense only with hard data-residency requirements or volumes large enough to amortize GPU costs.
Start on Model Studio — Capture the Chinese AI Model Dividend
Free credits for new users · OpenAI/Anthropic-compatible · One key for Qwen, DeepSeek, MiniMax
Claim Free Credits →
View ECS Plans
© 2026 Lieke Tech · About · Privacy · Contact
Sources: OpenRouter weekly data via public reporting (Sep 2026), Tencent Hunyuan official materials, Alibaba Cloud Model Studio official docs; live pricing on official campaign pages prevails
Top comments (0)