<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: liekeai</title>
    <description>The latest articles on DEV Community by liekeai (@liekeai).</description>
    <link>https://dev.to/liekeai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4099873%2F0523a9ef-d51e-4555-b720-9ac2704e8058.png</url>
      <title>DEV Community: liekeai</title>
      <link>https://dev.to/liekeai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/liekeai"/>
    <language>en</language>
    <item>
      <title>Qwen3.8-Flash: Full Review</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/qwen38-flash-full-review-4fde</link>
      <guid>https://dev.to/liekeai/qwen38-flash-full-review-4fde</guid>
      <description>&lt;p&gt;Qwen3.8-Flash Review: $0.15/M Input, 67% Cheaper Than DeepSeek | Lieke Tech&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;# Qwen3.8-Flash: Full Review&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Price cut just one day after launch — $0.15/M input tokens, 89% lower training cost, coding benchmarks match DeepSeek V4 Pro&lt;br&gt;
✅ Input $0.15/M tokens&lt;br&gt;
✅ Output $0.47/M tokens&lt;br&gt;
✅ Cache hit $0.016/M&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lieke-ai.com/r?lk=coupon&amp;amp;userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Claim $200 Free Credit →&lt;/a&gt;&lt;br&gt;
ECS from $4.50/mo&lt;/p&gt;

&lt;h2&gt;
  
  
  📌 Key Takeaways
&lt;/h2&gt;

&lt;p&gt;Launch then price cut: Qwen3.8-Flash launched on August 26 at $0.16/$0.47 per million tokens; Alibaba cut the China-side price on August 27 to ¥0.8 (input) and ¥2.7 (output), with cache hits at ¥0.1/M&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Architecture breakthrough: 125B-parameter MoE model activating only 6B per token, plus a 51B N-gram embedding layer — training cost is ~1/9 of Qwen3.7-Plus (89% reduction)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Competitive performance: Scores 62.5 on SWE-bench Pro (agentic coding), leads in 8 of 14 benchmarks, matching DeepSeek-V4-Flash and Claude Opus 4.6 on real-world tasks&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Mega context: Natively supports 262K tokens, extends to 1M with YaRN. Multimodal: text, image, and video input&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Cache economics: Cache hit at $0.016/M tokens — ideal for agents, RAG, and long conversations&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  💰 Latest Pricing (Effective August 27, 2026)
&lt;/h2&gt;

&lt;p&gt;Pricing via Alibaba Cloud Model Studio (international) and Bailian platform (China), per million tokens.&lt;/p&gt;

&lt;h3&gt;
  
  
  International Pricing (USD)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Billing Item&lt;/th&gt;
&lt;th&gt;Price (per 1M tokens)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;$0.15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;$0.47&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache hit&lt;/td&gt;
&lt;td&gt;$0.016&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  China Pricing (RMB, after Aug 27 cut)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Billing Item&lt;/th&gt;
&lt;th&gt;Launch Price (Aug 26)&lt;/th&gt;
&lt;th&gt;Current Price (Aug 27+)&lt;/th&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;¥1.0&lt;/td&gt;
&lt;td&gt;¥0.8&lt;/td&gt;
&lt;td&gt;↓20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;¥3.0&lt;/td&gt;
&lt;td&gt;¥2.7&lt;/td&gt;
&lt;td&gt;↓10%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache hit&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;¥0.1&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;💡 Real-world cost example:&lt;/strong&gt; An AI customer service app handling 1M requests/day (avg 2K input + 500 output tokens):&lt;/p&gt;

&lt;p&gt;Qwen3.8-Flash: ~$415/day total&lt;/p&gt;

&lt;p&gt;DeepSeek V4 Flash (peak): ~$1,260/day&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Savings: ~$300,000+ per year&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚔️ Competitive Pricing Comparison
&lt;/h2&gt;

&lt;p&gt;August 2026 pricing for major LLM APIs, USD per million tokens.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Active Params&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-Flash NEW&lt;/td&gt;
&lt;td&gt;$0.15&lt;/td&gt;
&lt;td&gt;$0.47&lt;/td&gt;
&lt;td&gt;6B (MoE)&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-Flash (≤128K)&lt;/td&gt;
&lt;td&gt;$0.11&lt;/td&gt;
&lt;td&gt;$0.67&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash&lt;/td&gt;
&lt;td&gt;$0.034&lt;/td&gt;
&lt;td&gt;$0.132&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;256K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-Max&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$6.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash (peak)&lt;/td&gt;
&lt;td&gt;$0.42&lt;/td&gt;
&lt;td&gt;$1.26&lt;/td&gt;
&lt;td&gt;13B&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash (off-peak 50%)&lt;/td&gt;
&lt;td&gt;$0.21&lt;/td&gt;
&lt;td&gt;$0.63&lt;/td&gt;
&lt;td&gt;13B&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Pro (peak)&lt;/td&gt;
&lt;td&gt;$1.26&lt;/td&gt;
&lt;td&gt;$3.78&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.3-Flash&lt;/td&gt;
&lt;td&gt;~$0.15+&lt;/td&gt;
&lt;td&gt;~$0.47+&lt;/td&gt;
&lt;td&gt;18B (MoE)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Key findings:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Qwen3.8-Flash output is 63% cheaper than DeepSeek V4 Flash peak ($0.47 vs $1.26)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Even against DeepSeek's off-peak pricing, Qwen3.8-Flash remains 25% cheaper on output&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;API price is 1/10 of GLM-5.3, and up to 1/20 during promotional periods&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Qwen3.7-Flash is cheaper but lacks the 1M context and multimodal capabilities of 3.8&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🏗️ Architecture Deep Dive
&lt;/h2&gt;

&lt;p&gt;Qwen3.8-Flash (codenamed Qwen3.8-Flash-Next during development) is an early preview of the Qwen4 architecture, with four core innovations:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Mixture-of-Experts (MoE)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;125B main parameters with only 6B activated per token (10 routed experts + 1 shared expert out of 512)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Additional 51B N-gram embedding layer stored in system RAM (not GPU), acting as a massive "phrase dictionary"&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;4B multi-token prediction parameters bring total footprint to ~180B&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. QSA Sparse Attention + GDN Hybrid
&lt;/h3&gt;

&lt;p&gt;Qwen Sparse Attention (QSA) works alongside GDN: GDN compresses historical context while QSA selects key information. In high-cache-hit 1M-token scenarios, this delivers 7.6x faster prefill and 4.9x faster decode.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Gated Residual Mechanism
&lt;/h3&gt;

&lt;p&gt;Splits the traditional single residual pathway into 4 parallel branches, dynamically gating information flow for better cross-layer communication and training stability.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Muon Optimizer
&lt;/h3&gt;

&lt;p&gt;Uses a refined Muon + AdamW hybrid strategy with refitted scaling laws, eliminating batch warm-up and significantly improving convergence efficiency and training throughput.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Result:&lt;/strong&gt; Training costs are approximately 1/9 of Qwen3.7-Plus (397B params, 17B activated) — an 89% reduction — while delivering superior coding and office task performance.&lt;/p&gt;

&lt;h2&gt;
  
  
  📊 Benchmark Performance
&lt;/h2&gt;

&lt;p&gt;Per Qwen's official technical report, Qwen3.8-Flash excels across multiple evaluation suites:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Domain&lt;/th&gt;
&lt;th&gt;Performance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench Pro&lt;/td&gt;
&lt;td&gt;Agentic coding&lt;/td&gt;
&lt;td&gt;62.5, leading peers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CoWorkBench&lt;/td&gt;
&lt;td&gt;Long-horizon office tasks&lt;/td&gt;
&lt;td&gt;Beats DeepSeek V4 Flash&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toolathlon Verified&lt;/td&gt;
&lt;td&gt;Real-world tool use&lt;/td&gt;
&lt;td&gt;Matches Claude Opus 4.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MathVision&lt;/td&gt;
&lt;td&gt;Visual math reasoning&lt;/td&gt;
&lt;td&gt;Strong multimodal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AndroidWorld&lt;/td&gt;
&lt;td&gt;Mobile agent tasks&lt;/td&gt;
&lt;td&gt;Embodied intelligence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ERQA&lt;/td&gt;
&lt;td&gt;Embodied reasoning&lt;/td&gt;
&lt;td&gt;Multimodal understanding&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Across 14 evaluations, the base model (6B activated) achieved the best results in 8. The fine-tuned version shows even stronger performance in coding, agents, and multimodal tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  🎯 Best Use Cases
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Qwen3.8-Flash is ideal for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;AI coding assistants: 62.5 on SWE-bench Pro, repository-level understanding with 1M context&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;AI agents / tool use: Native parallel tool calls, extremely low cost for long-horizon tasks&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Customer service / chatbots: $0.016/M cache hit makes multi-turn conversations nearly free&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Document analysis / RAG: 1M tokens processes hundreds of pages, cache speeds up 8x&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Multimodal apps: Text + image + video input, visual math and chart analysis&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;High-volume batch processing: 5,000 RPM / 5M TPM rate limits for enterprise scale&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🚀 How to Get Started
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Option 1: Alibaba Cloud Model Studio (International)
&lt;/h3&gt;

&lt;p&gt;Access Qwen3.8-Flash via Model Studio with OpenAI-compatible API. New users get free credit. The model serves on QwenCloud with 1M context by default.&lt;/p&gt;

&lt;p&gt;Claim $200 Free Credit →&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 2: Bailian Platform (China)
&lt;/h3&gt;

&lt;p&gt;Available on Alibaba Cloud's Bailian platform with new user free token quota. First to receive the latest Qwen releases.&lt;/p&gt;

&lt;p&gt;Bailian Console →&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 3: Open Weights (Self-hosted)
&lt;/h3&gt;

&lt;p&gt;The open-weight version, Qwen3.8-Flash-Next, is available on Hugging Face and ModelScope for download, fine-tuning, and local deployment. The production version with built-in tools and 1M default context is served via QwenCloud API.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;New user offers:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Alibaba Cloud international: up to $200 free trial credit&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Bailian platform: free token quota for new users (90 days)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;ECS cloud servers: from ~$4.50/month (2 vCPU, 2GB RAM)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  📰 September 19, 2026 update: Qwen3.8-27B generates working web apps from a single prompt
&lt;/h2&gt;

&lt;p&gt;Qwen3.8-27B keeps trending in developer circles this week. In a hands-on test published by QbitAI on September 19, a single prompt led Qwen3.8-27B to produce a 52KB offline data-analysis tool in about 5 minutes — automatically cleaning tables, aggregating revenue stats and rendering bar charts — and even a realistic 12306 ticket-booking page (a convincing front-end shell that, by default, does not connect to live data or payment backends). Engineer Alok paired Qwen3.8-27B with Cerebras to build an offline "AI computer desktop" that regenerates historical versions of Google, YouTube and other sites on the fly at roughly 1,950 tokens/s (individual benchmark). On September 18, the Qwen team also launched &lt;strong&gt;Qwen3.8-Omni-Flash&lt;/strong&gt;, a next-generation omnimodal model with text/image/audio/video inputs and a 1M-token context, cutting per-hour audio-input API pricing by more than 98% versus the previous generation. From local 27B deployment to cloud APIs, the Qwen3.8 family is making "one-prompt tool delivery" real — for enterprises, calling Qwen3.8-Flash, Qwen3.8-27B and Omni-Flash via Alibaba Cloud Model Studio / Bailian is the most cost-effective way to try this generation of models.&lt;/p&gt;

&lt;h2&gt;
  
  
  📰 September 8, 2026 update: Qwen3.8-27B quantization benchmark
&lt;/h2&gt;

&lt;p&gt;This week a Hacker News benchmark post by Quesma CTO Piotr Migdał gave a definitive answer for running Qwen3.8-27B locally: &lt;strong&gt;the 17GB Q4_K_M 4-bit quant matches the 55GB BF16 full model on Terminal-Bench 2.1&lt;/strong&gt; and fits on a 24GB consumer card (e.g. RTX 4090) with room for ~64K tokens of context. At 4-bit, GPQA Diamond, IFBench and Terminal-Bench 2.1 scores are statistically indistinguishable from the reference, at one-third the on-disk size. &lt;strong&gt;Compression hits a clear cliff at 1-bit&lt;/strong&gt;: the 6.2GB UD-IQ1_S falls to near-random on GPQA Diamond, and the xhigh reasoning effort actually makes it worse (the model often runs out of token budget without committing to an answer). Unsloth's stated minimum for agentic / tool-calling work remains the 9.8GB Q2_K_XL. For teams weighing a self-hosted 27B deployment, this is the most current, third-party-validated hardware and quantization guidance available.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/" rel="noopener noreferrer"&gt;Quesma: Benchmarking Qwen3.8 27B quantizations&lt;/a&gt; (202 HN points, 99 comments on Sept 8, 2026). The original pricing and architecture analysis above remain valid; this card only adds the local-deployment sizing reference.&lt;/p&gt;

&lt;h2&gt;
  
  
  ❓ FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Qwen3.8-Flash vs Qwen3.8-27B — what's the difference?
&lt;/h3&gt;

&lt;p&gt;Qwen3.8-Flash is a MoE model (125B total / 6B active) optimized for cost efficiency and high throughput. Qwen3.8-27B is a dense vision-language model (all 27B active) for deep reasoning at $0.42/$3.08 per million tokens. Choose Flash for scale, 27B for depth.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does the cache hit pricing work?
&lt;/h3&gt;

&lt;p&gt;When a request hits the context cache (e.g., repeated system prompts, long document prefixes), cached input tokens are billed at $0.016/M instead of $0.15/M — a 90% discount. This dramatically reduces costs for agents, RAG, and multi-turn conversations.&lt;/p&gt;

&lt;h3&gt;
  
  
  What input modalities are supported?
&lt;/h3&gt;

&lt;p&gt;Qwen3.8-Flash supports text and image input natively (input_modalities: ["text", "image"]), with video analysis in the production version. It handles chart analysis, document OCR, and visual math.&lt;/p&gt;

&lt;h3&gt;
  
  
  What are the rate limits?
&lt;/h3&gt;

&lt;p&gt;5,000 RPM (requests per minute) and 5,000,000 TPM (tokens per minute), suitable for enterprise-grade high-concurrency deployments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start Building with Qwen3.8-Flash Today
&lt;/h2&gt;

&lt;p&gt;$200 free credit + ECS from $4.50/month + open-weight models available&lt;/p&gt;

&lt;p&gt;Claim $200 Free Credit →&lt;br&gt;
View ECS Plans&lt;/p&gt;

&lt;p&gt;⏰ Alibaba Cloud September Deals · Exclusive Channel&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;New users: free trial credits + coupon bundle across ECS, databases and more&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Model Studio (Bailian) LLM platform: up to 1M free tokens per model for 90 days; up to 50% off for selected models during 22:00-08:00 (UTC+8) off-peak hours&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Overseas regions: Singapore / Hong Kong / US ECS with no ICP filing required&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Enterprise: annual ECS and GPU instances with up to 40% off, plus channel-only pricing on request&lt;/p&gt;

&lt;p&gt;🔥 Claim Free Credits →&lt;br&gt;
View ECS Plans →&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Offers subject to official campaign pages; new-user deals require identity verification.&lt;/p&gt;

&lt;p&gt;© 2026 Lieke Tech · &lt;a href="https://lieke-ai.com/en/about" rel="noopener noreferrer"&gt;About&lt;/a&gt; · &lt;a href="https://lieke-ai.com/en/privacy" rel="noopener noreferrer"&gt;Privacy&lt;/a&gt; · &lt;a href="https://lieke-ai.com/mailto:lieke-promo@coze.email" rel="noopener noreferrer"&gt;Contact&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Pricing sources: Alibaba Cloud Bailian announcement (Aug 27, 2026), Qwen official blog, Alibaba Cloud Model Studio pricing page&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>Qwen3 API Pricing: Complete Guide</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Sun, 20 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/qwen3-api-pricing-complete-guide-18gh</link>
      <guid>https://dev.to/liekeai/qwen3-api-pricing-complete-guide-18gh</guid>
      <description>&lt;p&gt;Latest Update (September 2026)&lt;/p&gt;

&lt;p&gt;Beyond text models, the &lt;strong&gt;Qwen3 voice ecosystem is heating up&lt;/strong&gt;: Alibaba Cloud Model Studio (Bailian) now serves official &lt;strong&gt;Qwen3-TTS-Flash&lt;/strong&gt; and &lt;strong&gt;Qwen3-ASR-Flash-Realtime&lt;/strong&gt; endpoints (17 expressive voices, multi-language, emotion recognition, dialect support). Independent inference providers are also pushing the open-weight Qwen3-TTS 1.7B model to the top of Coval's voice AI benchmarks - #1 latency for speech-to-text and #1 word-error-rate for text-to-speech as of mid-September 2026. For developers building voice agents, Qwen voice APIs combine &lt;strong&gt;leading open-source quality with low per-character cost&lt;/strong&gt;, and you can get started with the same free token allowance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qwen3 Model Pricing Overview
&lt;/h2&gt;

&lt;p&gt;Alibaba Cloud offers a range of Qwen3 models through DashScope/Model Studio. Prices below are per 1 million tokens, in USD for international access and CNY for China mainland.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input (per M tokens)&lt;/th&gt;
&lt;th&gt;Output (per M tokens)&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Free Tier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen Code CLI&lt;/td&gt;
&lt;td&gt;FREE&lt;/td&gt;
&lt;td&gt;FREE&lt;/td&gt;
&lt;td&gt;262K&lt;/td&gt;
&lt;td&gt;1,000-2,000 calls/day&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-Max (flagship)&lt;/td&gt;
&lt;td&gt;~$1.65 (¥12)&lt;/td&gt;
&lt;td&gt;~$4.95 (¥36)&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;70M tokens (new users)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Max&lt;/td&gt;
&lt;td&gt;~$0.83 (¥6)*&lt;/td&gt;
&lt;td&gt;~$2.48 (¥18)*&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;70M tokens (new users)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Plus&lt;/td&gt;
&lt;td&gt;~$0.55 (¥4)&lt;/td&gt;
&lt;td&gt;~$1.65 (¥12)&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;Free quota available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash (0-32K ctx)&lt;/td&gt;
&lt;td&gt;$0.03 (¥0.225)&lt;/td&gt;
&lt;td&gt;$0.13 (¥0.974)&lt;/td&gt;
&lt;td&gt;256K&lt;/td&gt;
&lt;td&gt;Free quota available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash (32-256K ctx)&lt;/td&gt;
&lt;td&gt;$0.10-$0.21&lt;/td&gt;
&lt;td&gt;$0.41-$0.83&lt;/td&gt;
&lt;td&gt;256K&lt;/td&gt;
&lt;td&gt;Free quota available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder-Plus&lt;/td&gt;
&lt;td&gt;~$0.83 (¥6)&lt;/td&gt;
&lt;td&gt;~$2.48 (¥18)&lt;/td&gt;
&lt;td&gt;262K&lt;/td&gt;
&lt;td&gt;Via Qwen Code free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder 480B (Vertex AI)&lt;/td&gt;
&lt;td&gt;$0.22&lt;/td&gt;
&lt;td&gt;$1.80&lt;/td&gt;
&lt;td&gt;262K&lt;/td&gt;
&lt;td&gt;Via Qwen Code free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-Omni (multimodal)&lt;/td&gt;
&lt;td&gt;~$0.25 (¥1.8)&lt;/td&gt;
&lt;td&gt;~$2.17 (¥15.8)&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;1M tokens free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;** Qwen3.7-Max prices shown are after the 50% limited-time discount (standard: ¥12/¥36). USD conversions approximate at 7.25 CNY/USD.*&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best Value:&lt;/strong&gt; Qwen3.7-Flash at $0.03/M input tokens is one of the cheapest production-grade LLMs available globally. For coding, Qwen Code CLI is completely free for up to 2,000 calls per day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free Tier Breakdown
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Option 1: Qwen Code CLI (Always Free)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Mainland China: 2,000 API calls/day via Qwen OAuth or ModelScope&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;International: 1,000 API calls/day via OpenRouter&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;No token limits within each call&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;60 requests/minute rate limit&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Uses Qwen3-Coder 480B model&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;npx @qwen-code/qwen-code@latest&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 2: New User Tokens
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;70 million+ free tokens across Qwen models&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Valid for 180 days (extended from 90 days)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;100 AI image generation credits&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;50 seconds of video generation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;200 CNY no-threshold coupon&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;No credit card required (international) / real-name verification (China)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Option 3: OpenRouter Free Models
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;qwen/qwen3-coder:free — 20 RPM, 200 requests/day&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;qwen3-32b — 30 RPM, 1,000 requests/day on GitHub Models&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  API Quick Start
&lt;/h2&gt;

&lt;p&gt;from openai import OpenAI&lt;/p&gt;

&lt;h1&gt;
  
  
  International endpoint
&lt;/h1&gt;

&lt;p&gt;client = OpenAI(&lt;br&gt;
    api_key="your-dashscope-key",&lt;br&gt;
    base_url="&lt;a href="https://dashscope-intl.aliyuncs.com/compatible-mode/v1" rel="noopener noreferrer"&gt;https://dashscope-intl.aliyuncs.com/compatible-mode/v1&lt;/a&gt;"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;response = client.chat.completions.create(&lt;br&gt;
    model="qwen3.7-max",&lt;br&gt;
    messages=[{"role": "user", "content": "Explain async/await in Python"}]&lt;br&gt;
)&lt;br&gt;
print(response.choices[0].message.content)&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost Comparison with Competitors
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input/M tokens&lt;/th&gt;
&lt;th&gt;Output/M tokens&lt;/th&gt;
&lt;th&gt;Free Tier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash&lt;/td&gt;
&lt;td&gt;$0.03&lt;/td&gt;
&lt;td&gt;$0.13&lt;/td&gt;
&lt;td&gt;Yes (70M tokens)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V3&lt;/td&gt;
&lt;td&gt;$0.14&lt;/td&gt;
&lt;td&gt;$0.56&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o Mini&lt;/td&gt;
&lt;td&gt;$0.15&lt;/td&gt;
&lt;td&gt;$0.60&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku&lt;/td&gt;
&lt;td&gt;$0.80&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini Flash&lt;/td&gt;
&lt;td&gt;$0.075&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen Code (CLI)&lt;/td&gt;
&lt;td&gt;FREE&lt;/td&gt;
&lt;td&gt;FREE&lt;/td&gt;
&lt;td&gt;1,000-2,000 calls/day&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Money-Saving Tips
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Use Flash models for simple tasks: Qwen3.7-Flash is 90% cheaper than Max and sufficient for classification, summarization, and simple chatbots.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Maximize free tier first: 70M tokens can handle approximately 70,000-100,000 typical API requests before needing to pay.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Use Qwen Code for development: 2,000 free coding calls per day replaces $10-20/month Copilot subscriptions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Context caching discount: Qwen3.7-Flash offers reduced pricing for cached tokens (¥0.225/M vs ¥0.749/M for 32K context).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Coding Plan Lite: Heavy coders can get unlimited Qwen3-Coder for 7.9 CNY (~$1.10) first month.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Qoder paid plans from $20/month: for AI agent deployments, Qoder Pro bundles 2,000 Credits with concurrent agents.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Start Building for Free
&lt;/h3&gt;

&lt;p&gt;Claim your 70M+ free tokens and $200 cloud credits today.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lieke-ai.com/r?lk=coupon&amp;amp;userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Get $200 Free Credits →&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No credit card required · 180-day AI token validity&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lieke-ai.com/en/articles/qwen-code-free-2026" rel="noopener noreferrer"&gt;→ Full Qwen Code setup guide&lt;/a&gt; · &lt;a href="https://lieke-ai.com/en/articles/free-cloud-hosting-2026" rel="noopener noreferrer"&gt;Free cloud hosting comparison&lt;/a&gt;&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>RAG in Production: Build a Document Q&amp;A Bot with Qwen + Bailian</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Sat, 19 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/rag-in-production-build-a-document-qa-bot-with-qwen-bailian-3he3</link>
      <guid>https://dev.to/liekeai/rag-in-production-build-a-document-qa-bot-with-qwen-bailian-3he3</guid>
      <description>&lt;h1&gt;
  
  
  RAG in Production: Build a Document Q&amp;amp;A Bot with Qwen + Bailian
&lt;/h1&gt;

&lt;p&gt;Everyone talks about RAG (retrieval-augmented generation). Fewer people show you the whole pipeline: how documents become chunks, how chunks become vectors, and how those vectors answer your users' questions. This post builds a production document Q&amp;amp;A bot with Qwen on Alibaba Cloud Bailian — the managed path that avoids running your own vector database.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why managed RAG for most teams
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;No vector database to operate (no pgvector tuning, no index rebuilds)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Built-in document parsing: PDF, Word, Markdown handled out of the box&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Managed embeddings + retrieval + LLM in one API surface&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Scales without you thinking about it&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The pipeline at a glance
&lt;/h2&gt;

&lt;p&gt;Documents → Parse → Chunk → Embed → Vector Index&lt;br&gt;
                                       ↓&lt;br&gt;
User question → Embed → Retrieve top-k → LLM (Qwen) → Answer + citations&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Create the knowledge base
&lt;/h2&gt;

&lt;p&gt;In Bailian, create a knowledge base and upload your documents. The console handles parsing and chunking. Practical tips:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Chunk size around 500 tokens works well for most business documents&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Keep a small overlap between chunks so sentence context survives the split&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Name chunks with source metadata — you will need it for citations&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 2: Wire retrieval + generation
&lt;/h2&gt;

&lt;p&gt;With the knowledge base ID, the application code is short:&lt;/p&gt;

&lt;p&gt;from dashscope import Application&lt;/p&gt;

&lt;p&gt;app = Application(&lt;br&gt;
    app_id=your_application_id,&lt;br&gt;
    api_key=api_key&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;response = app.call(&lt;br&gt;
    prompt="Summarize our refund policy in two sentences",&lt;br&gt;
    rag_options={"pipeline_ids": [knowledge_base_id]}&lt;br&gt;
)&lt;br&gt;
print(response.output.text)&lt;/p&gt;

&lt;p&gt;The response can include retrieved sources, which lets you render citations instead of hallucinating confidently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Production concerns
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Permissions: gate the bot behind your own auth layer; the knowledge base may contain sensitive docs&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Observability: log every question and its retrieved chunks — debugging a RAG app without this is painful&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Evaluation: build a small golden set of question/answer pairs and run it after every knowledge base update&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Cost control: cache common questions, use a smaller model for easy intents, and monitor token spend per session&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 4: When to move off the managed path
&lt;/h2&gt;

&lt;p&gt;Managed RAG is the right default. You only need to roll your own when you have very custom chunking, multi-tenant isolation at scale, or regulatory requirements around data residency that the managed service cannot meet. For those cases, the architecture is the same — you just own more of the pieces.&lt;/p&gt;

&lt;p&gt;I keep independent notes on building AI applications with Qwen and current cloud pricing at &lt;a href="https://lieke-ai.com/en" rel="noopener noreferrer"&gt;lieke-ai.com&lt;/a&gt;. The official Bailian documentation and free-token campaign live here: &lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Alibaba Cloud coupons&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>Qwen3-Coder vs GitHub Copilot vs Cursor</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Fri, 18 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/qwen3-coder-vs-github-copilot-vs-cursor-5ah0</link>
      <guid>https://dev.to/liekeai/qwen3-coder-vs-github-copilot-vs-cursor-5ah0</guid>
      <description>&lt;p&gt;Qwen3-Coder vs GitHub Copilot vs Cursor: 2026 AI Coding Showdown | Lieke Tech&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;# Qwen3-Coder vs GitHub Copilot vs Cursor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔥 New: &lt;a href="https://lieke-ai.com/en/articles/qoder-review-2026" rel="noopener noreferrer"&gt;Qoder Review&lt;/a&gt; — Alibaba's 6M-user agentic workbench with Claude/GPT/Gemini/Qwen. Free + BYOK, paid plans from $20/mo.&lt;/p&gt;

&lt;p&gt;The complete 2026 AI coding assistant comparison. Pricing, features, context window, and real-world performance — updated August 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick Verdict
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Your Situation&lt;/th&gt;
&lt;th&gt;Best Choice&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Budget-conscious developer&lt;/td&gt;
&lt;td&gt;Qwen Code CLI FREE&lt;/td&gt;
&lt;td&gt;1,000 requests/day, zero cost, 60/min rate limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power user / API-based&lt;/td&gt;
&lt;td&gt;Qwen3-Coder BEST VALUE&lt;/td&gt;
&lt;td&gt;$0.20/$1.00 per M tokens, 1M context, 480B model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub-centric team&lt;/td&gt;
&lt;td&gt;GitHub Copilot Pro&lt;/td&gt;
&lt;td&gt;Deep GitHub integration, $10/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI-native IDE preference&lt;/td&gt;
&lt;td&gt;Cursor Pro&lt;/td&gt;
&lt;td&gt;Multi-file editing, $20/month&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Pricing Comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Free Tier&lt;/th&gt;
&lt;th&gt;Paid Plan&lt;/th&gt;
&lt;th&gt;API Price (per M tokens)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen Code CLI&lt;/td&gt;
&lt;td&gt;1,000 req/day&lt;/td&gt;
&lt;td&gt;$0.99/mo (Coding Plan Lite)&lt;/td&gt;
&lt;td&gt;N/A (CLI tool)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder API&lt;/td&gt;
&lt;td&gt;70M tokens free&lt;/td&gt;
&lt;td&gt;Pay-as-you-go (from $0.20/M)&lt;/td&gt;
&lt;td&gt;$0.20 / $1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder Flash&lt;/td&gt;
&lt;td&gt;70M tokens free&lt;/td&gt;
&lt;td&gt;Pay-as-you-go&lt;/td&gt;
&lt;td&gt;$0.20 / $0.97&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Copilot&lt;/td&gt;
&lt;td&gt;2,000 completions/mo&lt;/td&gt;
&lt;td&gt;$10-$100/mo&lt;/td&gt;
&lt;td&gt;Credits-based&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cursor&lt;/td&gt;
&lt;td&gt;Limited (2-week trial)&lt;/td&gt;
&lt;td&gt;$20/mo (Pro)&lt;/td&gt;
&lt;td&gt;$20/mo includes usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;Trial only&lt;/td&gt;
&lt;td&gt;$20/mo (Pro)&lt;/td&gt;
&lt;td&gt;$3/$15 (Sonnet)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Key insight: Qwen3-Coder API costs 25x less than running GPT-4o through Copilot for equivalent token usage. For a developer processing 10M tokens/month, Qwen3-Coder costs ~$12 vs. $125+ for GPT-4o.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lieke-ai.com/r?lk=coupon&amp;amp;userCode=tzlh4rrj" rel="noopener noreferrer"&gt;🎁 Get 70 Million Free AI Tokens + $200 Cloud Credit — No Credit Card Needed →&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Feature Deep Dive
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Qwen3-Coder: The Open-Source Powerhouse
&lt;/h3&gt;

&lt;p&gt;Qwen3-Coder is a dedicated 480-billion-parameter coding model from Alibaba's Qwen team, now available through Alibaba Cloud's international API. It supports a 1-million-token context window — enough to process entire codebases in a single request.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Details&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model Size&lt;/td&gt;
&lt;td&gt;480B parameters (MoE architecture)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Window&lt;/td&gt;
&lt;td&gt;262K standard, up to 1M via API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Languages&lt;/td&gt;
&lt;td&gt;90+ programming languages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Free Tier&lt;/td&gt;
&lt;td&gt;70M tokens (180 days) + Qwen Code CLI free forever&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API Endpoint&lt;/td&gt;
&lt;td&gt;&lt;a href="https://dashscope-intl.aliyuncs.com/compatible-mode/v1" rel="noopener noreferrer"&gt;https://dashscope-intl.aliyuncs.com/compatible-mode/v1&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  GitHub Copilot: The Established Standard
&lt;/h3&gt;

&lt;p&gt;Copilot remains the most widely adopted tool with deep IDE integration (VS Code, JetBrains, Neovim, Xcode). In 2026, it shifted to credit-based billing with AI credits for chat/agent features. Code completions remain unlimited on paid plans.&lt;br&gt;
✓ Best IDE integration &amp;nbsp; ✓ Mature and stable &amp;nbsp; ✗ Credit system can get expensive&lt;/p&gt;

&lt;h3&gt;
  
  
  Cursor: The AI-Native IDE
&lt;/h3&gt;

&lt;p&gt;Cursor is a full VS Code fork with AI built in. It excels at multi-file refactoring and understands entire project structures. The dual-pool pricing (Auto + API) can surprise heavy users.&lt;br&gt;
✓ Best multi-file editing &amp;nbsp; ✓ Strong agent mode &amp;nbsp; ✗ Requires IDE migration &amp;nbsp; ✗ Premium models cost extra&lt;/p&gt;

&lt;h2&gt;
  
  
  Head-to-Head: Capabilities
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;Qwen3-Coder&lt;/th&gt;
&lt;th&gt;Copilot&lt;/th&gt;
&lt;th&gt;Cursor&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Code Completion&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-File Editing&lt;/td&gt;
&lt;td&gt;Good (API)&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Window&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;Model-dependent&lt;/td&gt;
&lt;td&gt;Project-level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent/Autonomous Mode&lt;/td&gt;
&lt;td&gt;Via Qwen Code&lt;/td&gt;
&lt;td&gt;Yes (Pro+)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open Source&lt;/td&gt;
&lt;td&gt;Yes (weights available)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-Host Option&lt;/td&gt;
&lt;td&gt;Yes (Qwen3-Coder 30B)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Offline Capability&lt;/td&gt;
&lt;td&gt;Yes (via Ollama)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly Cost (avg user)&lt;/td&gt;
&lt;td&gt;$0-$20&lt;/td&gt;
&lt;td&gt;$10-$39&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Getting Started with Qwen Code (Free Forever)
&lt;/h2&gt;

&lt;p&gt;Qwen Code is Alibaba's free AI coding CLI. It works alongside VS Code, JetBrains, or any editor — no migration required.&lt;/p&gt;

&lt;h3&gt;
  
  
  Installation
&lt;/h3&gt;

&lt;p&gt;Requires &lt;a href="https://nodejs.org/" rel="noopener noreferrer"&gt;Node.js 20+&lt;/a&gt;:&lt;br&gt;
npx @qwen-code/qwen-code@latest&lt;/p&gt;

&lt;h3&gt;
  
  
  What You Get Free
&lt;/h3&gt;

&lt;p&gt;1,000 requests per day (international) / 2,000 (China)&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;No token consumption — request-based, not token-based&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;60 requests per minute rate limit&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Powered by Qwen3-Coder model&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Terminal-based agent that can read files, run commands, and edit code&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Upgrade to Coding Plan Lite
&lt;/h3&gt;

&lt;p&gt;For power users: &lt;strong&gt;$0.99 first month&lt;/strong&gt; for unlimited Qwen3-Coder access. After trial, still significantly cheaper than alternatives.&lt;/p&gt;

&lt;p&gt;🚀 Start Building with Qwen3 — 70M Free Tokens (180 Days) →&lt;/p&gt;

&lt;h2&gt;
  
  
  When Each Tool Wins
&lt;/h2&gt;

&lt;h3&gt;
  
  
  🏆 Choose Qwen3-Coder if...
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;You want the lowest cost without sacrificing capability&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;You need a massive context window for large codebases&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;You prefer open-source models or self-hosting&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;You build AI-powered dev tools and need API access&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;You want a free CLI coding assistant forever&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Choose GitHub Copilot if...
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Your team lives in GitHub (PRs, Issues, Actions)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;You want zero-setup IDE integration&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Enterprise compliance and admin controls matter&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;You need broad model choice (GPT-5, Claude, Gemini)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Choose Cursor if...
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Multi-file refactoring is your daily work&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;You're comfortable switching to an AI-first IDE&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Agentic coding (autonomous multi-step tasks) is priority&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Cost Calculator: 10M Tokens/Month
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool/Model&lt;/th&gt;
&lt;th&gt;Input Cost&lt;/th&gt;
&lt;th&gt;Output Cost&lt;/th&gt;
&lt;th&gt;Est. Monthly&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash&lt;/td&gt;
&lt;td&gt;$0.03/M × 8M&lt;/td&gt;
&lt;td&gt;$0.13/M × 2M&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder Flash&lt;/td&gt;
&lt;td&gt;$0.20/M × 8M&lt;/td&gt;
&lt;td&gt;$0.97/M × 2M&lt;/td&gt;
&lt;td&gt;$3.54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder 480B&lt;/td&gt;
&lt;td&gt;$0.30/M × 8M&lt;/td&gt;
&lt;td&gt;$1.00/M × 2M&lt;/td&gt;
&lt;td&gt;$4.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V3&lt;/td&gt;
&lt;td&gt;$0.14/M × 8M&lt;/td&gt;
&lt;td&gt;$0.56/M × 2M&lt;/td&gt;
&lt;td&gt;$2.24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o (via Copilot)&lt;/td&gt;
&lt;td&gt;$2.50/M × 8M&lt;/td&gt;
&lt;td&gt;$10/M × 2M&lt;/td&gt;
&lt;td&gt;$40.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet&lt;/td&gt;
&lt;td&gt;$3/M × 8M&lt;/td&gt;
&lt;td&gt;$15/M × 2M&lt;/td&gt;
&lt;td&gt;$54.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Assumes 80% input / 20% output token split, typical for coding tasks. Qwen3-Coder 480B is 9x cheaper than GPT-4o and 12x cheaper than Claude Sonnet.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;Is Qwen Code really free?Yes. Qwen Code CLI offers 1,000 free requests per day internationally (2,000 in China) with no token limits. The free tier doesn't expire and doesn't require a credit card.&lt;br&gt;
Can Qwen3-Coder replace Copilot?For most developers, yes. Qwen3-Coder handles code completion, debugging, refactoring, and code review at a fraction of the cost. The main trade-off is less polished IDE integration compared to Copilot's native plugins.&lt;br&gt;
What is the Qwen3-Coder context window?Qwen3-Coder supports 262K tokens standard and up to 1 million tokens through the Alibaba Cloud API. This means it can process entire large codebases in a single conversation.&lt;br&gt;
How does the 70M free token trial work?New Alibaba Cloud international users receive 70 million AI tokens valid for 180 days, plus $200 in cloud credits, 2,000 image generations, and 1,650 seconds of video generation. No credit card required.&lt;br&gt;
Can I self-host Qwen3-Coder?Yes. The Qwen3-Coder 30B model is open-source and can run locally via Ollama (ollama run qwen3:7b for the smaller variant) or deployed on your own GPU infrastructure.&lt;/p&gt;

&lt;p&gt;🎁 Claim Your Free 70M Tokens + $200 Credit →&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Related:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://lieke-ai.com/en/articles/qwen3-api-pricing-2026" rel="noopener noreferrer"&gt;Qwen3 API Pricing Guide&lt;/a&gt; ·&lt;br&gt;
&lt;a href="https://lieke-ai.com/en/articles/free-cloud-hosting-2026" rel="noopener noreferrer"&gt;Free Cloud Hosting 2026&lt;/a&gt; ·&lt;br&gt;
&lt;a href="https://lieke-ai.com/en/articles/qwen-code-free-2026" rel="noopener noreferrer"&gt;Qwen Code Free Tutorial&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;⏰ Alibaba Cloud September Deals · Exclusive Channel&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;New users: free trial credits + coupon bundle across ECS, databases and more&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Model Studio (Bailian) LLM platform: up to 1M free tokens per model for 90 days; up to 50% off for selected models during 22:00-08:00 (UTC+8) off-peak hours&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Overseas regions: Singapore / Hong Kong / US ECS with no ICP filing required&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Enterprise: annual ECS and GPU instances with up to 40% off, plus channel-only pricing on request&lt;/p&gt;

&lt;p&gt;🔥 Claim Free Credits →&lt;br&gt;
View ECS Plans →&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Offers subject to official campaign pages; new-user deals require identity verification.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>Alibaba Cloud Makes Qwen Free: 2000 Daily AI Coding Calls + 70M Tokens</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Thu, 17 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/alibaba-cloud-makes-qwen-free-2000-daily-ai-coding-calls-70m-tokens-1adb</link>
      <guid>https://dev.to/liekeai/alibaba-cloud-makes-qwen-free-2000-daily-ai-coding-calls-70m-tokens-1adb</guid>
      <description>&lt;p&gt;Breaking News&lt;/p&gt;

&lt;h1&gt;
  
  
  Alibaba Cloud Makes Qwen Free: 2000 Daily AI Coding Calls + 70M Tokens
&lt;/h1&gt;

&lt;p&gt;Updated August 22, 2026&lt;/p&gt;

&lt;p&gt;ALL QWEN MODELS NOW FREE&lt;br&gt;
From flagship Qwen3.8-Max to Qwen3-Flash, not a limited trial&lt;br&gt;
Plus: 70M tokens, 100 AI images, 50s video, 200 CNY voucher&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lieke-ai.com/r?lk=coupon&amp;amp;userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Start Free Now&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Just Happened?
&lt;/h2&gt;

&lt;p&gt;In August 2026, Alibaba Cloud announced that its entire Qwen model family is now available for free. This includes the flagship Qwen3.8-Max, code models, and multimodal models. Combined with the new-user package of &lt;strong&gt;70 million free AI tokens&lt;/strong&gt;, Alibaba Cloud Model Studio (Bailian) is now one of the most generous free AI platforms available.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qwen Code: 2000 Free Daily AI Coding Runs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Qwen Code&lt;/strong&gt; is a terminal-based AI coding agent built on Qwen3-Coder models, competing with Google Gemini CLI, Claude Code, and Codex CLI.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Qwen Code Free Tier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Daily runs (mainland China)&lt;/td&gt;
&lt;td&gt;2,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Daily runs (international)&lt;/td&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token limits&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rate limit&lt;/td&gt;
&lt;td&gt;60 calls/minute&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Installation&lt;/td&gt;
&lt;td&gt;Single npm command&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Install Qwen Code
&lt;/h3&gt;

&lt;p&gt;Requirements: Node.js 20+.&lt;/p&gt;

&lt;p&gt;npx @qwen-code/qwen-code@latest&lt;/p&gt;

&lt;p&gt;Authenticate with your Alibaba Cloud account. China users get free tier automatically; international users via OpenRouter.&lt;/p&gt;

&lt;h3&gt;
  
  
  Capabilities
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Code understanding across large codebases beyond context windows&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Automated git workflows (PRs, rebases)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Code refactoring and optimization&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Documentation generation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Bug detection and autonomous terminal command execution&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Free API Access via Model Studio
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benefit&lt;/th&gt;
&lt;th&gt;Amount&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free AI tokens&lt;/td&gt;
&lt;td&gt;70 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI image generations&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI video generation&lt;/td&gt;
&lt;td&gt;50 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No-threshold voucher&lt;/td&gt;
&lt;td&gt;200 CNY (~$28)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Validity&lt;/td&gt;
&lt;td&gt;180 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credit card required&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Available Free Models
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-Max&lt;/td&gt;
&lt;td&gt;Flagship reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Max&lt;/td&gt;
&lt;td&gt;Agent workflows, 1M context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Plus&lt;/td&gt;
&lt;td&gt;Balanced performance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Flash&lt;/td&gt;
&lt;td&gt;High-speed tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder-Plus&lt;/td&gt;
&lt;td&gt;Code generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V3/V4-Pro&lt;/td&gt;
&lt;td&gt;Also available on platform&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  First API Call (Python)
&lt;/h3&gt;

&lt;p&gt;from openai import OpenAI&lt;br&gt;
client = OpenAI(&lt;br&gt;
    api_key="YOUR_API_KEY",&lt;br&gt;
    base_url="&lt;a href="https://dashscope.aliyuncs.com/compatible-mode/v1" rel="noopener noreferrer"&gt;https://dashscope.aliyuncs.com/compatible-mode/v1&lt;/a&gt;"&lt;br&gt;
)&lt;br&gt;
response = client.chat.completions.create(&lt;br&gt;
    model="qwen3.7-max",&lt;br&gt;
    messages=[{"role": "user", "content": "Hello!"}]&lt;br&gt;
)&lt;br&gt;
print(response.choices[0].message.content)&lt;/p&gt;

&lt;p&gt;OpenAI-compatible API for easy migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qwen Code vs Competitors
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Free Daily Calls&lt;/th&gt;
&lt;th&gt;Token Limit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen Code&lt;/td&gt;
&lt;td&gt;2,000/1,000&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini CLI&lt;/td&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Copilot CLI&lt;/td&gt;
&lt;td&gt;Paid only&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;Paid only&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Coding Plan Lite: 7.9 CNY First Month
&lt;/h2&gt;

&lt;p&gt;For power users, &lt;strong&gt;Coding Plan Lite&lt;/strong&gt; offers unlimited Qwen3-Coder calls at just &lt;strong&gt;7.9 CNY (~$1.10) first month&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Get Started
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Sign up at Alibaba Cloud International&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Verify with phone number (no credit card needed)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Claim 70M tokens automatically&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Install Qwen Code: npx @qwen-code/qwen-code@latest&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Build with the OpenAI-compatible API&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Start Building with Free AI Today&lt;br&gt;
Claim Your Free Access&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Qwen Code really free with no token limits?
&lt;/h3&gt;

&lt;p&gt;Yes. 2,000 daily runs in China, 1,000 internationally, no token consumption, 60 calls/min cap.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I use the free API commercially?
&lt;/h3&gt;

&lt;p&gt;Yes, within Alibaba Cloud terms. 180-day free quota for new Model Studio users.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need a credit card?
&lt;/h3&gt;

&lt;p&gt;No. Phone verification only for free tier.&lt;/p&gt;

&lt;p&gt;Disclaimer: Maintained by Lieke Cloud, authorized Alibaba Cloud channel partner. Terms may change. Verified August 2026.&lt;/p&gt;

&lt;p&gt;Contact: &lt;a href="https://lieke-ai.com/mailto:lieke-promo@coze.email" rel="noopener noreferrer"&gt;lieke-promo@coze.email&lt;/a&gt;&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>Qwen API Pricing 2026: The Complete Alibaba Cloud Model Studio Cost Guide</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Wed, 16 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/qwen-api-pricing-2026-the-complete-alibaba-cloud-model-studio-cost-guide-4mab</link>
      <guid>https://dev.to/liekeai/qwen-api-pricing-2026-the-complete-alibaba-cloud-model-studio-cost-guide-4mab</guid>
      <description>&lt;p&gt;&lt;strong&gt;📅 Data note (Sep 1, 2026):&lt;/strong&gt; All rates below are synced from the official &lt;a href="https://www.alibabacloud.com/help/en/model-studio/model-pricing?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Model Studio model inference pricing page&lt;/a&gt; (last updated Aug 31, 2026) and the Chinese &lt;a href="https://help.aliyun.com/zh/model-studio/model-pricing" rel="noopener noreferrer"&gt;百炼模型价格页&lt;/a&gt;. Standard list prices are shown; limited-time promotions (such as the Qwen3.7-Max 50% discount) are flagged. Always confirm final numbers in the console before production budgeting.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. How Model Studio Billing Works
&lt;/h2&gt;

&lt;p&gt;Alibaba Cloud Model Studio (the international name for &lt;strong&gt;Bailian / 百炼&lt;/strong&gt;) is Alibaba’s one-stop LLM platform: Qwen text models, DeepSeek, Qwen-VL vision models, QwQ reasoning models, embedding, rerank, speech and image generation, all behind an OpenAI-compatible API. Billing is straightforward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Pay-as-you-go by tokens. Fee = input tokens × input price + output tokens × output price. Prices are quoted per 1 million tokens.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tiered pricing on some models. For models with context tiers, the price is set by the total input tokens of that single request — and all tokens in the request are billed at that tier’s rate. A 100K-input request on a two-tier model (0–32K, 32K–128K) is billed entirely at the 32K–128K rate.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Batch inference = 50% off. Models supporting batch calls charge half price for both input and output, with results returned asynchronously.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Context caching = up to ~90% off cached input. Explicit cache creation costs 125% of the input price, but cache-hit input tokens cost about 10%. Repeated system prompts or long reference documents become dramatically cheaper.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Off-peak / night discounts. Some models carry automatic night discounts (22:00–08:00 Beijing time, UTC+8) with no signup — currently flagged on Qwen3.7 series list prices.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Free quota. New accounts get 1 million tokens per model (input and output each), valid for 90 days from activation — across dozens of models this totals roughly 70 million free tokens. On the international (Singapore) deployment, the free quota applies to the models listed there; mainland deployment (Beijing) has its own free-quota list.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. Qwen API Price Table — International Deployment (USD, per 1M tokens)
&lt;/h2&gt;

&lt;p&gt;These are the &lt;strong&gt;International scope&lt;/strong&gt; rates (Singapore region) most overseas developers use. Standard real-time prices; batch halves them where supported.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input (USD/1M)&lt;/th&gt;
&lt;th&gt;Output (USD/1M)&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-Max (flagship, GA Aug 3, 2026)&lt;/td&gt;
&lt;td&gt;$2.00 flat&lt;/td&gt;
&lt;td&gt;$6.00 flat&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Hardest agentic/coding tasks; one flat rate at any length&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Max&lt;/td&gt;
&lt;td&gt;$2.50 list (50% off → $1.25)&lt;/td&gt;
&lt;td&gt;$7.50 list (50% off → $3.75)&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Previous flagship — cheapest high-end while promo lasts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Max&lt;/td&gt;
&lt;td&gt;$1.20 → $2.40 → $3.00&lt;/td&gt;
&lt;td&gt;$6 → $12 → $15&lt;/td&gt;
&lt;td&gt;262K&lt;/td&gt;
&lt;td&gt;Reasoning &amp;amp; coding (tiers at 32K / 128K input)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Plus&lt;/td&gt;
&lt;td&gt;$0.48 (≤256K: $1.44)&lt;/td&gt;
&lt;td&gt;$1.92 (≤256K: $5.76)&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Multimodal mid-tier, production workhorse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen-Plus (qwen-plus-2025-12-01)&lt;/td&gt;
&lt;td&gt;from ~$0.40&lt;/td&gt;
&lt;td&gt;from ~$1.20&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;The default “start here” model for most apps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen-Flash&lt;/td&gt;
&lt;td&gt;from $0.05&lt;/td&gt;
&lt;td&gt;from $0.40&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;High-volume classification, tagging, extraction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash&lt;/td&gt;
&lt;td&gt;~$0.03–$0.07&lt;/td&gt;
&lt;td&gt;~$0.13–$0.20&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Ultra-cheap batch processing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen-Turbo (legacy)&lt;/td&gt;
&lt;td&gt;~$0.05&lt;/td&gt;
&lt;td&gt;~$0.20&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;No longer updated — use Qwen-Flash for new projects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QwQ-Plus (reasoning)&lt;/td&gt;
&lt;td&gt;~$0.82 (¥5.871 intl)&lt;/td&gt;
&lt;td&gt;~$2.47 (¥17.614 intl)&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;Deep thinking mode; free 1M tokens each&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen-Long (long context)&lt;/td&gt;
&lt;td&gt;~$0.07 (¥0.5 mainland)&lt;/td&gt;
&lt;td&gt;~$0.28 (¥2 mainland)&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Whole-document analysis on a budget&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Source: Alibaba Cloud Model Studio — Model inference pricing (Aug 31, 2026). CNY-converted rows marked with ¥ use the mainland rate at ~7.2 CNY/USD as an approximation; the international rate is billed in USD.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Qwen API Price Table — Mainland China Deployment (CNY, per 1M tokens)
&lt;/h2&gt;

&lt;p&gt;For China-facing products served from Beijing (North China 2). These rates include the models most commonly used by Chinese teams:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;模型 Model&lt;/th&gt;
&lt;th&gt;输入 Input (¥/1M)&lt;/th&gt;
&lt;th&gt;输出 Output (¥/1M)&lt;/th&gt;
&lt;th&gt;免费额度 Free quota&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;qwen-turbo / qwen-turbo-latest&lt;/td&gt;
&lt;td&gt;¥0.367&lt;/td&gt;
&lt;td&gt;¥1.468&lt;/td&gt;
&lt;td&gt;Batch 半价&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen-plus-2025-04-28 及早期版&lt;/td&gt;
&lt;td&gt;¥0.8&lt;/td&gt;
&lt;td&gt;¥2&lt;/td&gt;
&lt;td&gt;各 100 万 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3.6-plus (≤32K 档)&lt;/td&gt;
&lt;td&gt;¥2&lt;/td&gt;
&lt;td&gt;¥12&lt;/td&gt;
&lt;td&gt;各 100 万 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwq-plus (思考模式)&lt;/td&gt;
&lt;td&gt;¥1.6&lt;/td&gt;
&lt;td&gt;¥4&lt;/td&gt;
&lt;td&gt;各 100 万 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen-long&lt;/td&gt;
&lt;td&gt;¥0.5&lt;/td&gt;
&lt;td&gt;¥2&lt;/td&gt;
&lt;td&gt;各 100 万 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qvq-plus (视觉推理)&lt;/td&gt;
&lt;td&gt;¥2&lt;/td&gt;
&lt;td&gt;¥5&lt;/td&gt;
&lt;td&gt;各 100 万 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qvq-max&lt;/td&gt;
&lt;td&gt;¥8&lt;/td&gt;
&lt;td&gt;¥32&lt;/td&gt;
&lt;td&gt;各 100 万 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3-vl-flash (≤32K)&lt;/td&gt;
&lt;td&gt;¥0.15&lt;/td&gt;
&lt;td&gt;¥1.5&lt;/td&gt;
&lt;td&gt;Batch 半价&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Source: &lt;a href="https://help.aliyun.com/zh/model-studio/model-pricing" rel="noopener noreferrer"&gt;阿里云帮助中心 — 百炼模型价格&lt;/a&gt;. Mainland free quota: 1 million tokens each for input/output, valid 90 days after Bailian activation.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. How to Pick a Model (and What It Costs You)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;High-volume simple work (classification, tagging, extraction, chat triage): Qwen-Flash. At $0.05/$0.40 per million, 10 million input + 2 million output tokens cost roughly $1.30/day — under $40/month for a busy bot.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Mainstream app workhorse (summarization, drafting, RAG answers, coding assist): Qwen-Plus / Qwen3.7-Plus. 10M in + 2M out on Qwen3.7-Plus works out to about $8.6/month at standard rates — this is why Plus is the default for production.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Agentic and hard reasoning (multi-step tools, codebase agents, math): Qwen3.8-Max. Flat $2/$6 across the full 1M context means no long-prompt cliff; the same 10M+2M workload is about $32/month.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Batch/night jobs: batch API halves everything; night discounts on Qwen3.7 can reach 80% off list. Offline pipelines should never run at peak real-time price.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;DeepSeek is available on the same platform — DeepSeek-V4-Flash and V4-Pro are callable through Model Studio too, so you can A/B vendors without changing infrastructure.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Worked Examples: Real Monthly Bills
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Volume (monthly)&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Est. bill&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Support chatbot, 50K conversations&lt;/td&gt;
&lt;td&gt;20M input / 4M output&lt;/td&gt;
&lt;td&gt;Qwen-Flash&lt;/td&gt;
&lt;td&gt;~$2.6/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RAG knowledge assistant, 100K queries&lt;/td&gt;
&lt;td&gt;50M input / 10M output&lt;/td&gt;
&lt;td&gt;Qwen3.7-Plus&lt;/td&gt;
&lt;td&gt;~$43/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding agent, 5K heavy tasks&lt;/td&gt;
&lt;td&gt;30M input / 10M output&lt;/td&gt;
&lt;td&gt;Qwen3.8-Max&lt;/td&gt;
&lt;td&gt;~$120/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same coding agent, batch mode&lt;/td&gt;
&lt;td&gt;30M input / 10M output&lt;/td&gt;
&lt;td&gt;Qwen3.8-Max batch&lt;/td&gt;
&lt;td&gt;~$60/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Document processing, long-context&lt;/td&gt;
&lt;td&gt;100M input / 5M output&lt;/td&gt;
&lt;td&gt;Qwen-Long (mainland ¥)&lt;/td&gt;
&lt;td&gt;~¥60 (~$8)/month&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The takeaway: outside of heavy agent workloads, most production apps run on &lt;strong&gt;tens of dollars a month&lt;/strong&gt; — and the free quota covers the first ~70 million tokens of experimentation entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Free Tokens: What New Accounts Actually Get
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;1 million tokens per model, both input and output, for each model in the free-quota list.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Valid 90 days from Model Studio activation (or model release / application approval, whichever is later).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;International deployment: free quota is granted in the Singapore region; other international regions don’t carry it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Mainland deployment: free quota in the Beijing (North China 2) region.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The old unlimited free developer tier ended April 15, 2026 — the current program is this per-model trial pack.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  7. Three Ways to Cut the Bill Further
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Prompt caching. If your system prompt + retrieved context repeats across users, explicit cache hits drop input cost to ~10%. RAG apps routinely cut 60–80% of input spend this way.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Batch calls for non-interactive work. Evals, embeddings-style sweeps, nightly summaries: 50% off with no quality difference.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Model routing. Use Flash for triage and simple intents, escalate only the hard 5–10% of requests to Plus or Max. Most “Max-only” apps waste 80%+ of their budget.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  8. Getting Started in 10 Minutes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Register an Alibaba Cloud international account and claim the $200 starter credit (new users; approval ~3 business days where required).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Open the Model Studio console, activate the service — free tokens are granted automatically.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Create an API key and call the OpenAI-compatible endpoint (&lt;a href="https://dashscope-intl.aliyuncs.com/compatible-mode/v1" rel="noopener noreferrer"&gt;https://dashscope-intl.aliyuncs.com/compatible-mode/v1&lt;/a&gt;) with your existing OpenAI SDK — just change base URL and key.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Start on Qwen-Flash or Qwen-Plus, switch to Qwen3.8-Max only where quality demands it, and turn on caching once prompts stabilize.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Qwen API free?
&lt;/h3&gt;

&lt;p&gt;The hosted API is pay-per-token, but new accounts get ~1 million free tokens per model (roughly 70 million total) for 90 days. Most Qwen models are also open-weight, so you can self-host on GPU servers at zero per-token cost for steady high volume.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much does Qwen3.8-Max cost?
&lt;/h3&gt;

&lt;p&gt;$2.00 per million input tokens and $6.00 per million output tokens on international deployment, flat across the entire 1M-token context window. Batch calls halve this to $1/$3-ish effective rates.&lt;/p&gt;

&lt;h3&gt;
  
  
  Qwen vs DeepSeek — which is cheaper?
&lt;/h3&gt;

&lt;p&gt;DeepSeek-V4-Flash undercuts Qwen-Flash on output price; Qwen-Plus is slightly cheaper on input. Both are callable from the same Model Studio account, so run your own A/B — price differences are smaller than quality-fit differences for most workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do prices change often?
&lt;/h3&gt;

&lt;p&gt;Alibaba cuts prices aggressively as new models ship — Qwen3.7-Max is currently 50% off list and newer Flash tiers keep dropping. Budget against list prices for long-term planning and treat promotions as upside.&lt;/p&gt;

&lt;h3&gt;
  
  
  Start Building with Qwen Today
&lt;/h3&gt;

&lt;p&gt;$200 international credit · ~70 million free Qwen tokens · OpenAI-compatible API · Singapore + Beijing regions&lt;/p&gt;

&lt;p&gt;Claim $200 Free Credit&lt;br&gt;
Open Model Studio&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>Build a Free-Tier Image Pipeline with Alibaba Cloud OSS + CDN</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Tue, 15 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/build-a-free-tier-image-pipeline-with-alibaba-cloud-oss-cdn-25d2</link>
      <guid>https://dev.to/liekeai/build-a-free-tier-image-pipeline-with-alibaba-cloud-oss-cdn-25d2</guid>
      <description>&lt;h1&gt;
  
  
  Build a Free-Tier Image Pipeline with Alibaba Cloud OSS + CDN
&lt;/h1&gt;

&lt;p&gt;Every side project eventually hits the same wall: where do uploaded images live? Putting them on your app server bloats the disk, burns bandwidth, and makes backups painful. Object storage plus a CDN in front is the standard answer. This tutorial walks through the whole pipeline with Alibaba Cloud OSS and CDN — from bucket creation to a working upload endpoint — with cost numbers you can reason about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why object storage instead of the server disk
&lt;/h2&gt;

&lt;p&gt;Object storage (OSS) is built for exactly this job:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Pay per gigabyte stored and per request — no idle capacity&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Scale from 10 images to 10 million without touching infrastructure&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Integrates natively with a CDN for global edge delivery&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Your app server stays small and stateless&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 1: Create a bucket
&lt;/h2&gt;

&lt;p&gt;In the OSS console, create a bucket with these settings for a public-read, private-write pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Region: pick the one closest to your users (or the same region as your ECS instance to avoid cross-region traffic fees)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Access control: private (we will generate signed URLs or use a CDN origin)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Versioning: enable if images are user-generated and you want rollback safety&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 2: Upload with the SDK
&lt;/h2&gt;

&lt;p&gt;The Python SDK makes uploads a few lines:&lt;/p&gt;

&lt;p&gt;import oss2&lt;/p&gt;

&lt;p&gt;auth = oss2.Auth(access_key_id, access_key_secret)&lt;br&gt;
bucket = oss2.Bucket(auth, endpoint, bucket_name)&lt;/p&gt;

&lt;p&gt;bucket.put_object_from_file(&lt;br&gt;
    f"uploads/{uuid4().hex}.jpg",&lt;br&gt;
    "/tmp/source.jpg",&lt;br&gt;
    headers={"Content-Type": "image/jpeg"}&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;Never expose the access key to the browser. Upload from your backend, or use STS temporary credentials for client-side uploads — OSS supports both cleanly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Put a CDN in front
&lt;/h2&gt;

&lt;p&gt;Direct OSS URLs work, but for repeated views a CDN saves money and improves latency:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Create a CDN domain pointing at the OSS bucket as origin&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Enable image processing (resize, crop, webp conversion) — OSS Image Processing can transform images on the fly, so you store one original and serve many derivatives&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Set a reasonable cache TTL (7 days is fine for immutable content)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 4: What it costs
&lt;/h2&gt;

&lt;p&gt;Rough order of magnitude for a hobby project:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Storage: fractions of a cent per GB-month in standard tier&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;CDN traffic: cheaper than server egress in most regions&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Requests: pennies per million in the standard tier&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Exact pricing moves with region and tier, so check the official calculator before launch. The key insight is architectural: your app server stays at a fixed size no matter how viral your images go.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production notes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Validate uploads server-side: size limits, MIME whitelist, hash deduplication&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Use a lifecycle rule to move old originals to cold storage&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Log access via OSS logging or your CDN log service for abuse detection&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are comparing storage options or need a reference architecture for your project, I keep an independent summary of current cloud deals and setup notes at &lt;a href="https://lieke-ai.com/en" rel="noopener noreferrer"&gt;lieke-ai.com&lt;/a&gt;. Alibaba Cloud's official campaign page with current coupons: &lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Alibaba Cloud coupons&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>Deploy a Production Flask App on Alibaba Cloud ECS in 30 Minutes</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Mon, 14 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/deploy-a-production-flask-app-on-alibaba-cloud-ecs-in-30-minutes-88n</link>
      <guid>https://dev.to/liekeai/deploy-a-production-flask-app-on-alibaba-cloud-ecs-in-30-minutes-88n</guid>
      <description>&lt;h1&gt;
  
  
  Deploy a Production Flask App on Alibaba Cloud ECS in 30 Minutes
&lt;/h1&gt;

&lt;p&gt;Local development is easy. Production deployment is where beginners get stuck. This guide takes a Flask app from python app.py to a HTTPS-served, auto-restarting production service on a single Alibaba Cloud ECS instance — the whole path, not just the happy parts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Pick the right server
&lt;/h2&gt;

&lt;p&gt;For a small-to-medium Flask app, do not over-provision:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;2 vCPU / 4 GB RAM is comfortable for most API workloads&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Start with a 40 GB SSD system disk; move media to object storage instead of growing the disk&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Choose the region closest to your users to minimize latency&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are just starting and want minimal cost, a lightweight instance works fine for low traffic — you can migrate to full ECS later without re-architecting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: System setup
&lt;/h2&gt;

&lt;p&gt;sudo apt update &amp;amp;&amp;amp; sudo apt upgrade -y&lt;br&gt;
sudo apt install -y python3-pip nginx&lt;br&gt;
python3 -m venv /opt/app/venv&lt;/p&gt;

&lt;p&gt;Create a dedicated user for the app instead of running everything as root:&lt;/p&gt;

&lt;p&gt;sudo useradd -m -s /bin/bash appuser&lt;br&gt;
sudo chown -R appuser:appuser /opt/app&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Gunicorn + systemd
&lt;/h2&gt;

&lt;p&gt;Run Flask behind gunicorn, and let systemd keep it alive:&lt;/p&gt;

&lt;h1&gt;
  
  
  /etc/systemd/system/app.service
&lt;/h1&gt;

&lt;p&gt;[Unit]&lt;br&gt;
Description=Flask app&lt;br&gt;
After=network.target&lt;/p&gt;

&lt;p&gt;[Service]&lt;br&gt;
User=appuser&lt;br&gt;
WorkingDirectory=/opt/app&lt;br&gt;
ExecStart=/opt/app/venv/bin/gunicorn -w 2 -b 127.0.0.1:8000 app:app&lt;br&gt;
Restart=always&lt;/p&gt;

&lt;p&gt;[Install]&lt;br&gt;
WantedBy=multi-user.target&lt;/p&gt;

&lt;p&gt;Restart=always is the line that saves you at 3 AM. Enable it with systemctl enable app.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: nginx as reverse proxy
&lt;/h2&gt;

&lt;p&gt;server {&lt;br&gt;
    listen 80;&lt;br&gt;
    server_name api.example.com;&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;location / {
    proxy_pass http://127.0.0.1:8000;
    proxy_set_header Host $host;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header X-Forwarded-Proto $scheme;
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: HTTPS with Let's Encrypt
&lt;/h2&gt;

&lt;p&gt;sudo apt install -y certbot python3-certbot-nginx&lt;br&gt;
sudo certbot --nginx -d api.example.com&lt;/p&gt;

&lt;p&gt;Certbot rewrites the nginx config and wires up auto-renewal. Test it with certbot renew --dry-run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: The production checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Environment variables for secrets — never hardcode keys&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Uvicorn/gunicorn worker count = 2 × CPU cores + 1 (roughly)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Add a health endpoint and monitor it externally&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Set up log rotation for nginx and app logs&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Enable the security group firewall rules: only 80/443 open&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is genuinely it — a service that survives reboots, crashes, and certificate expiries, on hardware you control. For reference architectures and current cloud pricing notes I keep an independent summary at &lt;a href="https://lieke-ai.com/en" rel="noopener noreferrer"&gt;lieke-ai.com&lt;/a&gt;. If you are evaluating ECS pricing, the official campaign page with current offers is here: &lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Alibaba Cloud coupons&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>OpenClaw 2.0 Self-Hosting Guide: Deploy the Biggest Open-Source AI Agent Release on Your Own Cloud Server</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Sun, 13 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/openclaw-20-self-hosting-guide-deploy-the-biggest-open-source-ai-agent-release-on-your-own-cloud-16am</link>
      <guid>https://dev.to/liekeai/openclaw-20-self-hosting-guide-deploy-the-biggest-open-source-ai-agent-release-on-your-own-cloud-16am</guid>
      <description>&lt;p&gt;&lt;strong&gt;📰 Update (Aug 31, 2026):&lt;/strong&gt; OpenClaw has officially shipped &lt;strong&gt;version 2026.8.1, branded OpenClaw 2.0&lt;/strong&gt; — the largest release in the project’s history. According to the &lt;a href="https://openclaw.ai/blog/openclaw-2-accidentally" rel="noopener noreferrer"&gt;official OpenClaw blog&lt;/a&gt; and coverage by &lt;a href="https://github.blog/open-source/maintainers/openclaw-went-viral-meet-the-maintainers-building-and-securing-it/" rel="noopener noreferrer"&gt;GitHub’s official blog&lt;/a&gt;, the release was built by &lt;strong&gt;933 contributors (569 of them first-timers)&lt;/strong&gt; and merged &lt;strong&gt;more than 16,000 pull requests&lt;/strong&gt; — roughly half of the project’s entire merge history. The team went nearly seven weeks without shipping after 106 releases in 230 days, using the pause to rebuild the foundation. This guide explains what changed and how to run it on your own server, the right way.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What OpenClaw Is — and Why Everyone Is Talking About It
&lt;/h2&gt;

&lt;p&gt;OpenClaw is an open-source &lt;strong&gt;personal AI assistant runtime&lt;/strong&gt; that runs on your own machines and connects the models you already use (Claude, GPT, Gemini, Qwen, or local models) to the messaging channels you already live in — Telegram, Discord, Slack, WhatsApp, iMessage, and with community plugins, WeChat, Feishu, DingTalk and QQ. Started by Peter Steinberger as a weekend project in November 2025, it became the &lt;strong&gt;fastest-growing project in GitHub history&lt;/strong&gt;: roughly 388,000 stars and 81,000 forks by late August 2026.&lt;/p&gt;

&lt;p&gt;The architecture is simple in concept: a &lt;strong&gt;Gateway&lt;/strong&gt; service (default port 18789) brokers channels, agents, tools and policies; a &lt;strong&gt;Control UI&lt;/strong&gt; gives you a browser cockpit; paired &lt;strong&gt;nodes&lt;/strong&gt; (Mac, phone) extend reach; and everything runs under your own credentials and your own data.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. What’s New in OpenClaw 2.0 (v2026.8.1)
&lt;/h2&gt;

&lt;p&gt;This is not a point release. Almost every core module changed — Memory, Skills, Automations, Browser, Native App, Plugin system, Security, Cloud Workers and the agent permission model. The highlights that matter for self-hosters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Guided setup from credentials you already have. The installer detects existing Codex, ChatGPT or Claude CLI sign-ins, accepts an API key or a provider sign-in, and finds local Ollama / LM Studio models. It probes the chosen model with a live request before saving it. Fresh OpenAI setups default to GPT-5.6; local inference moved to a managed llama-server, with Gemma 4 as the RAM-gated default and a 64K default context.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Browser app is now the primary surface. The rebuilt Control UI opens straight into conversation, with docked panels for a workspace file editor, a git-backed Changes panel (PR status and CI summaries), a browser panel with element inspection and screenshot annotation, and a full-screen web terminal. In the team’s simulated test (mocked gateway, 50&amp;nbsp;ms latency), startup dropped from ~1.6&amp;nbsp;s to 575&amp;nbsp;ms and JavaScript requests from 140 to 45.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Sessions and transcripts move to SQLite. Faster, more durable history — with one caveat: downgrading back to a file-based release requires restoring archived legacy artifacts first, and sessions created after migration won’t appear in older releases. Take a verified backup before upgrading.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Shared cloud sessions — real multiplayer. A second person can join live work or take it over with full context; owners set read / suggest / draft / participate permissions. Work can start on your laptop and continue on a cloud worker. The docs are explicit about the ceiling: this is collaboration, not tenant isolation — single-operator and trusted-team deployments only.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Memory and skills that self-improve. Active Memory lets agents use past conversations in private sessions; Background Memory Consolidation writes long-term memory with source attribution. Self-learning converts proven task patterns into reusable Skills, curated through a Skill Workshop.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Automation that keeps working. Automations bind to their conversation context, /loop runs agents on a schedule or their own cadence, Workboard chains tasks, and the experimental Swarm lets one agent fan out parallel subagents.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Agents that reach further. Browser Agent reads pages and network requests; Desktop Control drives paired machines; Android Accessibility Control lets agents operate phones; official Teams and Zoom plugins join meetings as a browser guest; IMAP email can trigger tasks directly.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Security in 2.0: The Hardening You Actually Get
&lt;/h2&gt;

&lt;p&gt;OpenClaw had a rough first six months on the security side — a one-click RCE in the Control UI (CVE-2026-25253, fixed January), the “Claw Chain” credential-theft set (fixed April), and the ClawJacked injection flaw (fixed February). Version 2.0 is the team’s systematic answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Private Credential Requests: agents ask for secrets through a masked prompt; the value never enters the chat transcript or model context.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Secret egress host binding: shared-store secrets are tied to exact approved HTTPS destinations across CLI, Gateway RPC and Control UI; unbound substitution fails closed.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;One trust boundary per gateway: installs that would expose an instance to the network without authentication are blocked before any change is applied. The Gateway binds to loopback by default; unknown DM senders get a pairing-code challenge.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;openclaw security audit checks inbound access, tool blast radius, network exposure, browser-control exposure and plugin allowlists in one pass.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Plugin installs show source, version, capabilities and artifacts; code-bearing third-party plugins require explicit confirmation.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The team also published unusually honest red-team data: across a 2026 crowdsourced arena of 272,000 attacks in 41 agent scenarios, “harmful-and-hidden” success rates were 0.5% against Claude Opus 4.5, 1.0% against Sonnet 4.5, 1.3% against Haiku 4.5 and 8.5% against Gemini 2.5 Pro — while adaptive human attackers still exceed 80% against state-of-the-art defenses. Translation: &lt;strong&gt;model choice is the first layer, but tool policy, execution approvals and sandboxing remain the hard enforcement layer&lt;/strong&gt; — which is exactly why deployment hygiene matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Why Self-Host on a Cloud Server Instead of Your Laptop?
&lt;/h2&gt;

&lt;p&gt;OpenClaw runs fine on a personal machine, but the moment you want it to &lt;em&gt;actually work for you&lt;/em&gt;, a always-on server wins:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;24/7 availability: cron-driven automations, IMAP triggers and channel bots don’t stop when you close your laptop.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Stable network identity: pairing codes, webhooks and channel callbacks need a reachable, always-up endpoint.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Cloud workers for Swarm: 2.0’s cloud sessions hand off long-running work to a machine that can stay awake and scale.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Data ownership stays yours: the gateway, SQLite history and credentials live on infrastructure you control.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Chinese channels without ICP pain: a Singapore-region server is reachable globally and requires no Chinese filing — ideal for WeChat/Feishu/DingTalk gateway setups serving overseas or personal use.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Deployment Plan: Alibaba Cloud Singapore ECS
&lt;/h2&gt;

&lt;p&gt;We recommend an &lt;strong&gt;Alibaba Cloud ap-southeast-1 (Singapore) ECS&lt;/strong&gt;: low latency to both China and Southeast Asia, no ICP requirement, and straightforward security-group control.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.1 Sizing
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Minimum (gateway + API models): 2 vCPU / 2 GB RAM. OpenClaw’s Node gateway is light; if you use Claude/GPT/Qwen APIs (e.g. via &lt;a href="https://lieke-ai.com/r?lk=bailian&amp;amp;userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Alibaba Cloud Model Studio / Bailian&lt;/a&gt; for Qwen models), 2&amp;nbsp;GB is enough for personal use.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Recommended (comfortable + light local models): 2 vCPU / 4 GB RAM — headroom for the browser Control UI, SQLite under load, and small local models via llama-server.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Power users (local Gemma 4 / multi-agent Swarm): 4 vCPU / 8 GB or more; local model inference is RAM-bound.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5.2 Install steps (Ubuntu/Windows Server both supported)
&lt;/h3&gt;

&lt;p&gt;Runtime requirement is &lt;strong&gt;Node.js 24 (or 22.16+)&lt;/strong&gt;. On a fresh Linux host:&lt;/p&gt;

&lt;h1&gt;
  
  
  1. Install Node 24 (via nvm or nodesource), then:
&lt;/h1&gt;

&lt;p&gt;npm install -g openclaw@latest&lt;/p&gt;

&lt;h1&gt;
  
  
  2. Guided onboarding + background daemon (auto-start on boot)
&lt;/h1&gt;

&lt;p&gt;openclaw onboard --install-daemon&lt;/p&gt;

&lt;h1&gt;
  
  
  3. Verify the gateway
&lt;/h1&gt;

&lt;p&gt;openclaw gateway status --json&lt;br&gt;
openclaw gateway health&lt;/p&gt;

&lt;h1&gt;
  
  
  4. Run the built-in security audit BEFORE opening any channel
&lt;/h1&gt;

&lt;p&gt;openclaw security audit&lt;/p&gt;

&lt;p&gt;The 2.0 guided setup will discover any API keys or CLI sign-ins on the box; for a clean server, paste an API key (or use Qwen via Model Studio — significantly cheaper for high-volume agent traffic) and let the probe verify it.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Hardening Checklist (Do These Before You Connect a Channel)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Security group: do not open port 18789 to 0.0.0.0. The gateway binds loopback by default — keep it that way; use Tailscale or an SSH tunnel for remote Control UI access.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Run openclaw security audit after every upgrade and every new plugin.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Pairing codes: leave the unknown-sender pairing challenge on for every chat channel.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Approvals: keep host-execution approvals enabled for anything beyond a single trusted operator.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Backups before 2.0 upgrade: snapshot the VM (or back up ~/.openclaw) — the SQLite migration is effectively one-way.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Plugin hygiene: install code-bearing plugins only after reviewing source; prefer ClawHub entries that show a security audit.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;OS level: keep the ECS patched, disable password SSH (key-only), and restrict outbound where your channels allow.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  7. What It Costs
&lt;/h2&gt;

&lt;p&gt;OpenClaw itself is free and open source. Your costs are the server and the model tokens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Server (Alibaba Cloud China accounts): new-user burstable instances from 38&amp;nbsp;CNY/year for 2C2G3M; returning-user deals around 99&amp;nbsp;CNY/year; business 2C4G5M plans around 199&amp;nbsp;CNY/year. See the current hot deals page for live pricing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Server (international accounts): Alibaba Cloud International runs SMB coupon campaigns; a $200 free credit is available via application (use-case form, roughly 3 business days’ review — it is not instant). Check the international offers page.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tokens: using Qwen models through Model Studio (Bailian) can cut monthly agent API bills by an order of magnitude versus frontier closed models for routine automation.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  8. FAQ
&lt;/h2&gt;

&lt;p&gt;Is OpenClaw 2.0 safe to expose to the internet?&lt;/p&gt;

&lt;p&gt;Not directly. 2.0 explicitly blocks unauthenticated network exposure, binds to loopback by default, and challenges unknown senders with pairing codes. For remote access use Tailscale, a VPN, or an SSH tunnel — never open the gateway port to the world.&lt;/p&gt;

&lt;p&gt;Can I downgrade from 2.0 after upgrading?&lt;/p&gt;

&lt;p&gt;Carefully. Sessions now live in SQLite; before rolling back you must restore archived legacy transcript artifacts, and post-migration sessions won’t show in older releases. Take a VM snapshot or back up ~/.openclaw before upgrading.&lt;/p&gt;

&lt;p&gt;What server spec do I really need?&lt;/p&gt;

&lt;p&gt;2C2G runs the gateway plus API-based models comfortably for one user. 2C4G is the sweet spot if you use the browser Control UI heavily or run small local models. Local Gemma-class models need 8&amp;nbsp;GB+ RAM.&lt;/p&gt;

&lt;p&gt;Does OpenClaw work with WeChat and Feishu?&lt;/p&gt;

&lt;p&gt;Not in the official core, but mature community plugins exist for WeChat, Feishu, DingTalk, QQ and WeCom (search ClawHub/GitHub for openclaw-china). Run them through the same audit and approval hygiene as any plugin.&lt;/p&gt;

&lt;p&gt;Is shared cloud session mode suitable for a multi-tenant SaaS?&lt;/p&gt;

&lt;p&gt;No — the documentation states the controls are not tenant isolation or a security boundary. 2.0 multiplayer is for a single operator or a team that trusts each other.&lt;/p&gt;

&lt;p&gt;Sources: OpenClaw official blog “OpenClaw 2.0, Accidentally” (Aug 30, 2026); GitHub Blog “OpenClaw went viral” (Aug 27, 2026); OpenClaw v2026.8.1 release notes and gateway security docs. Prices and promotions verified Aug 31, 2026 — always confirm live pricing on the official pages before purchase.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>LLM API Pricing Comparison 2026</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Sat, 12 Sep 2026 07:00:22 +0000</pubDate>
      <link>https://dev.to/liekeai/llm-api-pricing-comparison-2026-1nga</link>
      <guid>https://dev.to/liekeai/llm-api-pricing-comparison-2026-1nga</guid>
      <description>&lt;p&gt;LLM API Pricing Comparison 2026: After DeepSeek Price Hike, Qwen Is the Value King | Lieke Tech&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;# LLM API Pricing Comparison 2026&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔥 Latest Update (Aug 27): &lt;a href="https://lieke-ai.com/en/articles/qwen38-flash-review-2026" rel="noopener noreferrer"&gt;Qwen3.8-Flash Deep Review&lt;/a&gt; — 125B MoE model just dropped 20%, now $0.15/M input, $0.47/M output, 73% cheaper than DeepSeek V4 Flash. Full pricing table, competitor comparison, and architecture analysis.&lt;/p&gt;

&lt;p&gt;After DeepSeek's 1100% price hike and industry-wide increases up to 80%, who offers the best value?&lt;br&gt;
✅ Qwen3.5-Flash only $0.07/M input&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Get $200 Free Credit →&lt;/a&gt;&lt;br&gt;
ECS from $4.50/mo&lt;/p&gt;

&lt;h2&gt;
  
  
  📌 Key Takeaways
&lt;/h2&gt;

&lt;p&gt;August 2026: The LLM API market has structurally shifted:&lt;/p&gt;

&lt;p&gt;DeepSeek raised prices twice: V4 Pro output went from $0.87 to $1.98/M (off-peak) and $3.96/M (peak), up 350%; cache-hit input surged 1100%&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Industry-wide increases: Morgan Stanley reports average Chinese LLM API input prices rose 48% YoY, output prices 80%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Qwen3.5-Flash holds its ground: $0.07/M input, $0.26/M output — 98% cheaper than category average, with 1M context and multimodal support&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;New user bonus: Alibaba Cloud Model Studio offers 70M free tokens + 100 AI images + 50 seconds of video generation, valid 180 days&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  💰 Global LLM API Price Comparison
&lt;/h2&gt;

&lt;p&gt;Prices in USD per million tokens, as of August 26, 2026. DeepSeek now uses peak/off-peak pricing (peak: weekdays 9:00-12:00, 14:00-18:00 Beijing time).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Total (1M+1M)&lt;/th&gt;
&lt;th&gt;vs Qwen3.5-Flash&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-Flash 🏆&lt;/td&gt;
&lt;td&gt;$0.07&lt;/td&gt;
&lt;td&gt;$0.26&lt;/td&gt;
&lt;td&gt;$0.33&lt;/td&gt;
&lt;td&gt;1x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash (off-peak)&lt;/td&gt;
&lt;td&gt;$0.22&lt;/td&gt;
&lt;td&gt;$0.66&lt;/td&gt;
&lt;td&gt;$0.88&lt;/td&gt;
&lt;td&gt;2.7x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash (peak)&lt;/td&gt;
&lt;td&gt;$0.44&lt;/td&gt;
&lt;td&gt;$1.32&lt;/td&gt;
&lt;td&gt;$1.76&lt;/td&gt;
&lt;td&gt;5.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Pro (off-peak)&lt;/td&gt;
&lt;td&gt;$0.66&lt;/td&gt;
&lt;td&gt;$1.98&lt;/td&gt;
&lt;td&gt;$2.64&lt;/td&gt;
&lt;td&gt;8x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Pro (peak)&lt;/td&gt;
&lt;td&gt;$1.32&lt;/td&gt;
&lt;td&gt;$3.96&lt;/td&gt;
&lt;td&gt;$5.28&lt;/td&gt;
&lt;td&gt;16x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;$2.80&lt;/td&gt;
&lt;td&gt;8.5x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.5 Flash&lt;/td&gt;
&lt;td&gt;$1.50&lt;/td&gt;
&lt;td&gt;$9.00&lt;/td&gt;
&lt;td&gt;$10.50&lt;/td&gt;
&lt;td&gt;32x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5 (current)&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;td&gt;36x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5 (after Sept)&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;td&gt;$18.00&lt;/td&gt;
&lt;td&gt;55x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;td&gt;$35.00&lt;/td&gt;
&lt;td&gt;106x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5 Pro&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;td&gt;$180.00&lt;/td&gt;
&lt;td&gt;$210.00&lt;/td&gt;
&lt;td&gt;636x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: Official pricing pages, DeepSeek API docs, Morgan Stanley research, GitHub state-of-llm-apis, August 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  🇨🇳 Chinese LLM API Prices (CNY per million tokens)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-Flash&lt;/td&gt;
&lt;td&gt;¥0.2&lt;/td&gt;
&lt;td&gt;¥2&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Multimodal, cache hit ¥0.02&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen-turbo&lt;/td&gt;
&lt;td&gt;¥0.3&lt;/td&gt;
&lt;td&gt;¥0.6&lt;/td&gt;
&lt;td&gt;131K&lt;/td&gt;
&lt;td&gt;Thinking mode output ¥3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-Plus&lt;/td&gt;
&lt;td&gt;¥0.8&lt;/td&gt;
&lt;td&gt;¥4.8&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;Enhanced reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash (off-peak)&lt;/td&gt;
&lt;td&gt;¥1.5&lt;/td&gt;
&lt;td&gt;¥4.5&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Weekends all off-peak&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash (peak)&lt;/td&gt;
&lt;td&gt;¥3&lt;/td&gt;
&lt;td&gt;¥9&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Up 200-350%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Pro (off-peak)&lt;/td&gt;
&lt;td&gt;¥4.5&lt;/td&gt;
&lt;td&gt;¥13.5&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Cache hit ¥0.15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Pro (peak)&lt;/td&gt;
&lt;td&gt;¥9&lt;/td&gt;
&lt;td&gt;¥27&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Cache hit up 1100%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;💡 Do the math:&lt;/strong&gt; Processing 1M input + 1M output tokens costs &lt;strong&gt;¥2.2&lt;/strong&gt; with Qwen3.5-Flash vs &lt;strong&gt;¥36&lt;/strong&gt; with DeepSeek V4 Pro at peak — a &lt;strong&gt;16x difference&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  🎬 New: Wan3.0 Video Generation API
&lt;/h2&gt;

&lt;p&gt;Launched August 24, 2026, Alibaba Cloud's Wan3.0 generates up to 30-second videos natively and is the first to accept documents (doc/xls/ppt/pdf/md) as input.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Resolution&lt;/th&gt;
&lt;th&gt;Price/sec&lt;/th&gt;
&lt;th&gt;Promo (Aug 24-Sep 23)&lt;/th&gt;
&lt;th&gt;30-sec video&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;480p&lt;/td&gt;
&lt;td&gt;$0.05&lt;/td&gt;
&lt;td&gt;$0.035&lt;/td&gt;
&lt;td&gt;$1.05&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;720p&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$0.07&lt;/td&gt;
&lt;td&gt;$2.10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$0.14&lt;/td&gt;
&lt;td&gt;$4.20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Compared to Google Veo 3.1 at $0.40/sec, Wan3.0 1080p costs only &lt;strong&gt;50% as much&lt;/strong&gt;. New users get 50 seconds of free video generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  📈 Why Are LLM APIs Raising Prices in 2026?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Soaring Compute Costs
&lt;/h3&gt;

&lt;p&gt;Nvidia AI servers up 15%+, HBM supply shortages, DRAM up 90-95% in Q1. GPU rental rates: H100 from $1.70 to $2.35/GPU/hour.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Agentic Workflows Spike Compute Demand
&lt;/h3&gt;

&lt;p&gt;GitHub's official announcement: "Agentic workflows have dramatically increased compute demands, with some single requests exceeding the cost of an entire plan." GitHub Copilot introduced session limits and 7-day token caps in April 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Explosive Inference Volume
&lt;/h3&gt;

&lt;p&gt;OpenRouter data: DeepSeek-V4-Flash processed 11.31 trillion tokens in a single week (Aug 3-7). OpenCode platform processed 8 trillion tokens in a single day.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Shift from Price War to Value Pricing
&lt;/h3&gt;

&lt;p&gt;Morgan Stanley's report is titled "Farewell to Price Wars, Hello to Intelligence Wars." Tencent Cloud raised prices twice this year; Zhipu AI three times.&lt;/p&gt;

&lt;h2&gt;
  
  
  ✅ Developer Cost-Saving Guide
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Strategy 1: Choose the right model&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Use Qwen3.5-Flash ($0.07/$0.26) for daily chat, content generation, and simple coding. Only upgrade to Plus or Max for complex reasoning. 90% of tasks work fine on Flash.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strategy 2: Maximize free credits&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Alibaba Cloud Model Studio gives new users 70M tokens (180 days) — enough for months of personal development. International users get $200 in cloud credits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strategy 3: Use caching and batch&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Qwen3.5-Flash cache hits cost only $0.003/M (90% off standard input). Batch File API: input $0.014/M, output $0.14/M — another 50% off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strategy 4: Schedule off-peak&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If using DeepSeek, weekends are all off-peak and weekday nights are half price. Schedule non-real-time tasks accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  ❓ FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Qwen3.5-Flash good enough?
&lt;/h3&gt;

&lt;p&gt;Qwen3.5-Flash supports 1M token context, multimodal input (text/image/video), Function Calling, and structured output. It scores 100/100 for pricing value in the coding category (lmmarketcap) and is 98% cheaper than the category average. It handles the vast majority of use cases.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I get 70M free tokens?
&lt;/h3&gt;

&lt;p&gt;Sign up for Alibaba Cloud and enter the Model Studio console — tokens are automatically credited, valid for 180 days. No credit card required for China region; international version offers $200 credit with a 30-day refund guarantee.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is DeepSeek still worth it after the hike?
&lt;/h3&gt;

&lt;p&gt;DeepSeek V4 Pro remains competitive for complex reasoning, but at peak pricing it approaches mid-tier international model costs. Use Qwen3.5-Flash daily and switch to V4 Pro only when you need its unique capabilities — preferably off-peak or on weekends.&lt;/p&gt;

&lt;h3&gt;
  
  
  What about international users?
&lt;/h3&gt;

&lt;p&gt;Alibaba Cloud International offers $200 free credits (30-day refund guarantee), with Qwen3.5-Flash at $0.07/$0.26 per M tokens across Singapore, Germany, and US nodes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start Building for Free
&lt;/h2&gt;

&lt;p&gt;New users: $200 cloud credit + 70M AI tokens + 80+ products, 30-day refund guarantee&lt;/p&gt;

&lt;p&gt;Claim $200 Free Credit →&lt;br&gt;
ECS from $4.50/mo&lt;/p&gt;

&lt;p&gt;China users: &lt;a href="https://bailian.console.aliyun.com/?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Get 70M free tokens on Bailian →&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;⏰ Alibaba Cloud September Deals · Exclusive Channel&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;New users: free trial credits + coupon bundle across ECS, databases and more&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Model Studio (Bailian) LLM platform: up to 1M free tokens per model for 90 days; up to 50% off for selected models during 22:00-08:00 (UTC+8) off-peak hours&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Overseas regions: Singapore / Hong Kong / US ECS with no ICP filing required&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Enterprise: annual ECS and GPU instances with up to 40% off, plus channel-only pricing on request&lt;/p&gt;

&lt;p&gt;🔥 Claim Free Credits →&lt;br&gt;
View ECS Plans →&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Offers subject to official campaign pages; new-user deals require identity verification.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>GPU Cloud Pricing 2026</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Fri, 11 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/gpu-cloud-pricing-2026-205p</link>
      <guid>https://dev.to/liekeai/gpu-cloud-pricing-2026-205p</guid>
      <description>&lt;h2&gt;
  
  
  1. The Price Hike Is Real: What Happened?
&lt;/h2&gt;

&lt;p&gt;On August 22, 2026, Bloomberg reported that Nvidia has notified its largest customers — including Microsoft, Google, and Oracle — that servers powered by its next-generation AI chips will see price increases of &lt;strong&gt;over 15%&lt;/strong&gt;, with some configurations nearing 17%. The new prices take effect for early 2027 shipments of Vera Rubin and Grace Blackwell systems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key figures:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;H100 one-year rental climbed from $1.70/GPU/hour (Oct 2025) to $2.35 — a 38% increase&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;H200 rental now at $3.50/GPU/hour&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;8x H200 server monthly rental: ~$11,500 USD&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;RTX 4090 scalped to ~$7,000 on secondary markets&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;On-demand GPU capacity essentially sold out globally; high-end cards backordered into 2028&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The primary driver is memory. Per TrendForce, Q1 2026 DRAM contract prices surged 90-95% quarter-over-quarter, with another 58-63% increase expected in Q2. The big three memory makers are shifting capacity to HBM, creating a severe shortage of standard DRAM. Memory now accounts for 62% of the Vera Rubin superchip's bill of materials.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Three Paths for Enterprises
&lt;/h2&gt;

&lt;p&gt;Amid rising compute costs, organizations have three options: &lt;strong&gt;buy GPU servers outright&lt;/strong&gt;, &lt;strong&gt;rent GPU cloud instances on demand&lt;/strong&gt;, or &lt;strong&gt;call LLM APIs directly&lt;/strong&gt;. Let's break down the costs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 1: Buy Your Own GPU Servers
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Hardware Price (est.)&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RTX 4090 single-GPU workstation&lt;/td&gt;
&lt;td&gt;$7,000-$8,500&lt;/td&gt;
&lt;td&gt;Inference, small-scale fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8x H200 server&lt;/td&gt;
&lt;td&gt;$140,000-$170,000&lt;/td&gt;
&lt;td&gt;Large model training, high-throughput inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8x H100 cluster&lt;/td&gt;
&lt;td&gt;$280,000+&lt;/td&gt;
&lt;td&gt;Large-scale training clusters&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Hidden costs include: colocation ($700-$2,100/month), electricity (8-GPU server ~$700/month), operations staff, and hardware depreciation (3-year refresh cycle).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Break-even point:&lt;/strong&gt; Industry experience shows self-hosting only makes sense when utilization exceeds 60% over 18-24 months. For short-term projects or variable workloads, buying hardware almost always loses money.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 2: Rent GPU Cloud Instances
&lt;/h3&gt;

&lt;p&gt;This is the choice for most enterprises — pay as you go, release when idle, zero upfront cost. Here are Alibaba Cloud's latest GPU prices:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;GPU Model&lt;/th&gt;
&lt;th&gt;Instance Type&lt;/th&gt;
&lt;th&gt;vCPU/Memory&lt;/th&gt;
&lt;th&gt;Monthly (discounted)&lt;/th&gt;
&lt;th&gt;On-Demand&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;T4 (16GB)&lt;/td&gt;
&lt;td&gt;ecs.gn6i-c4g1.xlarge&lt;/td&gt;
&lt;td&gt;4 vCPU / 15 GB&lt;/td&gt;
&lt;td&gt;~$260/month&lt;/td&gt;
&lt;td&gt;$0.50/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;T4 (16GB)&lt;/td&gt;
&lt;td&gt;ecs.gn6i-c8g1.2xlarge&lt;/td&gt;
&lt;td&gt;8 vCPU / 31 GB&lt;/td&gt;
&lt;td&gt;~$315/month&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V100 (16GB)&lt;/td&gt;
&lt;td&gt;ecs.gn6v-c8g1.2xlarge&lt;/td&gt;
&lt;td&gt;8 vCPU / 32 GB&lt;/td&gt;
&lt;td&gt;~$655/month&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A10 (24GB)&lt;/td&gt;
&lt;td&gt;ecs.gn7i series&lt;/td&gt;
&lt;td&gt;Various&lt;/td&gt;
&lt;td&gt;~$450/month&lt;/td&gt;
&lt;td&gt;$1.43/hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L20 (48GB)&lt;/td&gt;
&lt;td&gt;ecs.gn8is series&lt;/td&gt;
&lt;td&gt;Various&lt;/td&gt;
&lt;td&gt;~$970/month&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Source: Alibaba Cloud official pricing and authorized reseller quotes, August 2026. Prices may vary by region and promotion. Check the official website for current rates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost Saver: Spot / Best-Effort Instances&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Alibaba Cloud Container Service offers best-effort QoS GPU instances at roughly 40% of on-demand pricing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;T4 spot: $0.20/hour (vs $0.50 on-demand)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A10 spot: $0.57/hour (vs $1.43 on-demand)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Perfect for stateless inference, batch processing, and CI/CD workloads that tolerate interruption — a 60% cost reduction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 3: Call LLM APIs Directly (Cheapest)
&lt;/h3&gt;

&lt;p&gt;If your needs are AI coding, content generation, customer service chatbots, or similar applications, &lt;strong&gt;you don't need GPU servers at all&lt;/strong&gt; — just call an API. In August 2026, Alibaba Cloud's Model Studio (Bailian) announced that all Qwen models are free to call, with extremely competitive paid tiers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input Price&lt;/th&gt;
&lt;th&gt;Output Price&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Free Tier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder Flash&lt;/td&gt;
&lt;td&gt;$0.14/M tokens&lt;/td&gt;
&lt;td&gt;$0.56/M tokens&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;70M tokens for new users&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder Plus&lt;/td&gt;
&lt;td&gt;$1.40/M tokens&lt;/td&gt;
&lt;td&gt;$5.60/M tokens&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;All models free to call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash&lt;/td&gt;
&lt;td&gt;$0.04/M tokens&lt;/td&gt;
&lt;td&gt;$0.13/M tokens&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Let's do the math:&lt;/strong&gt; Using Qwen3-Coder Flash as an AI coding assistant at 1,000 requests/day (avg 2,000 input + 500 output tokens each):&lt;/p&gt;

&lt;p&gt;Input: 60M × $0.14 = $8.40 | Output: 15M × $0.56 = $8.40 | &lt;strong&gt;Total ~$16.80/month&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Renting an A10 GPU to run an open-source model costs ~$450/month minimum — the API approach is &lt;strong&gt;26x cheaper&lt;/strong&gt;, with zero ops overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Cost Comparison at a Glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factor&lt;/th&gt;
&lt;th&gt;Self-Hosted GPU&lt;/th&gt;
&lt;th&gt;Cloud GPU Rental&lt;/th&gt;
&lt;th&gt;LLM API&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Upfront cost&lt;/td&gt;
&lt;td&gt;$7,000-$280,000&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly (light use)&lt;/td&gt;
&lt;td&gt;$1,400-$2,800 (colo+power)&lt;/td&gt;
&lt;td&gt;$260/mo (T4 reserved)&lt;/td&gt;
&lt;td&gt;$0-$17 (within free tier)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly (heavy use)&lt;/td&gt;
&lt;td&gt;$4,200-$7,000 (multi-node)&lt;/td&gt;
&lt;td&gt;$970/mo (L20)&lt;/td&gt;
&lt;td&gt;$70-$700&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Elasticity&lt;/td&gt;
&lt;td&gt;❌ Fixed&lt;/td&gt;
&lt;td&gt;✅ Minute-level&lt;/td&gt;
&lt;td&gt;✅ Second-level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ops burden&lt;/td&gt;
&lt;td&gt;🔴 High&lt;/td&gt;
&lt;td&gt;🟡 Medium&lt;/td&gt;
&lt;td&gt;🟢 Zero&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data privacy&lt;/td&gt;
&lt;td&gt;🟢 Maximum&lt;/td&gt;
&lt;td&gt;🟡 Dedicated option&lt;/td&gt;
&lt;td&gt;🟡 Platform guaranteed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Hyperscale&lt;/td&gt;
&lt;td&gt;Mid-to-large&lt;/td&gt;
&lt;td&gt;Startups to enterprise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  💡 Lieke Tech Recommendation
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. If you're building AI apps, chatbots, content tools, or coding assistants:&lt;/strong&gt; Use the Bailian API. Qwen3-Coder Flash at $0.14/M tokens, 70M free tokens for new users, zero infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. If you need to run open-source models, fine-tune, or have custom inference:&lt;/strong&gt; Rent Alibaba Cloud GPU instances. T4 from $260/month, spot instances from $0.20/hour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. If you have 18+ months of steady high utilization and strict data residency requirements:&lt;/strong&gt; Then consider buying. In this price hike cycle, the lock-in value of short-term purchases is offset by supply shortages and extended lead times.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Act Now: Lock In Your Compute Costs
&lt;/h2&gt;

&lt;p&gt;Nvidia's price hike takes effect in early 2027, meaning now through year-end is the &lt;strong&gt;last window&lt;/strong&gt; at current pricing. Alibaba Cloud still offers promotional rates, and new users get free trials.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon" rel="noopener noreferrer"&gt;🌍 Claim $200 Free Credits&lt;/a&gt;&lt;br&gt;
&lt;a href="https://bailian.console.aliyun.com/?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;🤖 Try Bailian — 70M Free Tokens&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  New User Benefits
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;ECS Cloud Server: 2 vCPU / 2 GB RAM from $5.50/month (annual plan, renewal at same price)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;International Free Trial: $200 cloud credits + 70M AI tokens + 80+ products, no credit card required&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Bailian Platform: 70M tokens + 100 AI image generations + 50 seconds of video generation, valid 180 days&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;30-Day Money-Back Guarantee on eligible products&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Q: How does Alibaba Cloud GPU pricing compare to AWS/Azure?
&lt;/h3&gt;

&lt;p&gt;Alibaba Cloud has a significant price advantage in the Asia-Pacific region. For a V100-equivalent, Alibaba Cloud charges ~$655/month vs AWS p3.2xlarge at ~$2,200/month on-demand — roughly 70% cheaper. T4 instances are competitively priced at $0.50/hour with deeper monthly discounts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Will spot instances be reclaimed unexpectedly?
&lt;/h3&gt;

&lt;p&gt;Spot (best-effort) instances may be reclaimed during capacity shortages with a few minutes' notice. They're ideal for stateless, checkpoint-resumable workloads. Use reserved/monthly instances for production-critical services.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Are the API prices competitive after the free tier?
&lt;/h3&gt;

&lt;p&gt;Qwen3-Coder Flash at $0.14/M input and $0.56/M output is among the cheapest code models globally. Compared to OpenAI GPT-4o at $2.50/M input, it's over 90% cheaper, while delivering competitive coding benchmarks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Can I use Alibaba Cloud outside China?
&lt;/h3&gt;

&lt;p&gt;Absolutely. Alibaba Cloud International operates 30+ regions including Singapore, US (Virginia), Germany (Frankfurt), and Japan (Tokyo). International sign-ups get $200 in free credits and support credit cards and PayPal.&lt;/p&gt;

&lt;p&gt;Pricing data collected August 25, 2026 from Alibaba Cloud official pricing, Bloomberg, TrendForce, SemiAnalysis, and other sources. Cloud prices may change with promotions. Always verify current rates on the official Alibaba Cloud website. This article contains affiliate links; purchases through these links may earn Lieke Tech a commission at no additional cost to you.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
    <item>
      <title>Best Free Cloud Hosting in 2026</title>
      <dc:creator>liekeai</dc:creator>
      <pubDate>Thu, 10 Sep 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/liekeai/best-free-cloud-hosting-in-2026-5d89</link>
      <guid>https://dev.to/liekeai/best-free-cloud-hosting-in-2026-5d89</guid>
      <description>&lt;h2&gt;
  
  
  Free Cloud Hosting at a Glance
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Best Offer 2026:&lt;/strong&gt; Alibaba Cloud gives new users &lt;strong&gt;$200 in cloud credits&lt;/strong&gt;, &lt;strong&gt;70M+ AI tokens&lt;/strong&gt;, &lt;strong&gt;2,000 image generations&lt;/strong&gt;, and &lt;strong&gt;1,650 seconds of video generation&lt;/strong&gt; — all with a 30-day money-back guarantee.&lt;/p&gt;

&lt;h2&gt;
  
  
  Top Free Cloud Offers Compared
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Free Credits&lt;/th&gt;
&lt;th&gt;Duration&lt;/th&gt;
&lt;th&gt;AI Tokens&lt;/th&gt;
&lt;th&gt;Credit Card&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Alibaba Cloud&lt;/td&gt;
&lt;td&gt;$200&lt;/td&gt;
&lt;td&gt;30 days&lt;/td&gt;
&lt;td&gt;70M+&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AWS&lt;/td&gt;
&lt;td&gt;Free tier&lt;/td&gt;
&lt;td&gt;12 months&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Cloud&lt;/td&gt;
&lt;td&gt;$300&lt;/td&gt;
&lt;td&gt;90 days&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure&lt;/td&gt;
&lt;td&gt;$200&lt;/td&gt;
&lt;td&gt;30 days&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Oracle Cloud&lt;/td&gt;
&lt;td&gt;Always free&lt;/td&gt;
&lt;td&gt;Unlimited&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Alibaba Cloud is the only major provider that does not require a credit card for the free trial and includes generous AI token credits.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Can Do with $200 Free Credits
&lt;/h2&gt;

&lt;h3&gt;
  
  
  🚀 Cloud Computing (ECS)
&lt;/h3&gt;

&lt;p&gt;Deploy virtual servers with up to &lt;strong&gt;$90 in ECS credits&lt;/strong&gt;. Choose from multiple specifications and run multiple instances simultaneously.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;2 vCPU + 2GB RAM instances&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Linux or Windows Server&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Global data centers (Singapore, US, Europe)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;SSD block storage included&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  🤖 AI &amp;amp; Machine Learning
&lt;/h3&gt;

&lt;p&gt;Access &lt;strong&gt;Qwen3.8-Max&lt;/strong&gt;, the latest flagship model, plus 80+ AI tools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;70M+ free API tokens for Qwen models&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;2,000 AI image generations (Wan model)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;1,650 seconds of AI video generation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Qwen Code CLI: 1,000 free calls/day internationally&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;OpenAI-compatible API endpoint&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  💾 Storage &amp;amp; Database
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Object Storage Service (OSS) for files and media&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;ApsaraDB for RDS (MySQL, PostgreSQL, Redis)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Content Delivery Network (CDN)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Domain registration services&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to Claim Your Free Credits
&lt;/h2&gt;

&lt;p&gt;1&lt;/p&gt;

&lt;h3&gt;
  
  
  Visit the Campaign Page
&lt;/h3&gt;

&lt;p&gt;Go to the &lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon" rel="noopener noreferrer"&gt;Alibaba Cloud free trial page&lt;/a&gt; and click "Start for Free".&lt;/p&gt;

&lt;p&gt;2&lt;/p&gt;

&lt;h3&gt;
  
  
  Create Your Account
&lt;/h3&gt;

&lt;p&gt;Sign up with your email. Verification takes less than 2 minutes. &lt;strong&gt;No credit card required.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;3&lt;/p&gt;

&lt;h3&gt;
  
  
  Claim Credits
&lt;/h3&gt;

&lt;p&gt;Submit a short use-case description (about 100 words) for review; approved accounts typically receive credits within 3 business days. Cloud credits are valid 30 days; AI tokens up to 180 days.&lt;/p&gt;

&lt;p&gt;4&lt;/p&gt;

&lt;h3&gt;
  
  
  Start Building
&lt;/h3&gt;

&lt;p&gt;Deploy servers, call AI APIs, store files — your credits apply automatically to eligible services.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bonus: Qwen Code — Free AI Coding Assistant
&lt;/h2&gt;

&lt;p&gt;Beyond the free trial, Qwen Code CLI provides &lt;strong&gt;1,000 free API calls per day&lt;/strong&gt; for international users (2,000 in China) with &lt;strong&gt;no token limits&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;npx @qwen-code/qwen-code@latest&lt;/p&gt;

&lt;p&gt;Powered by Qwen3-Coder 480B (a 480-billion parameter MoE model with 262K context window), this competes directly with GitHub Copilot at &lt;strong&gt;zero cost&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lieke-ai.com/en/articles/qwen-code-free-2026" rel="noopener noreferrer"&gt;→ Full Qwen Code setup guide&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  After the Free Trial: Pay-as-You-Go Pricing
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th&gt;Starter Price&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ECS Cloud Server&lt;/td&gt;
&lt;td&gt;From $4.50/month&lt;/td&gt;
&lt;td&gt;2 vCPU, 2GB RAM, 3M bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.7-Flash API&lt;/td&gt;
&lt;td&gt;$0.03/M input tokens&lt;/td&gt;
&lt;td&gt;One of the cheapest LLMs globally&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3 Coder 480B&lt;/td&gt;
&lt;td&gt;$0.22/M input tokens&lt;/td&gt;
&lt;td&gt;Via Google Vertex AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Object Storage&lt;/td&gt;
&lt;td&gt;$0.018/GB/month&lt;/td&gt;
&lt;td&gt;Standard tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qoder Pro&lt;/td&gt;
&lt;td&gt;From $20/month&lt;/td&gt;
&lt;td&gt;2,000 Credits, AI agent workspace&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Compare with detailed ECS pricing on our &lt;a href="https://lieke-ai.com/en/articles/aliyun-ecs-price-2026-august" rel="noopener noreferrer"&gt;ECS pricing guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;Do I need a credit card for the free trial?No. Alibaba Cloud's free trial does not require a credit card. You can sign up and start using services immediately after email verification.&lt;/p&gt;

&lt;p&gt;How long are the free credits valid?Cloud credits ($200) are valid for 30 days. AI tokens (70M+) are valid for up to 180 days from activation. The 30-day money-back guarantee applies to paid services.&lt;/p&gt;

&lt;p&gt;Can I use the free credits for production workloads?Yes. The credits apply to standard cloud services including ECS, OSS, RDS, and CDN. Many startups use the free trial to launch MVPs and production applications.&lt;/p&gt;

&lt;p&gt;What happens after credits expire?Services continue running on a pay-as-you-go basis. You can set budget alerts or use the "stop when quota exhausted" feature to avoid unexpected charges. There is no auto-renewal trap.&lt;/p&gt;

&lt;p&gt;Is Qwen Code really free forever?Qwen Code provides 1,000 free API calls per day for international users (2,000 in China) with no token limits. This is not a trial — it is an ongoing free tier. For heavy usage, Qoder's coding plans start at low promotional prices on the China site (check qoder.cn for current offers).&lt;/p&gt;

&lt;h3&gt;
  
  
  Ready to Start?
&lt;/h3&gt;

&lt;p&gt;Join thousands of developers building on Alibaba Cloud's global infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon" rel="noopener noreferrer"&gt;Claim Your $200 Free Credits →&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No credit card required · 30-day money-back guarantee · 80+ products&lt;/p&gt;




&lt;p&gt;&lt;em&gt;More cloud deals and independent dev-tool guides: &lt;a href="https://lieke-ai.com" rel="noopener noreferrer"&gt;lieke-ai.com&lt;/a&gt; — &lt;a href="https://www.alibabacloud.com/en/campaign/smb-coupon?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;Alibaba Cloud international coupons&lt;/a&gt; · &lt;a href="https://www.aliyun.com/daily-act/ecs/activity_selection?userCode=tzlh4rrj" rel="noopener noreferrer"&gt;China new-user deals&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>ai</category>
      <category>programming</category>
      <category>alibaba</category>
    </item>
  </channel>
</rss>
