Verdict: Qwen 3.8 Max wins on price, Claude Opus 5.5 wins on agentic coding, and Kimi K3 wins on open weights. Qwen 3.8 Max lists at $2 per million input tokens and $6 per million output on Alibaba's Model Studio (published rate), against $3/$15 for Kimi K3 (API price card) and $4/$20 for Claude Opus 5.5 (model table). If you pay per token, Qwen is the pick. If you run long coding-agent sessions and cost is second, Claude Opus 5.5 posts the strongest agentic-coding score of the three (Anthropic's benchmark table). If you need frontier weights you can download and deploy, Kimi K3 is the only fully open 2.8T-class model here (weights on Hugging Face).
TL;DR
- The Issue: three frontier families now chase the same work - long-context reasoning, coding, and multimodal input - at list prices up to 3.3x apart on output tokens.
- The Insight: none of the three wins outright. Qwen wins the invoice, Claude wins the coding-agent run, Kimi wins the self-hosting case.
- The Solution: pick by your binding constraint. Price first: Qwen. Coding quality first: Claude. Control of the weights first: Kimi.
The three-way table
| Row | Qwen 3.8 Max | Claude Opus 5.5 | Kimi K3 |
|---|---|---|---|
| List price (input / output per 1M tokens) | $2 / $6 | $4 / $20 | $3 / $15 |
| Cached input per 1M | $0.25 | $0.20 cache read | writes $3 (5-min) / $6 (1-hr) |
| Context window | 1M hosted | 1M | 1,048,576 tokens |
| Max output | 131,072 tokens | 128K | not published in the model card |
| Total parameters | 2.4T, 95B active | not disclosed by Anthropic | 2.8T, 104B active |
| Open weights | yes, custom license, Aug 13 2026 | no | yes, Kimi K3 License, Jul 27 2026 |
| Native vision | hosted API: yes; open checkpoint: text-only | yes | yes, MoonViT-V2 401M, in the open weights |
| Reasoning control | low / medium / xhigh, default xhigh | adaptive thinking, effort controls | always on, effort adjustable |
| Independent index (Artificial Analysis) | 58 | not ranked in that comparison | 60 |
| Best for | high-volume API work on a budget | long agentic-coding sessions | open-weights self-hosting at frontier scale |
Where Qwen 3.8 Max loses
The open-weight release is the asterisk on Qwen's story. The published checkpoint, Qwen3.8-2.4T-A95B, ships under a custom license rather than Apache 2.0, and the open checkpoint is text-only with a native 256K context - the vision input and the default 1M context live in the hosted API, not in the download. If your plan is "get the weights, serve it yourself with everything the API has," that plan does not survive contact with the model card. And at launch, Alibaba's performance claims were largely internal rather than third-party verified.
Where Claude Opus 5.5 loses
Claude is the most expensive of the three on output tokens: $20 per million against Qwen's $6, a 3.3x gap at list price. It is also the only one of the three with no published weights, so there is no self-hosting route at any price (Anthropic's release). For anything where you need to own the stack - air-gapped deployments, fine-tuning, cost control at very high volume - Claude is structurally out.
Where Kimi K3 loses
Kimi K3's output tokens cost 2.5x Qwen's ($15 against $6 per million), so the open model is not the cheap model. And "open" here means open at datacenter scale: a 2.8T mixture-of-experts is a datacenter-scale deployment, not something you run on a workstation. Its headline coding scores are also partner-reported, using Moonshot's own harness, so read them as vendor benchmarks until independent tables catch up.
What our own readers open
One first-party datapoint, for what it is worth: across the articles on this site, pieces tagged kimi-k3 earn 2.84x the median article's views at the 30-day horizon (n=12 judged articles, measured 2026-10-07 from our own analytics store; method: per-article view counts against the site median). Kimi K3 is the subject our existing audience already reaches for, which is why this comparison exists as a page at all.
Honest takeaway
All three are good, and the differences are about constraints, not quality. The 3x price spread between Qwen and Claude on output tokens is real money on any serious workload - but so is the gap between a 66.4% agentic-coding score and everything below it (Anthropic's table). Our call: default to Qwen 3.8 Max for price-sensitive API work, reach for Claude Opus 5.5 when the deliverable is a long coding-agent run, and pick Kimi K3 only when you need the weights themselves.
FAQ
Q: Is Qwen cheaper than Claude and Kimi?
A: Yes, at list price. Qwen 3.8 Max is $2/$6 per million input/output tokens, against Kimi K3 at $3/$15 and Claude Opus 5.5 at $4/$20 (price sources: Qwen, Kimi, Claude).
Q: Can you self-host all three?
A: No. Kimi K3 publishes its full 2.8T weights (Hugging Face) and Qwen publishes the 2.4T-A95B checkpoint (Hugging Face) plus the Apache 2.0 27B sibling. Claude has no published weights (Anthropic).
Q: Which one is best for coding agents?
A: Claude Opus 5.5 posts the strongest published agentic-coding score of the three, 66.4% on Terminal-Bench 4.0. Kimi K3's 88.3% is on Terminal-Bench 2.1 with Moonshot's own harness - a different benchmark version, so the two numbers are not directly comparable.
Q: Which has the biggest context window?
A: A tie at about one million tokens: Qwen hosted, Claude, and Kimi all list 1M. The open Qwen checkpoint is the exception at 256K native.
Q: Are Qwen and Kimi actually open source?
A: Partly. Kimi K3 ships under the Kimi K3 License, and Qwen's 2.4T checkpoint under a custom license, while the smaller Qwen3.8-27B is Apache 2.0. None of the frontier-scale checkpoints here is plain Apache.
Q: What would a heavy month cost on each?
A: At the list prices cited above, 50M input plus 10M output tokens bills at roughly $160 on Qwen 3.8 Max, $300 on Kimi K3, and $400 on Claude Opus 5.5 (computed from $2/$6, $3/$15, $4/$20).
Related reading
- For running Qwen locally on a single GPU, see our Qwen 3.8 27B local setup guide.
- Our earlier Kimi K3 vs Claude Fable 5 coding comparison covers the coding case in depth.
- The wider field, including GPT-5.6 and Fable 5, is in our four-way frontier comparison.
- If you are weighing local options, the best local coding LLM guide adds GLM and Gemma to the mix.
Top comments (0)