Two sparse MoE models, same 1M-token window, same 128K output ceiling, same MIT license — and a ~3.1× price gap on ordinary tokens. The difference isn't "newer vs older." It's a genuine fork in design intent, and picking wrong costs you either money or completion rate.
Here's how I'd reason about it.
The fork: omni-modal vs agent-first
V2.5 takes text, images, video, and audio as native inputs. Xiaomi bolts a 729M-parameter vision encoder and a 261M-parameter audio encoder onto the language backbone, so visual and audio signal enter the same reasoning pass rather than being pre-summarized into text. That matters for:
- Screenshot-driven tool use (read the UI, then call the API)
- Video plus its audio track as one input
- Diagram and document-image extraction
- Visual-interface support agents
- Long video sequence reasoning
V2.5-Pro is specified as text-input only. Its published emphasis is deep reasoning, code development, and long-running tool orchestration.
Rule of thumb I use: if the workflow itself ingests pixels or waveforms, start with V2.5. Escalate to Pro only when the hard part is reasoning over text, source code, tools, or a long trajectory.
Architecture: what 310B/15B vs 1.02T/42B buys you
Both are sparse MoE with hybrid sliding-window + global attention and three Multi-Token Prediction layers. Pro is not a different species — it's the same recipe scaled up.
| Model card | V2.5 | Pro | Delta |
|---|---|---|---|
| Total params | 310B | 1.02T | ~3.3× |
| Active params | 15B | 42B | 2.8× |
| Hidden size | 4,096 | 6,144 | 1.5× |
| LLM layers | 48 | 70 | +22 |
| Attention heads | 64 | 128 | 2× |
| Routed experts | 256 | 384 | 1.5× |
| Experts per token | 8 | 8 | same |
| MTP layers | 3 | 3 | same |
Note the attention pattern shift: V2.5 runs a 5:1 local-to-global ratio, Pro runs 6:1. Xiaomi claims these cut KV-cache requirements by roughly 6× and 7× respectively versus full attention across every layer. If you're doing long-context inference, that's the number that determines whether you can stay on one node.
The 8-experts-per-token constant across both is worth internalizing. Pro's advantage is not "more experts fire per token" — it's a deeper, wider, more selective routing space.
Why per-step gains compound
A long agent run is a chain of dependent steps. One bad tool call early forces retries, poisons the context, or invalidates downstream work. So a small lift in per-step reliability can move end-to-end completion far more than the raw benchmark delta suggests. That's the entire case for Pro on agent workloads — and it's also why you should measure it on your trajectory length, not a leaderboard.
Benchmark deltas, read honestly
Xiaomi's base-model evaluation is the cleanest comparison available: same settings, same harness, no post-training confounds. Here's what the gap looks like by category.
General knowledge — modest.
- BBH +1.2
- MMLU +3.1
- MMLU-Pro +2.7
Math and science — where it opens up.
- GPQA-Diamond +8.6
- GSM8K +16.3
- MATH +18.5
Code — better, not uniformly better.
- HumanEval+ +4.3
- MBPP+ +3.2
- LiveCodeBench v6 +4.1
- SWE-Bench AgentLess +4.9
Pro is not 3× better. It costs ~3× more and gets progressively more valuable as reasoning difficulty and task duration climb. On BBH that's noise; on MATH it's a different tier.
Caveat worth stating: these are base-model rows. Production endpoints stack SFT, agentic RL, and Multi-Teacher On-Policy Distillation on top. Don't expect a hosted API to reproduce raw base-eval numbers, and definitely don't expect them to reproduce under your tool schemas and retry policy.
Long-horizon agents: the real separator
Post-training is where Pro earns its keep. Xiaomi's published V2.5 figures land at 56.1 SWE-bench Pro, 65.8 Terminal-Bench 2.0, and 62.1 Pass³ on the general portion of Claw-Eval. Pro reports 78.9% on SWE-bench Verified and is explicitly described as sustaining trajectories on the order of hundreds of tool calls.
There's a case study I keep coming back to: Pro writing a SysY compiler over 4.3 hours, 672 tool calls, all 233 tests passing. The takeaway isn't "use Pro for coding." It's that on a workload with hundreds of dependent actions, the marginal reliability difference is the difference between shipping and not.
Long context is a tie; long reasoning over context is not
Both expose 1M tokens of context and up to 128K output through Xiaomi's API, so raw capacity doesn't separate them.
- V2.5 wins when the long context is multimodal — video, doc images, visual agent traces.
- Pro is what you test when the context is the problem: a large repo, a long contract, a research corpus, a thousand-step trajectory.
A 1M window is a capacity claim, not a fidelity claim. Measure retrieval accuracy, instruction retention, and evidence grounding at the lengths you'll actually run at.
The economics
Overseas pay-as-you-go pricing, straight from the sheet:
| V2.5 | Pro | Ratio | |
|---|---|---|---|
| Uncached input / 1M | $0.14 | $0.435 | 3.11× |
| Cached input / 1M | $0.0028 | $0.0036 | 1.29× |
| Output / 1M | $0.28 | $0.87 | 3.11× |
| Context | 1M | 1M | — |
| Max output | 128K | 128K | — |
A single request with 1M uncached input tokens and 200K output tokens runs roughly $0.196 on V2.5 versus $0.609 on Pro — before cache hits, web-search charges, or provider billing quirks.
The interesting row is cache. For cache-hit input, Pro is ~29% more expensive, not 211%. If your agent has a stable system prompt, fixed tool definitions, or a repo prefix that stays resident, your effective price multiple is driven almost entirely by your cache-hit rate. Measure it before you conclude Pro is unaffordable.
The metric that settles arguments, though, is cost per successful task. A Pro run that completes a nasty workflow on the first attempt routinely beats three failed attempts on a cheaper model — you pay for the failures too.
Routing beats picking
Almost nobody needs one global answer. Route multimodal and routine traffic to V2.5; escalate only hard text reasoning, complex coding, and long-horizon execution.
| Dimension | Measure | Routing signal |
|---|---|---|
| Completion | successful tasks / attempts | escalate classes with low completion |
| Quality | human or rubric score | escalate high-stakes requests |
| Tool reliability | errors and retries | escalate long trajectories |
| Latency | time to accepted result (incl. retries) | keep interactive paths lightweight |
| Cost | spend per accepted task | cheapest model that still succeeds |
Workload-level guidance, condensed:
- Routine chat, high-volume generation, image/video/audio understanding, routine tool calling, cost-sensitive 1M context → V2.5
- Difficult math and science, repository-scale coding, hundreds of dependent tool calls → Pro
Testing both in one harness
Both endpoints are reachable through a single OpenAI-compatible layer — CometAPI exposes both MiMo variants behind one key, which is the practical way to A/B them on identical prompts, tools, and output limits:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["COMETAPI_KEY"],
base_url="https://api.cometapi.com/v1",
)
models = ["mimo-v2.5", "mimo-v2.5-pro"]
prompt = """
Review this implementation plan.
Identify hidden technical risks and propose the three highest-priority fixes.
"""
for model in models:
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_tokens=2000,
)
print(f"\n--- {model} ---")
print(response.choices[0].message.content)
Run a representative batch — trivial requests, hard reasoning, code edits, long-context pulls, agent tool loops — and log completion rate, token counts, latency, tool errors, retries, quality scores, and cost per success. That table, not a leaderboard, is your routing policy.
Where each one breaks
V2.5: the smaller backbone gives ground as reasoning difficulty rises. BBH shows almost nothing; GPQA-Diamond and MATH show a lot. And again — 1M context ≠ 1M-token reasoning fidelity.
Pro: ~3.1× standard input/output pricing, text-only input so it's not a drop-in for multimodal apps, and a 1.02T open-weight checkpoint that's a serious self-hosting project even with only 42B active per token.
Verdict
V2.5 is the default: multimodal, routine agents, cost-sensitive production. Pro is the escalation path: hard reasoning, complex software engineering, long autonomous runs. For mixed traffic, keep the bulk on V2.5 and pay for Pro only where difficulty or trajectory length justifies it.
FAQ
Is Pro strictly better than V2.5? No. It leads on hard text reasoning and coding, but it can't take image, video, or audio input and it's substantially more expensive.
Do both have a 1M context window? Yes — 1M tokens in, up to 128K out. Validate retrieval and reasoning at your operating lengths regardless.
What should a multimodal agent start on? V2.5, since it natively accepts visual and audio input. Pro currently targets text-input workflows.
When is Pro's price justified? When better per-step reliability changes whether a difficult math problem, a repo-scale coding task, or a many-step tool trajectory completes at all.
Should production standardize on one? Not necessarily — a routing layer keeps routine and multimodal traffic cheap while escalating the hard tail.
Originally published at cometapi.com
Top comments (0)