I've been running Xiaomi's MiMo-V2.5 through coding and document-analysis workloads for a while now, and the interesting thing isn't the benchmark sheet — it's that the standard model and the Pro model are optimized for genuinely different jobs. The standard variant is multimodal-first and cheap per token; Pro is text-only, a trillion parameters, and roughly 3.1× the price at Xiaomi's published rates. Picking between them is a routing decision, not a quality decision.
Here's what actually matters if you're about to wire one of these into a production path.
Architecture: what 310B/15B gets you
MiMo-V2.5 is a sparse MoE with 310B total parameters and ~15B active per token. The router picks 8 of 256 experts per token, so you get access to a much larger parameter pool than a plain 15B dense model without paying the full compute bill.
That does not mean self-hosting is cheap. You still need to store and serve all 310B weights, and expert routing plus interconnect bandwidth become the real constraints. Active params are a compute figure, not a memory figure.
The backbone is 48 layers: 1 dense plus 47 MoE. Attention interleaves sliding-window and global layers at a 5:1 ratio — 39 SWA layers with a 128-token window, 9 global layers. Xiaomi reports nearly a 6× KV-cache reduction versus an all-global design. That's the mechanism that makes the 1M context affordable, not a bonus.
On the perception side there's a 729M ViT (28 layers) for image/video and a 261M transformer (24 layers) for audio, both projected into the language backbone. Three Multi-Token Prediction modules (~329M params combined) support speculative decoding and RL efficiency. Training ran on roughly 48T tokens with context progressively extended to 1M in post-training.
Weights are MIT-licensed, public beta dated April 23, 2026.
Spec sheet
| Item | Value |
|---|---|
| Total / active params | 310B / 15B |
| Layers / routed experts | 48 (1 dense + 47 MoE) / 256, 8 active |
| Attention | 39 SWA (128-token window) + 9 global |
| Context / max output | 1M / 128K tokens |
| I/O | text, image, video, audio → text |
| Encoders | 729M ViT; 261M audio transformer |
| MTP | 3 modules, ~329M params |
| Training | ~48T tokens |
| API surface | tools, web search, streaming, structured output, context caching |
| Published rate limits | 100 RPM / 10M TPM |
| License | MIT |
Rate limits are provider-level and will differ by route and account tier.
Benchmarks: tool use is the strong signal
These are publisher-reported, so treat them as directional. Nothing here guarantees behavior on your repo or your media formats.
Coding and agents
| Benchmark | Score |
|---|---|
| MiMo Coding Bench | 71.8 |
| Claw-Eval Text | 62.3 |
| Terminal-Bench 2.0 | 65.8 |
| SWE-Bench Pro | 56.1 |
| Claw-Eval Multi-Turn | 63.2 |
| ResearchClawBench | 16.91 |
The 65.8 on Terminal-Bench 2.0 is the number I'd weight most for terminal-driven agents, and 56.1 on SWE-Bench Pro suggests real repository-level ability rather than autocomplete. ResearchClawBench at 16.91 is the outlier — if your workload is evidence gathering and citation, test that path separately, because this model is not obviously good at it.
Multimodal
| Benchmark | Score |
|---|---|
| CharXiv RQ | 81.0 |
| MMMU-Pro | 77.9 |
| HR-Bench 4K | 88.5 |
| OmniDocBench | 87.2 |
| Claw-Eval Multimodal | 23.8 |
| Video-MME | 87.7 |
| DailyOmni | 83.5 |
| VideoHolmes | 64.0 |
Document and high-res image understanding are clearly the strength. Claw-Eval Multimodal at 23.8 versus 87.2 on OmniDocBench is the gap that matters: perceiving media and acting on it through tools are different capabilities, and the second one needs end-to-end testing on your actual stack.
Standard vs Pro
| MiMo-V2.5 | MiMo-V2.5-Pro | |
|---|---|---|
| Total params | 310B | 1.02T |
| Active params | 15B | 42B |
| Layers / routed experts | 48 / 256 | 70 / 384 |
| Context ceiling | 1M | 1M |
| API max output | 128K | 128K |
| Native input | text, image, video, audio | text |
| SWE-Bench Pro | 56.1 | 57.2 |
| Terminal-Bench 2.0 | 65.8 | 68.4 |
Pro leads by 1.1 points on SWE-Bench Pro and 2.6 on Terminal-Bench 2.0 — a modest margin on the only two tests they share. Pro spends ~2.8× the active parameters to get there and drops multimodal input entirely. If media is a first-class input for you, Pro isn't a candidate at all.
Both model cards list 1M context and Xiaomi's API pages list 128K max output, but usable limits vary by endpoint, account, and request format.
Pricing
| Per MTok | Standard (official) | Pro (official) |
|---|---|---|
| Input, cache miss | $0.14 | $0.435 |
| Output | $0.28 | $0.87 |
| Input, cache hit | $0.0028 | $0.0036 |
Pro is ~3.1× standard on both uncached input and output. If you're routing across multiple vendors and want one key and one base URL, CometAPI lists both at 20% below the official uncached pairs ($0.112/$0.224 standard, $0.348/$0.696 Pro) — the relative gap between variants stays the same, so it doesn't change the routing math, only the absolute bill.
Price per token isn't cost per completed task. Tool retries, context bloat, output length, and failed-run recovery dominate real spend. Sample your own workload before you pick a default.
Calling it
Standard OpenAI-compatible client, nothing exotic:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["COMETAPI_KEY"],
base_url="https://api.cometapi.com/v1",
max_retries=0,
)
response = client.chat.completions.create(
model="mimo-v2.5",
max_tokens=256,
messages=[
{
"role": "user",
"content": "Summarize the main findings in this technical report.",
}
],
)
print(response.choices[0].message.content)
Verify the live endpoint and request schema before deploying — model IDs on third-party gateways move independently of the upstream platform.
Where I'd put it
Good fits: multimodal agents that mix screenshots, documents, video, audio, and tool calls; long-document analysis over repos, contracts, logs, and support history; media summarization and chart interpretation; coding and terminal agents; high-volume automation where sparse activation and low token price both matter.
Bad fits: anything requiring image, audio, or video generation — output is text only. And a 1M context ceiling is capacity, not uniform retrieval accuracy across the window.
Routing table
| Signal | Start with | Switch when |
|---|---|---|
| Image/video/audio is a first-class input | Standard | Only if multimodal preprocessing is acceptable |
| Heavy document or mixed-media volume | Standard | Pro only if reasoning failures dominate cost |
| Difficult repository engineering | Test both | Pro if completion gain beats 3.1× unit cost |
| Long autonomous terminal runs | Pro | Back to standard if the gain isn't measurable |
| High-volume routine automation | Standard | Escalate failed/high-value tasks only |
Default to standard, route only hard text-first coding and long-horizon agent work to Pro.
The deprecation clock
Xiaomi has announced that the official mimo-v2.5 and mimo-v2.5-pro API model names will be deprecated at 10:00 Beijing time on October 21, 2026, with no automatic replacement (deprecation notice). That applies to Xiaomi's own platform — check your gateway separately. Plan the migration and confirm the live route before you ship anything load-bearing against these IDs.
Quick answers
Open source? Yes, MIT — commercial use, modification, fine-tuning, redistribution.
Params? 310B total, 15B active per token.
Multimodal input? Standard: text, image, video, audio → text. Pro's official spec lists text input.
Context? Up to 1M on both model cards; Xiaomi's current API pages list 128K max output for both. Verify the deployed route.
Coding agents? Pro reports higher shared benchmark scores, but the margin is small — test on your repo and compare task completion, latency, and total cost.
Multimodal agents? Standard, without question — Pro is text-only.
Originally published at cometapi.com
Top comments (0)