DEV Community

Cover image for Xiaomi AI training, live: what the MiMo 2.6 RL dashboard shows

Xiaomi AI training, live: what the MiMo 2.6 RL dashboard shows

Xiaomi's AI lab is training its next model in public. The page at mimo.xiaomi.com/rl shows two reinforcement-learning runs for MiMo 2.6 "read directly from the trainer's logs": pass rates, data mix, restarts with the operators' reasons, and a cost counter that passed one million dollars on Thursday. If you have only ever seen a model through its launch blog, this is the part the launch blog leaves out.

TL;DR

  • Xiaomi is streaming two RL post-training runs: MiMo 2.6 Pro (about 1 trillion parameters, 42 billion active) and 2.6 Flash (309 billion, 15 billion active). Neither model is released.
  • The Pro run's counter read $1,048,236 on Thursday afternoon UTC, ticking at $5.71 a second. That is about $20,500 an hour, or roughly half a million dollars a day.
  • The Pro run restarted seven times in 51 hours, and the operators posted why, including "we removed the cyber dataset … since we observed some bad patterns in the rollout logs."
  • Pass rate went from 56.5 % to 61.5 % over 14 steps. Xiaomi's own DeepSWE score went from 58.4 to 65.8, which Hacker News immediately questioned as contamination.
  • Also in the episode: Neovim's forgotten Bitcoin, a Flock camera taken apart, and what one year of Servo donations paid for.

All dashboard numbers were read from its JSON endpoints between 13:20 and 13:45 UTC on September 17. They move.

What is the Xiaomi MiMo 2.6 live training dashboard?

The About text is short: "We are streaming our RL big runs. The mimo-v2.6 series is coming soon." The footer says "Open is what we value."

The Pro run started on September 15 at 10:32 UTC; the stream went public the next day at 04:00. By Wednesday night the Hacker News post had 480 points, more than most models get at release. It was found by the community, not announced: at 13:40 UTC on Thursday, @XiaomiMiMo's latest post on X was still a September 8 invite-only beta for a desktop app. Elie Bakouch of Hugging Face was one of the first to post it:

Elie Bakouch on X: Xiaomi is livestreaming the RL run of MiMo V2.6 Pro (1T, 42B active) and Flash (309B, 15B active)

The sizes follow the mixture-of-experts pattern: a trillion parameters stored, 42 billion used per token. The previous generation is on Hugging Face under an MIT license (MiMo-V2.5-Pro, 1.02 T parameters), and in May Xiaomi cut the V2.5 API price by "up to 99 %". So 2.6 will probably be cheap and open. The dashboard is a trailer for it.

The overview page of mimo.xiaomi.com/rl: live metrics for the Pro and Flash runs

How reinforcement learning post-training works, as the dashboard shows it

Pre-training teaches a model to predict text. Post-training with reinforcement learning teaches it to finish tasks: give it a problem, let it try, check the answer, and push the weights toward the attempts that worked. The dashboard exposes each part of that loop.

Per step, the Pro run samples 1,568 prompts and generates 16 attempts ("rollouts") for each: 25,088 sequences. The attempts run in sandboxes and get graded. By Thursday it had used almost two million sandboxes and trained on 30.2 billion tokens. The step-15 sampler drew from 24 data sources in five categories (code, general, cyber, visual, chat), and 1,061 of the 1,568 prompts were code: about 68 % of every batch. This is a coding-agent model being made.

A simplified sketch of one step, not Xiaomi's code:

# one RL step (illustrative sketch)
prompts = sample(datasets, n=1568)            # ~68 % code at step 15
rollouts = [model.generate(p) for p in prompts for _ in range(16)]
rewards = [sandbox_grade(r) for r in rollouts]  # pass / fail per attempt
pass_rate = mean(rewards)                     # the dashboard's headline line
model.update(rollouts, rewards)               # favour attempts that passed
Enter fullscreen mode Exit fullscreen mode

The headline metric, dynsam/avg@n, is that mean pass rate: the share of attempts that succeed. For Pro it went from 0.565 at step 1 to 0.615 at step 14. Five points in 14 steps is the sound of a model learning, or of a very expensive dice game. You cannot tell which from the line alone, which is why the other panels matter.

Why the MiMo 2.6 run keeps restarting

The Pro run restarted seven times between September 15 and 17. The operators post notices, and they read like an internal incident channel:

  • Sep 16, 21:14 UTC: "the mimo-v2.6-pro run is restarting due to a vram issue on one node."
  • Sep 17, 03:34 UTC: "we restarted the flash run from step 15. reason: a type of infra error on one of datasets was not correctly detected over the past ~3 hours."
  • Sep 17, 12:20 UTC: "there was a network connectivity issue between the pro training cluster and the grader deployment. we have restarted the run. we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs."

These are the normal failures of a big run. One bad GPU node can stall the whole synchronous job. The grader is a separate service, and if the network between trainer and grader breaks, rewards stop arriving. A data error that goes unnoticed for three hours means three hours of gradients built on wrong rewards, so you roll back to the last good checkpoint, as they did with Flash.

The cyber line is the interesting one. They did not say what the bad patterns were, so I won't guess. The general point stands: in RL a model learns whatever the reward pays for, and the rollout logs are where you catch it learning the wrong thing. Most labs would fix that quietly. Xiaomi posted it.

At $5.71 a second, each restart has a visible price. The Flash run, cheaper at $2.85 a second, stood at $475,440 with two restarts.

Is it benchmark contamination? The DeepSWE line

Xiaomi hand-fills a benchmark panel as checkpoints get evaluated ("a point appears when its step is on air"): DeepSWE v1.1 with mini-swe-agent, averaged over three runs. Pro went from 58.41 at step 1 to 65.78 at step 11. Flash went from 48.67 to 60.77.

DeepSWE v1.1 for mimo-v2.6-pro by training step, Xiaomi's own offline evaluation

On HN, liuliu asked the obvious question: "When you run benchmarks while training, isn't that the definition of contamination?" jampekka's version: the benchmark becomes the validation set, and you end up "benchmaxxing". lucrbvi replied that checkpoint evaluation is standard practice in big RL runs and is not training data.

Both are right, and the difference matters:

Contamination Checkpoint selection
Benchmark tasks in the training data Yes No
Benchmark score changes the weights directly Yes No
Benchmark score decides which checkpoint ships Maybe Yes
Result Score is meaningless Score is optimistic

Evaluating a checkpoint does not leak tasks into the weights. But if you pick the release checkpoint by the highest DeepSWE score, the published number is the best of many draws, and it will be higher than what you see on fresh tasks. Every lab does this. Xiaomi just does it where you can watch. Until a third party runs the released model, 65.78 is a number from the people training it. HN user esafak pointed to Artificial Analysis' DeepSWE chart as the frame to check it against later.

And the model is not a product yet. Users of today's V2.5 on HN were split: walrus01 called it "fast but makes basic mistakes". The stream is a promise.

Why is Xiaomi training its AI in public?

The thread had theories. rozab: "Why are they doing this? To try head off accusations about distillation?" thehamkercat: "chinese companies are more open than US or even EU companies". Nobody from Xiaomi said. My read: whatever the motive, the page is more useful to an engineer than any launch post, because it shows the failures. If you run training jobs, compare their restart notices with your own on-call log. The problems are the same at a trillion parameters.

Also today: Neovim, Flock, Servo

Neovim's untouched Bitcoin. Jake Manger checked the Bitcoin address in the neovim.io footer and posted it (HN). Per mempool.space, 10 BTC arrived on March 14, 2023, worth about $247,000 then and about $832,000 at Thursday's price. The last outgoing payment was in October 2019. Donations moved to Open Collective; the footer did not. The top worry in the thread was seymon's: "I hope they still have the private key." No maintainer had replied when we recorded.

A Flock camera, taken apart. A collective called stegan0gram removed a Flock license-plate reader from a pole in Wauwatosa, outside Milwaukee, and DDoSecrets published its partition images (Wired, HN). Micah Lee's analysis: Android 8.1 from 2017, a security patch level of June 2018, a Linux 3.18 kernel, on a build compiled in June 2025. Nineteen of the camera's 20 Flock apps share a library with a hard-coded API key, which the camera trades, with its MAC address, for credentials it then stores in plain text on an unencrypted partition. The media partition is encrypted, with the key stored next to it. Flock said: "We received no report through that [VDP] process."

John Scott-Railton on X: 1.6 million images logged in 21 days, and the key to decrypt files stored on the device itself

What one year of Servo donations bought. Servo, the Rust browser engine now at Linux Foundation Europe, used its monthly Open Collective and GitHub Sponsors money to pay one long-time maintainer, Josh Bowman-Matthews, part-time. The report (HN): 1,150 pull requests reviewed, 8 new maintainers, 114 newcomer issues, 92 % of them fixed. That is what donations do when someone spends them. After Neovim, a pointed remark.

Verdict: SHIP IT

I stamped it SHIP IT. Seven public restarts and a note that says "we observed some bad patterns" and pulled a dataset are worth more than one polished launch post. The caveat is the benchmark line: it is still graded by the people training the model, and a checkpoint picked by its score will look better than it is. Judge MiMo 2.6 when someone else runs it.

FAQ

What is Xiaomi MiMo?
Xiaomi's family of large language models. MiMo-V2.5-Pro is on Hugging Face under an MIT license; MiMo 2.6 Pro and Flash are in training and not yet released.

How much does it cost to train an AI model like MiMo 2.6 Pro?
Xiaomi's own counter showed about $1.05 million for the Pro RL run after 51 hours, at $5.71 per second. That counts post-training only; pre-training is outside it.

Is running a benchmark during training contamination?
Not by itself: evaluating checkpoints does not put benchmark tasks into the weights. It does make the published score optimistic if the release checkpoint is chosen by that score.

Where can I watch the MiMo training run?
At mimo.xiaomi.com/rl, while the runs are live.

Sources


This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Top comments (1)

Collapse
 
aifrontierpost profile image
AI Frontier Post •

Pulling the cyber dataset over "some bad patterns" was the most consequential event in the whole 51-hour log, and it is the one detail the operators declined to explain. A transparency dashboard that redacts exactly what went wrong in training is doing marketing, not transparency.