Xiaomi Just Open-Sourced an Embodied-AI Foundation Model — and 29 Seconds of Video Is Now Enough to Teach a Robot a New Skill
Subtitle: While everyone argues about parameter counts, the embodied-AI field just quietly changed its most important metric: data efficiency. Xiaomi open-sourced Xiaomi-Robotics-1 (100K hours pretraining, full pipeline), CyberOne started "internship" at its EV factory, and Tsinghua + BIT released HOST, which teaches a robot a new skill from a 29-second video.
Here's the thing nobody tells you about embodied AI: the bottleneck was never the model. It was the data. Collecting robot teleoperation data costs millions of dollars and thousands of human-hours per skill. So this week, two stories landed that reframe the entire field — and developers should care more about them than about any single benchmark number.
What actually happened
Xiaomi open-sourced Xiaomi-Robotics-1 on Aug 5. Not a teaser, not a blog post — the full pipeline. Pretrained on 100,000+ hours of UMI (Universal Manipulation Interface) data, post-trained on 10,000+ hours of cross-embodiment data, with everything from real-robot post-training to deployment plus benchmark evaluation code released on GitHub (XiaomiRobotics/Xiaomi-Robotics-1) and Hugging Face.
The architecture is a two-stage "pretrain + post-train" paradigm: pretraining learns general action generation — predicting action sequences that move a scene from current state to target state given visual observations and language descriptions. That's the "foundation model" part: it doesn't know one robot, it knows how to move scenes toward goals.
Almost simultaneously, CNBC reported Xiaomi's CyberOne humanoid is working at its EV factory as an "intern." Open source for the foundation, factories for the scenarios. Goldman Sachs is bullish, maintaining Buy with a HK$40 target price.
The 29-second video
Then the story that made me re-read the headline twice: Tsinghua and Beijing Institute of Technology co-open-sourced HOST, which teaches a robot a new skill from a single 29-second video.
Think about what that means for a moment. The traditional embodied-AI data pipeline looks like: hours of teleoperation → label → train → deploy. If HOST-class approaches scale, the cost of teaching a robot a new behavior collapses from "weeks of data collection" to "film it once, show it, done." That's not an incremental improvement — that's a change in who can compete in this field at all. When data-collection costs compress, small teams can play in foundation-model territory.
Why the capital market is telling the same story
It's easy to dismiss this as academic noise. But look at where the money went in July: 39 funding rounds, ¥11B+ in the humanoid sector, and the heaviest concentration wasn't in complete-robot startups — it was in core components: reducers, dexterous hands, tactile sensors. A reducer maker supplying Foxconn just closed a nine-figure round. A 22-year-old HIT undergraduate's team, building bipedal robots "from a dorm room," raised ¥500M.
When capital spreads from finished robots into bottleneck components, an industry is moving from concept validation to supply-chain positioning — the classic signature of a sector on the eve of mass production.
And the IPO window is opening: Unitree's STAR Market bookbuilding implies a valuation of up to ¥55 billion (~US$7.4B), clearing regulatory review in a record 73 days, with public subscription on Aug 10. The "first humanoid-robot stock" premium is setting the benchmark for every embodied-intelligence listing that follows.
Meanwhile, the policy side: the FCC's new rule swept in robot vacuums, and Shark, iRobot, and Eufy issued collective responses. Existing models are unaffected; new foreign-made models face restrictions. For Chinese makers like Roborock, Dreame, and Ecovacs, "localized production" just shifted from a cost option to a compliance necessity.
The throughline
Open source and mass production are moving toward each other. Xiaomi open-sources the foundation model to court developers while putting CyberOne in its own factory to close the real-scenario data loop. The training paradigm for embodied intelligence is being rewritten — and the metric that matters is no longer parameter scale, it's learning speed.
As a developer, this is the part I find genuinely interesting: the field just gave away its foundation layer. If you've been waiting to build on top of embodied AI without owning a robotics lab, this week was the moment the barrier started falling. The open-source camp (Xiaomi, HOST) is now offering a route contrast with the closed-source in-house camp (Unitree, Zhiyuan) — and the next 12 months will show which one compounds faster.
What's your take — does open-sourcing the foundation model actually accelerate embodied AI, or does it just commoditize the wrong layer? And would you bet on data-efficiency research (29-second video learning) over scale for the next big jump?
Based on SinoBot Daily Pulse #57 — Aug 6, 2026. Sources: ITHome, TechNode, CNBC, Caixin Global, 36Kr, Mashable, Vacuum Wars, EET-China, Zhidx, Sina Finance, TMTPost.
Top comments (0)