Today's Highlights
· NVIDIA's DreamDojo trains a general-purpose robot world model on 44,000 hours of first-person human video
· Europe's first humanoid unicorn is born: Humanoid raises $152 million, with Bosch as contract manufacturer and Schaeffler placing orders
· China tightens rules on autonomous-driving "blue indicator lights": newly filed vehicle models are banned from factory installation, and national standard GB 4785 revision launched in parallel
· Disney's accelerator backs two physical AI foundation-model companies simultaneously for the first time: Physical Intelligence and FieldAI
· Seoul Gangnam's robotaxi service tops 1,000 rides over two-plus months since switching to paid fares, with only two vehicles in operation
I. Research Progress
DreamDojo: feeding 44,000 hours of human video into a world model, then distilling it to real-time on a single GPU · world-model
A longstanding problem with world models is that they can "watch but not be commanded" — human video without action labels can teach physical common sense but not action control. NVIDIA's GEAR team's DreamDojo sidesteps this with continuous latent actions: a 700M-parameter spatiotemporal Transformer VAE ingests two adjacent frames and compresses the action that occurred between them into a 32-dimensional vector as a unified proxy label, bridging human video and robot data. The training set totals 44,711 hours of first-person video, covering 9,869 scene categories, 6,015 tasks, and 43,237 objects — by the team's account, its variety of skills is 96 times larger than the previous largest comparable dataset. After a three-stage pipeline of human-video pretraining, post-training on target embodiments (Fourier GR-1, Unitree G1), and four-step denoising distillation, the model runs at 10.81 frames per second on a single H100 and can be teleoperated in real time with a VR controller; ablations show latent-action pretraining nearly matches ground-truth action pretraining that requires extra hardware (in-lab PSNR 20.83 vs. 21.02).
NVIDIA GEAR · arXiv 2602.06949 source · Analysis: AI随想录 source (WeChat, CN)
Xiaomi-Robotics-1: a mobile-manipulation foundation model built on 100,000 hours of UMI trajectories · vla
Xiaomi Robotics swapped out the pretraining data source for its VLA — instead of real-robot teleoperation, it used a UMI handheld gripper plus a first-person camera to collect over 100,000 hours of real manipulation trajectories across homes, supermarkets, industrial sites, offices, and outdoor settings. The key to scaling this up was annotation: the team cut trajectories into equal-length segments, then used Qwen3.5-27B to automatically generate descriptions of "how the scene state changed," labeling the entire dataset in about two weeks via a producer-consumer pipeline. The model uses a MoT architecture combining Qwen3-VL with a diffusion Transformer; the paper reports performance continuing to improve as data volume and parameter count increase, and stronger pretraining correlating with better real-robot performance in unseen environments during post-training.
Xiaomi Robotics · Analysis: 大语言模型和具身智体及自动驾驶 source (WeChat, CN)
Qwen-AgentWorld: bringing world models from embodied AI into language agents · world-model
Current LLM agent research almost exclusively optimizes the policy side (state → action), while the other half of the interaction loop — how the environment responds — has long gone seriously unmodeled. The Qwen team trained a "language world model" that unifies simulation of seven environment types — MCP, Search, Terminal, SWE, Android, Web, and OS — through a three-stage CPT→SFT→RL pipeline, with GUI scenarios represented as text via accessibility trees instead of pixels. The paper claims it can serve either as a scalable, perturbation-injectable simulator to replace real environments for agentic RL (reportedly even outperforming training on real environments), or as a "warm-up" that internalizes next-state prediction into meta-reasoning. Two model sizes were released simultaneously: 35B-A3B and 397B-A17B.
Qwen team · arXiv 2606.24597 source · Analysis: 小胡学 LLM source (WeChat, CN)
Survey on latent world models for autonomous driving: the metric shouldn't be how realistic the frames look, but how much you can trust the decision · autonomy
This survey argues that although video generation, occupancy prediction, latent planning, and VLA reasoning produce different outputs, they increasingly share latent representations as a common computational substrate, making them comparable within a single framework. The authors accordingly group methods into four categories — neural simulation, latent planning and reinforcement learning, generative data synthesis, and cognitive reasoning — and propose three conceptual metrics: closed-loop safety gap, temporal coherence, and cautious-reasoning cost, pushing world models to move from "generating plausible-looking futures" toward "generating futures that can support safe decisions." The paper is also candid about a point of contention: methods that merely turn a sunny-day image into a rainy-night one, without learning temporal evolution conditioned on action, are closer to conditional generative models than to world models in the strict sense.
Rongxiang Zeng, Yongqi Dong (RWTH Aachen University · Delft University of Technology) · Submitted to IEEE T-ITS, under review · Analysis: 飞行汽车研究 source (WeChat, CN)
Open Source · Tools · Benchmarks
· WorldScore leaderboard: a unified benchmark for "world generation" that puts 3D, 4D, and video generation models under a single protocol; the test set has 3,000 samples (2,000 static / 1,000 dynamic), scored on 3 controllability metrics + 4 quality metrics + 3 dynamics metrics, normalized. Hosted on Hugging Face and still open for submissions, the current snapshot includes 34 model entries source (WeChat, CN)
II. Funding and Deals
Humanoid (UK's SKL Robotics) | Series A | $152 million | post-money valuation $1.35 billion · humanoid ⚠️ Single-source account
Led by Prime Movers Lab, with participation from Schaeffler, Bosch, Fubon Financial Holding Ventures, and Aglaé Ventures; the company's cumulative funding now stands at $270 million, and by its own account it is Europe's first humanoid-focused unicorn. What matters here isn't the valuation but the industrial structure of this round: Schaeffler has already signed a commercial partnership agreement and plans to deploy thousands of units across its manufacturing facilities, while Bosch will act as a contract manufacturing partner handling hardware design, production, and supply chain — the startup owns the model and the robot body, the industrial groups build and deploy it, a path distinct from the US "Big Tech plus venture capital" model. The product line also deliberately downplays fixation on form factor: the core platform, HMND-01 Alpha, is planned in both wheeled and bipedal versions, and part of this round's funding is earmarked for scaling up production of the wheeled humanoid, with plans to deploy Beta units at manufacturing, logistics, and retail customer sites in Q4 2026. The company disclosed in May roughly 34,000 pre-orders corresponding to about $2.4 billion in projected annual recurring revenue, but these figures come from management statements — pre-orders are not deliveries, and intent-of-purchase amounts should not be treated as confirmed revenue.Source: 机器人电报 source (WeChat, CN)
DeepRobotics AI (深度机智) | New round | Over RMB 100 million (third round in four months) · world-model ⚠️ Single-source account
State-owned capital platforms, industrial capital, and leading financial institutions jointly participated, with existing shareholders following on; the company says the next round is already being finalized. This follows an over-RMB-100-million round in May (Zhongguancun Capital, Chengtong Sci-Tech Innovation Fund, etc.) and a several-hundred-million-RMB round in June (led by China Life Yangtze River Delta Sci-Tech Innovation Fund). The company was co-incubated in 2025 by Beijing Zhongguancun Academy and the Zhongguancun Institute of Artificial Intelligence, and takes a "human learning" approach — driving its embodied foundation model PhysBrain directly from first-person human data rather than fitting robot trajectories. Per its own disclosures, Z-WM_v1 scored 64.96 on WorldArena, surpassing the previous top score of 64.24; on SimplerEnv's WidowX test its average success rate was 80.2% (vs. 57.1% for π0.5), and 98.8% average on LIBERO; the company says it has secured tens of millions of RMB in real orders. All of these benchmark and order figures are from company disclosures and have not been independently reproduced.Source: 具身智能最前沿 source (WeChat, CN)
Zhongke Yuandongli (中科原动力) | Series B2 | Several hundred million RMB | post-money valuation over RMB 1 billion · industrial
Following last week's disclosure of this round, the investor and valuation details have now been confirmed: the round was funded exclusively by Anhui Chery Intelligent Technology, a unit of Chery Holding Group, making it the largest round the company has disclosed to date, with a post-money valuation exceeding RMB 1 billion. This marks the company's 10th funding round.Source: Robo百科 source (WeChat, CN)
Unitree Robotics | STAR Market IPO | Bookbuilding on August 5 · humanoid
Following the previously set August 10 subscription date, the offering timeline has advanced further: multiple media outlets report that bookbuilding — setting the offering price — will take place on August 5. The prospectus had already directly addressed the FCC covered list, stating that models currently on sale have completed certification, while new models face the risk of being unable to obtain the certification needed to enter the US market.Source: Sohu IPO周报 source
III. Commercialization and Deployment
Seoul Gangnam's robotaxi tops 1,000 rides in two-plus months since going paid, with only two vehicles in operation · autonomy
Kakao Mobility disclosed on August 2 that its autonomous taxi service in Seoul's Gangnam district has carried over 1,000 passengers cumulatively since switching to paid fares in April of this year — a period of just over two months. The denominator here is also worth noting: only 2 vehicles are currently in operation, working out to roughly 12 rides per day for the whole fleet, or about 6 rides per vehicle. Operating hours are limited to weekdays, 10 p.m. to 5 a.m., covering Gangnam Station, Apgujeong, Seolleung, Bongeunsa, Daechi, Gaepo-dong, and Yangjae. The company's interpretation is that charging a fare itself changed user perception — shifting from "trying it once out of curiosity" to "a transportation service I'm paying for." It plans to negotiate with the city of Seoul to gradually expand operating routes and coverage, while raising the share of its end-to-end approach within its core autonomous-driving technology, "AI Planner," and putting it through real-vehicle road testing this year.Source: Seoul Economic Daily source
DoorDash pays couriers to load delivery robots — about $5 for 5 minutes · adjacent
Delivery robots that can travel miles on their own are stuck at the last few meters between a restaurant's pickup counter and the curb. Over the past month, couriers in the Phoenix area have periodically received app tasks to "load a Dot robot": one courier in Mesa, Arizona, drove about 2 miles to a restaurant to pick up an order and place it into a robot waiting in the parking lot — a task taking 5 minutes for about $5 in pay. Her question was straightforward: why not have restaurant staff on-site do this instead. Dot is a delivery robot DoorDash launched last September, about stroller-sized with a 30-pound payload capacity, able to travel on roads and sidewalks; the company describes this as a limited pilot to support merchants during peak hours, emphasizing that couriers will still handle the vast majority of deliveries. This kind of "human patching robot" task isn't new — in February, couriers had already received tasks to go close Waymo car doors. The cost of handing off the last stretch of automation is, for now, being shifted onto independent contractors paid per task.Source: Business Insider source
Magic Atom (魔法原子) breaks ground on Wuxi headquarters, plans annual output of 10,000 quadrupeds · humanoid ⚠️ Planning figures
On August 1, Magic Atom's headquarters facility broke ground in Wuxi's Liangxi Science and Technology City, with plans for 4 joint-module production lines and 2 finished-robot production lines; at full capacity, the facility is projected to produce 9,000 small quadruped robots and 1,000 large quadruped robots annually. These are planned figures, not current delivery volumes. The company was founded in January 2024 and, in July, unveiled a full-size humanoid, MagicBot X1, and an industrial wheeled humanoid, MagicBot D1, at WAIC. What is already deployed is its industrial wheeled humanoid, which has entered Dreame's (Chinese smart-home appliance maker) smart manufacturing plant to handle production tasks, while its traffic-management humanoid "Xiaomai" underwent a real-world trial at the Wuxi Marathon.Source: 观点网 source
Shiyue Technology (拾玥科技) signs technical framework agreement with a subsidiary of CRRC · embodied ⚠️ Statement of intent
The two sides plan to collaborate on adapting dexterous manipulation technology for high-end equipment manufacturing scenarios, along with operating-condition validation and technical iteration; the partnership is currently still at the scenario-adaptation and technical-validation stage, with no orders or deliveries involved yet. Shiyue's technology stack includes the Lumi-Mind dexterous manipulation foundation model, the Lumi-Tac visuo-tactile sensor, the Lumi-Dex dexterous hand, and the Lumi-Mod compact motor module.Source: AI云资讯 source
IV. Industry Developments
China's autonomous-driving "blue indicator light" hits pause: newly filed models banned from factory installation, national standard revision launched in parallel · autonomy ⚠️ Reported unofficially
According to a report circulating since July 29, because the practice does not comply with the current mandatory national standard GB 4785-2019, starting from the Ministry of Industry and Information Technology's 410th batch of product announcement filings, passenger vehicles are no longer permitted to come from the factory with an external blue indicator light signaling autonomous-driving status, and inspection agencies must individually check every vehicle submitted for testing. No official document has been published yet, but several automakers have confirmed receiving the notice and stated they will comply. The compliance issue at the root is straightforward: the standard permits only white, red, amber, and yellow external signal-light colors, blue is not on the legally permitted list, and the standard doesn't even include a lighting category for "autonomous-driving indicator light." The China Automotive Technology and Research Center flagged two practical risks the same evening — nighttime glare interfering with the judgment of nearby drivers, and cars fitted with the light being more likely to have other drivers cut in front of them or test the limits of the autonomous-driving system at close range. The feature was first mass-produced as standard equipment by Li Auto's L9 in August 2022, and was subsequently adopted by Xpeng, AITO, Xiaomi, and NIO. Enforcement is currently focused on the new-vehicle admission stage; there is no requirement yet to remove the light from vehicles already on the road, and industry experts believe disabling the blue light via an OTA update would be sufficient to comply. China's national automotive standards committee has launched a revision of GB 4785 (project number 20262535-Q-339), and the GRE and GRVA working groups under the UN's WP.29 are also researching the safety benefits and potential risks of this type of indicator light — in other words, this looks more like tightening admission first and then drawing a compliance boundary through the standard, not necessarily a permanent blanket ban.Source: 时代周报 source
Disney's accelerator backs two physical AI foundation-model companies simultaneously for the first time · embodied
Among the five companies named to the 12th cohort of the Disney Accelerator, announced July 31, all five are AI-related, and two are physical AI foundation-model companies: Physical Intelligence and FieldAI. The division of labor between the two is clear — PI's π-series general-purpose models address "what the robot should do," while FieldAI's Field Foundation Models (FFMs), built around a "belief world model" prediction engine, address how to make reliable decisions in dynamic, unstructured environments without maps, GPS, or preset paths, with the system deployed at the edge without relying on the cloud. FieldAI emerged publicly in August 2025 with a $405 million funding round, backed by a fund tied to Jeff Bezos, Intel Capital, NVIDIA, Bill Gates' Breakthrough Energy Ventures, and Samsung. For Disney, this fills the last missing piece — it already has, at the base layer, the open-source physics engine Newton co-released with NVIDIA and Google DeepMind, plus its own in-house simulator Kamino; at the character layer it already has the BDX and Olaf robots deployed in its parks. What's been missing is the middle layer: general-purpose task execution and autonomous navigation in uncertain environments. The gap is just as clear on the deployment side — Olaf's debut at Disneyland Paris was limited to performing a pre-choreographed routine, and the company has explicitly said visitors shouldn't yet expect extended one-on-one interaction with it; PI's generalization demos have been concentrated on a small number of indoor tasks with outdoor stability unverified, and FieldAI's deployments have mostly been in controlled, physically isolated professional settings — neither company has yet tackled the kind of high-density human-robot environment a theme park represents.Source: Sohu (source: DeepTech深科技) source
San Francisco mayor demands robotaxis "prove it, then deploy," lays out four emergency tests · autonomy
Mayor Daniel Lurie sent a letter last Wednesday to California Transportation Secretary Toks Omishakin, demanding that robotaxi operators demonstrate their vehicles can handle major disruptions before scaling up deployment, laying out four specific tests: quickly clearing broken-down vehicles from traffic lanes, dynamically rerouting in emergencies, sharing operational data with local agencies in real time, and demonstrating through testing that the system can withstand major surges in traffic and demand. The letter was triggered by two incidents: a July 4th fireworks event that drew more than 100,000 people to the waterfront, during which Waymo vehicles got stuck in gridlock, blocked lanes, and trapped a Muni shuttle bus carrying thousands of passengers; and a citywide power outage in December 2025. Lurie had previously backed Waymo's expansion upon taking office, but has now reversed course and thrown his support behind Congressman Kevin Mullin's federal bill; San Francisco's transportation and public safety departments have also voiced support. Waymo responded that it welcomes the mayor's input and said its fleet has successfully supported multiple large-scale events, including World Cup matches.Source: News Anyway source
Hyundai merges AVP headquarters and 42dot into one integrated team, to unveil autonomous-driving roadmap later this month · autonomy
Hyundai Motor Group's two core autonomous-driving R&D organizations — the Advanced Vehicle Platform (AVP) headquarters and 42dot — have formed an integrated team with overlapping executive appointments: Vice President Kwon Jeong-hyun now heads both AVP's Autonomous Driving Development Center and 42dot's Autonomous Driving Division; Director Jeremy Ma now heads both AVP's Silicon Valley office and 42dot's Silicon Valley division; and President Park Min-woo represents both organizations, forming a triangular collaboration structure. The stated division of labor is "42dot internalizes frontier technology, AVP handles mass-production execution." Group Executive Chairman Chung Eui-sun already announced, at an AI summit in San Francisco on July 24, the group's transformation into a physical AI solutions company, with a path extending from autonomous vehicles and robots to AI-defined factories and ultimately city-scale intelligence; the group owns Boston Dynamics. The full autonomous-driving technology roadmap will be unveiled at the 2026 CEO Investor Day on Yeouido, Seoul, on the 26th of this month.Source: Asia Today source
Stardust Intelligence (星尘智能) unveils Lumo-2, releasing 20+ real-robot household chore videos at once · world-model ⚠️ Demo/self-reported
The approach behind Stardust Intelligence's second-generation embodied foundation model, Lumo-2, is "predict the world first, then generate actions" — predicting the future in a lightweight latent dynamics space, which is faster than explicit text-based reasoning and less compute-intensive than directly generating future video; the company describes it as the industry's first implicit world-action model built for home settings. On the engineering side, the company reports that chunked autoregressive decoding makes end-to-end inference 2.71 times faster than standard autoregressive decoding without loss of accuracy. Alongside the release came over 20 real-robot videos covering tasks like flipping an egg in a pan, weighing out 500 grams of millet, catching a ball rolling down from a height, grinding beans and making coffee, tying a bow, and folding clothes. Worth distinguishing: these are demos and self-evaluations — the company claims overall performance across four real, complex tasks surpassing π0.5 and Fast-WAM, but this is a vendor claim, and the success rates and stability over continuous operation have not been independently verified.Source: 机器之心Pro source
SK Telecom convenes 8 Korean physical AI startups, starting with training data built in digital twins · adjacent
SK Telecom held a "Physical AI and Robotics Startup Conference" at SK Searin in Jongno, Seoul, on July 30, attended by executives from SK Hynix, SK AX, and the group's AI committee, along with the CEOs of 8 startups: Gauss Labs, Real World, MakinaRocks, C-Mes Robotics, Wiro Robotics, Uilo Robotics, Config Intelligence, and Holiday Robotics, spanning technology from humanoids and wearable robots to industrial collaborative robots and robot foundation models. The most closely watched topic was connecting SK's AI chips, AI data centers, and digital twin capabilities with the startups' physical AI technology to generate foundation-model training data within virtually replicated industrial sites, and to run joint proofs-of-concept in manufacturing, logistics, and service scenarios.Source: 코리아스타트업포스트 source
NVIDIA partners with Kawasaki Heavy Industries on shipbuilding robots · industrial
The two companies will jointly develop AI robots for shipbuilding in Japan, aiming to build a next-generation digital shipyard combining robotics, AI, and simulation technology, covering welding, painting, inspection, and material handling. This extends NVIDIA's compute and simulation stack into a heavy-industry setting with an extremely long equipment lifecycle and demanding precision requirements — another point of expansion from data centers into industrial robotics and digital twins.Source: Yahoo Finance source
Unitree unveils wheeled UGV "AS2-W," previously spotted in a China-Mongolia joint military exercise · adjacent ⚠️ Vendor-reported
Unitree has listed a four-wheeled unmanned ground vehicle, the AS2-W, on its official site: it weighs 25 kg and can carry up to 150 kg while stationary; unloaded range is about 30 km, dropping to about 25 km with a 16 kg payload. It can autonomously climb an 80 cm obstacle and a 45° slope, moves at over 6 meters per second, has a joint peak torque of about 95 N·m, and is equipped with a 648Wh battery and a 64–128-line LiDAR, with 150 TOPS of onboard compute (Unitree did not specify who manufactures this processor, though it is widely assumed to be an NVIDIA solution). All of the above specifications come from the manufacturer. An armed version of this platform has previously appeared in the China-Mongolia "Steppe Partner 2026" joint exercise, with Chinese military reports stating that similar robot platforms carried out reconnaissance and fire-support tasks.Source: TechRadar source
Hardware · Supply Chain
· Bill-of-materials (BOM) cost for a full unit: two cost estimates diverge — one puts the drop from about $130,000 in 2023 to about $30,000 in July 2026 (a 70% decline), the other puts it at a drop from $131,000 to $46,000 (a decline of nearly two-thirds); the two estimates also diverge on the localization rate for core components in China (over 75% vs. already at 85% in the first half of the year). The consensus is that the main cost-reduction battleground is actuators, lead screws, and reducers, and that the next price step-down awaits scaled-up Chinese production of lead screws source (WeChat, CN)source (WeChat, CN)
· Planetary roller screws: the component with the highest barrier to entry and the widest divergence in estimates — accounting for about 22% of BOM, with Chinese localization rates variously reported at under 10% or 40–50%; imported units cost RMB 5,000–10,000 each, and localized production could bring that down to around RMB 2,000, bottlenecked by 0.5-micron-level machining precision and yield ramp-up from 60% to 85% source (WeChat, CN)source (WeChat, CN)
· Harmonic reducers: LEADERDRIVE (绿的谐波) has signed a three-year exclusive supply agreement with Unitree covering 110,000 units of finished-robot supply; humanoid-related orders in Q1 2026 reached RMB 582 million, up 286% year-over-year. Average unit price has already fallen from about RMB 1,900 in 2017 to about RMB 1,000, with a longer-term target of RMB 400–500 once scaled production is reached source (WeChat, CN)
· Coreless motors: Zhongshen Dynamics (灵巧驱控) says its on-hand orders for the first half of the year reached 1 million units, versus 120,000 for all of last year; Leadshine Technology (雷赛智能) delivered 120,000 units in 2025, a 20-fold year-over-year increase. These are industry-reported figures, and on-hand orders are not the same as deliveries source (WeChat, CN)
· Dahuan UDH-3-7 three-fingered dexterous hand: the full hand weighs 900 grams with 7 active degrees of freedom, fingertip force of 50N, maximum lift/pull load of 15 kg, and thumb parallel-swing support of 0–180°; at WAIC it was mounted on 7 general-purpose humanoids from Geek+ (Chinese warehouse-robotics maker) to complete tote-box handling and irregular-item sorting source (WeChat, CN)
· LinkerHand (灵心巧手): in just over a year, the company has raised seven funding rounds totaling over RMB 2 billion by its own disclosures, and its products cover all three drivetrain approaches — tendon-driven, direct-drive, and linkage-based — the leader hasn't converged on one, suggesting the downstream market has not yet settled on a unified hand architecture; third-party data puts total dexterous-hand sales across the industry in 2025 at about 19,200 units, with Fourier Intelligence (Chinese robotics maker) leading with over 10,000 units source (WeChat, CN)
Top comments (0)