A General World Model Is Still Far. Next-State Prediction Already Runs.
2026-10-08
World Labs put the phrase back in circulation. Marble turns a picture into a 3D scene you can walk through. Atlas, released in September 2026, is billed as a next-generation world model: text, images, video, and 3D go into one spatial context, and the model generates what comes next. It is meant to power later versions of Marble, and to feed robot simulation.
It looks as if machines are close to understanding the physical world.
A general world model is still far away.
That claim has to be split. Left in one piece, it is either empty pessimism, or the reverse exaggeration: a trajectory predictor in a car, called a finished foundation model.
1. The name covers three things
Generative world models. Sora, Genie, Marble, Atlas. A picture or a sentence goes in. A video, a 3D scene, or a new camera view comes out. Genie takes actions. Atlas does more than paint the next pixel. It puts several inputs into one spatial representation.
Looking right is still not understanding. Gravity breaks. Objects vanish. Over a long stretch the scene stops matching itself. The model learns what usually comes next. A larger scale might grow some physical intuition. There is not enough evidence to treat a demo as an ability. There is also no evidence that scaling cannot produce it. Leave both open.
Next-state prediction already running in engineering. This layer is old, and it often does not use the name. Trajectory prediction, occupancy networks, and ego dynamics in driving. Grasp and motion prediction in robots. Cloud simulators that manufacture rare cases and run closed-loop tests. The job is the same: given the state and an action, estimate what happens next.
Draw the line. Under this definition a bicycle model, a Kalman filter, and the physics engine inside a simulator all count. "World models are already in use" means this function is old. It does not mean a new foundation model is done.
Occupancy networks and trajectory predictors enter decisions because they write the world as grids, trajectories, and occupancy, and those can be scored. A collision is measurable. A miss is measurable. A generative model leaves the state in the picture, and a mistake often only looks a little odd. Both learn statistics. Both can break physics. The difference is where the state is written, and whether a mistake can be caught.
The general world model. This is the one Fei-Fei Li and Yann LeCun are after. One representation that holds 3D space and time, objects, cause, and interaction. It still works in another scene and on another body. It can answer what would have happened if the action had been different. Planning that uses it beats a special-purpose module in a closed loop.
That one is far.
2. "What was seen" does not give "what if I act"
The usual explanation is that data and compute are short, and enough of both will grow the ability. Atlas's own bet is the same: more training compute, more ability.
That bet is not falsified. It should not be treated as already falsified. Hard only means it cannot be done today.
The gap that can be stated is narrower. Next-frame training learns what usually follows what was seen. Planning asks what the world does if I take an action that never happened.
Records of what already happened do not separate the two. A cup falls because I pushed it, or because it slipped. Correlation can look convincing. Telling the cause takes an action, then the result. More samples make "what usually happens" more stable. If the data has no record of acting, intervention and counterfactuals do not appear on their own.
A few concrete difficulties follow.
Representation. Geometry, meaning, material, dynamics, and intent still live in different tools. One shared context means the model can read them together. Reading them together is not a representation you can intervene on.
Interaction. The model has to separate a change I caused from a change the world made on its own. Action has to be a condition, and the actions cannot be only the ones people already took.
Closed loop. On a robot, in a car, or in AR/MR, trying it yourself is expensive, risky, and full of rare cases.
Evaluation. Accurate pixels are not correct physics. Physics tests and closed-loop tests exist. They do not line up with each other, and they do not predict whether a system can ship. The missing piece is a standard that matches real action.
Latency. A car and a robot need low delay, and a failure that can be explained. Today's generative models do not sit in that seat.
3. Driving is the cleanest test
Some people say autonomous driving does not use a world model. On the vehicle it is perception, prediction, and planning. That is half right.
The vehicle will not run an explicit generative world model. Compute, latency, safety checks, and the need to explain a failure all forbid it. In that sense, driving does not run that kind of world model.
The vehicle still cannot do without "what happens next." Trajectory prediction estimates how other vehicles and people move. An occupancy network estimates which space will be taken. A dynamics model estimates how the vehicle responds after an action. Cloud simulation uses the same kind of method to build rare cases and run closed-loop tests. These are narrow predictors, and they can be checked.
When planning uses dynamic programming or model predictive control, this term is that predictor:
V(s) = min over a of [ R(s, a) + γ Σ P(s' | s, a) V(s') ]
A model-free method can skip writing down P(s' | s, a) and learn the action directly. Closed-loop control depends on how the environment changes, either as an explicit probability or welded into the policy.
An end-to-end network that drives shows that it reacts to the road. It does not show that a piece can be pulled out and asked, "what if I change lanes now?" Reacting and being available to ask are different things.
What the car already has is a special-purpose predictor on a narrow state: lanes, vehicles, occupancy. The argument is about another road, another kind of road user, an action never seen, and whether that answer is more reliable than the modules already there.
4. AR/MR supplies samples, not the model
AR/MR is often pulled into the same picture. Its value is not a few extra labels. In real 3D space it records, continuously and with position, what the scene was, what the person did, and what the scene became.
Camera pose, depth, meshes, and point clouds record space. Object class, material, and whether something can be handled record attributes. Gesture, gaze, speech, and grasp trajectories record action. The scene before, during, and after the action records order in time.
The samples are useful. They are still mostly trajectories of what was seen. Force at contact, mass, and friction are often missing from the picture. AR/MR is a data factory and a proving ground. It is not a general world model.
5. How to tell it is getting close
"Far" has to be something that can be proven wrong.
Long consistency. For minutes, objects remain, physics still holds, and the scene still matches itself.
Intervention. Given an action, the model predicts how the state changes. A next frame that merely looks reasonable does not count.
Counterfactuals. What if the action had been different. A harder check: two models can tie on "does the next frame look right" and disagree on "what if I had not done that." A high prediction score does not pin down the counterfactual.
Closed loop. On a robot or a driving task, planning with the model beats the special-purpose method already in use. This is the ledger. If the earlier items fail, this one does not arrive.
Physics. Gravity, collision, friction, and occlusion hold steadily, rather than by an occasional lucky guess.
Cause. It separates occurring together from causing. It supports both "I act" and "what if I had not."
None of these is stably met for the general case. "Still far" has criteria.
6. What is actually true
Next-state prediction, as a function, already runs in engineering. The thing called a world model, and expected to be a third foundation model after language models and vision-language models, is early. Across scenes. An action never seen. A counterfactual. A closed loop where it beats a special-purpose module.
On the vehicle, the stack is still perception, prediction, and planning. Large models carry meaning and reasoning. Next-state prediction sits in a special-purpose module, or is welded into an end-to-end policy. A generative model does not run there on its own.
In cloud simulation the method is already useful: rare cases, closed-loop tests, synthetic data.
AR/MR is a source of data and a place to check.
Robots are the closest proving ground, and the least mature. Here the system has to act and then see the result. "What was seen" and "I acted" start to come apart.
7.
Two extremes are both wrong in the same way. Each one flattens the word.
One says a world model is about to change everything. A demo is treated as an ability. Generation is treated as understanding.
The other says a world model is hype and useless, and throws out the predictors already running in cars and simulators.
The direction is right. Representation, prediction, and action in the physical world are a main line after language models and vision-language models. The work will enter driving, robots, and AR/MR in pieces, by domain, sometimes as its own module and sometimes welded into a policy. It will not arrive one day as a single general model and rewrite the rest.
Maturity is three questions. Can the world be written as a state a program can use. Can change be predicted reliably. Can the system act safely on that prediction.
On a narrow task, a lane or a grasp, usable versions of all three already exist. Change the body, the physics, or the action to one never seen, and all three break.
The third question is the watershed. The first two can stop at "looks real." The third requires the answer to be right, because a wrong action causes harm. In the general sense, that watershed is still far.
Top comments (0)