There are two ways to make a robot act, and the field hasn't picked one yet.
A VLA takes what the camera sees, what you tell it, and the robot's own state, and maps it straight to an action. A trained reflex, learned from huge robot datasets. π0.5 and SmolVLA work like this. It's fast and simple because almost nothing sits between seeing and moving. The weakness is that it only knows what it learned. Push it far enough outside its training and it can fail, because it carries no model of what happens next.
A world model goes the other way. It learns to predict how the world changes, then uses that learned simulator to test moves and imagine outcomes before the robot commits. Dreamer is built this way, and DayDreamer put it on real robots. The strength is foresight: it can reason over possible futures instead of repeating past actions. The cost is compute and its own mistakes, because every imagined future has to be predicted, and a bad prediction means a bad move.
So the trade is reflex against foresight. They aren't rivals though. World models are increasingly used to generate experience and training data for the action models. The likely future is hybrid, where the imagination teaches the reflex.