The researchers took the basic Veo-2 and fine-tuned it based on the first frame + the robot's actions to generate future consistent frames from its 4 cameras. This is called action-conditioned rollout and, in essence, allows for inexpensive and safe evaluation of the robot's policy using just the world model.
🟢 Why is this cooler than regular simulation?
Strict physical simulators work well if the situation is simple and predictable. AI simulation can be scaled to non-trivial worlds. Moreover, every object in a physical simulator requires clear assets and manual tuning + heavy calculations. You can't go far with that. Here, on the other hand, you can add new objects and cases as much as you want - just write a prompt or edit the initial frame with Nano Banana.
Of course, there are also downsides. They mainly concern the quality of strict modeling, especially of fine physics. But there's still a lot to come.
Google has learned to fairly decently evaluate the robot's policy using Veo (see the last graph). Add a policy update, and you'll already get reinforcement learning. So far, they're not doing this consciously, again due to the lack of accuracy in the World model.
The resulting fine-tuning was cleverly named Veo (Robotics). It's A Big Step
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
