Why do AI-generated videos often look good but break physics?
Google researchers with University of Copenhagen and Oxford introduced VGGRPO a new method that trains a latent geometry model to directly read scene depth and motion from a video diffusion model’s internal representations.
Then it uses reinforcement learning to reward smooth camera paths and consistent 3D geometry.
Result: VGGRPO outperforms prior approaches in camera stability, geometry consistency, and overall quality while eliminating expensive VAE decoding.
Post #4388
826
- ❤🔥 4
- 👍 4
- 🥰 2