🔥 Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
💡 The paper presents Lift4D, a test-time optimization framework for reconstructing dynamic non-rigid objects from monocular video. The problem addressed is the difficulty in reconstructing 4D representations of dynamic objects from single-view video due to the scarcity of 4D training data and the limitations of prior approaches that either directly predict 4D representations or initialize a 3D representation and refine it based on video evidence.
The method involves adapting a single-view 3D reconstruction model to yield temporally consistent per-frame predictions, which provides a coherent initialization for a deformable 3D Gaussian Splatting representation. This representation is then optimized to match the input video through an occlusion-aware optimization that recovers visible surface details and completes unobserved regions using a view-conditioned diffusion prior.
The results show that Lift4D improves over prior 4D reconstruction methods, particularly on challenging in-the-wild sequences with severe occlusions and non-rigid motion. The framework effectively handles complex scenarios by integrating visual cues from direct observations with data-driven priors over geometry and appearance, making it a significant contribution to the field of 4D reconstruction from monocular video.
📅 Published on Jun 22
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.23688
• PDF: https://arxiv.org/pdf/2606.23688
• Project Page: https://lift4d.github.io/
━━━━━━━━━━━━━━━━━━━━━━━━
@Machine_learn
Post #4009
2.76K