Сотрудники нашего Центра регулярно обмениваются полезными статьями. А мы решили поделиться этими материалами с вами!
В сегодняшней подборке статьи на тему CV:
🟡Collaborative Dynamic 3D Scene Graphs for Open-Vocabulary Urban Scene Understanding
In this work, we present CURB-OSG, an open-vocabulary dynamic 3D scene graph engine that generates hierarchical decompositions of urban driving scenes via multi-agent collaboration.
🟡SAPFormer: Shape-aware propagation Transformer for point clouds
In this paper, we propose the Shape-Aware Propagation Transformer (SAPFormer), which flexibly captures the semantic information of point clouds in geometric space and effectively extracts the contextual geometric space information.
🟡SplatTalk: 3D VQA with Gaussian Splatting
In this work, we introduce SplatTalk, a novel method that uses a generalizable 3D Gaussian Splatting (3DGS) framework to produce 3D tokens suitable for direct input into a pretrained LLM, enabling effective zero-shot 3D visual question answering (3D VQA) for scenes with only posed images.
🟡Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
In this paper, we exploratively propose a novel method called TPCNet, the first outdoor 3D visual grounding model upon the paradigm of prompt-guided point cloud sensor combination, including both LiDAR and radar contexts.🟡CQVPR: Landmark-aware Contextual Queries for Visual Place Recognition
We propose the Contextual Query VPR (CQVPR), which integrates contextual information with detailed pixel-level visual features. By leveraging a set of learnable contextual queries, our method automatically learns the high-level contexts with respect to landmarks and their surrounding areas.
#CV