TGViewer
Channel Public Channel
iMapDAY

iMapDAY

@imapday

Сделал канал для размещения новостей от меня @yuddim и моей команды, занимающейся трехмерным компьютерным зрением роботов и автомобилей. Также давно хотелось собирать в одном месте интересные для меня научные публикации и технологические заметки.
Subscribers
358
Photos
630
Videos
74
Links
176

Showing posts older than #723 · Back to latest

Older Posts 4 shown
Post #722 221
И продолжение:
• Task-Aware Bimanual Affordance Prediction via VLM-Guided Semantic-Geometric Reasoning (https://openreview.net/attachment?id=05Zf4KJrRn&name=pdf)
• Semantic–Geometric Task Representations for Bimanual Manipulation from Human Demonstrations to Robot Action Planning (https://openreview.net/attachment?id=R2I0asfyZT&name=pdf)
• Towards Task Planning for Proactive Safety in Home Service Robots: A Scene Graph-Augmented LLM Approach (https://openreview.net/attachment?id=juTZvTVQvT&name=pdf)
• PEACE: A Planner–Executor Agent with Constraint Enforcement for UAVs (https://openreview.net/attachment?id=inQPnS3D75&name=pdf)
• Energy-Based Action Heads Know When They Don’t Know (https://openreview.net/attachment?id=PF9lhBseXQ&name=pdf)
• M2H-MX: Multi-Task Semantic and Geometric Perception for Real-Time Monocular 3D Scene Graph Construction (https://openreview.net/attachment?id=KG9QYhLIBV&name=pdf)
• Uncertainty-Aware Symbolic State Monitoring with Vision-Language Models and Gaussian Naive Bayes (https://openreview.net/attachment?id=9vqwDyFhDX&name=pdf)
• From Obstacles to Etiquette: Robot Social Navigation with VLM-Informed Path Selection (https://openreview.net/attachment?id=gnpxmJ3S2j&name=pdf)
• LiftNav: Path Planning via Semantic Lifting in TSDF-Guided Gaussian Splatting (https://openreview.net/attachment?id=s6rxtvo20o&name=pdf)
• R5DGS: Semantic-Aware 4D Gaussian Splatting with Rigid Body Constraints for Efficient Dynamic Scene Reconstruction (https://openreview.net/attachment?id=LFHg6tcn7e&name=pdf)
• Tool-Augmented VLM Agents for Zero-Shot 3D Visual Grounding on Point Clouds (https://openreview.net/attachment?id=Ft2t57khSZ&name=pdf)
• LLM Tool Workflows for Robot Explainability and Natural Language Commanding (https://openreview.net/attachment?id=NUu9P1LwbT&name=pdf)
• Unified Point Cloud Corruption-Aware Reasoning Engine (https://openreview.net/attachment?id=63dSwa83gC&name=pdf)
• Where Did I Leave My Glasses? Open-Vocabulary Semantic Exploration in Real-World Semi-Static Environments (https://openreview.net/attachment?id=4SSJXzXnxM&name=pdf)
  • 🔥 1
Post #721 192
Из-за высокой релевантности этого воркшопа приведу полный список принятых докладов на него (там есть два доклада от коллег из ИТМО), ссылки работают если авторизоваться в openreview:
• Distributional Semantics for Robust Global Localization in Cluttered, Geometrically Aliased Environments (https://openreview.net/attachment?id=FjjDqommMK&name=pdf)
• IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models (https://openreview.net/attachment?id=oMOpXYZJ8v&name=pdf)
• Introducing a framework to reduce foundation models hallucinations in robotics application (https://openreview.net/attachment?id=WGyXVAfAqR&name=pdf)
• HERMES: Habit- and Episode-aware Retrieval Memory for Embodied Systems (https://openreview.net/attachment?id=IkUhq9v0JX&name=pdf)
• 4D Latent Mapping for Mobile Manipulation Policy Learning (https://openreview.net/attachment?id=1aaWTejVLX&name=pdf)
• Efficient world models with tree-structured sparsity (https://openreview.net/attachment?id=xXrDiTMLH5&name=pdf)
• Occupancy-Aware Reasoning for Safe Quadrotor Navigation: Perception-Aware MPPI (https://openreview.net/attachment?id=aOrQSYmp9L&name=pdf)
• IFG: Internet-Scale Guidance for Functional Grasping Generation (https://openreview.net/attachment?id=rYOqxbLwyh&name=pdf)
• PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution (https://openreview.net/attachment?id=PqZOaoqtso&name=pdf)
• BlabberSeg: Semantic Perception for Reliable Open-Vocabulary UAV Safe Landing (https://openreview.net/attachment?id=2oMctsz534&name=pdf)
• EVE: A Generator-Verifier System for Generative Policies (https://openreview.net/attachment?id=hKPbg5pu4q&name=pdf)
• NaviTrace: Evaluating Embodied Navigation of Vision-Language Models (https://openreview.net/attachment?id=rv6VhBMW7I&name=pdf)
• Ontology-Guided Reasoning for Affordance-Based Explanations of Robot Navigation (https://openreview.net/attachment?id=ZzABFRPCr9&name=pdf)
• From Exploration to Reuse: An Embodied Agent Framework for Manipulation Skill Learning (https://openreview.net/attachment?id=Dwya5QWvED&name=pdf)
• Affordance-based Robot Manipulation with Flow Matching (https://openreview.net/attachment?id=MmtkpTDfwN&name=pdf)
• Learning Compositional Symbolic Task Rules from Demonstrations with Inductive Logic Programming (https://openreview.net/attachment?id=OjVJRVwGS7&name=pdf)
• Flying with Style: Learning Agile Aerial Cinematography from Labeled Egocentric Video (https://openreview.net/attachment?id=Fjhagih2vV&name=pdf)
• Dynamic Control Barrier Function Regulation with Vision-Language Models for Safe, Adaptive, and Realtime Visual Navigation (https://openreview.net/attachment?id=uMoX27VOyY&name=pdf)
• Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow (https://openreview.net/attachment?id=NMQ9Qw5j3i&name=pdf)
• Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning (https://openreview.net/attachment?id=MqiKNCxZOu&name=pdf)
• RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots (https://openreview.net/attachment?id=6qkdDhfhWn&name=pdf)
• Integrating Topological Object Recognition into Semantic SLAM for Unseen Cluttered Environments (https://openreview.net/attachment?id=hnAp82m78m&name=pdf)
• Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection (https://openreview.net/attachment?id=G7NAwtCjdV&name=pdf)
• GovernorVLA: Adaptive Embodied Reasoning for Long-tail Autonomous Driving (https://openreview.net/attachment?id=tgZFh6O1xW&name=pdf)
• Towards Contextual Robot Intent Explanation for Hierarchical Vision–Language–Action Collaboration (https://openreview.net/attachment?id=NtTlOzgUkc&name=pdf)
• Latent Dynamics Modulation Guidance for Semantic Obstacle Avoidance in Diffusion Policies (https://openreview.net/attachment?id=w4rhh0VpES&name=pdf)
• TIGR: a Mixture-of-Foundation-Model-Experts for 3D-informed Task-Aware Grasping (https://openreview.net/attachment?id=XEB7wJSFbQ&name=pdf)
  • 🔥 2
Post #711 224
А также много времени провел на воркшопе по извлечению и использованию семантической информации для повышения автономии роботов https://www.dynsyslab.org/icra2026-workshop-on-semantics-for-reliable-robot-autonomy/

На этом воркшопе выступал Lukas Scmid (H-index 15, https://scholar.google.com/citations?hl=en&user=r79fGI0AAAAJ) автор полезного метода построения графов 4D-сцены DAAAM (Describe Anything Anywhere at Any Moment) (проект с кодом https://nicolasgorlo.com/DAAAM_25), а также метода Gaussian Mapping for Evolving Scenes https://arxiv.org/abs/2506.06909. Он достаточно активен и судя по аффилиациям, ушел из MIT и возглавил команду в Германии. Также он прорекламировал свежую книгу SLAM Handbook - https://github.com/SLAM-Handbook-contributors/slam-handbook-public-release/tree/main - она выглядит современной и полезной. Он также предствил масштабный онлайн-архив материалов по 3D-графам сцены https://3dscenegraphs.com (обязательно загляните!)

Аспирант Kush Hari и его руководитель Ken Goldberg рассказали о их свежем датасете и бенчмарке Robo2VLM (https://berkeleyautomation.github.io/robo2vlm/) и подходе RoboSQ: Semantic Queries for Task-Aligned Robot Training Data (https://autolab.berkeley.edu/assets/publications/media/2026_ICRA_Robo_SQ_final_version.pdf). Сам Ken Goldberg повторил кусок своей пленарной лекции про Code-as-Policy (https://arxiv.org/abs/2603.22435) и Graph-as-Policy для качественного и интерпретируемого управления роботами c помощью кодовых агентов.

Повторил кусок своего доклада и Manolis Savva из Simon Fraser University (H-index 53, https://scholar.google.com/citations?user=4D2vsdYAAAAJ&hl=en&oi=ao).

Выступала также Masha Itkina (H-index, https://scholar.google.com/citations?user=JAmTk5gAAAAJ), рассказывала про создание отказоустойчивых систем управления роботами, например, проект FAIL-Detect (https://cxu-tri.github.io/FAIL-Detect-Website/). Также она рассказывала про свои авторитетные публикации https://toyotaresearchinstitute.github.io/lbm1/, https://co-training-lbm.github.io, https://tri-ml.github.io/vla_foundry/
Post #701 149
В заключительный день 5 июня удалось посетить два воркшопа.

Побывал на воркшопе по локализации роботов на основе построенных заранее карт https://sites.google.com/view/icra2026-priors-map-workshop/home

На нем запомнились вот такие постеры:
-vS-Graphs: Integrating Visual SLAM and Situational Graphs (https://arxiv.org/abs/2503.01783) (проект с датасетом https://snt-arg.github.io/vsgraphs-results/) (код https://github.com/snt-arg/visual_sgraphs)
-Informed Visual S-Graphs (ivS-Graphs) - BIM Informed Visual SLAM for Construction Monitoring (https://arxiv.org/abs/2509.13972v1)
-COMPASS: COmpact Multi-channel Prior-map And Scene Signature for Floor-Plan-Based Visual Localization (https://arxiv.org/abs/2604.25388)
-OsmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning https://arxiv.org/html/2507.12753v1 (код https://anonymous.4open.science/r/osmAG-LLM)
-OSM-BKI: OpenStreetMap-Guided Bayesian Kernel Inference for Domain-Robust LiDAR Semantic Mapping (код https://github.com/arpg/OSM-BKI)
-INHerit-SG: Incremental Hierarchical Semantic Scene Graphs with RAG-Style Retrieval (https://arxiv.org/abs/2602.12971) (проект с датасетом, но пока без кода https://fangyuktung.github.io/INHeritSG.github.io/)
- CAD Assemblies as Object-Level Priors for Robotic Operations Planning (https://drive.google.com/file/d/15PoyRPDi-dIDrGhh97lHirCgJRZ_x5R6/view?usp=drivesdk )
- RoboBIM: Scaling Robot Semantic Foundations through Agentic Extraction of BIM Priors (код https://show2instruct.github.io/RoboBIM/)
- Sat-RoMa: Robust Dense Satellite Geo-Registration for GPS-Denied Navigation
- Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation (https://arxiv.org/abs/2605.18184)
-Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment (https://arxiv.org/abs/2603.01613)
-Towards Zero-Shot Global Localization via Dense Embedding Maps Based on Satellite Priors (код https://github.com/ctu-mrs/mrs_uav_system ) (симулятор https://mrs.fel.cvut.cz/flight-forge)
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →