摘要
3D open-vocabulary semantic segmentation has shown great potential in applications such as autonomous driving and mixed reality. However, achieving accurate segmentation in dynamic environments remains challenging due to motion-induced inconsistencies. To address this issue, we incorporate scene flow as temporal information into a static semantic backbone to enhance semantic consistency and accuracy over time. Our method captures inter-frame motion cues from point cloud sequences and leverages them, together with a local clustering mechanism, to refine semantic label consistency in consecutive frames. Furthermore, we introduce a two-way scene flow-based data augmentation strategy that exploits both forward and backward motion to jointly train the model in bidirectional temporal contexts. On the large-scale nuScenes autonomous driving dataset, our method achieves a 0.4% overall improvement in hIoU and a 2.37% gain under high-motion scenes. On the synthetic object-centric dataset, it achieves a 4.53% overall hIoU improvement and a 6.09% gain in high-motion scenes, while reducing the ID switch rate by 0.5%.
| 原文 | 英語 |
|---|---|
| 頁(從 - 到) | 8140-8147 |
| 頁數 | 8 |
| 期刊 | IEEE Robotics and Automation Letters |
| 卷 | 11 |
| 發行號 | 7 |
| DOIs | |
| 出版狀態 | 已發佈 - 2026 7月 1 |
ASJC Scopus subject areas
- 控制與系統工程
- 生物醫學工程
- 人機介面
- 機械工業
- 電腦視覺和模式識別
- 電腦科學應用
- 控制和優化
- 人工智慧
指紋
深入研究「Open-Vocabulary Semantic Segmentation for Dynamic 3D Scenes Using Scene Flow Estimation」主題。共同形成了獨特的指紋。引用此
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS