Abstract
3D open-vocabulary semantic segmentation has shown great potential in applications such as autonomous driving and mixed reality. However, achieving accurate segmentation in dynamic environments remains challenging due to motion-induced inconsistencies. To address this issue, we incorporate scene flow as temporal information into a static semantic backbone to enhance semantic consistency and accuracy over time. Our method captures inter-frame motion cues from point cloud sequences and leverages them, together with a local clustering mechanism, to refine semantic label consistency in consecutive frames. Furthermore, we introduce a two-way scene flow-based data augmentation strategy that exploits both forward and backward motion to jointly train the model in bidirectional temporal contexts. On the large-scale nuScenes autonomous driving dataset, our method achieves a 0.4% overall improvement in hIoU and a 2.37% gain under high-motion scenes. On the synthetic object-centric dataset, it achieves a 4.53% overall hIoU improvement and a 6.09% gain in high-motion scenes, while reducing the ID switch rate by 0.5%.
| Original language | English |
|---|---|
| Pages (from-to) | 8140-8147 |
| Number of pages | 8 |
| Journal | IEEE Robotics and Automation Letters |
| Volume | 11 |
| Issue number | 7 |
| DOIs | |
| Publication status | Published - 2026 Jul 1 |
Keywords
- Semantic scene understanding
- deep learning for visual perception
ASJC Scopus subject areas
- Control and Systems Engineering
- Biomedical Engineering
- Human-Computer Interaction
- Mechanical Engineering
- Computer Vision and Pattern Recognition
- Computer Science Applications
- Control and Optimization
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Open-Vocabulary Semantic Segmentation for Dynamic 3D Scenes Using Scene Flow Estimation'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS