LiDAR semantic segmentation is a crucial task in autonomous driving and robotics, where real-time performance is essential for online decision-making. Recent trends exploit range images and Vision Transformers, using the self-attention mechanism. However, these approaches often lack explicit spatial priors and involve a large number of parameters. To tackle these limitations, we propose a novel method, adapting the Retentive Network architecture from the Natural Language Processing (NLP) field, for its efficient sequence modeling capabilities, directly operating on the range-view representation. Our approach incorporates a circular retention (CiR) mechanism that explicitly captures spatial relationships and continual circular property of the range image while modeling long-range dependencies and preserving the receptive field. In addition, we introduce a new set of range-view augmentations, adapted from 3D techniques, to improve generalization and mitigate class imbalance. Extensive experiments on three large-scale datasets, as SemanticKITTI, PandaSet and Semantic-POSS demonstrate that our method achieve state-of-the-art performance among range-view approaches on two out of three datasets, while satisfying real-time constraints. The code is available at https://github.com/SiMoM0/RangeRet.

Revisiting Retentive Networks for Fast Range-View 3D LiDAR Semantic Segmentation

Mosco, Simone
;
Li, Wanmeng;Pretto, Alberto
2026

Abstract

LiDAR semantic segmentation is a crucial task in autonomous driving and robotics, where real-time performance is essential for online decision-making. Recent trends exploit range images and Vision Transformers, using the self-attention mechanism. However, these approaches often lack explicit spatial priors and involve a large number of parameters. To tackle these limitations, we propose a novel method, adapting the Retentive Network architecture from the Natural Language Processing (NLP) field, for its efficient sequence modeling capabilities, directly operating on the range-view representation. Our approach incorporates a circular retention (CiR) mechanism that explicitly captures spatial relationships and continual circular property of the range image while modeling long-range dependencies and preserving the receptive field. In addition, we introduce a new set of range-view augmentations, adapted from 3D techniques, to improve generalization and mitigate class imbalance. Extensive experiments on three large-scale datasets, as SemanticKITTI, PandaSet and Semantic-POSS demonstrate that our method achieve state-of-the-art performance among range-view approaches on two out of three datasets, while satisfying real-time constraints. The code is available at https://github.com/SiMoM0/RangeRet.
2026
Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11577/3616958
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 1
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact