期刊文献+

基于Transformer的通用和鲁棒体素特征学习的目标检测 被引量:1

General and robust voxel feature learning with Transformer for 3D object detection
下载PDF
导出
摘要 自注意力网络和Transformer主导了机器翻译和自然语言处理领域,并在诸如图像分类和目标检测等图像视觉任务中显示出巨大潜力。受到Transformer在2D图像视觉任务中取得的巨大进步的启发,提出了一种基于传统Transformer的新颖和鲁棒的体素特征编码器。首先,探究自注意力对序列数据的排列不变性,并将其应用于点云数据处理。其次,基于自注意力构造体素特征层,根据体素内所有点之间的空间关系和上下文信息交换自适应地学习体素的局部和鲁棒上下文。最后,构建了以体素特征层为核心的通用3D目标检测框架。VFT(voxel feature learning with Transformer)是通用的体素特征提取器,可以嵌入任何其他基于体素方法的3D物体检测框架中。在KITTI数据集上进行的实验结果表明,本方法在3D目标检测方面表现出优越的性能。 The self-attention networks and Transformer have dominated machine translation and natural language processing fields,and shown great potential in image vision tasks such as image classification and object detection.Inspired by the great progress of Transformer,we propose a novel general and robust voxel feature encoder for 3D object detection based on the traditional Transformer.We first investigate the permutation invariance of sequence data of the self-attention and apply it to point cloud processing.Then we construct a voxel feature layer based on the self-attention to adaptively learn local and robust context of a voxel according to the spatial relationship and context information exchanging between all points within the voxel.Lastly,we construct a general voxel feature learning framework with the voxel feature layer as the core for 3D object detection.The voxel feature with Transformer(VFT)can be plugged into any other voxel-based 3D object detection framework easily,and serves as the backbone for voxel feature extractor.Experiments results on the KITTI dataset demonstrate that our method achieves the state-of-the-art performance on 3D object detection.
作者 李阳 葛洪伟 LI Yang;GE Hongwei(Jiangsu Provincial Engineering Laboratory of Pattern Recognition and Computational Intelligence, Wuxi 214122, China;School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi 214122, China)
出处 《Journal of Measurement Science and Instrumentation》 CAS CSCD 2022年第1期51-60,共10页 测试科学与仪器(英文版)
基金 National Natural Science Foundation of China(No.61806006) Innovation Program for Graduate of Jiangsu Province(No.KYLX160-781) University Superior Discipline Construction Project of Jiangsu Province。
关键词 3D目标检测 自注意力网络 基于Transformer的体素特征学习 点云 编码解码器 3D object detection self-attention networks voxel feature with Transformer(VFT) point cloud encoder-decoder
  • 相关文献

同被引文献4

引证文献1

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部