跳到主要导航 跳到搜索 跳到主要内容

Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection

  • Xin Li
  • , Botian Shi
  • , Yuenan Hou
  • , Xingjiao Wu
  • , Tianlong Ma
  • , Yikang Li
  • , Liang He*
  • *此作品的通讯作者
  • East China Normal University
  • Shanghai AI Laboratory
  • Fudan University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D points and dense 2D pixels. Recent approaches either fuse the image features with the point cloud features that are projected onto the 2D image plane or combine the sparse point cloud with dense image pixels. These fusion approaches often suffer from severe information loss, thus causing sub-optimal performance. To address these problems, we construct the homogeneous structure between the point cloud and images to avoid projective information loss by transforming the camera features into the LiDAR 3D space. In this paper, we propose a homogeneous multi-modal feature fusion and interaction method (HMFI) for 3D object detection. Specifically, we first design an image voxel lifter module (IVLM) to lift 2D image features into the 3D space and generate homogeneous image voxel features. Then, we fuse the voxelized point cloud features with the image features from different regions by introducing the self-attention based query fusion mechanism (QFM). Next, we propose a voxel feature interaction module (VFIM) to enforce the consistency of semantic information from identical objects in the homogeneous point cloud and image voxel representations, which can provide object-level alignment guidance for cross-modal feature fusion and strengthen the discriminative ability in complex backgrounds. We conduct extensive experiments on the KITTI and Waymo Open Dataset, and the proposed HMFI achieves better performance compared with the state-of-the-art multi-modal methods. Particularly, for the 3D detection of cyclist on the KITTI benchmark, HMFI surpasses all the published algorithms by a large margin.

源语言英语
主期刊名Computer Vision – ECCV 2022 - 17th European Conference, Proceedings
编辑Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, Tal Hassner
出版商Springer Science and Business Media Deutschland GmbH
691-707
页数17
ISBN(印刷版)9783031198380
DOI
出版状态已出版 - 2022
活动17th European Conference on Computer Vision, ECCV 2022 - Tel Aviv, 以色列
期限: 23 10月 202227 10月 2022

出版系列

姓名Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
13698 LNCS
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议17th European Conference on Computer Vision, ECCV 2022
国家/地区以色列
Tel Aviv
时期23/10/2227/10/22

学术指纹

探究 'Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection' 的科研主题。它们共同构成独一无二的学术指纹。

引用此