跳到主要导航 跳到搜索 跳到主要内容

3D Dense Captioning via Prototypical Momentum Distillation

  • Jinpeng Mi
  • , Ying Wang
  • , Shaofei Jin
  • , Shiming Zhang
  • , Xian Wei
  • , Jianwei Zhang
  • University of Shanghai for Science and Technology
  • University of Hamburg

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

3D dense captioning aims to describe the crucial regions in 3D visual scenes in the form of natural language. Recent prevailing approaches achieve promising results by leveraging complicated structures incorporated with large-scale models, which necessitate abundant parameters and pose challenges regarding its practical applications. Besides, with limited training data, 3D dense captioners are often susceptible to overfitting, directly degrading caption generation performance. Drawing inspiration from the recent advancements in knowledge distillation, we propose a novel approach termed Prototypical Momentum Distillation (PMD) to prompt the model to generate more detailed captions. PMD incorporates Momentum Distillation (MD) with an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to transfer knowledge by considering the uncertainty of the teacher knowledge. Specifically, we employ the original captioner as the student model and maintain an Exponential Moving Average (EMA) copy of the captioner as the teacher model to impart knowledge as the auxiliary supervision of the student. To abate the misleading caused by uncertain knowledge, we present an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to cluster the distilled knowledge according to its confidence. We then transfer the rearranged knowledge from the teacher to guide the training route of the student. We conduct extensive experiments and ablation studies on two widely used benchmark datasets, ScanRefer and Nr3D. Experimental results demonstrate that PMD outperforms all state-of-the-art approaches on the benchmarks with MLE training, highlighting its effectiveness.

源语言英语
主期刊名2025 IEEE International Conference on Robotics and Automation, ICRA 2025
编辑Christian Ott, Henny Admoni, Sven Behnke, Stjepan Bogdan, Aude Bolopion, Youngjin Choi, Fanny Ficuciello, Nicholas Gans, Clement Gosselin, Kensuke Harada, Erdal Kayacan, H. Jin Kim, Stefan Leutenegger, Zhe Liu, Perla Maiolino, Lino Marques, Takamitsu Matsubara, Anastasia Mavromatti, Mark Minor, Jason O'Kane, Hae Won Park, Hae-Won Park, Ioannis Rekleitis, Federico Renda, Elisa Ricci, Laurel D. Riek, Lorenzo Sabattini, Shaojie Shen, Yu Sun, Pierre-Brice Wieber, Katsu Yamane, Jingjin Yu
出版商Institute of Electrical and Electronics Engineers Inc.
1444-1450
页数7
ISBN(电子版)9798331541392
DOI
出版状态已出版 - 2025
活动2025 IEEE International Conference on Robotics and Automation, ICRA 2025 - Atlanta, 美国
期限: 19 5月 202523 5月 2025

丛书

姓名Proceedings - IEEE International Conference on Robotics and Automation
ISSN(印刷版)1050-4729

会议

会议2025 IEEE International Conference on Robotics and Automation, ICRA 2025
国家/地区美国
Atlanta
时期19/05/2523/05/25

学术指纹

探究 '3D Dense Captioning via Prototypical Momentum Distillation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此