跳到主要导航 跳到搜索 跳到主要内容

VME-Transformer: Enhancing Visual Memory Encoding for Navigation in Interactive Environments

  • Jiwei Shen
  • , Pengjie Lou
  • , Liang Yuan
  • , Shujing Lyu*
  • , Yue Lu
  • *此作品的通讯作者
  • East China Normal University
  • Beijing University of Chemical Technology

科研成果: 期刊稿件文章同行评审

摘要

The efficiency of a robotic system is primarily determined by its ability to navigate complex and interactive environments. In real-world scenarios, cluttered surroundings are common, requiring a robot to navigate diverse spaces and displace objects to pave a path towards its objective. Consequently, 'Visual Interactive Navigation' presents several challenges, including how to retain historical exploration information from partially observable visual signals, and how to utilize sparse rewards in reinforcement learning to simultaneously learn a latent representation and a control policy. Addressing these challenges, we introduce a Transformer-based Visual Memory Encoder (VME-Transformer), capable of embedding both recent and long-term exploration information into memory. Additionally, we explicitly estimate the robot's next pose, conditioned on the impending action, to bootstrap the learning process of the high-capacity VME-Transformer. We further regularize the value function by introducing input perturbations, thereby enhancing its generalization capabilities in previously unseen environments. In the Visual Interactive Navigation tasks within the iGibson environment, the VME-Transformer demonstrates superior performance compared to state-of-the-art methods, underlining its effectiveness.

源语言英语
页(从-至)643-650
页数8
期刊IEEE Robotics and Automation Letters
9
1
DOI
出版状态已出版 - 1 1月 2024

指纹

探究 'VME-Transformer: Enhancing Visual Memory Encoding for Navigation in Interactive Environments' 的科研主题。它们共同构成独一无二的指纹。

引用此