跳到主要导航 跳到搜索 跳到主要内容

Zero-shot visual grounding via coarse-to-fine representation learning

  • Jinpeng Mi*
  • , Shaofei Jin
  • , Zhiqian Chen
  • , Dan Liu
  • , Xian Wei
  • , Jianwei Zhang
  • *此作品的通讯作者
  • University of Shanghai for Science and Technology
  • University of Hamburg

科研成果: 期刊稿件文章同行评审

摘要

Visual grounding (VG) locates target objects in visual scenes by understanding given natural language queries. Current methods for VG mainly focus on grounding referring expressions or noun phrases covered in the labeled training samples. Despite their grounding prowess, these approaches struggle in grounding novel query–image pairs excluded from the training data. This shortage is usually caused by the deficiency of discriminative representation learning both in images and queries. To address these issues, we propose a one-stage coarse-to-fine framework for zero-shot VG to ground novel query-image samples. Specifically, in the coarse stage, we mine the global context information in the visual features and query embeddings by employing a multi-head self-attention block, strengthening the intra-modality relations in the visual and textual features. In the fine stage, we first learn the query-aware visual representations based on the acquired global context information via a multi-modal relation-enhanced transformer block, which explores the informative information by modeling the cross-modal interaction among visual and textual domains. We further excavate target-oriented discriminative representations from the acquired query-aware visual representations by a noun phrase-guided multi-modal interaction network, which augments the interaction between target-related phrases and the obtained query-aware visual representations to enhance the distinction of target regions, enhancing the subsequent referred target grounding and generalizing. In order to validate the proposed approach, we implement extensive experiments and ablation studies on public benchmark datasets, including RefCOCO, RefCOCO+, RefCOCOg, Flickr30K Entity, Flickr-Split-0 and Flickr-Split-1. Experimental results demonstrate that our approach substantially improves the grounding accuracy and achieves new state-of-the-art performance under single-stage training and testing.

源语言英语
文章编号128621
期刊Neurocomputing
610
DOI
出版状态已出版 - 28 12月 2024

学术指纹

探究 'Zero-shot visual grounding via coarse-to-fine representation learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此