跳到主要导航 跳到搜索 跳到主要内容

BiMAC: Bidirectional Multimodal Alignment in Contrastive Learning

  • Shanghai Jiao Tong University
  • University of York

科研成果: 期刊稿件会议文章同行评审

摘要

Achieving robust performance in vision-language tasks requires strong multimodal alignment, where textual and visual data interact seamlessly. Existing frameworks often combine contrastive learning with image captioning to unify visual and textual representations. However, reliance on global representations and unidirectional information flow from images to text limits their ability to reconstruct visual content accurately from textual descriptions. To address this limitation, we propose BiMAC, a novel framework that enables bidirectional interactions between images and text at both global and local levels. BiMAC employs advanced components to simultaneously reconstruct visual content from textual cues and generate textual descriptions guided by visual features. By integrating a text-region alignment mechanism, BiMAC identifies and selects relevant image patches for precise cross-modal interaction, reducing information noise and enhancing mapping accuracy. BiMAC achieves state-of-the-art performance across diverse vision-language tasks, including image-text retrieval, captioning, and classification.

源语言英语
页(从-至)22290-22298
页数9
期刊Proceedings of the AAAI Conference on Artificial Intelligence
39
21
DOI
出版状态已出版 - 11 4月 2025
活动39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025 - Philadelphia, 美国
期限: 25 2月 20254 3月 2025

学术指纹

探究 'BiMAC: Bidirectional Multimodal Alignment in Contrastive Learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此