摘要
Achieving robust performance in vision-language tasks requires strong multimodal alignment, where textual and visual data interact seamlessly. Existing frameworks often combine contrastive learning with image captioning to unify visual and textual representations. However, reliance on global representations and unidirectional information flow from images to text limits their ability to reconstruct visual content accurately from textual descriptions. To address this limitation, we propose BiMAC, a novel framework that enables bidirectional interactions between images and text at both global and local levels. BiMAC employs advanced components to simultaneously reconstruct visual content from textual cues and generate textual descriptions guided by visual features. By integrating a text-region alignment mechanism, BiMAC identifies and selects relevant image patches for precise cross-modal interaction, reducing information noise and enhancing mapping accuracy. BiMAC achieves state-of-the-art performance across diverse vision-language tasks, including image-text retrieval, captioning, and classification.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 22290-22298 |
| 页数 | 9 |
| 期刊 | Proceedings of the AAAI Conference on Artificial Intelligence |
| 卷 | 39 |
| 期 | 21 |
| DOI | |
| 出版状态 | 已出版 - 11 4月 2025 |
| 活动 | 39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025 - Philadelphia, 美国 期限: 25 2月 2025 → 4 3月 2025 |
学术指纹
探究 'BiMAC: Bidirectional Multimodal Alignment in Contrastive Learning' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver