跳到主要导航 跳到搜索 跳到主要内容

[Figure presented]MARS: Multimodal-Assisted Refined Semantic Alignment

  • Junjie Xu
  • , Xingjiao Wu*
  • , Zihao Zhang
  • , Shuwen Yang
  • , Tianlong Ma
  • , Daoguo Dong
  • , Liang He
  • *此作品的通讯作者
  • East China Normal University

科研成果: 期刊稿件文章同行评审

摘要

Audio-to-image generation (AIG) faces challenges in fine-grained semantic alignment, particularly semantic semantic misalignment, and loss of visual detail. To address these issues, we proposed MARS [Figure presented] (Multimodal-Assisted Refined Semantic alignment), a novel framework leveraging a Mamba-based audio encoder to manage the complexity of long audio sequences, coupled with a fine-grained multimodal alignment strategy using visual descriptions from multimodal large language models. We enhanced semantic coherence and aesthetic quality by fine-tuning an image generator using an image aesthetic perception generator. Furthermore, we validated MARS on VGGSound and VEGAS benchmarks, comprising 37,250 and 9,500 records, respectively. The results suggest that MARS significantly outperforms existing methods, achieving average improvements of 28.73% in semantic relevance and 127.35% in aesthetic scores compared with the best AIG generation baseline. In addition, cross-domain evaluations on the AudioCaps and Clotho datasets confirmed the robustness and generalization capability of MARS, with an average improvement of 73.9% on the V2A metric.

源语言英语
文章编号104292
期刊Information Processing and Management
63
1
DOI
出版状态已出版 - 1月 2026

学术指纹

探究 '[Figure presented]MARS: Multimodal-Assisted Refined Semantic Alignment' 的科研主题。它们共同构成独一无二的学术指纹。

引用此