跳到主要导航 跳到搜索 跳到主要内容

VAMF: Variance-guided attention modulation framework for infrared and visible image fusion

  • Hafiz Tayyab Mustafa
  • , Mujtaba Asad
  • , He Jiang
  • , Zhonglong Zheng*
  • , Pourya Shamsolmoali
  • , Jie Yang*
  • *此作品的通讯作者
  • Zhejiang Normal University
  • Zhejiang Institute of Photoelectronics
  • Shanghai Jiao Tong University
  • China University of Mining and Technology

科研成果: 期刊稿件文章同行评审

摘要

Infrared and visible image fusion aims to generate a single composite image that preserves the thermal radiation of targets from the infrared modality and the rich texture details from the visible modality. However, current methods often struggle to balance this exchange adaptively. They typically rely either on computationally intensive attention mechanisms to capture global context or on standard convolutions that fail to model long-range dependencies. Crucially, they rarely provide a principled mechanism for modality-specific feature discrimination. Consequently, important thermal targets and fine textures are frequently over-smoothed, resulting in inconsistent emphasis across modalities. This paper presents a novel variance-guided attention modulation framework for infrared and visible image fusion that explicitly decouples global context modeling from local detail extraction while prioritizing modality-specific feature discrimination. The proposed architecture introduces a shared shallow feature extractor followed by multiple adaptive feature modulation modules (AFMMs), which integrate variance-guided global context aggregation with local feature enhancement to jointly capture global contextual information and fine spatial details in a modality-aware manner. Each AFMM incorporates a multiscale contextual feature aggregation (MCFA) block and a gated feature refinement network (GFRN). In particular, the MCFA block leverages a variance-guided attention modulation block to perform variance-based global context modeling, effectively distinguishing and enhancing thermal radiation and texture details according to their unique statistical characteristics. The fusion module further enables effective information exchange through a two-stage process: an efficient channel attention mechanism refines modality-specific channels, followed by a cross-modal feature interaction module that facilitates bidirectional cross-modal guidance, allowing features from one modality to adaptively enhance the other. Comprehensive experiments on benchmark datasets demonstrate that our approach outperforms existing and recent advanced fusion methods.

源语言英语
文章编号132653
期刊Expert Systems with Applications
327
DOI
出版状态已出版 - 25 9月 2026

指纹

探究 'VAMF: Variance-guided attention modulation framework for infrared and visible image fusion' 的科研主题。它们共同构成独一无二的指纹。

引用此