TY - JOUR
T1 - VAMF
T2 - Variance-guided attention modulation framework for infrared and visible image fusion
AU - Mustafa, Hafiz Tayyab
AU - Asad, Mujtaba
AU - Jiang, He
AU - Zheng, Zhonglong
AU - Shamsolmoali, Pourya
AU - Yang, Jie
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/9/25
Y1 - 2026/9/25
N2 - Infrared and visible image fusion aims to generate a single composite image that preserves the thermal radiation of targets from the infrared modality and the rich texture details from the visible modality. However, current methods often struggle to balance this exchange adaptively. They typically rely either on computationally intensive attention mechanisms to capture global context or on standard convolutions that fail to model long-range dependencies. Crucially, they rarely provide a principled mechanism for modality-specific feature discrimination. Consequently, important thermal targets and fine textures are frequently over-smoothed, resulting in inconsistent emphasis across modalities. This paper presents a novel variance-guided attention modulation framework for infrared and visible image fusion that explicitly decouples global context modeling from local detail extraction while prioritizing modality-specific feature discrimination. The proposed architecture introduces a shared shallow feature extractor followed by multiple adaptive feature modulation modules (AFMMs), which integrate variance-guided global context aggregation with local feature enhancement to jointly capture global contextual information and fine spatial details in a modality-aware manner. Each AFMM incorporates a multiscale contextual feature aggregation (MCFA) block and a gated feature refinement network (GFRN). In particular, the MCFA block leverages a variance-guided attention modulation block to perform variance-based global context modeling, effectively distinguishing and enhancing thermal radiation and texture details according to their unique statistical characteristics. The fusion module further enables effective information exchange through a two-stage process: an efficient channel attention mechanism refines modality-specific channels, followed by a cross-modal feature interaction module that facilitates bidirectional cross-modal guidance, allowing features from one modality to adaptively enhance the other. Comprehensive experiments on benchmark datasets demonstrate that our approach outperforms existing and recent advanced fusion methods.
AB - Infrared and visible image fusion aims to generate a single composite image that preserves the thermal radiation of targets from the infrared modality and the rich texture details from the visible modality. However, current methods often struggle to balance this exchange adaptively. They typically rely either on computationally intensive attention mechanisms to capture global context or on standard convolutions that fail to model long-range dependencies. Crucially, they rarely provide a principled mechanism for modality-specific feature discrimination. Consequently, important thermal targets and fine textures are frequently over-smoothed, resulting in inconsistent emphasis across modalities. This paper presents a novel variance-guided attention modulation framework for infrared and visible image fusion that explicitly decouples global context modeling from local detail extraction while prioritizing modality-specific feature discrimination. The proposed architecture introduces a shared shallow feature extractor followed by multiple adaptive feature modulation modules (AFMMs), which integrate variance-guided global context aggregation with local feature enhancement to jointly capture global contextual information and fine spatial details in a modality-aware manner. Each AFMM incorporates a multiscale contextual feature aggregation (MCFA) block and a gated feature refinement network (GFRN). In particular, the MCFA block leverages a variance-guided attention modulation block to perform variance-based global context modeling, effectively distinguishing and enhancing thermal radiation and texture details according to their unique statistical characteristics. The fusion module further enables effective information exchange through a two-stage process: an efficient channel attention mechanism refines modality-specific channels, followed by a cross-modal feature interaction module that facilitates bidirectional cross-modal guidance, allowing features from one modality to adaptively enhance the other. Comprehensive experiments on benchmark datasets demonstrate that our approach outperforms existing and recent advanced fusion methods.
KW - Channel attention
KW - Cross-modal
KW - Deep learning
KW - Image fusion
UR - https://www.scopus.com/pages/publications/105038779372
U2 - 10.1016/j.eswa.2026.132653
DO - 10.1016/j.eswa.2026.132653
M3 - 文章
AN - SCOPUS:105038779372
SN - 0957-4174
VL - 327
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 132653
ER -