Skip to main navigation Skip to search Skip to main content

MAFFUNet: A Multimodal Attention and Feature Fusion Framework for Remote Sensing Image Integration

  • Songjia Zhang
  • , Yi Lin*
  • , Qingli Li
  • , Xiaonan Yang
  • *Corresponding author for this work
  • Tongji University
  • East China Normal University
  • Shanghai Information Technology Research Institute

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal image fusion in remote sensing integrates hyperspectral and multispectral data to enhance applications, such as land cover classification and environmental monitoring. Challenges, including modality heterogeneity, spectral degradation, and cross-domain variability, require advanced fusion algorithms to capture nonlinear correlations and ensure robust generalization. We propose MAFFUNet, a lightweight U-Net-based framework featuring a five-layer encoder-decoder architecture with depthwise separable convolutions for computational efficiency and support for temporal data through flattening. MAFFUNet incorporates three novel modules: adversarial feature fusion to maximize mutual information, hybrid attention and multiscale fusion to optimize spectral and spatial features, and cross-domain regularization to enhance adaptability across domains. Evaluations on the Chikusei, PaviaC, and PaviaU datasets at 4×, 8×, and 16× scaling factors demonstrate MAFFUNet's superior performance over traditional and deep learning methods in metrics, such as root mean square error, peak signal-to-noise ratio, erreur relative globale adimensionnelle de synthèse, and spectral angle mapper, achieving excellent spectral fidelity and spatial detail preservation. Its low computational complexity makes it ideal for resource-constrained environments, offering an efficient and robust solution for hyperspectral-multispectral image fusion.

Original languageEnglish
Pages (from-to)20563-20573
Number of pages11
JournalIEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
Volume19
DOIs
StatePublished - 2026

Keywords

  • Attention mechanism
  • cross-domain generalization
  • deep learning
  • multimodal fusion
  • remote sensing image

Fingerprint

Dive into the research topics of 'MAFFUNet: A Multimodal Attention and Feature Fusion Framework for Remote Sensing Image Integration'. Together they form a unique fingerprint.

Cite this