TY - GEN
T1 - Multimodal DeepFake Detection via Audio-Visual Feature Fusion and Temporal Attention
AU - Li, Xiang
AU - Ma, Zeyu
AU - Zhang, Zhizhong
AU - Chen, Mingang
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - With the rapid development of AI-generated content, DeepFake videos have become increasingly realistic, posing significant challenges to the authenticity of information dissemination and the security of personal identity. As a result, detecting DeepFake videos has emerged as a critical task in the field of digital media forensics. Despite extensive research efforts, existing methods still face two major limitations: (i) insufficient capacity in modeling long-range temporal dependencies, which hinders the detection of complex dynamic inconsistencies; and (ii) reliance on a single modality, making it difficult to fully exploit the complementary information between audio and visual streams. To address these issues, we propose a novel multimodal DeepFake detection framework that jointly leverages audio and visual information for video detection. Specifically, we employ a visual Transformer with temporal self-attention to model dynamic inconsistencies across sampled video frames, and adopt an audio encoder to extract discriminative features from Mel-spectrogram representations. The extracted visual and audio features are then globally aggregated and fused through a lightweight classification head. Extensive experiments demonstrate that our approach effectively captures cross-modal forgery cues and achieves strong generalization capability across diverse and challenging Deep-Fake scenarios.
AB - With the rapid development of AI-generated content, DeepFake videos have become increasingly realistic, posing significant challenges to the authenticity of information dissemination and the security of personal identity. As a result, detecting DeepFake videos has emerged as a critical task in the field of digital media forensics. Despite extensive research efforts, existing methods still face two major limitations: (i) insufficient capacity in modeling long-range temporal dependencies, which hinders the detection of complex dynamic inconsistencies; and (ii) reliance on a single modality, making it difficult to fully exploit the complementary information between audio and visual streams. To address these issues, we propose a novel multimodal DeepFake detection framework that jointly leverages audio and visual information for video detection. Specifically, we employ a visual Transformer with temporal self-attention to model dynamic inconsistencies across sampled video frames, and adopt an audio encoder to extract discriminative features from Mel-spectrogram representations. The extracted visual and audio features are then globally aggregated and fused through a lightweight classification head. Extensive experiments demonstrate that our approach effectively captures cross-modal forgery cues and achieves strong generalization capability across diverse and challenging Deep-Fake scenarios.
KW - Audio-visual fusion
KW - DeepFake detection
KW - DeepFake video
KW - Multimodal learning
UR - https://www.scopus.com/pages/publications/105033147277
U2 - 10.1109/ICBAIE66852.2025.11326621
DO - 10.1109/ICBAIE66852.2025.11326621
M3 - 会议稿件
AN - SCOPUS:105033147277
T3 - 2025 6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering, ICBAIE 2025
SP - 390
EP - 393
BT - 2025 6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering, ICBAIE 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering, ICBAIE 2025
Y2 - 17 October 2025 through 19 October 2025
ER -