跳到主要导航 跳到搜索 跳到主要内容

Multimodal DeepFake Detection via Audio-Visual Feature Fusion and Temporal Attention

  • Xiang Li
  • , Zeyu Ma
  • , Zhizhong Zhang
  • , Mingang Chen*
  • *此作品的通讯作者
  • Shanghai Second Polytechnic University
  • Shanghai Key Laboratory of Computer Software Evaluating and Testing
  • Shanghai Development Center of Computer Software Technology
  • Shanghai Normal University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

With the rapid development of AI-generated content, DeepFake videos have become increasingly realistic, posing significant challenges to the authenticity of information dissemination and the security of personal identity. As a result, detecting DeepFake videos has emerged as a critical task in the field of digital media forensics. Despite extensive research efforts, existing methods still face two major limitations: (i) insufficient capacity in modeling long-range temporal dependencies, which hinders the detection of complex dynamic inconsistencies; and (ii) reliance on a single modality, making it difficult to fully exploit the complementary information between audio and visual streams. To address these issues, we propose a novel multimodal DeepFake detection framework that jointly leverages audio and visual information for video detection. Specifically, we employ a visual Transformer with temporal self-attention to model dynamic inconsistencies across sampled video frames, and adopt an audio encoder to extract discriminative features from Mel-spectrogram representations. The extracted visual and audio features are then globally aggregated and fused through a lightweight classification head. Extensive experiments demonstrate that our approach effectively captures cross-modal forgery cues and achieves strong generalization capability across diverse and challenging Deep-Fake scenarios.

源语言英语
主期刊名2025 6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering, ICBAIE 2025
出版商Institute of Electrical and Electronics Engineers Inc.
390-393
页数4
ISBN(电子版)9798331538873
DOI
出版状态已出版 - 2025
活动6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering, ICBAIE 2025 - Shanghai, 中国
期限: 17 10月 202519 10月 2025

出版系列

姓名2025 6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering, ICBAIE 2025

会议

会议6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering, ICBAIE 2025
国家/地区中国
Shanghai
时期17/10/2519/10/25

指纹

探究 'Multimodal DeepFake Detection via Audio-Visual Feature Fusion and Temporal Attention' 的科研主题。它们共同构成独一无二的指纹。

引用此