Evaluating multimodal fusion strategies for audio–visual deepfake detection
File(s) 671-WorldS4 2026 (1).pdf (203.57 KB)
Accepted version
Author(s)
Loh, Caitlin
Vekkot, Susmitha
Shukla, Pancham
Type
Conference Paper
Abstract
Deepfakes generated using modern machine learning techniques pose growing risks to digital trust by enabling realistic manipulation of both audio and video content. Many existing detection approaches rely on a single modality, limiting robustness when confronted with increasingly sophisticated forgeries. This paper evaluates a multi-modal deepfake detection framework that integrates audio and visual information using deep learning. The proposed system combines state-of-the-art audio and visual encoders within a modular architecture and conducts a systematic comparison of three fusion strategies: early fusion, late fusion, and cross-attention. Experiments are conducted on the PolyGlotFake dataset, a multilingual benchmark containing synthetic and authentic audio–visual media. Results show that multimodal approaches substantially outperform unimodal baselines, with late fusion achieving an AUROC of 0.955 and cross-attention models reaching accuracies of up to 0.996. These findings provide a controlled comparison of fusion strategies and demonstrate that multimodal fusion significantly improves detection performance and highlights its potential for building more robust deepfake detection systems.
Date Acceptance
2026-06-30
Citation
Lecture Notes in Networks and Systems
ISSN
2367-3370
Publisher
Springer
Journal / Book Title
Lecture Notes in Networks and Systems
Copyright Statement
Subject to copyright. This paper is embargoed until publication. Once published the author’s accepted manuscript will be made available under a CC-BY License in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy).
License URL
Source
WorldS4 2026
Publication Status
Accepted
Start Date
2026-07-28
Finish Date
2026-07-30
Coverage Spatial
London, UK
