Learning from talking faces: from deepfake detection to speech recognition
File(s)
Author(s)
Haliassos, Alex
Type
Thesis
Abstract
Humans are highly sensitive to facial movements and expressions, as well as vocal subtleties, all of which play a critical role in human interactions. Analogously, the advent of deep learning has enabled machines to adeptly extract and interpret information from talking faces, offering potential for numerous applications. This thesis explores how large datasets of speaking faces can be used to learn deep representations related to facial behaviour, aiming to enhance the performance of two specific tasks: DeepFake detection and audio-visual speech recognition. We begin by proposing a DeepFake detector, dubbed LipForensics, that performs well on samples generated by unseen forgery types - a central goal in forgery detection literature - while at the same time withstanding low-level corruptions, which other approaches struggle with. LipForensics transfers the knowledge gained by a lipreading model to force the detector to ascertain whether a given video is fake or not by focusing on high-level information related to anomalous mouth movements, which are ubiquitous in fake videos. In our second main contribution, we address some of the limitations of LipForensics. Most notably, we remove the need for supervised pre-training through a cross-modal, self-supervised method that exploits the correspondence between the visual and auditory modalities to target inconsistencies related to lexical content, emotion, and identity. Next, turning our focus to speech recognition, we propose RAVEn, a self-supervised method that uses a combination of masked prediction and cross- and within-modal losses to learn visual and auditory speech representations entirely from raw data. Finally, we make several modifications to RAVEn to address the asymmetries between the auditory and visual modalities, leading to substantial performance improvements and state-of-the-art results among similar approaches in various challenging settings.
Version
Open Access
Date Issued
2024-04-11
Date Awarded
01/01/2025
License URL
Advisor
Pantic, Maja
Sponsor
Imperial College London
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
