Towards expressive audio-driven facial animation
File(s)
Author(s)
Bigata Casademunt, Antoni
Type
Thesis
Abstract
The way we communicate with our devices is constantly evolving, leading to the rapid development of novel Human-Computer Interaction (HCI) modalities. One such modality is audio-driven facial animation, where a static image is brought to life by animating it to match a person's speech. However, current methods are often limited in their ability to generate high-quality, realistic animations that capture the full nuance of human expression. This thesis highlights the need for more natural animation by integrating Non-Speech Vocalizations (NSVs) and emotions into the generation process, an aspect overlooked by most state-of-the-art methods. This work initially demonstrates that even subtle cues like head motion must match the speaker's emotional state to be perceived as natural. The thesis then introduces a model that, for the first time, can generate realistic laughter sequences from audio, underscoring the importance of NSVs in facial animation. This research is subsequently extended to create a unified framework that simultaneously handles speech, a wide range of NSVs, and continuous emotional expressions. A core contribution is a new pipeline that allows for the generation of high-quality, temporally consistent videos of arbitrary length. It is the first model to successfully integrate speech and NSVs, producing realistic animations that accurately reflect the subtleties of human communication. Finally, this framework is adapted to address the specific task of lip synchronisation (lip-sync), where only the mouth region of an existing video is modified to match new audio. This task introduces a unique set of challenges, as the model must avoid being influenced by the original video's expressions and robustly handle any facial occlusions that may arise. Throughout this work, a strong emphasis is placed on developing a suite of new evaluation metrics alongside these new tasks, providing valuable tools to guide future research in the field.
Version
Open Access
Date Issued
2025-07-18
Date Awarded
01/01/2026
License URL
Advisor
Pantic, Maja
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
