Machine learning methods for facial attribute editing
File(s)
Author(s)
Ververas, Evangelos
Type
Thesis
Abstract
In this thesis, we explore the problem of facial attribute analysis and editing in images, from the scope of two major fields of Machine Learning, those of Components Analysis (CA) and Deep Learning (DL).
First, we present a CA method for analysing and editing facial data.
Then, we present a DL algorithm for animating facial images according to expressions and speech.
Finally, we present a method for improving gaze estimation generalisation to unseen image domains and showcase applications to eye gaze editing.
Although CA methods are able to capture only linear relationships in data, they can still be useful with well-aligned data, such as UV maps of facial texture.
In this Thesis, we propose robust extensions of the Joint and Individual Variance Explained (JIVE) method, for the recovery of joint and individual components in visual facial data, captured in unconstrained conditions and possibly containing sparse non-Gaussian errors and missing data.
We demonstrate the effectiveness of the proposed methods to several computer vision applications, namely facial expression synthesis and 2D and 3D face age progression in-the-wild.
CA methods usually fall short in image generation, as they fail to generate details.
On the contrary, Image-to-image (i2i) translation, which is the problem of translating images between image domains, has recently seen remarkable progress since the advent of DL and Generative Adversarial Networks (GANs).
In this Thesis, we study the problem of i2i translation, under a set of continuous parameters that correspond to statistical blendshape models of facial motion.
We show that it is possible to edit facial images according to expression and speech blendshapes using ``sliders'', which are more flexible than discrete expressions or action units.
Lastly, realistically animating gaze is crucial for achieving high quality facial animations.
To this end, large datasets of faces with gaze annotations are required for training.
In this Thesis, we present a weakly-supervised method for improving gaze estimation generalization to unseen domains, by harnessing arbitrary unlabelled ``in-the-wild'' face images.
Unlike previous methods, we tackle gaze estimation as end-to-end, dense 3D reconstruction of eyes and experimentally validate the benefits of this choice.
Particularly, we show improvements in semi-supervised and cross-dataset gaze estimation.
Finally, we showcase how our methods can be employed for training efficient models for gaze editing.
First, we present a CA method for analysing and editing facial data.
Then, we present a DL algorithm for animating facial images according to expressions and speech.
Finally, we present a method for improving gaze estimation generalisation to unseen image domains and showcase applications to eye gaze editing.
Although CA methods are able to capture only linear relationships in data, they can still be useful with well-aligned data, such as UV maps of facial texture.
In this Thesis, we propose robust extensions of the Joint and Individual Variance Explained (JIVE) method, for the recovery of joint and individual components in visual facial data, captured in unconstrained conditions and possibly containing sparse non-Gaussian errors and missing data.
We demonstrate the effectiveness of the proposed methods to several computer vision applications, namely facial expression synthesis and 2D and 3D face age progression in-the-wild.
CA methods usually fall short in image generation, as they fail to generate details.
On the contrary, Image-to-image (i2i) translation, which is the problem of translating images between image domains, has recently seen remarkable progress since the advent of DL and Generative Adversarial Networks (GANs).
In this Thesis, we study the problem of i2i translation, under a set of continuous parameters that correspond to statistical blendshape models of facial motion.
We show that it is possible to edit facial images according to expression and speech blendshapes using ``sliders'', which are more flexible than discrete expressions or action units.
Lastly, realistically animating gaze is crucial for achieving high quality facial animations.
To this end, large datasets of faces with gaze annotations are required for training.
In this Thesis, we present a weakly-supervised method for improving gaze estimation generalization to unseen domains, by harnessing arbitrary unlabelled ``in-the-wild'' face images.
Unlike previous methods, we tackle gaze estimation as end-to-end, dense 3D reconstruction of eyes and experimentally validate the benefits of this choice.
Particularly, we show improvements in semi-supervised and cross-dataset gaze estimation.
Finally, we showcase how our methods can be employed for training efficient models for gaze editing.
Version
Open Access
Date Issued
2022-04
Date Awarded
2023-08
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Zafeiriou, Stefanos
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
