Dynamic face parsing in the wild
File(s)
Author(s)
Wang, Yujiang
Type
Thesis
Abstract
Landmark-based facial descriptors are widely utilised in face video analysis to obtain useful facial dynamics. Such a sparse facial representation is not capable of constructing the full dynamics of each facial component like eyes and mouths, while such dynamics can be essential to the recognition of higher-level features like facial expressions, emotions, identity, and so on. A dense facial descriptor such as the segmentation masks generated by face parsing, however, can effectively overcome those limitations by providing pixel-wise semantic information that is generally more discriminative and more desirable to facial analysis tasks. Recently, Deep Convolutional Neural Networks (DCNNs) have made impressive progress in semantic image segmentation, a task that performs per-pixel classifications. Those deep segmentation models can naturally generate pixel-level predictions for facial images, however, face parsing in the wild is still a challenging task. The model's ability to accurately segment different facial regions is crucial to generate high-quality face masks. Besides, how to adapt the segmentation models designed for static images to the continuous environment of face videos also requires consideration. To satisfy the real-time requirement under realistic scenarios, the acceleration problem needs to be resolved. This thesis investigates different aspects of in-the-wild face parsing and proposes several novel approaches of constructing robust face segmentation masks. To increase the robustness of eye segmentation against low-quality video scenarios, we encode the shape priors of eyes into the training procedure of deep segmentation model. Additionally, the segmentation model's sensitivity to semantic facial contours is enhanced by introducing the Dilated Convolutions with Lateral Inhibitions, which is a convolutional operator biologically inspired by human visual systems. To exploit information from both temporal and spatial domains in face videos, we propose a ConvLSTM-FCN model to generate temporal-smoothed face segmentation masks which are more tolerant to video variations. Eventually, we consider to accelerate the process of dynamic face parsing via Reinforcement Learning to learn a globally-optimised key scheduler. This thesis contributes towards in-the-wild face parsing from different aspects such as improving the fundamental network architectures and optimising the performance under realistic scenarios. It can benefit downstream tasks that require detailed facial dynamics such as facial expression recognition and lip-reading. It can also be inspiring to future works on semantic image/video segmentation, and to other works pursuing face parsing with higher visual qualities or with better working efficiency.
Version
Open Access
Date Issued
2020-09
Date Awarded
2021-02
Copyright Statement
Creative Commons Attribution-NonCommercial 4.0 International Licence
License URL
Advisor
Pantic, Maja
Sponsor
China Scholarship Council
Engineering and Physical Sciences Research Council
Grant Number
201708060212; EP/N007743/1
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)