Deep face recognition in the wild
File(s)
Author(s)
Deng, Jiankang
Type
Thesis
Abstract
Face recognition is the most prominent bio-metric technique for identity authentication and is widely used in applications, such as access control, finance, law enforcement, and public security. Even though face biometrics is widely deployed, the current face recognition system is still limited to applications where strict control is imposed on the process of face image capture.
This thesis targets on develop unconstrained face recognition technology, which is robust to a range of appearance variations and degradation factors, based on recent advances in deep neural networks and sophisticated prior information conveyed by 3D face models.We develop a complete pipeline for deep face recognition in real-life, naturalistic conditions (in-the-wild), covering the main steps (e.g. face detection, face alignment and face feature embedding).
First, we innovatively propose RetinaFace, which unifies face box prediction, 2D facial landmark localisation and 3D vertices regression under one common target: point regression on the image plane. The proposed 3D mesh regression branch can be easily incorporated, without any optimisation difficulty, in parallel with the existing box and 2D landmark regression branches during joint training. RetinaFace can achieve state-of-the-art face detection while being efficient through single-shot inference.
Then, we propose stacked dense U-Nets for face alignment under large pose variations. Our method employs a novel scale aggregation network topology structure and a channel aggregation building block to improve the model's capacity without sacrificing the computational complexity and model size. Extensive experiments on many in-the-wild datasets, validate the robustness of the proposed full-pose facial landmark localisation method.
Through sampling the image using the fitted 3D face model, a facial UV can be created. Unfortunately, due to self-occlusion, such a UV map is always incomplete. We devise a meticulously designed architecture that combines local and global adversarial DCNNs to learn an identity-preserving facial UV completion model. We demonstrate that by attaching the completed UV to the fitted mesh and generating instances of arbitrary poses, we can increase pose variations for training deep face recognition models, and minimise pose discrepancy during testing, which lead to pose-invariant face recognition.
For large-scale face recognition, we propose an Additive Angular Margin Loss (ArcFace) to obtain highly discriminative features. The proposed ArcFace has a clear geometric interpretation due to its exact correspondence to geodesic distance on a hypersphere. Besides, ArcFace only adds negligible computational complexity during training and current servers can easily support millions of identities. ArcFace achieves state-of-the-art performance on many face recognition benchmarks including large-scale image and video datasets.
Based on the proposed efficient face localisation method (RetinaFace) and effective face feature embedding method (ArcFace), we achieve state-of-the-art performance in the global Face Recognition Vendor Test (FRVT). We have also released our solution to facilitate future research.
This thesis targets on develop unconstrained face recognition technology, which is robust to a range of appearance variations and degradation factors, based on recent advances in deep neural networks and sophisticated prior information conveyed by 3D face models.We develop a complete pipeline for deep face recognition in real-life, naturalistic conditions (in-the-wild), covering the main steps (e.g. face detection, face alignment and face feature embedding).
First, we innovatively propose RetinaFace, which unifies face box prediction, 2D facial landmark localisation and 3D vertices regression under one common target: point regression on the image plane. The proposed 3D mesh regression branch can be easily incorporated, without any optimisation difficulty, in parallel with the existing box and 2D landmark regression branches during joint training. RetinaFace can achieve state-of-the-art face detection while being efficient through single-shot inference.
Then, we propose stacked dense U-Nets for face alignment under large pose variations. Our method employs a novel scale aggregation network topology structure and a channel aggregation building block to improve the model's capacity without sacrificing the computational complexity and model size. Extensive experiments on many in-the-wild datasets, validate the robustness of the proposed full-pose facial landmark localisation method.
Through sampling the image using the fitted 3D face model, a facial UV can be created. Unfortunately, due to self-occlusion, such a UV map is always incomplete. We devise a meticulously designed architecture that combines local and global adversarial DCNNs to learn an identity-preserving facial UV completion model. We demonstrate that by attaching the completed UV to the fitted mesh and generating instances of arbitrary poses, we can increase pose variations for training deep face recognition models, and minimise pose discrepancy during testing, which lead to pose-invariant face recognition.
For large-scale face recognition, we propose an Additive Angular Margin Loss (ArcFace) to obtain highly discriminative features. The proposed ArcFace has a clear geometric interpretation due to its exact correspondence to geodesic distance on a hypersphere. Besides, ArcFace only adds negligible computational complexity during training and current servers can easily support millions of identities. ArcFace achieves state-of-the-art performance on many face recognition benchmarks including large-scale image and video datasets.
Based on the proposed efficient face localisation method (RetinaFace) and effective face feature embedding method (ArcFace), we achieve state-of-the-art performance in the global Face Recognition Vendor Test (FRVT). We have also released our solution to facilitate future research.
Version
Open Access
Date Issued
2020-08
Date Awarded
2021-05
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Zafeiriou, Stefanos
Sponsor
Engineering and Physical Sciences Research Council
Grant Number
EP/N007743/1
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
