Mnemonic Descent Method: A recurrent process applied for end-to-end face alignment
File(s) trigeorgis2016mnemonic.pdf (5.41 MB)
Accepted version
Author(s)
Trigeorgis, G
Snape, P
Nicolaou, MA
Antonakos, E
Zafeiriou, S
Type
Conference Paper
Abstract
Cascaded regression has recently become the method of
choice for solving non-linear least squares problems such
as deformable image alignment. Given a sizeable training
set, cascaded regression learns a set of generic rules that
are sequentially applied to minimise the least squares problem.
Despite the success of cascaded regression for problems
such as face alignment and head pose estimation, there
are several shortcomings arising in the strategies proposed
thus far. Specifically, (a) the regressors are learnt independently,
(b) the descent directions may cancel one another
out and (c) handcrafted features (e.g., HoGs, SIFT etc.) are
mainly used to drive the cascade, which may be sub-optimal
for the task at hand. In this paper, we propose a combined
and jointly trained convolutional recurrent neural network
architecture that allows the training of an end-to-end to
system that attempts to alleviate the aforementioned drawbacks.
The recurrent module facilitates the joint optimisation
of the regressors by assuming the cascades form a nonlinear
dynamical system, in effect fully utilising the information
between all cascade levels by introducing a memory
unit that shares information across all levels. The convolutional
module allows the network to extract features that
are specialised for the task at hand and are experimentally
shown to outperform hand-crafted features. We show that
the application of the proposed architecture for the problem
of face alignment results in a strong improvement over the
current state-of-the-art.
choice for solving non-linear least squares problems such
as deformable image alignment. Given a sizeable training
set, cascaded regression learns a set of generic rules that
are sequentially applied to minimise the least squares problem.
Despite the success of cascaded regression for problems
such as face alignment and head pose estimation, there
are several shortcomings arising in the strategies proposed
thus far. Specifically, (a) the regressors are learnt independently,
(b) the descent directions may cancel one another
out and (c) handcrafted features (e.g., HoGs, SIFT etc.) are
mainly used to drive the cascade, which may be sub-optimal
for the task at hand. In this paper, we propose a combined
and jointly trained convolutional recurrent neural network
architecture that allows the training of an end-to-end to
system that attempts to alleviate the aforementioned drawbacks.
The recurrent module facilitates the joint optimisation
of the regressors by assuming the cascades form a nonlinear
dynamical system, in effect fully utilising the information
between all cascade levels by introducing a memory
unit that shares information across all levels. The convolutional
module allows the network to extract features that
are specialised for the task at hand and are experimentally
shown to outperform hand-crafted features. We show that
the application of the proposed architecture for the problem
of face alignment results in a strong improvement over the
current state-of-the-art.
Date Issued
2016-06-26
Date Acceptance
2016-03-02
Publisher
Computer Vision Foundation (CVF)
Copyright Statement
© the authors
Sponsor
Engineering & Physical Science Research Council (EPSRC)
Commission of the European Communities
Grant Number
EP/J017787/1
688520
Source
International Conference on Computer Vision and Pattern Recognition
Publication Status
Accepted
Start Date
2016-06-26
Finish Date
2016-07-01
Coverage Spatial
Las Vegas, Nevada, USA
