Speech-driven facial animation using polynomial fusion of features
File(s)1912.05833v1.pdf (500.21 KB)
Working paper
Author(s)
Type
Working Paper
Abstract
Speech-driven facial animation involves using a speech signal to generate
realistic videos of talking faces. Recent deep learning approaches to facial
synthesis rely on extracting low-dimensional representations and concatenating
them, followed by a decoding step of the concatenated vector. This accounts for
only first-order interactions of the features and ignores higher-order
interactions. In this paper we propose a polynomial fusion layer that models
the joint representation of the encodings by a higher-order polynomial, with
the parameters modelled by a tensor decomposition. We demonstrate the the
suitability of this approach through experiments on generated videos evaluated
on a range of metrics on video quality, audiovisual synchronisation and
generation of blinks.
realistic videos of talking faces. Recent deep learning approaches to facial
synthesis rely on extracting low-dimensional representations and concatenating
them, followed by a decoding step of the concatenated vector. This accounts for
only first-order interactions of the features and ignores higher-order
interactions. In this paper we propose a polynomial fusion layer that models
the joint representation of the encodings by a higher-order polynomial, with
the parameters modelled by a tensor decomposition. We demonstrate the the
suitability of this approach through experiments on generated videos evaluated
on a range of metrics on video quality, audiovisual synchronisation and
generation of blinks.
Date Issued
2020-02-19
Citation
2020
Publisher
arXiv
Copyright Statement
© 2020 The Author(s)
Identifier
http://arxiv.org/abs/1912.05833v1
Subjects
cs.LG
cs.LG
eess.AS
stat.ML
Publication Status
Published