Realistic speech-driven facial animation with GANs
File(s)Vougioukas2020_Article_RealisticSpeech-DrivenFacialAn.pdf (4.8 MB)
Published version
Author(s)
Vougioukas, Konstantinos
Petridis, Stavros
Pantic, Maja
Type
Journal Article
Abstract
Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals. The majority of work in this domain creates a mapping from audio features to visual features. This approach often requires post-processing using computer graphics techniques to produce realistic albeit subject dependent results. We present an end-to-end system that generates videos of a talking head, using only a still image of a person and an audio clip containing speech, without relying on handcrafted intermediate features. Our method generates videos which have (a) lip movements that are in sync with the audio and (b) natural facial expressions such as blinks and eyebrow movements. Our temporal GAN uses 3 discriminators focused on achieving detailed frames, audio-visual synchronization, and realistic expressions. We quantify the contribution of each component in our model using an ablation study and we provide insights into the latent representation of the model. The generated videos are evaluated based on sharpness, reconstruction quality, lip-reading accuracy, synchronization as well as their ability to generate natural blinks.
Date Issued
2019-10-13
Date Acceptance
2019-10-01
Citation
International Journal of Computer Vision, 2019, 128, pp.1398-1413
ISSN
0920-5691
Publisher
Springer Verlag
Start Page
1398
End Page
1413
Journal / Book Title
International Journal of Computer Vision
Volume
128
Copyright Statement
© The Author(s) 2019. This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.
Identifier
http://gateway.webofknowledge.com/gateway/Gateway.cgi?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000490090600001&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=1ba7043ffcc86c417c072aa74d649202
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Computer Science
Generative modelling
Face generation
Speech-driven animation
HEAD
Publication Status
Published
Date Publish Online
2019-10-13