End-to-end visual speech recognition with LSTMS
File(s)petridislipantic_icassp2017.pdf (335.85 KB)
Accepted version
Author(s)
Petridis, S
Li, Z
Pantic, M
Type
Conference Paper
Abstract
Traditional visual speech recognition systems consist of two stages, feature extraction and classification. Recently, several deep learning approaches have been presented which automatically extract features from the mouth images and aim to replace the feature extraction stage. However, research on joint learning of features and classification is very limited. In this work, we present an end-to-end visual speech recognition system based on Long-Short Memory (LSTM) networks. To the best of our knowledge, this is the first model which simultaneously learns to extract features directly from the pixels and perform classification and also achieves state-of-the-art performance in visual speech classification. The model consists of two streams which extract features directly from the mouth and difference images, respectively. The temporal dynamics in each stream are modelled by an LSTM and the fusion of the two streams takes place via a Bidirectional LSTM (BLSTM). An absolute improvement of 9.7% over the base line is reported on the OuluVS2 database, and 1.5% on the CUAVE database when compared with other methods which use a similar visual front-end.
Date Issued
2017-06-19
Date Acceptance
2017-03-05
Citation
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2017, pp.2592-2596
ISBN
9781509041176
ISSN
1520-6149
Publisher
IEEE
Start Page
2592
End Page
2596
Journal / Book Title
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
Copyright Statement
© 2017 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Sponsor
Commission of the European Communities
Commission of the European Communities
Grant Number
645094
688835
Source
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Publication Status
Published
Start Date
2017-03-05
Finish Date
2017-03-09