Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network
File(s)
Author(s)
Type
Conference Paper
Abstract
The automatic recognition of spontaneous emotions from speech is a challenging task. On the one hand, acoustic features need to be robust enough to capture the emotional content for various styles of speaking, and while on the other, machine learning algorithms need to be insensitive to outliers while being able to model the context. Whereas the latter has been tackled by the use of Long Short-Term Memory (LSTM) networks, the former is still under very active investigations, even though more than a decade of research has provided a large set of acoustic descriptors. In this paper, we propose a solution to the problem of `context-aware' emotional relevant feature extraction, by combining Convolutional Neural Networks (CNNs) with LSTM networks, in order to automatically learn the best representation of the speech signal directly from the raw time representation. In this novel work on the so-called end-to-end speech emotion recognition, we show that the use of the proposed topology significantly outperforms the traditional approaches based on signal processing techniques for the prediction of spontaneous and natural emotions on the RECOLA database.
Date Issued
2016-05-19
Date Acceptance
2015-12-21
Citation
41st IEEE International Conference on Acoustics, Speech, and Signal Processing, 2016
Publisher
IEEE
Journal / Book Title
41st IEEE International Conference on Acoustics, Speech, and Signal Processing
Copyright Statement
This article is under embargo until publication
Sponsor
Commission of the European Communities
Identifier
https://ieeexplore.ieee.org/document/7472669
Grant Number
645378
Source
41st IEEE International Conference on Acoustics, Speech, and Signal Processing
Subjects
Science & Technology
Technology
Acoustics
Engineering, Electrical & Electronic
Engineering
end-to-end learning
raw waveform
emotion recognition
deep learning
CNN
LSTM
NEURAL-NETWORKS
Publication Status
Published
Start Date
2016-03-20
Coverage Spatial
Shanghai, China
Date Publish Online
2016-05-19
