Speaker-independent emotion recognition exploiting a psychologically-inspired binary cascade classification schema
File(s)IJSP_2012_Margarita_Kotti.pdf (369.66 KB)
Accepted version
Author(s)
Kotti, Margarita
Paternò, Fabio
Type
Journal Article
Abstract
In this paper, a psychologically-inspired binary cascade classification schema is proposed for speech emotion recognition. Performance is enhanced because commonly confused pairs of emotions are distinguishable from one another. Extracted features are related to statistics of pitch, formants, and energy contours, as well as spectrum, cepstrum, perceptual and temporal features, autocorrelation, MPEG-7 descriptors, Fujisakis model parameters, voice quality, jitter, and shimmer. Selected features are fed as input to K nearest neighborhood classifier and to support vector machines. Two kernels are tested for the latter: Linear and Gaussian radial basis function. The recently proposed speaker-independent experimental protocol is tested on the Berlin emotional speech database for each gender separately. The best emotion recognition accuracy, achieved by support vector machines with linear kernel, equals 87.7%, outperforming state-of-the-art approaches. Statistical analysis is first carried out with respect to the classifiers error rates and then to evaluate the information expressed by the classifiers confusion matrices. © Springer Science+Business Media, LLC 2011.
Date Issued
2012-06
Citation
International Journal of Speech Technology, 2012, 15 (2), pp.131-150
ISSN
1381-2416
Publisher
Springer
Start Page
131
End Page
150
Journal / Book Title
International Journal of Speech Technology
Volume
15
Issue
2
Copyright Statement
© 2012 Springer Science+Business Media, LLC. The original publication is available at www.springerlink.com
Description
06.08.13 KB. Ok to add accepted version to spiral, embargo period expired. Springer