Sorry, you need to enable JavaScript to visit this website.

facebooktwittermailshare

UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION

Abstract: 

The scarcity of emotional speech data is a bottleneck of developing automatic speech emotion recognition (ASER) systems. One way to alleviate this issue is to use unsupervised feature learning techniques to learn features from the widely available general speech and use these features to train emotion classifiers. These unsupervised methods, such as denoising autoencoder (DAE), variational autoencoder (VAE), adversarial autoencoder (AAE) and adversarial variational Bayes (AVB), can capture the intrinsic structure of the data distribution in the learned feature representation. In this work, we systematically investigate four kinds of unsupervised feature learning methods for improving speaker-independent ASER. We show that all methods improve the performance regarding unweighted accuracy rating (UAR) and F1-score over methods that use hand-crafted features or that do not perform feature learning on external datasets. We also show that VAE, AAE and AVB methods, which control the distribution of the latent representation, outperform DAE that does not control such distribution. This suggests the benefits of using variational inference methods to learn features from general speech for the speech tasks such as ASER that has very limited labeled data.

up
0 users have voted:

Paper Details

Authors:
Sefik Emre Eskimez, Zhiyao Duan, Wendi Heinzelman
Submitted On:
19 April 2018 - 4:01pm
Short Link:
Type:
Poster
Event:
Presenter's Name:
Sefik Emre Eskimez
Paper Code:
SP-P1.5
Document Year:
2018
Cite

Document Files

icassp-2018-poster.pdf

(253)

Subscribe

[1] Sefik Emre Eskimez, Zhiyao Duan, Wendi Heinzelman, "UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION", IEEE SigPort, 2018. [Online]. Available: http://sigport.org/3017. Accessed: Aug. 10, 2020.
@article{3017-18,
url = {http://sigport.org/3017},
author = {Sefik Emre Eskimez; Zhiyao Duan; Wendi Heinzelman },
publisher = {IEEE SigPort},
title = {UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION},
year = {2018} }
TY - EJOUR
T1 - UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION
AU - Sefik Emre Eskimez; Zhiyao Duan; Wendi Heinzelman
PY - 2018
PB - IEEE SigPort
UR - http://sigport.org/3017
ER -
Sefik Emre Eskimez, Zhiyao Duan, Wendi Heinzelman. (2018). UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION. IEEE SigPort. http://sigport.org/3017
Sefik Emre Eskimez, Zhiyao Duan, Wendi Heinzelman, 2018. UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION. Available at: http://sigport.org/3017.
Sefik Emre Eskimez, Zhiyao Duan, Wendi Heinzelman. (2018). "UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION." Web.
1. Sefik Emre Eskimez, Zhiyao Duan, Wendi Heinzelman. UNSUPERVISED LEARNING APPROACH TO FEATURE ANALYSIS FOR AUTOMATIC SPEECH EMOTION RECOGNITION [Internet]. IEEE SigPort; 2018. Available from : http://sigport.org/3017