Sorry, you need to enable JavaScript to visit this website.

facebooktwittermailshare

Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks

Abstract: 

Head movement is an integral part of face-to-face communications. It is important to investigate methodologies to generate naturalistic movements for conversational agents (CAs). The predominant method for head movement generation is using rules based on the meaning of the message. However, the variations of head movements by these methods are bounded by the predefined dictionary of gestures. Speech-driven methods offer an alternative approach, learning the relationship between speech and head movements from real recordings. However, previous studies do not generate novel realizations for a repeated speech signal. Conditional generative adversarial network (GAN) provides a framework to generate multiple realizations of head movements for each speech segment by sampling from a conditioned distribution. We build a conditional GAN with bidirectional long-short term memory (BLSTM), which is suitable for capturing the long-short term dependencies of time-continuous signals. This model learns the distribution of head movements conditioned on speech prosodic features. We compare this model with a dynamic Bayesian network (DBN) and BLSTM models optimized to reduce mean squared error (MSE) or to increase concordance correlation. The objective evaluations and subjective evaluations of the results showed better performance for the conditional GAN model compared with these baseline systems.

up
0 users have voted:

Paper Details

Authors:
Najmeh Sadoughi, Carlos Busso
Submitted On:
1 May 2018 - 8:43pm
Short Link:
Type:
Poster
Event:
Document Year:
2018
Cite

Document Files

Sadoughi_2018-poster.pdf

(19 downloads)

Subscribe

[1] Najmeh Sadoughi, Carlos Busso, "Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks", IEEE SigPort, 2018. [Online]. Available: http://sigport.org/3198. Accessed: May. 20, 2018.
@article{3198-18,
url = {http://sigport.org/3198},
author = {Najmeh Sadoughi; Carlos Busso },
publisher = {IEEE SigPort},
title = {Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks},
year = {2018} }
TY - EJOUR
T1 - Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks
AU - Najmeh Sadoughi; Carlos Busso
PY - 2018
PB - IEEE SigPort
UR - http://sigport.org/3198
ER -
Najmeh Sadoughi, Carlos Busso. (2018). Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks. IEEE SigPort. http://sigport.org/3198
Najmeh Sadoughi, Carlos Busso, 2018. Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks. Available at: http://sigport.org/3198.
Najmeh Sadoughi, Carlos Busso. (2018). "Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks." Web.
1. Najmeh Sadoughi, Carlos Busso. Novel Realizations of Speech-driven Head Movements with Generative Adversarial Networks [Internet]. IEEE SigPort; 2018. Available from : http://sigport.org/3198