Sorry, you need to enable JavaScript to visit this website.

facebooktwittermailshare

A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition

Abstract: 

We propose a novel speaker-dependent (SD) approach to joint training of deep neural networks (DNNs) with an explicit speech separation structure for multi-talker speech recognition in a single-channel setting. First, a multi-condition training strategy is designed for a SD-DNN recognizer in multi-talker scenarios, which can significantly reduce the decoding runtime and improve the recognition accuracy over the approaches that use speaker-independent DNN models with a complicated joint decoding framework. In addition, a SD regression DNN for mapping the acoustic features of mixed speech to the speech features of a target speaker is jointly trained with the SD recognition DNN for acoustic modeling. Our experiments on the Speech Separation Challenge (SSC) task show that the proposed SD recognition system under multi-condition training achieves an average word error rate (WER) of 3.8\%, yielding a relative WER reduction of 65.1\% from the proposed DNN pre-processing approach under clean-condition training \cite{Tu15}. Furthermore, the jointly trained DNN system generates a relative WER reduction of 13.2\% from the state-of-the-art systems under multi-condition training.

up
0 users have voted:

Paper Details

Authors:
Yan-Hui Tu, Jun Du, Li-Rong Dai, Chin-Hui Lee
Submitted On:
15 October 2016 - 2:48am
Short Link:
Type:
Presentation Slides
Event:
Presenter's Name:
Yan-Hui Tu
Paper Code:
106
Document Year:
2016
Cite

Document Files

Yanhui_ISCSLP2016_oral.pdf

(386)

Subscribe

[1] Yan-Hui Tu, Jun Du, Li-Rong Dai, Chin-Hui Lee, "A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition", IEEE SigPort, 2016. [Online]. Available: http://sigport.org/1216. Accessed: Sep. 27, 2020.
@article{1216-16,
url = {http://sigport.org/1216},
author = {Yan-Hui Tu; Jun Du; Li-Rong Dai; Chin-Hui Lee },
publisher = {IEEE SigPort},
title = {A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition},
year = {2016} }
TY - EJOUR
T1 - A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition
AU - Yan-Hui Tu; Jun Du; Li-Rong Dai; Chin-Hui Lee
PY - 2016
PB - IEEE SigPort
UR - http://sigport.org/1216
ER -
Yan-Hui Tu, Jun Du, Li-Rong Dai, Chin-Hui Lee. (2016). A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition. IEEE SigPort. http://sigport.org/1216
Yan-Hui Tu, Jun Du, Li-Rong Dai, Chin-Hui Lee, 2016. A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition. Available at: http://sigport.org/1216.
Yan-Hui Tu, Jun Du, Li-Rong Dai, Chin-Hui Lee. (2016). "A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition." Web.
1. Yan-Hui Tu, Jun Du, Li-Rong Dai, Chin-Hui Lee. A Speaker-Dependent Deep Learning Approach to Joint Speech Separation and Acoustic Modeling for Multi-Talker Automatic Speech Recognition [Internet]. IEEE SigPort; 2016. Available from : http://sigport.org/1216