Sorry, you need to enable JavaScript to visit this website.

facebooktwittermailshare

SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING

Abstract: 

In this paper we propose speaker characterization using time delay neural networks and long short-term memory neural networks (TDNN-LSTM) speaker embedding. Three types of front-end feature extraction are investigated to find good features for speaker embedding. Three kinds of data augmentation are used to increase the amount and diversity of the training data. The proposed methods are evaluated with the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) tasks. Experimental results show that the proposed methods achieve a decision cost of 0.400 with the pooled SRE 2018 development set with a single system. In addition, by applying simple average score combination on the outputs of 12 systems, the proposed methods achieve an equal error rate (EER) of 5.56% and a minimum decision cost function of 0.423 with the SRE 2016 evaluation set.

up
1 user has voted: CHIH-TING YEH

Paper Details

Authors:
Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang
Submitted On:
7 May 2019 - 11:01pm
Short Link:
Type:
Poster
Event:

Document Files

ICASSP2019_poster_A0.pdf

(36)

Subscribe

[1] Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang, "SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING", IEEE SigPort, 2019. [Online]. Available: http://sigport.org/3997. Accessed: Sep. 16, 2019.
@article{3997-19,
url = {http://sigport.org/3997},
author = {Chia-Ping Chen; Su-Yu Zhang; Chih-Ting Yeh; Jia-Ching Wang; Tenghui Wang; Chien-Lin Huang },
publisher = {IEEE SigPort},
title = {SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING},
year = {2019} }
TY - EJOUR
T1 - SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING
AU - Chia-Ping Chen; Su-Yu Zhang; Chih-Ting Yeh; Jia-Ching Wang; Tenghui Wang; Chien-Lin Huang
PY - 2019
PB - IEEE SigPort
UR - http://sigport.org/3997
ER -
Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang. (2019). SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING. IEEE SigPort. http://sigport.org/3997
Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang, 2019. SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING. Available at: http://sigport.org/3997.
Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang. (2019). "SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING." Web.
1. Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang. SPEAKER CHARACTERIZATION USING TDNN-LSTM BASED SPEAKER EMBEDDING [Internet]. IEEE SigPort; 2019. Available from : http://sigport.org/3997