Sorry, you need to enable JavaScript to visit this website.

facebooktwittermailshare

Small energy masking for improved neural network training for end-to-end speech recognition

Abstract: 

In this paper, we present a Small Energy Masking (SEM) algorithm, which masks inputs having values below a certain threshold. More specifically, a time-frequency bin is masked if the filterbank energy in this bin is less than a certain energy threshold. A uniform distribution is employed to randomly generate the ratio of this energy threshold to the peak filterbank energy of each utterance in decibels. The unmasked feature elements are scaled so that the total sum of the feature values remain the same through this masking procedure. This very simple algorithm shows relatively 11.2% and 13.5% Word Error Rate (WER) improvements on the standard LibriSpeech test-clean and test-other sets over the baseline end-to-end speech recognition system. Additionally, compared to the input dropout algorithm, SEM algorithm shows relatively 7.7% and 11.6% improvements on the same LibriSpeech test-clean and test-other sets. With a modified shallow-fusion technique with a Transformer LM, we obtained a 2.62% WER on the Lib-riSpeech test-clean set and a 7.87% WER on the LibriSpeech test-other set.

up
0 users have voted:

Paper Details

Authors:
Chanwoo Kim, Kwangyoun Kim, Sathish Reddy Indurthi
Submitted On:
5 May 2020 - 5:27pm
Short Link:
Type:
Presentation Slides
Event:
Presenter's Name:
Chanwoo Kim
Document Year:
2020
Cite

Document Files

20200508_icassp_small_energy_masking_paper_3965_presentation.pdf

(21)

Subscribe

[1] Chanwoo Kim, Kwangyoun Kim, Sathish Reddy Indurthi, "Small energy masking for improved neural network training for end-to-end speech recognition", IEEE SigPort, 2020. [Online]. Available: http://sigport.org/5125. Accessed: Jun. 06, 2020.
@article{5125-20,
url = {http://sigport.org/5125},
author = {Chanwoo Kim; Kwangyoun Kim; Sathish Reddy Indurthi },
publisher = {IEEE SigPort},
title = {Small energy masking for improved neural network training for end-to-end speech recognition},
year = {2020} }
TY - EJOUR
T1 - Small energy masking for improved neural network training for end-to-end speech recognition
AU - Chanwoo Kim; Kwangyoun Kim; Sathish Reddy Indurthi
PY - 2020
PB - IEEE SigPort
UR - http://sigport.org/5125
ER -
Chanwoo Kim, Kwangyoun Kim, Sathish Reddy Indurthi. (2020). Small energy masking for improved neural network training for end-to-end speech recognition. IEEE SigPort. http://sigport.org/5125
Chanwoo Kim, Kwangyoun Kim, Sathish Reddy Indurthi, 2020. Small energy masking for improved neural network training for end-to-end speech recognition. Available at: http://sigport.org/5125.
Chanwoo Kim, Kwangyoun Kim, Sathish Reddy Indurthi. (2020). "Small energy masking for improved neural network training for end-to-end speech recognition." Web.
1. Chanwoo Kim, Kwangyoun Kim, Sathish Reddy Indurthi. Small energy masking for improved neural network training for end-to-end speech recognition [Internet]. IEEE SigPort; 2020. Available from : http://sigport.org/5125