標題: A new duration modeling approach for Mandarin speech
作者: Chen, SH
Lai, WH
Wang, YR
電信工程研究所
Institute of Communications Engineering
關鍵字: duration modeling;Mandarin;text-to-speech
公開日期: 1-Jul-2003
摘要: In this paper, a new duration modeling approach for Mandarin speech is proposed. It explicitly takes several major affecting factors as multiplicative companding factors (CFs) and estimates all model parameters by an EM algorithm. Besides, the three basic Tone 3 patterns (i.e., full tone, half tone and sandhi tone) are also properly considered via using three different Us to separate their affections on syllable duration. Experimental results showed that the variance of the syllable duration was greatly reduced from 180.17 to 2.52 frame(2) (1 frame =5 ms) by the syllable duration modeling to eliminate effects from those affecting factors. Moreover, the estimated Us of those affecting factors agreed well to our prior linguistic knowledge. Two extensions of the duration modeling method are also performed. One is the use of the same technique to model initial and final durations. The other is to replace the multiplicative model with an additive one. Lastly, a preliminary study of applying the proposed model to predict syllable duration for TTS is also performed. Experimental results showed that it outperformed the conventional regressive prediction method.
URI: http://dx.doi.org/10.1109/TSA.2003.814377
http://hdl.handle.net/11536/27760
ISSN: 1063-6676
DOI: 10.1109/TSA.2003.814377
期刊: IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
Volume: 11
Issue: 4
起始頁: 308
結束頁: 320
Appears in Collections:Articles