Improving Phoneme Sequence Recognition using Phoneme Duration Information in DNN-HSMM

M. Asadolahzade Kermanshahi; M. M. Homayounpour

Improving Phoneme Sequence Recognition using Phoneme Duration Information in DNN-HSMM

محل انتشار: مجله هوش مصنوعی و داده کاوی، دوره: 7، شماره: 1

سال انتشار: 1398

نوع سند: مقاله ژورنالی

زبان: انگلیسی

مشاهده: 523

فایل این مقاله در 11 صفحه با فرمت PDF قابل دریافت می باشد

دریافت فایل کامل مقاله

صدور گواهی نمایه سازی
من نویسنده این مقاله هستم

این مقاله در بخشهای موضوعی زیر دسته بندی شده است:

هوش مصنوعی > شبکه عصبی

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

https://civilica.com/doc/894078

شناسه ملی سند علمی:

JR_JADM-7-1_012

تاریخ نمایه سازی: 19 تیر 1398

چکیده مقاله:

Improving phoneme recognition has attracted the attention of many researchers due to its applications in various fields of speech processing. Recent research achievements show that using deep neural network (DNN) in speech recognition systems significantly improves the performance of these systems. There are two phases in DNN-based phoneme recognition systems including training and testing. Most previous research attempted to improve training phase such as training algorithms, different types of network, network architecture, feature type, etc. But in this study, we focus on test phase which is related to generate phoneme sequence that is also essential to achieve good phoneme recognition accuracy. Past research used Viterbi algorithm on hidden Markov model (HMM) to generate phoneme sequences. We address an important problem associated with this method. To deal with the problem of considering geometric distribution of state duration in HMM, we use real duration probability distribution for each phoneme with the aid of hidden semi-Markov model (HSMM). We also represent each phoneme with only one state to simply use phonemes duration information in HSMM. Furthermore, we investigate the performance of a post-processing method, which corrects the phoneme sequence obtained from the neural network, based on our knowledge about phonemes. The experimental results using the Persian FarsDat corpus show that using extended Viterbi algorithm on HSMM achieves phoneme recognition accuracy improvements of 2.68% and 0.56% over conventional methods using Gaussian mixture model-hidden Markov models (GMM-HMMs) and Viterbi on HMM, respectively. The post-processing method also increases the accuracy compared to before its application.

کلیدواژه ها:

Phoneme Recognition ، Deep Neural Network ، Hidden Markov Model ، Hidden Semi-Markov Model ، Extended Viterbi Algorithm

نویسندگان

M. Asadolahzade Kermanshahi

Computer Engineering and IT Department, Amirkabir University of Technology, Tehran, Iran

M. M. Homayounpour

Computer Engineering and IT Department, Amirkabir University of Technology, Tehran, Iran