Minimum Classification Error Training with Speech Synthesis-Based Regularization for Speech Recognition

International Conference on Signal Processing and Machine Learning Pub Date : 2019-11-27 DOI:10.1145/3372806.3372819

Naoto Umezaki, Takumi Okubo, Hideyuki Watanabe, S. Katagiri, M. Ohsaki

{"title":"Minimum Classification Error Training with Speech Synthesis-Based Regularization for Speech Recognition","authors":"Naoto Umezaki, Takumi Okubo, Hideyuki Watanabe, S. Katagiri, M. Ohsaki","doi":"10.1145/3372806.3372819","DOIUrl":null,"url":null,"abstract":"To increase the utility of Regularization, which is a common framework for avoiding the underestimation of ideal Bayes error, for speech recognizer training, we propose a new classifier training concept that incorporates a regularization term that represents the speech synthesis ability of classifier parameters. To implement our new concept, we first introduce a speech recognizer that embeds Line Spectral Pairs-Conjugate Structure-Algebraic Code Excited Linear Prediction (LSP-CS-ACELP) in a Multi-Prototype State-Transition-Model (MP-STM) classifier, define a regularization term that represents the speech synthesis ability by the distance between a training sample and its nearest MP-STM word model, and formalize a new Minimum Classification Error (MCE) training method for jointly minimizing a conventional smooth classification error count loss and the newly defined regularization term. We evaluated the proposed training method in an isolated-word, closed-vocabulary, and speaker-independent speech recognition task whose Bayes error is estimated to be about 20% and found that our method successfully produced an estimate of Bayes error (about 18.4%) with a single training run over a training dataset without such data resampling as Cross-Validation or the assumptions of sample distribution. Moreover, we investigated the quality of the synthesized speech using LSP parameters derived from the trained prototypes and found that the quality of the Bayes error estimation is clearly supported by the speech synthesis ability preserved in the training.","PeriodicalId":340004,"journal":{"name":"International Conference on Signal Processing and Machine Learning","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-11-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"International Conference on Signal Processing and Machine Learning","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3372806.3372819","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

Abstract

To increase the utility of Regularization, which is a common framework for avoiding the underestimation of ideal Bayes error, for speech recognizer training, we propose a new classifier training concept that incorporates a regularization term that represents the speech synthesis ability of classifier parameters. To implement our new concept, we first introduce a speech recognizer that embeds Line Spectral Pairs-Conjugate Structure-Algebraic Code Excited Linear Prediction (LSP-CS-ACELP) in a Multi-Prototype State-Transition-Model (MP-STM) classifier, define a regularization term that represents the speech synthesis ability by the distance between a training sample and its nearest MP-STM word model, and formalize a new Minimum Classification Error (MCE) training method for jointly minimizing a conventional smooth classification error count loss and the newly defined regularization term. We evaluated the proposed training method in an isolated-word, closed-vocabulary, and speaker-independent speech recognition task whose Bayes error is estimated to be about 20% and found that our method successfully produced an estimate of Bayes error (about 18.4%) with a single training run over a training dataset without such data resampling as Cross-Validation or the assumptions of sample distribution. Moreover, we investigated the quality of the synthesized speech using LSP parameters derived from the trained prototypes and found that the quality of the Bayes error estimation is clearly supported by the speech synthesis ability preserved in the training.

查看原文本刊更多论文

基于语音合成的正则化最小分类误差训练用于语音识别

正则化是避免理想贝叶斯误差低估的常用框架，为了提高正则化在语音识别器训练中的效用，我们提出了一种新的分类器训练概念，该概念包含了代表分类器参数语音合成能力的正则化项。为了实现我们的新概念，我们首先在多原型状态转换模型(MP-STM)分类器中引入了一个语音识别器，该识别器嵌入了线谱对共轭结构代数码激发线性预测(LSP-CS-ACELP)，定义了一个正则化项，该正则化项通过训练样本与其最近的MP-STM单词模型之间的距离来表示语音合成能力。并形式化了一种新的最小分类误差(MCE)训练方法，该方法将传统的平滑分类误差计数损失和新定义的正则化项联合最小化。我们在一个孤立词、封闭词汇和说话人独立的语音识别任务中评估了所提出的训练方法，该任务的贝叶斯误差估计约为20%，并发现我们的方法在一个训练数据集上运行一次训练就成功地产生了贝叶斯误差估计(约18.4%)，而没有交叉验证或样本分布假设等数据重新采样。此外，我们使用从训练原型中获得的LSP参数研究了合成语音的质量，发现贝叶斯误差估计的质量明显得到了训练中保留的语音合成能力的支持。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

International Conference on Signal Processing and Machine Learning

自引率

0.00%

发文量