Perceptual MVDR-based unsupervised built-in speaker normalization for Kazakh speech recognition

2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT) Pub Date : 2014-10-01 DOI:10.1109/ICAICT.2014.7035914

Zhandos Yessenbayev, U. Yapanel

引用次数: 0

Abstract

In this work we present a novel approach to unsupervised speaker normalization on top of the Perceptual MVDR-based Built-in Speaker Normalization technique. We showed that the proposed method can be efficient for the task of phonetic recognition on TIMIT and then applied it to Kazakh speech recognition. From the experiments, we see that this method is able to improve the relative performance of ASR systems up to 20% The analysis of the optimal warp factor selection by the algorithm revealed a nice gender separation ability which may be used for gender/speaker classification tasks.

查看原文本刊更多论文

基于感知mvdr的哈萨克语语音识别的无监督内置说话人归一化

在这项工作中，我们在基于感知mvdr的内置说话人归一化技术的基础上提出了一种新的无监督说话人归一化方法。结果表明，该方法可以有效地完成语音识别任务，并将其应用于哈萨克语语音识别中。实验结果表明，该方法可将ASR系统的相对性能提高20%以上。通过对算法的最优扭曲因子选择的分析，表明该算法具有良好的性别分离能力，可用于性别/说话人分类任务。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT)

自引率

0.00%

发文量