Hybrid acoustic models for distant and multichannel large vocabulary speech recognition

2013 IEEE Workshop on Automatic Speech Recognition and Understanding Pub Date : 2013-12-01 DOI:10.1109/ASRU.2013.6707744

P. Swietojanski, Arnab Ghoshal, S. Renals

引用次数: 112

Abstract

We investigate the application of deep neural network (DNN)-hidden Markov model (HMM) hybrid acoustic models for far-field speech recognition of meetings recorded using microphone arrays. We show that the hybrid models achieve significantly better accuracy than conventional systems based on Gaussian mixture models (GMMs). We observe up to 8% absolute word error rate (WER) reduction from a discriminatively trained GMM baseline when using a single distant microphone, and between 4-6% absolute WER reduction when using beamforming on various combinations of array channels. By training the networks on audio from multiple channels, we find the networks can recover significant part of accuracy difference between the single distant microphone and beamformed configurations. Finally, we show that the accuracy of a network recognising speech from a single distant microphone can approach that of a multi-microphone setup by training with data from other microphones.

查看原文本刊更多论文

远距离和多通道大词汇语音识别的混合声学模型

研究了深度神经网络(DNN)-隐马尔可夫模型(HMM)混合声学模型在麦克风阵列录制会议远场语音识别中的应用。结果表明，混合模型比基于高斯混合模型(GMMs)的传统系统具有更好的精度。我们观察到，当使用单个远距离麦克风时，从鉴别训练的GMM基线来看，绝对字错误率(WER)降低了8%，当在各种阵列信道组合上使用波束形成时，绝对字错误率(WER)降低了4-6%。通过对多声道音频进行训练，我们发现网络可以恢复单远端麦克风和波束形成配置之间的很大一部分精度差异。最后，我们表明，通过使用来自其他麦克风的数据进行训练，网络识别来自单个远端麦克风的语音的准确性可以接近多麦克风设置的准确性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2013 IEEE Workshop on Automatic Speech Recognition and Understanding

自引率

0.00%

发文量