The IBM 2007 speech transcription system for European parliamentary speeches

2007 IEEE Workshop on Automatic Speech Recognition & Understanding (ASRU) Pub Date : 2007-12-01 DOI:10.1109/ASRU.2007.4430158

B. Ramabhadran, O. Siohan, A. Sethy

引用次数: 43

Abstract

TC-STAR is an European Union funded speech to speech translation project to transcribe, translate and synthesize European Parliamentary Plenary Speeches (EPPS). This paper describes IBM's English speech recognition system submitted to the TC-STAR 2007 Evaluation. Language model adaptation based on clustering and data selection using relative entropy minimization provided significant gains in the 2007 evaluation. The additional advances over the 2006 system that we present in this paper include unsupervised training of acoustic and language models; a system architecture that is based on cross-adaptation across complementary systems and system combination through generation of an ensemble of systems using randomized decision tree state-tying. These advances reduced the error rate by 30% relative over the best-performing system in the TC-STAR 2006 evaluation on the 2006 English development and evaluation test sets, and produced one of the best performing systems on the 2007 evaluation in English with a word error rate of 7.1%.

查看原文本刊更多论文

欧洲议会演讲的IBM 2007语音转录系统

TC-STAR是欧盟资助的演讲到演讲翻译项目，用于转录、翻译和合成欧洲议会全体会议演讲(EPPS)。本文介绍了IBM公司提交TC-STAR 2007评估的英语语音识别系统。基于聚类的语言模型自适应和使用相对熵最小化的数据选择在2007年的评估中取得了显著的进展。我们在本文中介绍的2006年系统的其他进步包括声学和语言模型的无监督训练;一种系统架构，它基于互补系统之间的交叉适应和系统组合，通过使用随机决策树状态绑定生成系统集合。这些进步将错误率相对于2006年英语发展和评估测试集的TC-STAR 2006评估中表现最好的系统降低了30%，并产生了2007年英语评估中表现最好的系统之一，单词错误率为7.1%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2007 IEEE Workshop on Automatic Speech Recognition & Understanding (ASRU)

自引率

0.00%

发文量