Voice activity detection using harmonic frequency components in likelihood ratio test

2010 IEEE International Conference on Acoustics, Speech and Signal Processing Pub Date : 2010-03-14 DOI:10.1109/ICASSP.2010.5495611

L. Tan, B. J. Borgstrom, A. Alwan

引用次数: 44

Abstract

This paper proposes a new statistical model-based likelihood ratio test (LRT) VAD to obtain reliable speech / non-speech decisions. In the proposed method, the likelihood ratio (LR) is calculated differently for voiced frames, as opposed to unvoiced frames: only DFT bins containing harmonic spectral peaks are selected for LR computation. To evaluate the new VAD's effectiveness in improving the noise-robustness of ASR, its decisions are applied to pre-processing techniques such as non-linear spectral subtraction, minimum mean square error short-time spectral amplitude estimator, and frame dropping. From the ASR experiments conducted on the Aurora2 database, the proposed harmonic frequency-based LRTs give better results than conventional LRT-based VADs and the standard G.729B and ETSI AMR VADs.

查看原文本刊更多论文

似然比检验中谐波频率分量的语音活动检测

本文提出了一种新的基于统计模型的似然比检验(LRT) VAD来获得可靠的语音/非语音决策。在该方法中，对浊音帧和非浊音帧的似然比(LR)进行不同的计算:仅选择包含谐波谱峰的DFT箱进行LR计算。为了评估新的VAD在提高ASR噪声鲁棒性方面的有效性，将其决策应用于非线性谱减法、最小均方误差短时谱幅估计和降帧等预处理技术。在Aurora2数据库上进行的ASR实验表明，基于谐波频率的lrt比传统的lrt VADs和标准的G.729B和ETSI AMR VADs具有更好的效果。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2010 IEEE International Conference on Acoustics, Speech and Signal Processing

自引率

0.00%

发文量