Voice activity detection based on frequency modulation of harmonics

2013 IEEE International Conference on Acoustics, Speech and Signal Processing Pub Date : 2013-10-21 DOI:10.1109/ICASSP.2013.6638954

Chung-Chien Hsu, Tse-En Lin, Jian-Hueng Chen, T. Chi

引用次数: 12

Abstract

In this paper, we propose a voice activity detection (VAD) algorithm based on spectro-temporal modulation structures of input sounds. A multi-resolution spectro-temporal analysis framework is used to inspect prominent speech structures. By comparing with an adaptive threshold, the proposed VAD distinguishes speech from non-speech based on the energy of the frequency modulation of harmonics. Compared with three standard VADs, ITU-T G.729B, ETSI AMR1 and AMR2, our proposed VAD significantly outperforms them in non-stationary noises in terms of the receiver operating characteristic (ROC) curves and the recognition rates from a practical distributed speech recognition (DSR) system.

查看原文本刊更多论文

基于谐波调频的语音活动检测

本文提出了一种基于输入声音的频谱-时间调制结构的语音活动检测(VAD)算法。使用多分辨率频谱-时间分析框架来检测突出的语音结构。通过与自适应阈值的比较，本文提出的VAD基于谐波的调频能量来区分语音和非语音。与ITU-T G.729B、ETSI AMR1和AMR2三种标准VAD相比，本文提出的VAD在非平稳噪声条件下，在接收机工作特性曲线(ROC)和实际分布式语音识别(DSR)系统的识别率方面都明显优于它们。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2013 IEEE International Conference on Acoustics, Speech and Signal Processing

自引率

0.00%

发文量