Sound Source Distance Estimation in Rooms based on Statistical Properties of Binaural Signals

IEEE Transactions on Audio Speech and Language Processing Pub Date : 2013-08-01 DOI:10.1109/TASL.2013.2260155

Eleftheria Georganti, T. May, S. Par, J. Mourjopoulos

引用次数: 36

Abstract

A novel method for the estimation of the distance of a sound source from binaural speech signals is proposed. The method relies on several statistical features extracted from such signals and their binaural cues. Firstly, the standard deviation of the difference of the magnitude spectra of the left and right binaural signals is used as a feature for this method. In addition, an extended set of additional statistical features that can improve distance detection is extracted from an auditory front-end which models the peripheral processing of the human auditory system. The method incorporates the above features into two classification frameworks based on Gaussian mixture models and Support Vector Machines and the relative merits of those frameworks are evaluated. The proposed method achieves distance detection when tested in various acoustical environments and performs well in unknown environments. Its performance is also compared to an existing binaural distance detection method.

查看原文本刊更多论文

基于双耳信号统计特性的室内声源距离估计

提出了一种从双耳语音信号中估计声源距离的新方法。该方法依赖于从这些信号及其双耳线索中提取的几个统计特征。该方法首先利用左右双耳信号的星等谱差的标准差作为特征;此外，从模拟人类听觉系统外围处理的听觉前端提取了一组扩展的附加统计特征，可以改进距离检测。该方法将上述特征融合到基于高斯混合模型和支持向量机的两种分类框架中，并对两种框架的优劣进行了比较。该方法在各种声环境下均能实现距离检测，在未知环境下也能取得良好的效果。并与现有的双耳距离检测方法进行了比较。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

IEEE Transactions on Audio Speech and Language Processing 工程技术-工程：电子与电气

自引率

0.00%

发文量

审稿时长

24.0 months

期刊介绍： The IEEE Transactions on Audio, Speech and Language Processing covers the sciences, technologies and applications relating to the analysis, coding, enhancement, recognition and synthesis of audio, music, speech and language. In particular, audio processing also covers auditory modeling, acoustic modeling and source separation. Speech processing also covers speech production and perception, adaptation, lexical modeling and speaker recognition. Language processing also covers spoken language understanding, translation, summarization, mining, general language modeling, as well as spoken dialog systems.