Audiovisual arrays for untethered spoken interfaces

Proceedings. Fourth IEEE International Conference on Multimodal Interfaces Pub Date : 2002-10-14 DOI:10.1109/ICMI.2002.1167026

K. Wilson, Vibhav Rangarajan, N. Checka, Trevor Darrell

引用次数: 9

Abstract

When faced with a distant speaker at a known location in a noisy environment, a microphone array can provide a significantly improved audio signal for speech recognition. Estimating the location of a speaker in a reverberant environment from audio information alone can be quite difficult, so we use an array of video cameras to aid localization. Stereo processing techniques are used on pairs of cameras, and foreground 3-D points are grouped to estimate the trajectory of people as they move in an environment. These trajectories are used to guide a microphone array beamformer. Initial results using this system for speech recognition demonstrate increased recognition rates compared to non-array processing techniques.

查看原文本刊更多论文

用于无系绳语音接口的视听阵列

当在嘈杂的环境中面对一个在已知位置的远距离说话者时，麦克风阵列可以为语音识别提供显着改进的音频信号。仅从音频信息估计扬声器在混响环境中的位置是相当困难的，所以我们使用一系列摄像机来帮助定位。立体处理技术用于成对的摄像机，前景3d点被分组，以估计人们在环境中移动时的轨迹。这些轨迹用于引导麦克风阵列波束形成器。使用该系统进行语音识别的初步结果表明，与非阵列处理技术相比，识别率有所提高。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings. Fourth IEEE International Conference on Multimodal Interfaces

自引率

0.00%

发文量