用于感知用户界面的音视频阵列源分离

Workshop on Perceptive User Interfaces Pub Date : 2001-11-15 DOI:10.1145/971478.971500

K. Wilson, N. Checka, D. Demirdjian, Trevor Darrell

{"title":"用于感知用户界面的音视频阵列源分离","authors":"K. Wilson, N. Checka, D. Demirdjian, Trevor Darrell","doi":"10.1145/971478.971500","DOIUrl":null,"url":null,"abstract":"Steerable microphone arrays provide a flexible infrastructure for audio source separation. In order for them to be used effectively in perceptual user interfaces, there must be a mechanism in place for steering the focus of the array to the sound source. Audio-only steering techniques often perform poorly in the presence of multiple sound sources or strong reverberation. Video-only techniques can achieve high spatial precision but require that the audio and video subsystems be accurately calibrated to preserve this precision. We present an audio-video localization technique that combines the benefits of the two modalities. We implement our technique in a test environment containing multiple stereo cameras and a room-sized microphone array. Our technique achieves an 8.9 dB improvement over a single far-field microphone and a 6.7 dB improvement over source separation based on video-only localization.","PeriodicalId":416822,"journal":{"name":"Workshop on Perceptive User Interfaces","volume":"12 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2001-11-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"20","resultStr":"{\"title\":\"Audio-video array source separation for perceptual user interfaces\",\"authors\":\"K. Wilson, N. Checka, D. Demirdjian, Trevor Darrell\",\"doi\":\"10.1145/971478.971500\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Steerable microphone arrays provide a flexible infrastructure for audio source separation. In order for them to be used effectively in perceptual user interfaces, there must be a mechanism in place for steering the focus of the array to the sound source. Audio-only steering techniques often perform poorly in the presence of multiple sound sources or strong reverberation. Video-only techniques can achieve high spatial precision but require that the audio and video subsystems be accurately calibrated to preserve this precision. We present an audio-video localization technique that combines the benefits of the two modalities. We implement our technique in a test environment containing multiple stereo cameras and a room-sized microphone array. Our technique achieves an 8.9 dB improvement over a single far-field microphone and a 6.7 dB improvement over source separation based on video-only localization.\",\"PeriodicalId\":416822,\"journal\":{\"name\":\"Workshop on Perceptive User Interfaces\",\"volume\":\"12 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2001-11-15\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"20\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Workshop on Perceptive User Interfaces\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/971478.971500\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Workshop on Perceptive User Interfaces","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/971478.971500","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 20

摘要

可操纵麦克风阵列为音频源分离提供了灵活的基础设施。为了使它们能够有效地用于感知用户界面，必须有一种机制将阵列的焦点转向声源。在存在多个声源或强烈混响的情况下，纯音频转向技术通常表现不佳。仅视频技术可以实现高空间精度，但需要精确校准音频和视频子系统以保持这种精度。我们提出了一种结合了这两种方式的优点的音频-视频定位技术。我们在一个包含多个立体摄像机和一个房间大小的麦克风阵列的测试环境中实现了我们的技术。我们的技术比单远场麦克风提高了8.9 dB，比基于视频定位的源分离提高了6.7 dB。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Audio-video array source separation for perceptual user interfaces

Steerable microphone arrays provide a flexible infrastructure for audio source separation. In order for them to be used effectively in perceptual user interfaces, there must be a mechanism in place for steering the focus of the array to the sound source. Audio-only steering techniques often perform poorly in the presence of multiple sound sources or strong reverberation. Video-only techniques can achieve high spatial precision but require that the audio and video subsystems be accurately calibrated to preserve this precision. We present an audio-video localization technique that combines the benefits of the two modalities. We implement our technique in a test environment containing multiple stereo cameras and a room-sized microphone array. Our technique achieves an 8.9 dB improvement over a single far-field microphone and a 6.7 dB improvement over source separation based on video-only localization.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Workshop on Perceptive User Interfaces

自引率

0.00%

发文量