Cooperative Speech Separation With a Microphone Array and Asynchronous Wearable Devices

Interspeech Pub Date : 2022-09-18 DOI:10.21437/interspeech.2022-11025

R. Corey, Manan Mittal, Kanad Sarkar, A. Singer

{"title":"Cooperative Speech Separation With a Microphone Array and Asynchronous Wearable Devices","authors":"R. Corey, Manan Mittal, Kanad Sarkar, A. Singer","doi":"10.21437/interspeech.2022-11025","DOIUrl":null,"url":null,"abstract":"We consider the problem of separating speech from several talkers in background noise using a fixed microphone array and a set of wearable devices. Wearable devices can provide reliable information about speech from their wearers, but they typically cannot be used directly for multichannel source separation due to network delay, sample rate offsets, and relative motion. Instead, the wearable microphone signals are used to compute the speech presence probability for each talker at each time-frequency index. Those parameters, which are robust against small sample rate offsets and relative motion, are used to track the second-order statistics of the speech sources and background noise. The fixed array then separates the speech signals using an adaptive linear time-varying multichannel Wiener filter. The proposed method is demonstrated using real-room recordings from three human talkers with binaural earbud microphones and an eight-microphone tabletop array. but are useful for distin-guishing between different sources because of their known positions relative to the talkers. The proposed system uses the wearable devices to estimate SPP values, which are then used to learn the second-order statistics for each source at the microphones of the fixed array. The array separates the sources using an adaptive linear time-varying spatial filter suitable for real-time applications. This work combines the cooperative ar-chitecture of [19], the distributed SPP method of [18], and the motion-robust modeling of [15]. The system is implemented adaptively and demonstrated using live human talkers.","PeriodicalId":73500,"journal":{"name":"Interspeech","volume":"1 1","pages":"5398-5402"},"PeriodicalIF":0.0000,"publicationDate":"2022-09-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Interspeech","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.21437/interspeech.2022-11025","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

We consider the problem of separating speech from several talkers in background noise using a fixed microphone array and a set of wearable devices. Wearable devices can provide reliable information about speech from their wearers, but they typically cannot be used directly for multichannel source separation due to network delay, sample rate offsets, and relative motion. Instead, the wearable microphone signals are used to compute the speech presence probability for each talker at each time-frequency index. Those parameters, which are robust against small sample rate offsets and relative motion, are used to track the second-order statistics of the speech sources and background noise. The fixed array then separates the speech signals using an adaptive linear time-varying multichannel Wiener filter. The proposed method is demonstrated using real-room recordings from three human talkers with binaural earbud microphones and an eight-microphone tabletop array. but are useful for distin-guishing between different sources because of their known positions relative to the talkers. The proposed system uses the wearable devices to estimate SPP values, which are then used to learn the second-order statistics for each source at the microphones of the fixed array. The array separates the sources using an adaptive linear time-varying spatial filter suitable for real-time applications. This work combines the cooperative ar-chitecture of [19], the distributed SPP method of [18], and the motion-robust modeling of [15]. The system is implemented adaptively and demonstrated using live human talkers.

查看原文本刊更多论文

基于麦克风阵列和异步可穿戴设备的协同语音分离

我们考虑了在背景噪声中使用固定麦克风阵列和一组可穿戴设备从多个说话者中分离语音的问题。可穿戴设备可以提供有关其佩戴者的语音的可靠信息，但由于网络延迟、采样率偏移和相对运动，它们通常不能直接用于多通道源分离。相反，使用可穿戴麦克风信号来计算每个说话者在每个时频指数下的语音存在概率。这些参数对小采样率偏移和相对运动具有鲁棒性，用于跟踪语音源和背景噪声的二阶统计量。固定阵列然后使用自适应线性时变多通道维纳滤波器分离语音信号。该方法通过使用双耳耳塞麦克风和8个麦克风桌面阵列的三个人的真实房间录音进行了演示。但是对于区分不同的来源是有用的，因为它们相对于说话者的已知位置。该系统使用可穿戴设备来估计SPP值，然后使用SPP值来学习固定阵列麦克风处每个源的二阶统计量。该阵列使用适合于实时应用的自适应线性时变空间滤波器分离源。本工作结合[19]的协作架构、[18]的分布式SPP方法和[15]的运动鲁棒建模。该系统是自适应实现的，并使用真人说话者进行演示。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Interspeech

自引率

0.00%

发文量