Kalamkas Zhagyparova, Ruslan Zhagypar, A. Zollanvari, M. Akhtar
{"title":"Supervised Learning-based Sound Source Distance Estimation Using Multivariate Features","authors":"Kalamkas Zhagyparova, Ruslan Zhagypar, A. Zollanvari, M. Akhtar","doi":"10.1109/TENSYMP52854.2021.9551007","DOIUrl":null,"url":null,"abstract":"This paper introduces the use of supervised machine learning methods with a combination of several sound source distance-dependent features to tackle the problem of distance-of-arrival (DisOA) estimation. The DisOA estimation is approached as a classification problem, which aims to classify a recorded audio signal into one of the predefined four DisOA classes regardless of the orientation angle. The datasets for both training and testing purposes are simulated by convolving appropriate room impulse responses with anechoic speech signals. The performance of three conventional and efficient classifiers was examined along with various subsets of four extracted features including: 1) Diffuseness (DIFF); 2) Binaural spectral magnitude difference standard deviation (BSMD-STD); 3) Magnitude squared coherence (MSC); and 4) Direct-to-reverberant ratio (DRR). The simulations consider the use of different source signals as well as varying directions-of-arrival and the room sizes. Our empirical results show that the use of a single univariate feature, namely, MSC, along with K-nearest neighbor (KNN) could potentially lead to an accurate DisOA classification rule.","PeriodicalId":137485,"journal":{"name":"2021 IEEE Region 10 Symposium (TENSYMP)","volume":"27 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-08-23","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2021 IEEE Region 10 Symposium (TENSYMP)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/TENSYMP52854.2021.9551007","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 1
Abstract
This paper introduces the use of supervised machine learning methods with a combination of several sound source distance-dependent features to tackle the problem of distance-of-arrival (DisOA) estimation. The DisOA estimation is approached as a classification problem, which aims to classify a recorded audio signal into one of the predefined four DisOA classes regardless of the orientation angle. The datasets for both training and testing purposes are simulated by convolving appropriate room impulse responses with anechoic speech signals. The performance of three conventional and efficient classifiers was examined along with various subsets of four extracted features including: 1) Diffuseness (DIFF); 2) Binaural spectral magnitude difference standard deviation (BSMD-STD); 3) Magnitude squared coherence (MSC); and 4) Direct-to-reverberant ratio (DRR). The simulations consider the use of different source signals as well as varying directions-of-arrival and the room sizes. Our empirical results show that the use of a single univariate feature, namely, MSC, along with K-nearest neighbor (KNN) could potentially lead to an accurate DisOA classification rule.