基于多分辨率HMM的声事件分类

2018 26th European Signal Processing Conference (EUSIPCO) Pub Date : 2018-09-01 DOI:10.23919/EUSIPCO.2018.8553131

P. Baggenstoss

{"title":"基于多分辨率HMM的声事件分类","authors":"P. Baggenstoss","doi":"10.23919/EUSIPCO.2018.8553131","DOIUrl":null,"url":null,"abstract":"Real-world acoustic events span a wide range of time and frequency resolutions, from short clicks to longer tonals. This is a challenge for the hidden Markov model (HMM), which uses a fixed segmentation and feature extraction, forcing a compromise between time and frequency resolution. The multiresolution HMM (MR-HMM) is an extension of the HMM that assumes not only an underlying (hidden) random state sequence, but also an underlying random segmentation, with segments spanning a wide range of sizes and processed using a variety of feature extraction methods. It is shown that the MR-HMM alone, as an acoustic event classifier, has performance comparable to state of the art discriminative classifiers on three open data sets. However, as a generative classifier, the MR-HMM models the underlying data generation process and can generate synthetic data, allowing weaknesses of individual class models to be discovered and corrected. To demonstrate this point, the MR-HMM is combined with auxiliary features that capture temporal information, resulting in significantly improved performance.","PeriodicalId":303069,"journal":{"name":"2018 26th European Signal Processing Conference (EUSIPCO)","volume":"107 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2018-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"7","resultStr":"{\"title\":\"Acoustic Event Classification Using Multi-Resolution HMM\",\"authors\":\"P. Baggenstoss\",\"doi\":\"10.23919/EUSIPCO.2018.8553131\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Real-world acoustic events span a wide range of time and frequency resolutions, from short clicks to longer tonals. This is a challenge for the hidden Markov model (HMM), which uses a fixed segmentation and feature extraction, forcing a compromise between time and frequency resolution. The multiresolution HMM (MR-HMM) is an extension of the HMM that assumes not only an underlying (hidden) random state sequence, but also an underlying random segmentation, with segments spanning a wide range of sizes and processed using a variety of feature extraction methods. It is shown that the MR-HMM alone, as an acoustic event classifier, has performance comparable to state of the art discriminative classifiers on three open data sets. However, as a generative classifier, the MR-HMM models the underlying data generation process and can generate synthetic data, allowing weaknesses of individual class models to be discovered and corrected. To demonstrate this point, the MR-HMM is combined with auxiliary features that capture temporal information, resulting in significantly improved performance.\",\"PeriodicalId\":303069,\"journal\":{\"name\":\"2018 26th European Signal Processing Conference (EUSIPCO)\",\"volume\":\"107 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2018-09-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"7\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2018 26th European Signal Processing Conference (EUSIPCO)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.23919/EUSIPCO.2018.8553131\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2018 26th European Signal Processing Conference (EUSIPCO)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.23919/EUSIPCO.2018.8553131","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 7

摘要

现实世界的声学事件跨越了广泛的时间和频率分辨率，从短的点击到较长的音调。这对隐马尔可夫模型(HMM)来说是一个挑战，隐马尔可夫模型使用固定的分割和特征提取，迫使在时间和频率分辨率之间做出妥协。多分辨率HMM (MR-HMM)是HMM的扩展，它不仅假设底层的(隐藏的)随机状态序列，而且假设底层的随机分割，其中的片段跨越了广泛的大小范围，并使用各种特征提取方法进行处理。结果表明，作为一种声学事件分类器，MR-HMM在三个开放数据集上的性能可与最先进的判别分类器相媲美。然而，作为一种生成分类器，MR-HMM对底层数据生成过程进行建模，可以生成合成数据，从而发现和纠正单个类模型的弱点。为了证明这一点，MR-HMM与捕获时间信息的辅助特征相结合，从而显著提高了性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Acoustic Event Classification Using Multi-Resolution HMM

Real-world acoustic events span a wide range of time and frequency resolutions, from short clicks to longer tonals. This is a challenge for the hidden Markov model (HMM), which uses a fixed segmentation and feature extraction, forcing a compromise between time and frequency resolution. The multiresolution HMM (MR-HMM) is an extension of the HMM that assumes not only an underlying (hidden) random state sequence, but also an underlying random segmentation, with segments spanning a wide range of sizes and processed using a variety of feature extraction methods. It is shown that the MR-HMM alone, as an acoustic event classifier, has performance comparable to state of the art discriminative classifiers on three open data sets. However, as a generative classifier, the MR-HMM models the underlying data generation process and can generate synthetic data, allowing weaknesses of individual class models to be discovered and corrected. To demonstrate this point, the MR-HMM is combined with auxiliary features that capture temporal information, resulting in significantly improved performance.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2018 26th European Signal Processing Conference (EUSIPCO)

自引率

0.00%

发文量