基于长短期记忆神经网络的城市声音分类

2019 Federated Conference on Computer Science and Information Systems (FedCSIS) Pub Date : 2019-09-26 DOI:10.15439/2019F185

Yurij Lezhenin, N. Bogach, Evgeny Pyshkin

{"title":"基于长短期记忆神经网络的城市声音分类","authors":"Yurij Lezhenin, N. Bogach, Evgeny Pyshkin","doi":"10.15439/2019F185","DOIUrl":null,"url":null,"abstract":"Environmental sound classification has received more attention in recent years. Analysis of environmental sounds is difficult because of its unstructured nature. However, the presence of strong spectro-temporal patterns makes the classification possible. Since LSTM neural networks are efficient at learning temporal dependencies we propose and examine a LSTM model for urban sound classification. The model is trained on magnitude mel-spectrograms extracted from UrbanSound8K dataset audio. The proposed network is evaluated using 5-fold cross-validation and compared with the baseline CNN. It is shown that the LSTM model outperforms a set of existing solutions and is more accurate and confident than the CNN.","PeriodicalId":168208,"journal":{"name":"2019 Federated Conference on Computer Science and Information Systems (FedCSIS)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-09-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"24","resultStr":"{\"title\":\"Urban Sound Classification using Long Short-Term Memory Neural Network\",\"authors\":\"Yurij Lezhenin, N. Bogach, Evgeny Pyshkin\",\"doi\":\"10.15439/2019F185\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Environmental sound classification has received more attention in recent years. Analysis of environmental sounds is difficult because of its unstructured nature. However, the presence of strong spectro-temporal patterns makes the classification possible. Since LSTM neural networks are efficient at learning temporal dependencies we propose and examine a LSTM model for urban sound classification. The model is trained on magnitude mel-spectrograms extracted from UrbanSound8K dataset audio. The proposed network is evaluated using 5-fold cross-validation and compared with the baseline CNN. It is shown that the LSTM model outperforms a set of existing solutions and is more accurate and confident than the CNN.\",\"PeriodicalId\":168208,\"journal\":{\"name\":\"2019 Federated Conference on Computer Science and Information Systems (FedCSIS)\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-09-26\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"24\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2019 Federated Conference on Computer Science and Information Systems (FedCSIS)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.15439/2019F185\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2019 Federated Conference on Computer Science and Information Systems (FedCSIS)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.15439/2019F185","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 24

摘要

环境声分类近年来受到越来越多的关注。分析环境声音是困难的，因为它是非结构化的。然而，强烈的光谱-时间模式的存在使得分类成为可能。由于LSTM神经网络在学习时间依赖性方面是有效的，我们提出并检验了一个用于城市声音分类的LSTM模型。该模型是在UrbanSound8K数据集音频提取的震级谱图上进行训练的。使用5倍交叉验证对所提出的网络进行评估，并与基线CNN进行比较。结果表明，LSTM模型优于一组现有的解决方案，并且比CNN更准确和更有信心。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Urban Sound Classification using Long Short-Term Memory Neural Network

Environmental sound classification has received more attention in recent years. Analysis of environmental sounds is difficult because of its unstructured nature. However, the presence of strong spectro-temporal patterns makes the classification possible. Since LSTM neural networks are efficient at learning temporal dependencies we propose and examine a LSTM model for urban sound classification. The model is trained on magnitude mel-spectrograms extracted from UrbanSound8K dataset audio. The proposed network is evaluated using 5-fold cross-validation and compared with the baseline CNN. It is shown that the LSTM model outperforms a set of existing solutions and is more accurate and confident than the CNN.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2019 Federated Conference on Computer Science and Information Systems (FedCSIS)

自引率

0.00%

发文量