基于生成扩散模型的语音增强

IF 0.5 Q4 COMPUTER SCIENCE, INFORMATION SYSTEMS
O. V. Girfanov, A. G. Shishkin
{"title":"基于生成扩散模型的语音增强","authors":"O. V. Girfanov,&nbsp;A. G. Shishkin","doi":"10.3103/S0005105523050035","DOIUrl":null,"url":null,"abstract":"<p>An alternative approach to speech denoising using generative diffusion models that model the distribution of training data is proposed. In recent years, such models have led to promising results to be obtained in the field of generating signals of various kinds, and these are superior in many ways to previous generative models, such as variational autoencoders. However, diffusion models have not yet found wide application in the field of speech denoising. A new diffusion model is presented, which can be used to denoise real speech signals using a deep neural network. Our own data set, with more than 150 h of pure speech in Russian, has been created. The obtained results, estimated using the metrics scale invariant signal to distortion ratio and perceptual evaluation of speech quality, are comparable or superior to the results of the best discriminative models.</p>","PeriodicalId":42995,"journal":{"name":"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS","volume":null,"pages":null},"PeriodicalIF":0.5000,"publicationDate":"2023-11-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Speech Enhancement with Generative Diffusion Models\",\"authors\":\"O. V. Girfanov,&nbsp;A. G. Shishkin\",\"doi\":\"10.3103/S0005105523050035\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<p>An alternative approach to speech denoising using generative diffusion models that model the distribution of training data is proposed. In recent years, such models have led to promising results to be obtained in the field of generating signals of various kinds, and these are superior in many ways to previous generative models, such as variational autoencoders. However, diffusion models have not yet found wide application in the field of speech denoising. A new diffusion model is presented, which can be used to denoise real speech signals using a deep neural network. Our own data set, with more than 150 h of pure speech in Russian, has been created. The obtained results, estimated using the metrics scale invariant signal to distortion ratio and perceptual evaluation of speech quality, are comparable or superior to the results of the best discriminative models.</p>\",\"PeriodicalId\":42995,\"journal\":{\"name\":\"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS\",\"volume\":null,\"pages\":null},\"PeriodicalIF\":0.5000,\"publicationDate\":\"2023-11-24\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://link.springer.com/article/10.3103/S0005105523050035\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q4\",\"JCRName\":\"COMPUTER SCIENCE, INFORMATION SYSTEMS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS","FirstCategoryId":"1085","ListUrlMain":"https://link.springer.com/article/10.3103/S0005105523050035","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q4","JCRName":"COMPUTER SCIENCE, INFORMATION SYSTEMS","Score":null,"Total":0}
引用次数: 0

摘要

提出了一种使用生成扩散模型对训练数据的分布进行建模的语音去噪方法。近年来,这种模型在生成各种类型的信号方面取得了可喜的成果,并且在许多方面优于以前的生成模型,如变分自编码器。然而,扩散模型在语音去噪领域还没有得到广泛的应用。提出了一种新的扩散模型,该模型可以利用深度神经网络对真实语音信号进行降噪。我们已经创建了自己的数据集,其中有超过150小时的俄语纯语音。使用尺度不变的信号失真比和语音质量的感知评价来估计所获得的结果,与最佳判别模型的结果相当或优于。
本文章由计算机程序翻译,如有差异,请以英文原文为准。

Speech Enhancement with Generative Diffusion Models

Speech Enhancement with Generative Diffusion Models

An alternative approach to speech denoising using generative diffusion models that model the distribution of training data is proposed. In recent years, such models have led to promising results to be obtained in the field of generating signals of various kinds, and these are superior in many ways to previous generative models, such as variational autoencoders. However, diffusion models have not yet found wide application in the field of speech denoising. A new diffusion model is presented, which can be used to denoise real speech signals using a deep neural network. Our own data set, with more than 150 h of pure speech in Russian, has been created. The obtained results, estimated using the metrics scale invariant signal to distortion ratio and perceptual evaluation of speech quality, are comparable or superior to the results of the best discriminative models.

求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS
AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS COMPUTER SCIENCE, INFORMATION SYSTEMS-
自引率
40.00%
发文量
18
期刊介绍: Automatic Documentation and Mathematical Linguistics  is an international peer reviewed journal that covers all aspects of automation of information processes and systems, as well as algorithms and methods for automatic language analysis. Emphasis is on the practical applications of new technologies and techniques for information analysis and processing.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信