基于生成扩散模型的语音增强

IF 0.5 Q4 COMPUTER SCIENCE, INFORMATION SYSTEMS

AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS Pub Date : 2023-11-24 DOI:10.3103/S0005105523050035

O. V. Girfanov, A. G. Shishkin

{"title":"基于生成扩散模型的语音增强","authors":"O. V. Girfanov, A. G. Shishkin","doi":"10.3103/S0005105523050035","DOIUrl":null,"url":null,"abstract":"<p>An alternative approach to speech denoising using generative diffusion models that model the distribution of training data is proposed. In recent years, such models have led to promising results to be obtained in the field of generating signals of various kinds, and these are superior in many ways to previous generative models, such as variational autoencoders. However, diffusion models have not yet found wide application in the field of speech denoising. A new diffusion model is presented, which can be used to denoise real speech signals using a deep neural network. Our own data set, with more than 150 h of pure speech in Russian, has been created. The obtained results, estimated using the metrics scale invariant signal to distortion ratio and perceptual evaluation of speech quality, are comparable or superior to the results of the best discriminative models.</p>","PeriodicalId":42995,"journal":{"name":"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS","volume":"57 5","pages":"249 - 257"},"PeriodicalIF":0.5000,"publicationDate":"2023-11-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Speech Enhancement with Generative Diffusion Models\",\"authors\":\"O. V. Girfanov, A. G. Shishkin\",\"doi\":\"10.3103/S0005105523050035\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<p>An alternative approach to speech denoising using generative diffusion models that model the distribution of training data is proposed. In recent years, such models have led to promising results to be obtained in the field of generating signals of various kinds, and these are superior in many ways to previous generative models, such as variational autoencoders. However, diffusion models have not yet found wide application in the field of speech denoising. A new diffusion model is presented, which can be used to denoise real speech signals using a deep neural network. Our own data set, with more than 150 h of pure speech in Russian, has been created. The obtained results, estimated using the metrics scale invariant signal to distortion ratio and perceptual evaluation of speech quality, are comparable or superior to the results of the best discriminative models.</p>\",\"PeriodicalId\":42995,\"journal\":{\"name\":\"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS\",\"volume\":\"57 5\",\"pages\":\"249 - 257\"},\"PeriodicalIF\":0.5000,\"publicationDate\":\"2023-11-24\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://link.springer.com/article/10.3103/S0005105523050035\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q4\",\"JCRName\":\"COMPUTER SCIENCE, INFORMATION SYSTEMS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS","FirstCategoryId":"1085","ListUrlMain":"https://link.springer.com/article/10.3103/S0005105523050035","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q4","JCRName":"COMPUTER SCIENCE, INFORMATION SYSTEMS","Score":null,"Total":0}

引用次数: 0

摘要

提出了一种使用生成扩散模型对训练数据的分布进行建模的语音去噪方法。近年来，这种模型在生成各种类型的信号方面取得了可喜的成果，并且在许多方面优于以前的生成模型，如变分自编码器。然而，扩散模型在语音去噪领域还没有得到广泛的应用。提出了一种新的扩散模型，该模型可以利用深度神经网络对真实语音信号进行降噪。我们已经创建了自己的数据集，其中有超过150小时的俄语纯语音。使用尺度不变的信号失真比和语音质量的感知评价来估计所获得的结果，与最佳判别模型的结果相当或优于。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

Speech Enhancement with Generative Diffusion Models

查看原文本刊更多论文

Speech Enhancement with Generative Diffusion Models

An alternative approach to speech denoising using generative diffusion models that model the distribution of training data is proposed. In recent years, such models have led to promising results to be obtained in the field of generating signals of various kinds, and these are superior in many ways to previous generative models, such as variational autoencoders. However, diffusion models have not yet found wide application in the field of speech denoising. A new diffusion model is presented, which can be used to denoise real speech signals using a deep neural network. Our own data set, with more than 150 h of pure speech in Russian, has been created. The obtained results, estimated using the metrics scale invariant signal to distortion ratio and perceptual evaluation of speech quality, are comparable or superior to the results of the best discriminative models.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS COMPUTER SCIENCE, INFORMATION SYSTEMS-

自引率

40.00%

发文量

期刊介绍： Automatic Documentation and Mathematical Linguistics is an international peer reviewed journal that covers all aspects of automation of information processes and systems, as well as algorithms and methods for automatic language analysis. Emphasis is on the practical applications of new technologies and techniques for information analysis and processing.