Design of single-channel speech enhancement algorithm in noisy acoustic environments

IF 3 3区 计算机科学 Q2 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE
Computer Speech and Language Pub Date : 2026-10-01 Epub Date: 2026-02-12 DOI:10.1016/j.csl.2026.101955
Yi-Fu Zhao, Guang-Hui Dong, Nan Liu
{"title":"Design of single-channel speech enhancement algorithm in noisy acoustic environments","authors":"Yi-Fu Zhao,&nbsp;Guang-Hui Dong,&nbsp;Nan Liu","doi":"10.1016/j.csl.2026.101955","DOIUrl":null,"url":null,"abstract":"<div><div>In speech enhancement, Transformers and Self-Attention-based denoising networks are widely used and perform well, and speech enhancement serves as a valuable front-end for speech recognition. However, existing dual-branch architectures lack sufficient natural speech phase extraction due to the phase spectrum’s sensitivity and easy compensation, and traditional dilated convolution architectures are unsuitable for resource-constrained devices, creating an urgent need for lightweight alternatives. Thus, this paper proposes TFEM-PHASEN-MINI, a discrete dual-branch phase extraction architecture based on the Base and Detail Feature Modules. It uses DilatedReparamBlock to replace the Dense Encoder’s dilated convolution module, balancing computational efficiency and performance by fusing Convolutional Neural Networks and Transformers. It also designs a time-frequency feature extraction module to verify integrating speech recognition modules into speech enhancement, and adds a Phase Enhancement Module to address insufficient phase-spectrum speech phoneme feature extraction (caused by magnitude spectrum over-compensation) via parallel phase estimation. On the VoiceBank+DEMAND dataset, it achieves scores of 3.44, 4.72, 4.18, 17.13, 2.10, and 0.96 for PESQ, CSIG, COVL, FWSSNR, CEPS, and STOI, respectively. On the DNS-Challenge dataset, it attains scores of 3.20 and 3.57 for WB-PESQ and NB-PESQ, respectively. On the EARS-WHAM testset and its blind testset, it improves the metrics of PESQ, CSIG, CBAK, COVL, SSNR, FWSSNR, CEPS, and STOI by 0.56, 1.00, 0.94, 0.83, 8.42, 5.26, 0.21, and 0.15 respectively, and achieves non-intrusive metrics (Overall Quality of 3.80, Noisiness of 4.18, Discontinuity of 4.32, Coloration of 3.85, Loudness of 3.45), showing optimal generalization. Though it has relatively lower CBAK and SSNR on the VoiceBank+DEMAND dataset, it remains overall advanced. Computational complexity and device inference tests verify the balance between its computational efficiency and accuracy.</div></div>","PeriodicalId":50638,"journal":{"name":"Computer Speech and Language","volume":"100 ","pages":"Article 101955"},"PeriodicalIF":3.0000,"publicationDate":"2026-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Computer Speech and Language","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0885230826000185","RegionNum":3,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/2/12 0:00:00","PubModel":"Epub","JCR":"Q2","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0

Abstract

In speech enhancement, Transformers and Self-Attention-based denoising networks are widely used and perform well, and speech enhancement serves as a valuable front-end for speech recognition. However, existing dual-branch architectures lack sufficient natural speech phase extraction due to the phase spectrum’s sensitivity and easy compensation, and traditional dilated convolution architectures are unsuitable for resource-constrained devices, creating an urgent need for lightweight alternatives. Thus, this paper proposes TFEM-PHASEN-MINI, a discrete dual-branch phase extraction architecture based on the Base and Detail Feature Modules. It uses DilatedReparamBlock to replace the Dense Encoder’s dilated convolution module, balancing computational efficiency and performance by fusing Convolutional Neural Networks and Transformers. It also designs a time-frequency feature extraction module to verify integrating speech recognition modules into speech enhancement, and adds a Phase Enhancement Module to address insufficient phase-spectrum speech phoneme feature extraction (caused by magnitude spectrum over-compensation) via parallel phase estimation. On the VoiceBank+DEMAND dataset, it achieves scores of 3.44, 4.72, 4.18, 17.13, 2.10, and 0.96 for PESQ, CSIG, COVL, FWSSNR, CEPS, and STOI, respectively. On the DNS-Challenge dataset, it attains scores of 3.20 and 3.57 for WB-PESQ and NB-PESQ, respectively. On the EARS-WHAM testset and its blind testset, it improves the metrics of PESQ, CSIG, CBAK, COVL, SSNR, FWSSNR, CEPS, and STOI by 0.56, 1.00, 0.94, 0.83, 8.42, 5.26, 0.21, and 0.15 respectively, and achieves non-intrusive metrics (Overall Quality of 3.80, Noisiness of 4.18, Discontinuity of 4.32, Coloration of 3.85, Loudness of 3.45), showing optimal generalization. Though it has relatively lower CBAK and SSNR on the VoiceBank+DEMAND dataset, it remains overall advanced. Computational complexity and device inference tests verify the balance between its computational efficiency and accuracy.
噪声环境下单通道语音增强算法设计
在语音增强中,变压器和基于自注意的去噪网络得到了广泛的应用,并且表现良好,语音增强是语音识别的一个有价值的前端。然而,由于相位谱的灵敏度和易于补偿,现有的双分支架构缺乏足够的自然语音相位提取,传统的扩展卷积架构不适合资源受限的设备,因此迫切需要轻量级的替代方案。因此,本文提出了一种基于基本特征模块和细节特征模块的离散双支路相位提取体系结构TFEM-PHASEN-MINI。它使用DilatedReparamBlock来取代Dense Encoder的扩展卷积模块,通过融合卷积神经网络和变压器来平衡计算效率和性能。设计时频特征提取模块,验证将语音识别模块集成到语音增强中;增加相位增强模块,通过并行相位估计解决相谱语音音素特征提取不足(幅度谱过补偿)的问题。在VoiceBank+DEMAND数据集上,PESQ、CSIG、COVL、FWSSNR、CEPS和STOI的得分分别为3.44、4.72、4.18、17.13、2.10和0.96。在DNS-Challenge数据集上,WB-PESQ和NB-PESQ的得分分别为3.20和3.57。在ear - wham测试集及其盲测试集上,将PESQ、CSIG、CBAK、COVL、SSNR、FWSSNR、CEPS和STOI指标分别提高了0.56、1.00、0.94、0.83、8.42、5.26、0.21和0.15,实现了非侵入性指标(综合质量为3.80、噪声为4.18、不连续性为4.32、显色性为3.85、响度为3.45),呈现出最佳泛化效果。虽然它在VoiceBank+DEMAND数据集上的CBAK和sssnr相对较低,但总体上仍处于领先地位。计算复杂度和设备推理测试验证了其计算效率和精度之间的平衡。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Computer Speech and Language
Computer Speech and Language 工程技术-计算机:人工智能
CiteScore
11.30
自引率
4.70%
发文量
80
审稿时长
22.9 weeks
期刊介绍: Computer Speech & Language publishes reports of original research related to the recognition, understanding, production, coding and mining of speech and language. The speech and language sciences have a long history, but it is only relatively recently that large-scale implementation of and experimentation with complex models of speech and language processing has become feasible. Such research is often carried out somewhat separately by practitioners of artificial intelligence, computer science, electronic engineering, information retrieval, linguistics, phonetics, or psychology.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书