{"title":"Design of single-channel speech enhancement algorithm in noisy acoustic environments","authors":"Yi-Fu Zhao, Guang-Hui Dong, Nan Liu","doi":"10.1016/j.csl.2026.101955","DOIUrl":null,"url":null,"abstract":"<div><div>In speech enhancement, Transformers and Self-Attention-based denoising networks are widely used and perform well, and speech enhancement serves as a valuable front-end for speech recognition. However, existing dual-branch architectures lack sufficient natural speech phase extraction due to the phase spectrum’s sensitivity and easy compensation, and traditional dilated convolution architectures are unsuitable for resource-constrained devices, creating an urgent need for lightweight alternatives. Thus, this paper proposes TFEM-PHASEN-MINI, a discrete dual-branch phase extraction architecture based on the Base and Detail Feature Modules. It uses DilatedReparamBlock to replace the Dense Encoder’s dilated convolution module, balancing computational efficiency and performance by fusing Convolutional Neural Networks and Transformers. It also designs a time-frequency feature extraction module to verify integrating speech recognition modules into speech enhancement, and adds a Phase Enhancement Module to address insufficient phase-spectrum speech phoneme feature extraction (caused by magnitude spectrum over-compensation) via parallel phase estimation. On the VoiceBank+DEMAND dataset, it achieves scores of 3.44, 4.72, 4.18, 17.13, 2.10, and 0.96 for PESQ, CSIG, COVL, FWSSNR, CEPS, and STOI, respectively. On the DNS-Challenge dataset, it attains scores of 3.20 and 3.57 for WB-PESQ and NB-PESQ, respectively. On the EARS-WHAM testset and its blind testset, it improves the metrics of PESQ, CSIG, CBAK, COVL, SSNR, FWSSNR, CEPS, and STOI by 0.56, 1.00, 0.94, 0.83, 8.42, 5.26, 0.21, and 0.15 respectively, and achieves non-intrusive metrics (Overall Quality of 3.80, Noisiness of 4.18, Discontinuity of 4.32, Coloration of 3.85, Loudness of 3.45), showing optimal generalization. Though it has relatively lower CBAK and SSNR on the VoiceBank+DEMAND dataset, it remains overall advanced. Computational complexity and device inference tests verify the balance between its computational efficiency and accuracy.</div></div>","PeriodicalId":50638,"journal":{"name":"Computer Speech and Language","volume":"100 ","pages":"Article 101955"},"PeriodicalIF":3.0000,"publicationDate":"2026-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Computer Speech and Language","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0885230826000185","RegionNum":3,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/2/12 0:00:00","PubModel":"Epub","JCR":"Q2","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0
Abstract
In speech enhancement, Transformers and Self-Attention-based denoising networks are widely used and perform well, and speech enhancement serves as a valuable front-end for speech recognition. However, existing dual-branch architectures lack sufficient natural speech phase extraction due to the phase spectrum’s sensitivity and easy compensation, and traditional dilated convolution architectures are unsuitable for resource-constrained devices, creating an urgent need for lightweight alternatives. Thus, this paper proposes TFEM-PHASEN-MINI, a discrete dual-branch phase extraction architecture based on the Base and Detail Feature Modules. It uses DilatedReparamBlock to replace the Dense Encoder’s dilated convolution module, balancing computational efficiency and performance by fusing Convolutional Neural Networks and Transformers. It also designs a time-frequency feature extraction module to verify integrating speech recognition modules into speech enhancement, and adds a Phase Enhancement Module to address insufficient phase-spectrum speech phoneme feature extraction (caused by magnitude spectrum over-compensation) via parallel phase estimation. On the VoiceBank+DEMAND dataset, it achieves scores of 3.44, 4.72, 4.18, 17.13, 2.10, and 0.96 for PESQ, CSIG, COVL, FWSSNR, CEPS, and STOI, respectively. On the DNS-Challenge dataset, it attains scores of 3.20 and 3.57 for WB-PESQ and NB-PESQ, respectively. On the EARS-WHAM testset and its blind testset, it improves the metrics of PESQ, CSIG, CBAK, COVL, SSNR, FWSSNR, CEPS, and STOI by 0.56, 1.00, 0.94, 0.83, 8.42, 5.26, 0.21, and 0.15 respectively, and achieves non-intrusive metrics (Overall Quality of 3.80, Noisiness of 4.18, Discontinuity of 4.32, Coloration of 3.85, Loudness of 3.45), showing optimal generalization. Though it has relatively lower CBAK and SSNR on the VoiceBank+DEMAND dataset, it remains overall advanced. Computational complexity and device inference tests verify the balance between its computational efficiency and accuracy.
期刊介绍:
Computer Speech & Language publishes reports of original research related to the recognition, understanding, production, coding and mining of speech and language.
The speech and language sciences have a long history, but it is only relatively recently that large-scale implementation of and experimentation with complex models of speech and language processing has become feasible. Such research is often carried out somewhat separately by practitioners of artificial intelligence, computer science, electronic engineering, information retrieval, linguistics, phonetics, or psychology.