{"title":"Robust Speaker Verification Using Improved PNCC Based on GMM-UBM","authors":"Xin-Xing Jing, Bing Xiang, Haiyan Yang, Ping Zhou","doi":"10.14355/IJAPE.2015.04.003","DOIUrl":null,"url":null,"abstract":"Focused on the issue that the robustness of traditional Mel Frequency Cepstral Coefficient (MFCC) feature degrades drastically in speaker verification in noisy environments, a kind of suitable extraction method for low SNR environments based on Gaussian Mixture Model‐Universal Background Model (GMM‐UBM) and improved Power Normalized Cepstral Coefficient (PNCC) is proposed. First, the PNCC feature is extracted after the Voice Activity Detection (VAD), which uses long term analysis to remove the effect of background noise. Then, Cepstral Mean Variance Normalization (CMVN), Feature Warping and other methods are used to improve PNCC. Finally, GMM‐UBM‐MAP is set as the baseline system for speaker verification test with TIMIT speech database, the robustness of four different features (MFCC, GFCC, PNCC and improved PNCC) are analyzed and compared in different noisy conditions. The experimental results indicate that MFCC has achieved the highest recognition rate under the environment of clean speech. By mixing the test speech with sine noises, the improved PNCC is more robust against different low‐SNR noises than other original features and its Equal Error Rate (EER) reduce significantly in low‐SNR noise environments.","PeriodicalId":270763,"journal":{"name":"International Journal of Automation and Power Engineering","volume":"10 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"International Journal of Automation and Power Engineering","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.14355/IJAPE.2015.04.003","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 3
Abstract
Focused on the issue that the robustness of traditional Mel Frequency Cepstral Coefficient (MFCC) feature degrades drastically in speaker verification in noisy environments, a kind of suitable extraction method for low SNR environments based on Gaussian Mixture Model‐Universal Background Model (GMM‐UBM) and improved Power Normalized Cepstral Coefficient (PNCC) is proposed. First, the PNCC feature is extracted after the Voice Activity Detection (VAD), which uses long term analysis to remove the effect of background noise. Then, Cepstral Mean Variance Normalization (CMVN), Feature Warping and other methods are used to improve PNCC. Finally, GMM‐UBM‐MAP is set as the baseline system for speaker verification test with TIMIT speech database, the robustness of four different features (MFCC, GFCC, PNCC and improved PNCC) are analyzed and compared in different noisy conditions. The experimental results indicate that MFCC has achieved the highest recognition rate under the environment of clean speech. By mixing the test speech with sine noises, the improved PNCC is more robust against different low‐SNR noises than other original features and its Equal Error Rate (EER) reduce significantly in low‐SNR noise environments.