基于谱图图像的说话人识别

2021 IEEE International Conference on Smart Information Systems and Technologies (SIST) Pub Date : 2021-04-28 DOI:10.1109/SIST50301.2021.9465954

S. Kadyrov, Cemil Turan, Altynbek Amirzhanov, Cemal Ozdemir

{"title":"基于谱图图像的说话人识别","authors":"S. Kadyrov, Cemil Turan, Altynbek Amirzhanov, Cemal Ozdemir","doi":"10.1109/SIST50301.2021.9465954","DOIUrl":null,"url":null,"abstract":"Speaker identification is used to identify the owner of the voice among many people based on the uniqueness of everyone’s speech style. In this paper, we combine Convolutional Neural Network with Recurrent Neural Network using Long Short-Term Memory models for speaker recognition and implement the deep learning architecture on our dataset of spectrogram images for 77 different non-native speakers reading the same texts in Turkish. Usage of identical text reading eliminates the possible variations and diversities on spectrograms depending on vocabularies. Experiments show that the used method is very effective on recognition rate with satisfying performance and over 98% accuracy.","PeriodicalId":318915,"journal":{"name":"2021 IEEE International Conference on Smart Information Systems and Technologies (SIST)","volume":"10 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-04-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"6","resultStr":"{\"title\":\"Speaker Recognition from Spectrogram Images\",\"authors\":\"S. Kadyrov, Cemil Turan, Altynbek Amirzhanov, Cemal Ozdemir\",\"doi\":\"10.1109/SIST50301.2021.9465954\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Speaker identification is used to identify the owner of the voice among many people based on the uniqueness of everyone’s speech style. In this paper, we combine Convolutional Neural Network with Recurrent Neural Network using Long Short-Term Memory models for speaker recognition and implement the deep learning architecture on our dataset of spectrogram images for 77 different non-native speakers reading the same texts in Turkish. Usage of identical text reading eliminates the possible variations and diversities on spectrograms depending on vocabularies. Experiments show that the used method is very effective on recognition rate with satisfying performance and over 98% accuracy.\",\"PeriodicalId\":318915,\"journal\":{\"name\":\"2021 IEEE International Conference on Smart Information Systems and Technologies (SIST)\",\"volume\":\"10 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2021-04-28\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"6\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2021 IEEE International Conference on Smart Information Systems and Technologies (SIST)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/SIST50301.2021.9465954\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2021 IEEE International Conference on Smart Information Systems and Technologies (SIST)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SIST50301.2021.9465954","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 6

摘要

说话人识别是根据每个人说话风格的独特性，在众多人群中识别声音的所有者。在本文中，我们将卷积神经网络和循环神经网络结合使用长短期记忆模型进行说话人识别，并在我们的频谱图图像数据集上实现深度学习架构，该数据集包含77个不同的非母语人士阅读相同的土耳其语文本。使用相同的文本阅读消除了谱图因词汇不同而可能出现的变化和多样性。实验表明，该方法具有良好的识别率，准确率达到98%以上。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Speaker Recognition from Spectrogram Images

Speaker identification is used to identify the owner of the voice among many people based on the uniqueness of everyone’s speech style. In this paper, we combine Convolutional Neural Network with Recurrent Neural Network using Long Short-Term Memory models for speaker recognition and implement the deep learning architecture on our dataset of spectrogram images for 77 different non-native speakers reading the same texts in Turkish. Usage of identical text reading eliminates the possible variations and diversities on spectrograms depending on vocabularies. Experiments show that the used method is very effective on recognition rate with satisfying performance and over 98% accuracy.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2021 IEEE International Conference on Smart Information Systems and Technologies (SIST)

自引率

0.00%

发文量