K. Trakulsuk, A. Suchato, P. Punyabukkana, C. Wutiwiwatchai
{"title":"基于几何模型的色调自然度感知预测","authors":"K. Trakulsuk, A. Suchato, P. Punyabukkana, C. Wutiwiwatchai","doi":"10.1109/JCSSE.2014.6841845","DOIUrl":null,"url":null,"abstract":"Naturalness is an important issue in the Text-To-Speech (TTS) system. To support arbitrarily defined pitch contours for any synthesized syllables, a TTS should be able to maintain the naturalness of the synthetic speech. This work proposed an automatic evaluation of pitch contours in order to determine the level of naturalness of synthesized syllables when perceived by human listeners. By analyzing results, tone perception experiments conducted on human listeners in this work, a syllable tone naturalness prediction model based on the midpoint and endpoint of the syllable's rhyme part was proposed. The model was then used for developing a tone naturalness prediction algorithm using geometric models of pitch contours. The evaluation of the tone naturalness prediction algorithm involved human listeners perceiving the naturalness of syllables with 45 pitch contour patterns, each of which with 2 repetitions. The proposed algorithm achieved approximately 80% consistency rate compared against human listeners' decisions on tone naturalness of the syllables.","PeriodicalId":331610,"journal":{"name":"2014 11th International Joint Conference on Computer Science and Software Engineering (JCSSE)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2014-05-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Prediction of tone naturalness perception using geometric model\",\"authors\":\"K. Trakulsuk, A. Suchato, P. Punyabukkana, C. Wutiwiwatchai\",\"doi\":\"10.1109/JCSSE.2014.6841845\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Naturalness is an important issue in the Text-To-Speech (TTS) system. To support arbitrarily defined pitch contours for any synthesized syllables, a TTS should be able to maintain the naturalness of the synthetic speech. This work proposed an automatic evaluation of pitch contours in order to determine the level of naturalness of synthesized syllables when perceived by human listeners. By analyzing results, tone perception experiments conducted on human listeners in this work, a syllable tone naturalness prediction model based on the midpoint and endpoint of the syllable's rhyme part was proposed. The model was then used for developing a tone naturalness prediction algorithm using geometric models of pitch contours. The evaluation of the tone naturalness prediction algorithm involved human listeners perceiving the naturalness of syllables with 45 pitch contour patterns, each of which with 2 repetitions. The proposed algorithm achieved approximately 80% consistency rate compared against human listeners' decisions on tone naturalness of the syllables.\",\"PeriodicalId\":331610,\"journal\":{\"name\":\"2014 11th International Joint Conference on Computer Science and Software Engineering (JCSSE)\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2014-05-14\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2014 11th International Joint Conference on Computer Science and Software Engineering (JCSSE)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/JCSSE.2014.6841845\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2014 11th International Joint Conference on Computer Science and Software Engineering (JCSSE)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/JCSSE.2014.6841845","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Prediction of tone naturalness perception using geometric model
Naturalness is an important issue in the Text-To-Speech (TTS) system. To support arbitrarily defined pitch contours for any synthesized syllables, a TTS should be able to maintain the naturalness of the synthetic speech. This work proposed an automatic evaluation of pitch contours in order to determine the level of naturalness of synthesized syllables when perceived by human listeners. By analyzing results, tone perception experiments conducted on human listeners in this work, a syllable tone naturalness prediction model based on the midpoint and endpoint of the syllable's rhyme part was proposed. The model was then used for developing a tone naturalness prediction algorithm using geometric models of pitch contours. The evaluation of the tone naturalness prediction algorithm involved human listeners perceiving the naturalness of syllables with 45 pitch contour patterns, each of which with 2 repetitions. The proposed algorithm achieved approximately 80% consistency rate compared against human listeners' decisions on tone naturalness of the syllables.