基于语速的汉语方言TTS分层韵律模型的自适应研究

2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE) Pub Date : 2015-10-01 DOI:10.1109/ICSDA.2015.7357862

Chen-Yu Chiang

{"title":"基于语速的汉语方言TTS分层韵律模型的自适应研究","authors":"Chen-Yu Chiang","doi":"10.1109/ICSDA.2015.7357862","DOIUrl":null,"url":null,"abstract":"This paper presents a new approach to developing a speaking rate (SR)-dependent hierarchical prosodic model (SR-HPM) to be utilized in a SR-controlled TTS for Taiwanese (Min-Nan) language, a resource-limited Chinese dialect. The main issue is to conquer the difficulty of building the SR-HPM directly from a Taiwanese database with sparse coverage of linguistic context, prosody and SR. By using the property that Taiwanese and Mandarin Chinese share the same linguistic characteristics, we propose an adaptation approach to constructing Taiwanese SR-HPM from a small Taiwanese corpus of fast SR with the help of an existing Mandarin SRHPM which is well-trained from a large Mandarin corpus with utterances covering a wide range of SR. The proposed method includes two parts: adaptation of normalization functions (NFs) and adaptive prosody labeling and modeling algorithm (PLM). Both of these two parts are formulated based on MAP estimations with the existing Mandarin SR-HPM serving as an informative prior. Effectiveness of the proposed approach was evaluated by an experiment of prosody generation for Taiwanese TTS using a small corpus of fast speech with SR in 4.5-6.8 syllables/sec. Experimental results showed that the generated prosody sounded quite natural for SR in a wide range of 3.4-6.8 syllables/sec.","PeriodicalId":290790,"journal":{"name":"2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE)","volume":"100 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2015-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":"{\"title\":\"A study on adaptation of speaking rate-dependent hierarchical prosodic model for Chinese dialect TTS\",\"authors\":\"Chen-Yu Chiang\",\"doi\":\"10.1109/ICSDA.2015.7357862\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper presents a new approach to developing a speaking rate (SR)-dependent hierarchical prosodic model (SR-HPM) to be utilized in a SR-controlled TTS for Taiwanese (Min-Nan) language, a resource-limited Chinese dialect. The main issue is to conquer the difficulty of building the SR-HPM directly from a Taiwanese database with sparse coverage of linguistic context, prosody and SR. By using the property that Taiwanese and Mandarin Chinese share the same linguistic characteristics, we propose an adaptation approach to constructing Taiwanese SR-HPM from a small Taiwanese corpus of fast SR with the help of an existing Mandarin SRHPM which is well-trained from a large Mandarin corpus with utterances covering a wide range of SR. The proposed method includes two parts: adaptation of normalization functions (NFs) and adaptive prosody labeling and modeling algorithm (PLM). Both of these two parts are formulated based on MAP estimations with the existing Mandarin SR-HPM serving as an informative prior. Effectiveness of the proposed approach was evaluated by an experiment of prosody generation for Taiwanese TTS using a small corpus of fast speech with SR in 4.5-6.8 syllables/sec. Experimental results showed that the generated prosody sounded quite natural for SR in a wide range of 3.4-6.8 syllables/sec.\",\"PeriodicalId\":290790,\"journal\":{\"name\":\"2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE)\",\"volume\":\"100 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2015-10-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"3\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICSDA.2015.7357862\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICSDA.2015.7357862","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 3

摘要

本文提出了一种基于语速的分层韵律模型(SR- hpm)，并将其应用于资源有限的闽南语的分层韵律控制语音系统中。主要问题是克服直接从语言上下文、韵律和sr的稀疏覆盖的台湾数据库中构建SR-HPM的困难。利用台语和普通话具有相同的语言特征，本文提出了一种自适应方法，利用已有的普通话语料库，从一个小的台湾语料库中构建台湾SR- hpm。该方法包括两个部分:自适应归一化函数(NFs)和自适应韵律标记和建模算法(PLM)。这两个部分都是基于MAP估计而制定的，现有的普通话SR-HPM作为信息先验。本文使用一个小语料库对台湾TTS进行韵律生成实验，该语料库包含4.5-6.8个音节/秒的快速语音。实验结果表明，生成的韵律在3.4-6.8个音节/秒的大范围内听起来非常自然。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

A study on adaptation of speaking rate-dependent hierarchical prosodic model for Chinese dialect TTS

This paper presents a new approach to developing a speaking rate (SR)-dependent hierarchical prosodic model (SR-HPM) to be utilized in a SR-controlled TTS for Taiwanese (Min-Nan) language, a resource-limited Chinese dialect. The main issue is to conquer the difficulty of building the SR-HPM directly from a Taiwanese database with sparse coverage of linguistic context, prosody and SR. By using the property that Taiwanese and Mandarin Chinese share the same linguistic characteristics, we propose an adaptation approach to constructing Taiwanese SR-HPM from a small Taiwanese corpus of fast SR with the help of an existing Mandarin SRHPM which is well-trained from a large Mandarin corpus with utterances covering a wide range of SR. The proposed method includes two parts: adaptation of normalization functions (NFs) and adaptive prosody labeling and modeling algorithm (PLM). Both of these two parts are formulated based on MAP estimations with the existing Mandarin SR-HPM serving as an informative prior. Effectiveness of the proposed approach was evaluated by an experiment of prosody generation for Taiwanese TTS using a small corpus of fast speech with SR in 4.5-6.8 syllables/sec. Experimental results showed that the generated prosody sounded quite natural for SR in a wide range of 3.4-6.8 syllables/sec.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE)

自引率

0.00%

发文量