Deep learning-based speaking rate-dependent hierarchical prosodie model for Mandarin TTS

2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) Pub Date : 2017-12-01 DOI:10.1109/APSIPA.2017.8282228

Yen-Ting Lin, Chen-Yu Chiang

{"title":"Deep learning-based speaking rate-dependent hierarchical prosodie model for Mandarin TTS","authors":"Yen-Ting Lin, Chen-Yu Chiang","doi":"10.1109/APSIPA.2017.8282228","DOIUrl":null,"url":null,"abstract":"Speaking Rate-dependent Hierarchical Prosodie Model (SR-HPM) is a syllable-based statistical prosodie model and has been successfully served as a prosody generation model in a speaking rate-controlled text-to-speech system for Mandarin, and two Chinese dialects: Taiwan Min and Si-Xian Hakka. Excited by the success of utilizing deep learning (DL) techniques in parametric speech synthesis based on the HMM-based speech synthesis system, this study aims to improve the performance of the SR-HPM in prosody generation by replacing the conventional cascaded statistical sub-models with DL-based models, i.e. the DL-based SR-HPM. Each of the sub-model is first independently realized by a specially designed DL-based model based on its input-output characteristics. Then, all sub-models are cascaded and unified as one deep neural structure with their parameters being obtained by an end-to-end (linguistic feature-to-prosodic acoustic feature) optimization manner. The subjective and objective tests show that the DL-based SR-HPM performs better than the conventional statistical SR-HPM in prosody generation.","PeriodicalId":142091,"journal":{"name":"2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)","volume":"6 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2017-12-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/APSIPA.2017.8282228","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Speaking Rate-dependent Hierarchical Prosodie Model (SR-HPM) is a syllable-based statistical prosodie model and has been successfully served as a prosody generation model in a speaking rate-controlled text-to-speech system for Mandarin, and two Chinese dialects: Taiwan Min and Si-Xian Hakka. Excited by the success of utilizing deep learning (DL) techniques in parametric speech synthesis based on the HMM-based speech synthesis system, this study aims to improve the performance of the SR-HPM in prosody generation by replacing the conventional cascaded statistical sub-models with DL-based models, i.e. the DL-based SR-HPM. Each of the sub-model is first independently realized by a specially designed DL-based model based on its input-output characteristics. Then, all sub-models are cascaded and unified as one deep neural structure with their parameters being obtained by an end-to-end (linguistic feature-to-prosodic acoustic feature) optimization manner. The subjective and objective tests show that the DL-based SR-HPM performs better than the conventional statistical SR-HPM in prosody generation.

查看原文本刊更多论文

基于深度学习的汉语TTS语速分层韵律模型

基于语速的分层韵律模型(SR-HPM)是一种基于音节的韵律统计模型，已成功应用于普通话、台湾闽话和泗县客家方言的语速控制文本-语音系统中。基于深度学习技术在基于hmm的参数化语音合成系统中的成功应用，本研究旨在用基于DL的模型(即基于DL的SR-HPM)取代传统的级联统计子模型，从而提高SR-HPM在韵律生成方面的性能。每个子模型首先由一个专门设计的基于dl的模型根据其输入输出特性独立实现。然后，将所有子模型级联统一为一个深度神经结构，并通过端到端(语言特征到韵律声学特征)优化方式获得其参数。主观和客观测试表明，基于dl的SR-HPM在韵律生成方面优于传统的统计SR-HPM。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)

自引率

0.00%

发文量