Mix-MaxETTS: A text-to-emotional speech synthesis model based on a deep encoder–decoder structure for the transfer of secondary emotions

IF 2 4区 计算机科学 Q3 ENGINEERING, ELECTRICAL & ELECTRONIC
ETRI Journal Pub Date : 2026-08-18 Epub Date: 2025-11-26 DOI:10.4218/etrij.2025-0058
Seyyed Mahdi Hassani, Mohammad Reza Kangavari
{"title":"Mix-MaxETTS: A text-to-emotional speech synthesis model based on a deep encoder–decoder structure for the transfer of secondary emotions","authors":"Seyyed Mahdi Hassani,&nbsp;Mohammad Reza Kangavari","doi":"10.4218/etrij.2025-0058","DOIUrl":null,"url":null,"abstract":"<p>Given the importance of emotions in social interactions, emotional speech synthesis has attracted significant attention in the field of human–computer interaction. Remarkable advancements have been made in emotional text-to-speech synthesis, but most previous studies have concentrated on imitating styles associated with a specific primary emotion, neglecting secondary emotions that arise from mixtures of primary emotions. Therefore, there is a need to leverage both primary and secondary emotions in speech synthesis to facilitate more engaging, realistic, and natural interactions among artificial social agents. To address this gap, we propose a text-to-emotional speech synthesis model designed to generate nuanced mixtures of emotions that effectively convey secondary emotions during interactions. By adjusting the values of each basic emotion, we can control the mix of emotions in the synthetic speech. Our proposed method distinguishes between primary emotions and variations in mixed emotions while learning emotional styles. The effectiveness of the proposed framework was validated through both objective and subjective evaluations.</p>","PeriodicalId":11901,"journal":{"name":"ETRI Journal","volume":"48 4","pages":"693-710"},"PeriodicalIF":2.0000,"publicationDate":"2026-08-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.4218/etrij.2025-0058","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"ETRI Journal","FirstCategoryId":"94","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.4218/etrij.2025-0058","RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/11/26 0:00:00","PubModel":"Epub","JCR":"Q3","JCRName":"ENGINEERING, ELECTRICAL & ELECTRONIC","Score":null,"Total":0}
引用次数: 0

Abstract

Given the importance of emotions in social interactions, emotional speech synthesis has attracted significant attention in the field of human–computer interaction. Remarkable advancements have been made in emotional text-to-speech synthesis, but most previous studies have concentrated on imitating styles associated with a specific primary emotion, neglecting secondary emotions that arise from mixtures of primary emotions. Therefore, there is a need to leverage both primary and secondary emotions in speech synthesis to facilitate more engaging, realistic, and natural interactions among artificial social agents. To address this gap, we propose a text-to-emotional speech synthesis model designed to generate nuanced mixtures of emotions that effectively convey secondary emotions during interactions. By adjusting the values of each basic emotion, we can control the mix of emotions in the synthetic speech. Our proposed method distinguishes between primary emotions and variations in mixed emotions while learning emotional styles. The effectiveness of the proposed framework was validated through both objective and subjective evaluations.

Mix-MaxETTS:一种基于深度编码器-解码器结构的文本-情感语音合成模型,用于次要情感的传递
鉴于情感在社会互动中的重要性,情感语音合成在人机交互领域受到了广泛关注。在情感文本到语音合成方面已经取得了显著的进步,但大多数先前的研究都集中在模仿与特定主要情绪相关的风格上,而忽略了由主要情绪混合产生的次要情绪。因此,有必要在语音合成中利用主要和次要情感,以促进人工社会代理之间更有吸引力、更现实、更自然的互动。为了解决这一差距,我们提出了一种文本到情感的语音合成模型,旨在生成微妙的情感混合,在交互过程中有效地传达次要情感。通过调整每个基本情绪的值,我们可以控制合成语音中的情绪混合。我们提出的方法在学习情绪风格时区分主要情绪和混合情绪的变化。通过客观和主观评价验证了所提出框架的有效性。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
ETRI Journal
ETRI Journal 工程技术-电信学
CiteScore
4.00
自引率
7.10%
发文量
98
审稿时长
6.9 months
期刊介绍: ETRI Journal is an international, peer-reviewed multidisciplinary journal published bimonthly in English. The main focus of the journal is to provide an open forum to exchange innovative ideas and technology in the fields of information, telecommunications, and electronics. Key topics of interest include high-performance computing, big data analytics, cloud computing, multimedia technology, communication networks and services, wireless communications and mobile computing, material and component technology, as well as security. With an international editorial committee and experts from around the world as reviewers, ETRI Journal publishes high-quality research papers on the latest and best developments from the global community.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书