{"title":"Mix-MaxETTS: A text-to-emotional speech synthesis model based on a deep encoder–decoder structure for the transfer of secondary emotions","authors":"Seyyed Mahdi Hassani, Mohammad Reza Kangavari","doi":"10.4218/etrij.2025-0058","DOIUrl":null,"url":null,"abstract":"<p>Given the importance of emotions in social interactions, emotional speech synthesis has attracted significant attention in the field of human–computer interaction. Remarkable advancements have been made in emotional text-to-speech synthesis, but most previous studies have concentrated on imitating styles associated with a specific primary emotion, neglecting secondary emotions that arise from mixtures of primary emotions. Therefore, there is a need to leverage both primary and secondary emotions in speech synthesis to facilitate more engaging, realistic, and natural interactions among artificial social agents. To address this gap, we propose a text-to-emotional speech synthesis model designed to generate nuanced mixtures of emotions that effectively convey secondary emotions during interactions. By adjusting the values of each basic emotion, we can control the mix of emotions in the synthetic speech. Our proposed method distinguishes between primary emotions and variations in mixed emotions while learning emotional styles. The effectiveness of the proposed framework was validated through both objective and subjective evaluations.</p>","PeriodicalId":11901,"journal":{"name":"ETRI Journal","volume":"48 4","pages":"693-710"},"PeriodicalIF":2.0000,"publicationDate":"2026-08-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.4218/etrij.2025-0058","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"ETRI Journal","FirstCategoryId":"94","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.4218/etrij.2025-0058","RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/11/26 0:00:00","PubModel":"Epub","JCR":"Q3","JCRName":"ENGINEERING, ELECTRICAL & ELECTRONIC","Score":null,"Total":0}
引用次数: 0
Abstract
Given the importance of emotions in social interactions, emotional speech synthesis has attracted significant attention in the field of human–computer interaction. Remarkable advancements have been made in emotional text-to-speech synthesis, but most previous studies have concentrated on imitating styles associated with a specific primary emotion, neglecting secondary emotions that arise from mixtures of primary emotions. Therefore, there is a need to leverage both primary and secondary emotions in speech synthesis to facilitate more engaging, realistic, and natural interactions among artificial social agents. To address this gap, we propose a text-to-emotional speech synthesis model designed to generate nuanced mixtures of emotions that effectively convey secondary emotions during interactions. By adjusting the values of each basic emotion, we can control the mix of emotions in the synthetic speech. Our proposed method distinguishes between primary emotions and variations in mixed emotions while learning emotional styles. The effectiveness of the proposed framework was validated through both objective and subjective evaluations.
期刊介绍:
ETRI Journal is an international, peer-reviewed multidisciplinary journal published bimonthly in English. The main focus of the journal is to provide an open forum to exchange innovative ideas and technology in the fields of information, telecommunications, and electronics.
Key topics of interest include high-performance computing, big data analytics, cloud computing, multimedia technology, communication networks and services, wireless communications and mobile computing, material and component technology, as well as security.
With an international editorial committee and experts from around the world as reviewers, ETRI Journal publishes high-quality research papers on the latest and best developments from the global community.