What makes this voice sound so bad? A multidimensional analysis of state-of-the-art text-to-speech systems

2012 IEEE Spoken Language Technology Workshop (SLT) Pub Date : 2012-12-01 DOI:10.1109/SLT.2012.6424229

Florian Hinterleitner, C. Norrenbrock, S. Möller, U. Heute

引用次数: 14

Abstract

This paper presents research on perceptual quality dimensions of synthetic speech. We generated 57 stimuli from 16/19 female/male German text-to-speech systems (TTS) and asked listeners to judge the perceptual distances between them in a sorting task. Through a subsequent multidimensional scaling algorithm, we extracted three dimensions. Via expert listening and a comparison to ratings gathered on 16 attribute scales, the three dimensions can be assigned to naturalness of voice, temporal distortions and calmness. These dimensions are discussed in detail and compared to the perceptual quality dimensions from previous multidimensional analyses. Moreover, the results are analyzed depending on the type of TTS system. The identified dimensions will be used in the future to build a dimension-based quality predictor for synthetic speech.

查看原文本刊更多论文

这声音怎么这么难听?最先进的文本到语音系统的多维分析

本文对合成语音的感知质量维度进行了研究。我们从16/19个女性/男性德语文本到语音系统(TTS)中产生57个刺激，并要求听众在分类任务中判断它们之间的感知距离。通过随后的多维缩放算法，我们提取了三个维度。通过专家的倾听，并与16个属性量表收集的评分进行比较，这三个维度可以被分配到声音的自然度、时间扭曲和冷静。对这些维度进行了详细的讨论，并与以前多维分析中的感知质量维度进行了比较。此外，根据TTS系统的类型对结果进行了分析。识别的维度将在未来用于构建基于维度的合成语音质量预测器。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2012 IEEE Spoken Language Technology Workshop (SLT)

自引率

0.00%

发文量