自发和配音语音的鲁棒时间对齐及其在自动对话替换中的应用

2010 18th European Signal Processing Conference Pub Date : 2010-08-23 DOI:10.5281/ZENODO.42077

Pieter Soens, W. Verhelst

{"title":"自发和配音语音的鲁棒时间对齐及其在自动对话替换中的应用","authors":"Pieter Soens, W. Verhelst","doi":"10.5281/ZENODO.42077","DOIUrl":null,"url":null,"abstract":"In this paper, we present a robust system for the temporal alignment of 2 renditions of the same speech utterance. The system operates in 2 steps: during analysis, the timing relationships between the speech segments of the utterance that serves as a timing reference and the corresponding speech segments in the replacement utterance are measured by means of a dedicated dynamic time warping algorithm. The obtained warping paths are then processed and used to synthesize a high-quality speech utterance that is time-aligned with the reference. Subjective audio-visual listening tests performed within the context of a difficult Automatic Dialogue Replacement task demonstrated that the proposed system achieves a significant improvement compared to the industry-standard benchmark, both in terms of achieved lip-synchronization accuracy as well as in overall sound quality of the synthesized utterances.","PeriodicalId":409817,"journal":{"name":"2010 18th European Signal Processing Conference","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2010-08-23","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":"{\"title\":\"Robust temporal alignment of spontaneous and dubbed speech and its application for automatic dialogue replacement\",\"authors\":\"Pieter Soens, W. Verhelst\",\"doi\":\"10.5281/ZENODO.42077\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this paper, we present a robust system for the temporal alignment of 2 renditions of the same speech utterance. The system operates in 2 steps: during analysis, the timing relationships between the speech segments of the utterance that serves as a timing reference and the corresponding speech segments in the replacement utterance are measured by means of a dedicated dynamic time warping algorithm. The obtained warping paths are then processed and used to synthesize a high-quality speech utterance that is time-aligned with the reference. Subjective audio-visual listening tests performed within the context of a difficult Automatic Dialogue Replacement task demonstrated that the proposed system achieves a significant improvement compared to the industry-standard benchmark, both in terms of achieved lip-synchronization accuracy as well as in overall sound quality of the synthesized utterances.\",\"PeriodicalId\":409817,\"journal\":{\"name\":\"2010 18th European Signal Processing Conference\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2010-08-23\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"2\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2010 18th European Signal Processing Conference\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.5281/ZENODO.42077\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2010 18th European Signal Processing Conference","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.5281/ZENODO.42077","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 2

摘要

在本文中，我们提出了一个鲁棒的系统，用于同一语音的两种再现的时间对齐。系统工作分为两个步骤:在分析时，通过专用的动态时间规整算法测量作为时间参考的话语的语音段与替代话语中相应语音段之间的时间关系。然后对获得的扭曲路径进行处理并用于合成与参考时间对齐的高质量语音。在一个困难的自动对话替换任务的背景下进行的主观视听听力测试表明，与行业标准基准相比，所提出的系统在实现的唇部同步精度和合成话语的整体音质方面都取得了显着改善。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Robust temporal alignment of spontaneous and dubbed speech and its application for automatic dialogue replacement

In this paper, we present a robust system for the temporal alignment of 2 renditions of the same speech utterance. The system operates in 2 steps: during analysis, the timing relationships between the speech segments of the utterance that serves as a timing reference and the corresponding speech segments in the replacement utterance are measured by means of a dedicated dynamic time warping algorithm. The obtained warping paths are then processed and used to synthesize a high-quality speech utterance that is time-aligned with the reference. Subjective audio-visual listening tests performed within the context of a difficult Automatic Dialogue Replacement task demonstrated that the proposed system achieves a significant improvement compared to the industry-standard benchmark, both in terms of achieved lip-synchronization accuracy as well as in overall sound quality of the synthesized utterances.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2010 18th European Signal Processing Conference

自引率

0.00%

发文量