A corpus of audio-visual recordings of linguistically balanced, Danish sentences for speech-in-noise experiments

IF 2.4 3区计算机科学 Q2 ACOUSTICS

Speech Communication Pub Date : 2024-10-09 DOI:10.1016/j.specom.2024.103141

Abigail Anne Kressner , Kirsten Maria Jensen-Rico , Johannes Kizach , Brian Kai Loong Man , Anja Kofoed Pedersen , Lars Bramsløw , Lise Bruun Hansen , Laura Winther Balling , Brent Kirkwood , Tobias May

{"title":"A corpus of audio-visual recordings of linguistically balanced, Danish sentences for speech-in-noise experiments","authors":"Abigail Anne Kressner , Kirsten Maria Jensen-Rico , Johannes Kizach , Brian Kai Loong Man , Anja Kofoed Pedersen , Lars Bramsløw , Lise Bruun Hansen , Laura Winther Balling , Brent Kirkwood , Tobias May","doi":"10.1016/j.specom.2024.103141","DOIUrl":null,"url":null,"abstract":"<div><div>A typical speech-in-noise experiment in a research and development setting can easily contain as many as 20 conditions, or even more, and often requires at least two test points per condition. A sentence test with enough sentences to make this amount of testing possible without repetition does not yet exist in Danish. Thus, a new corpus has been developed to facilitate the creation of a sentence test that is large enough to address this need. The corpus itself is made up of audio and audio-visual recordings of 1200 linguistically balanced sentences, all of which are spoken by two female and two male talkers. The sentences were constructed using a novel, template-based method that facilitated control over both word frequency and sentence structure. The sentences were evaluated linguistically in terms of phonemic distributions, naturalness, and connotation, and thereafter, recorded, postprocessed, and rated on their audio, visual, and pronunciation qualities. This paper describes in detail the methodology employed to create and characterize this corpus.</div></div>","PeriodicalId":49485,"journal":{"name":"Speech Communication","volume":"165 ","pages":"Article 103141"},"PeriodicalIF":2.4000,"publicationDate":"2024-10-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Speech Communication","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0167639324001122","RegionNum":3,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"ACOUSTICS","Score":null,"Total":0}

引用次数: 0

Abstract

A typical speech-in-noise experiment in a research and development setting can easily contain as many as 20 conditions, or even more, and often requires at least two test points per condition. A sentence test with enough sentences to make this amount of testing possible without repetition does not yet exist in Danish. Thus, a new corpus has been developed to facilitate the creation of a sentence test that is large enough to address this need. The corpus itself is made up of audio and audio-visual recordings of 1200 linguistically balanced sentences, all of which are spoken by two female and two male talkers. The sentences were constructed using a novel, template-based method that facilitated control over both word frequency and sentence structure. The sentences were evaluated linguistically in terms of phonemic distributions, naturalness, and connotation, and thereafter, recorded, postprocessed, and rated on their audio, visual, and pronunciation qualities. This paper describes in detail the methodology employed to create and characterize this corpus.

查看原文本刊更多论文

用于噪声语音实验的语言均衡的丹麦语句子视听记录语料库

在研发环境中，一个典型的噪声语音实验很容易包含多达 20 个条件，甚至更多，而且通常每个条件至少需要两个测试点。在丹麦语中，还没有一个句子测试包含足够多的句子，可以在不重复的情况下进行如此大量的测试。因此，我们开发了一个新的语料库，以帮助创建一个足够大的句子测试来满足这一需求。语料库本身由 1200 个语言平衡句子的音频和视听录音组成，所有句子均由两名女性和两名男性说话者说出。这些句子是用一种新颖的、基于模板的方法构建的，便于控制词频和句子结构。这些句子从音位分布、自然度和内涵等方面进行了语言学评估，然后进行录音、后处理，并根据其音频、视觉和发音质量进行评分。本文详细介绍了创建和描述该语料库的方法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Speech Communication 工程技术-计算机：跨学科应用

CiteScore

6.80

自引率

6.20%

发文量

审稿时长

19.2 weeks

期刊介绍： Speech Communication is an interdisciplinary journal whose primary objective is to fulfil the need for the rapid dissemination and thorough discussion of basic and applied research results. The journal''s primary objectives are: • to present a forum for the advancement of human and human-machine speech communication science; • to stimulate cross-fertilization between different fields of this domain; • to contribute towards the rapid and wide diffusion of scientifically sound contributions in this domain.