台湾建立的车内语音资料库之收集与注释

Int. J. Comput. Linguistics Chin. Lang. Process. Pub Date : 2005-07-01 DOI:10.30019/IJCLCLP.200507.0005

Hsien-Chang Wang, Chung-Hsien Yang, Jhing-Fa Wang, Chung-Hsien Wu, Jen-Tzung Chien

{"title":"台湾建立的车内语音资料库之收集与注释","authors":"Hsien-Chang Wang, Chung-Hsien Yang, Jhing-Fa Wang, Chung-Hsien Wu, Jen-Tzung Chien","doi":"10.30019/IJCLCLP.200507.0005","DOIUrl":null,"url":null,"abstract":"This paper describes a project that aims to create a Mandarin speech database for the automobile setting (TAICAR). A group of researchers from several universities and research institutes in Taiwan have participated in the project. The goal is to generate a corpus for the development and testing of various speech-processing techniques. There are six recording sites in this project. Various words, sentences, and spontaneously queries uttered in the vehicular navigation setting have been collected in this project. A preliminary corpus of utterances from 192 speakers was created from utterances generated in different vehicles. The database contains more than 163,000 files, occupying 16.8 gigabytes of disk space.","PeriodicalId":436300,"journal":{"name":"Int. J. Comput. Linguistics Chin. Lang. Process.","volume":"47 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2005-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"8","resultStr":"{\"title\":\"TAICAR-The Collection and Annotation of an In-Car Speech Database Created in Taiwan\",\"authors\":\"Hsien-Chang Wang, Chung-Hsien Yang, Jhing-Fa Wang, Chung-Hsien Wu, Jen-Tzung Chien\",\"doi\":\"10.30019/IJCLCLP.200507.0005\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper describes a project that aims to create a Mandarin speech database for the automobile setting (TAICAR). A group of researchers from several universities and research institutes in Taiwan have participated in the project. The goal is to generate a corpus for the development and testing of various speech-processing techniques. There are six recording sites in this project. Various words, sentences, and spontaneously queries uttered in the vehicular navigation setting have been collected in this project. A preliminary corpus of utterances from 192 speakers was created from utterances generated in different vehicles. The database contains more than 163,000 files, occupying 16.8 gigabytes of disk space.\",\"PeriodicalId\":436300,\"journal\":{\"name\":\"Int. J. Comput. Linguistics Chin. Lang. Process.\",\"volume\":\"47 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2005-07-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"8\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Int. J. Comput. Linguistics Chin. Lang. Process.\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.30019/IJCLCLP.200507.0005\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Int. J. Comput. Linguistics Chin. Lang. Process.","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.30019/IJCLCLP.200507.0005","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 8

摘要

本文描述了一个旨在为汽车设置(TAICAR)创建普通话语音数据库的项目。来自台湾几所大学和研究机构的一组研究人员参与了该项目。目标是生成一个语料库，用于开发和测试各种语音处理技术。在这个项目中有六个录音地点。在这个项目中收集了车辆导航设置中发出的各种单词、句子和自发查询。从不同的载体中产生的话语创建了来自192个说话人的初步语料库。该数据库包含超过16.3万个文件，占用16.8 gb的磁盘空间。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

TAICAR-The Collection and Annotation of an In-Car Speech Database Created in Taiwan

This paper describes a project that aims to create a Mandarin speech database for the automobile setting (TAICAR). A group of researchers from several universities and research institutes in Taiwan have participated in the project. The goal is to generate a corpus for the development and testing of various speech-processing techniques. There are six recording sites in this project. Various words, sentences, and spontaneously queries uttered in the vehicular navigation setting have been collected in this project. A preliminary corpus of utterances from 192 speakers was created from utterances generated in different vehicles. The database contains more than 163,000 files, occupying 16.8 gigabytes of disk space.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Int. J. Comput. Linguistics Chin. Lang. Process.

自引率

0.00%

发文量