面向知识交流的濒危语言计算机口语语料库构建研究

Zihui Xiao, Junjun Fan, Wei-Nan Gao
{"title":"面向知识交流的濒危语言计算机口语语料库构建研究","authors":"Zihui Xiao, Junjun Fan, Wei-Nan Gao","doi":"10.1145/3568739.3568801","DOIUrl":null,"url":null,"abstract":"This Corpus/corpora is a computer database that stores language materials. Corpus in the world is dominated by lingua franca corpus. However, these corpora have limited samples and are difficult to meet the needs of machine artificial intelligence. The paper proposes to build a Knowledge-Communication-Oriented spoken corpus for Endangered Languages to enrich the content of corpus research. A Knowledge-Communication-Oriented Spoken Corpus was divided into three sub-corporas, the words sub-corpu, the sentences sub-corpus, and the narrative discourses sub-corpus. This paper mainly introduces the methods of constructing the spoken corpus of endangered languages from the aspects of corpus collection, corpus arrangement and corpus annotation.","PeriodicalId":200698,"journal":{"name":"Proceedings of the 6th International Conference on Digital Technology in Education","volume":"41 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-09-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"On the Construction of Knowledge-Communication-Oriented Computer Spoken Corpus for Endangered Languages\",\"authors\":\"Zihui Xiao, Junjun Fan, Wei-Nan Gao\",\"doi\":\"10.1145/3568739.3568801\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This Corpus/corpora is a computer database that stores language materials. Corpus in the world is dominated by lingua franca corpus. However, these corpora have limited samples and are difficult to meet the needs of machine artificial intelligence. The paper proposes to build a Knowledge-Communication-Oriented spoken corpus for Endangered Languages to enrich the content of corpus research. A Knowledge-Communication-Oriented Spoken Corpus was divided into three sub-corporas, the words sub-corpu, the sentences sub-corpus, and the narrative discourses sub-corpus. This paper mainly introduces the methods of constructing the spoken corpus of endangered languages from the aspects of corpus collection, corpus arrangement and corpus annotation.\",\"PeriodicalId\":200698,\"journal\":{\"name\":\"Proceedings of the 6th International Conference on Digital Technology in Education\",\"volume\":\"41 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-09-16\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the 6th International Conference on Digital Technology in Education\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/3568739.3568801\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 6th International Conference on Digital Technology in Education","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3568739.3568801","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0

摘要

这个语料库是一个存储语言材料的计算机数据库。世界语料库以通用语语料库为主。然而,这些语料库样本有限,难以满足机器人工智能的需求。本文提出构建面向知识交流的濒危语言口语语料库,以丰富语料库研究的内容。知识交际型口语语料库分为三个子语料库:词子语料库、句子子语料库和叙事语子语料库。本文主要从语料库采集、语料库整理和语料库标注等方面介绍了构建濒危语言口语语料库的方法。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
On the Construction of Knowledge-Communication-Oriented Computer Spoken Corpus for Endangered Languages
This Corpus/corpora is a computer database that stores language materials. Corpus in the world is dominated by lingua franca corpus. However, these corpora have limited samples and are difficult to meet the needs of machine artificial intelligence. The paper proposes to build a Knowledge-Communication-Oriented spoken corpus for Endangered Languages to enrich the content of corpus research. A Knowledge-Communication-Oriented Spoken Corpus was divided into three sub-corporas, the words sub-corpu, the sentences sub-corpus, and the narrative discourses sub-corpus. This paper mainly introduces the methods of constructing the spoken corpus of endangered languages from the aspects of corpus collection, corpus arrangement and corpus annotation.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
自引率
0.00%
发文量
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信