数据驱动的俄语网络爬行语料库拉丁短语识别方法

V. Benko, K. Rausova
{"title":"数据驱动的俄语网络爬行语料库拉丁短语识别方法","authors":"V. Benko, K. Rausova","doi":"10.17586/2541-9781-2020-4-11-20","DOIUrl":null,"url":null,"abstract":"Latin phrases are an integral part of the language of educated speakers in many (European) languages. Besides lexical units of Latin origin that have been already adapted to the orthography of the respective host language and calques, phrases retaining the original form and orthography can also be found in many texts. Due to the rather low frequency of the phenomenon, however, any systematic attempt of its analysis was a real challenge before the advent of very large (multi-Gigaword) corpora. Our paper presents a method of semi-automatic detection of Latin phrases in a Russian web corpus based on applying a Latin tagger and a series of filtrations performed by standard Linux utilities. The preliminary analysis of the resulting candidate list is shown in the concluding part of the paper.","PeriodicalId":226779,"journal":{"name":"Intelligent Memory Systems","volume":"30 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Data-Driven Approach to Identification of Latin Phrases in Russian Web-Crawled Corpora\",\"authors\":\"V. Benko, K. Rausova\",\"doi\":\"10.17586/2541-9781-2020-4-11-20\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Latin phrases are an integral part of the language of educated speakers in many (European) languages. Besides lexical units of Latin origin that have been already adapted to the orthography of the respective host language and calques, phrases retaining the original form and orthography can also be found in many texts. Due to the rather low frequency of the phenomenon, however, any systematic attempt of its analysis was a real challenge before the advent of very large (multi-Gigaword) corpora. Our paper presents a method of semi-automatic detection of Latin phrases in a Russian web corpus based on applying a Latin tagger and a series of filtrations performed by standard Linux utilities. The preliminary analysis of the resulting candidate list is shown in the concluding part of the paper.\",\"PeriodicalId\":226779,\"journal\":{\"name\":\"Intelligent Memory Systems\",\"volume\":\"30 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Intelligent Memory Systems\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.17586/2541-9781-2020-4-11-20\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Intelligent Memory Systems","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.17586/2541-9781-2020-4-11-20","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0

摘要

在许多(欧洲)语言中,拉丁短语是受过教育的人语言的一个组成部分。除了拉丁起源的词汇单位已经适应了各自的宿主语言和方言的正字法之外,在许多文本中也可以找到保留原始形式和正字法的短语。然而,由于这种现象的频率相当低,在超大型(多千兆字)语料库出现之前,任何对其进行系统分析的尝试都是一个真正的挑战。本文提出了一种半自动检测俄语网络语料库中的拉丁短语的方法,该方法基于一个拉丁标注器和一系列由标准Linux实用程序执行的过滤。本文的结语部分对最终候选名单进行了初步分析。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
Data-Driven Approach to Identification of Latin Phrases in Russian Web-Crawled Corpora
Latin phrases are an integral part of the language of educated speakers in many (European) languages. Besides lexical units of Latin origin that have been already adapted to the orthography of the respective host language and calques, phrases retaining the original form and orthography can also be found in many texts. Due to the rather low frequency of the phenomenon, however, any systematic attempt of its analysis was a real challenge before the advent of very large (multi-Gigaword) corpora. Our paper presents a method of semi-automatic detection of Latin phrases in a Russian web corpus based on applying a Latin tagger and a series of filtrations performed by standard Linux utilities. The preliminary analysis of the resulting candidate list is shown in the concluding part of the paper.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
自引率
0.00%
发文量
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信