数据驱动的俄语网络爬行语料库拉丁短语识别方法

Intelligent Memory Systems Pub Date : 1900-01-01 DOI:10.17586/2541-9781-2020-4-11-20

V. Benko, K. Rausova

{"title":"数据驱动的俄语网络爬行语料库拉丁短语识别方法","authors":"V. Benko, K. Rausova","doi":"10.17586/2541-9781-2020-4-11-20","DOIUrl":null,"url":null,"abstract":"Latin phrases are an integral part of the language of educated speakers in many (European) languages. Besides lexical units of Latin origin that have been already adapted to the orthography of the respective host language and calques, phrases retaining the original form and orthography can also be found in many texts. Due to the rather low frequency of the phenomenon, however, any systematic attempt of its analysis was a real challenge before the advent of very large (multi-Gigaword) corpora. Our paper presents a method of semi-automatic detection of Latin phrases in a Russian web corpus based on applying a Latin tagger and a series of filtrations performed by standard Linux utilities. The preliminary analysis of the resulting candidate list is shown in the concluding part of the paper.","PeriodicalId":226779,"journal":{"name":"Intelligent Memory Systems","volume":"30 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Data-Driven Approach to Identification of Latin Phrases in Russian Web-Crawled Corpora\",\"authors\":\"V. Benko, K. Rausova\",\"doi\":\"10.17586/2541-9781-2020-4-11-20\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Latin phrases are an integral part of the language of educated speakers in many (European) languages. Besides lexical units of Latin origin that have been already adapted to the orthography of the respective host language and calques, phrases retaining the original form and orthography can also be found in many texts. Due to the rather low frequency of the phenomenon, however, any systematic attempt of its analysis was a real challenge before the advent of very large (multi-Gigaword) corpora. Our paper presents a method of semi-automatic detection of Latin phrases in a Russian web corpus based on applying a Latin tagger and a series of filtrations performed by standard Linux utilities. The preliminary analysis of the resulting candidate list is shown in the concluding part of the paper.\",\"PeriodicalId\":226779,\"journal\":{\"name\":\"Intelligent Memory Systems\",\"volume\":\"30 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Intelligent Memory Systems\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.17586/2541-9781-2020-4-11-20\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Intelligent Memory Systems","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.17586/2541-9781-2020-4-11-20","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

在许多(欧洲)语言中，拉丁短语是受过教育的人语言的一个组成部分。除了拉丁起源的词汇单位已经适应了各自的宿主语言和方言的正字法之外，在许多文本中也可以找到保留原始形式和正字法的短语。然而，由于这种现象的频率相当低，在超大型(多千兆字)语料库出现之前，任何对其进行系统分析的尝试都是一个真正的挑战。本文提出了一种半自动检测俄语网络语料库中的拉丁短语的方法，该方法基于一个拉丁标注器和一系列由标准Linux实用程序执行的过滤。本文的结语部分对最终候选名单进行了初步分析。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Data-Driven Approach to Identification of Latin Phrases in Russian Web-Crawled Corpora

Latin phrases are an integral part of the language of educated speakers in many (European) languages. Besides lexical units of Latin origin that have been already adapted to the orthography of the respective host language and calques, phrases retaining the original form and orthography can also be found in many texts. Due to the rather low frequency of the phenomenon, however, any systematic attempt of its analysis was a real challenge before the advent of very large (multi-Gigaword) corpora. Our paper presents a method of semi-automatic detection of Latin phrases in a Russian web corpus based on applying a Latin tagger and a series of filtrations performed by standard Linux utilities. The preliminary analysis of the resulting candidate list is shown in the concluding part of the paper.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Intelligent Memory Systems

自引率

0.00%

发文量