Using transitivity information for morphological and syntactic disambiguation of pronouns in Ukrainian

Vìsnik Nacìonalʹnogo unìversitetu "Lʹvìvsʹka polìtehnìka". Serìâ Ìnformacìjnì sistemi ta merežì Pub Date : 2019-06-10 DOI:10.23939/sisn2019.01.101

N. Kotsyba, Bohdan Moskalevskyi

{"title":"Using transitivity information for morphological and syntactic disambiguation of pronouns in Ukrainian","authors":"N. Kotsyba, Bohdan Moskalevskyi","doi":"10.23939/sisn2019.01.101","DOIUrl":null,"url":null,"abstract":"The paper presents a short introduction to several electronic resources for Ukrainian language, namely, two treebanks: the Gold standard (ab. 130 thousand tokens), manually annotated in the Universal Dependencies flavour (https://universaldependencies.org/), which comprises the training data for a machine-trained syntactic parser, and a big (near 3 billion tokens), automatically annotated General Treebank (also known as Zvidusil), as well as a valency dictionary, developed by the Institute for Ukrainian, NGO (Kyiv) in 2015-2019 (https://mova.institute/). We also describe an experimental usage of the valency dictionary information to boost the performance of the syntactic parser. As a proof of concept, we discuss the case of syntactic and morphological ambiguity of frequently used Ukrainian pronouns його, її, їх ‘his, her, their’ and ways of improving the syntactic parser’s performance using the supervised machine learning techniques with a theoretical linguistic support. Apart from the multiple morphological ambiguity (24+ possible tags for each of these forms), one of the challenges connected with the presented linguistic phenomenon, is that its correct disambiguation involves anaphora resolution and semantic roles identification. On the one hand, this makes the disambiguation process much more complicated, given the followed annotation design, on the other hand, by resolving a seemingly low-level (morphological) problem we gain a bonus in the form of significant textual analysis hints which can be later used in various NLP applications for Ukrainian. The present article is a practical follow-up of its more theoretical predecessor (Kotsyba, Moskalevskyi 2018 [11]), where the linguistic underpinnings of the syntactic and morphological interpretation of the pronouns його, її, їх in comparison with other Slavic languages are presented in greater detail.","PeriodicalId":444399,"journal":{"name":"Vìsnik Nacìonalʹnogo unìversitetu \"Lʹvìvsʹka polìtehnìka\". Serìâ Ìnformacìjnì sistemi ta merežì","volume":"33 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-06-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Vìsnik Nacìonalʹnogo unìversitetu \"Lʹvìvsʹka polìtehnìka\". Serìâ Ìnformacìjnì sistemi ta merežì","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.23939/sisn2019.01.101","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

The paper presents a short introduction to several electronic resources for Ukrainian language, namely, two treebanks: the Gold standard (ab. 130 thousand tokens), manually annotated in the Universal Dependencies flavour (https://universaldependencies.org/), which comprises the training data for a machine-trained syntactic parser, and a big (near 3 billion tokens), automatically annotated General Treebank (also known as Zvidusil), as well as a valency dictionary, developed by the Institute for Ukrainian, NGO (Kyiv) in 2015-2019 (https://mova.institute/). We also describe an experimental usage of the valency dictionary information to boost the performance of the syntactic parser. As a proof of concept, we discuss the case of syntactic and morphological ambiguity of frequently used Ukrainian pronouns його, її, їх ‘his, her, their’ and ways of improving the syntactic parser’s performance using the supervised machine learning techniques with a theoretical linguistic support. Apart from the multiple morphological ambiguity (24+ possible tags for each of these forms), one of the challenges connected with the presented linguistic phenomenon, is that its correct disambiguation involves anaphora resolution and semantic roles identification. On the one hand, this makes the disambiguation process much more complicated, given the followed annotation design, on the other hand, by resolving a seemingly low-level (morphological) problem we gain a bonus in the form of significant textual analysis hints which can be later used in various NLP applications for Ukrainian. The present article is a practical follow-up of its more theoretical predecessor (Kotsyba, Moskalevskyi 2018 [11]), where the linguistic underpinnings of the syntactic and morphological interpretation of the pronouns його, її, їх in comparison with other Slavic languages are presented in greater detail.

查看原文本刊更多论文

及物性信息在乌克兰语代词形态和句法消歧中的应用

本文简要介绍了乌克兰语的几个电子资源，即两个树库:黄金标准(约13万个代币)，以通用依赖风格手工注释(https://universaldependencies.org/)，其中包括机器训练语法解析器的训练数据，以及一个大的(近30亿个代币)，自动注释的General Treebank(也称为Zvidusil)，以及由乌克兰研究所开发的价字典，NGO(基辅)在2015-2019年(https://mova.institute/)。我们还描述了一种使用价字典信息来提高语法解析器性能的实验方法。作为概念证明，我们讨论了经常使用的乌克兰代词його， її， їх ' his, her, their '的句法和形态歧义的情况，以及在理论语言学支持下使用监督机器学习技术提高句法解析器性能的方法。除了多种形态歧义(每种形式都有超过24种可能的标签)之外，与所提出的语言现象相关的挑战之一是其正确的歧义消除涉及回指解决和语义角色识别。一方面，考虑到下面的注释设计，这使得消歧过程变得更加复杂，另一方面，通过解决一个看似低级的(形态学)问题，我们获得了重要的文本分析提示，这些提示可以稍后在乌克兰语的各种NLP应用程序中使用。本文是其更具理论性的前身(Kotsyba, Moskalevskyi 2018[11])的实际后续，其中更详细地介绍了与其他斯拉夫语言相比，代词його， її， їх的句法和形态解释的语言学基础。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Vìsnik Nacìonalʹnogo unìversitetu "Lʹvìvsʹka polìtehnìka". Serìâ Ìnformacìjnì sistemi ta merežì

自引率

0.00%

发文量