Ananya -僧伽罗语命名实体识别(NER)系统

2016 Moratuwa Engineering Research Conference (MERCon) Pub Date : 2016-04-05 DOI:10.1109/MERCON.2016.7480111

S. A. P. M. Manamini, A. F. Ahamed, R. Rajapakshe, G. H. A. Reemal, Sanath Jayasena, G. Dias, Surangika Ranathunga

{"title":"Ananya -僧伽罗语命名实体识别(NER)系统","authors":"S. A. P. M. Manamini, A. F. Ahamed, R. Rajapakshe, G. H. A. Reemal, Sanath Jayasena, G. Dias, Surangika Ranathunga","doi":"10.1109/MERCON.2016.7480111","DOIUrl":null,"url":null,"abstract":"Named-Entity-Recognition (NER) is one of the major tasks under Natural Language Processing, which is widely used in the fields of Computer Science and Computational Linguistics. However, the amount of prior research done on NER for Sinhala is very minimal. In this paper, we present data-driven techniques to detect Named Entities in Sinhala text, with the use of Conditional Random Fields (CRF) and Maximum Entropy (ME) statistical modeling methods. Results obtained from experiments indicate that CRF, which provided the highest accuracy for the same task for other languages outperforms ME in Sinhala NER as well. Furthermore, we identify different linguistic features such as orthographic word level and contextual information that are effective with both CRF and ME Algorithms.","PeriodicalId":184790,"journal":{"name":"2016 Moratuwa Engineering Research Conference (MERCon)","volume":"147 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2016-04-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"9","resultStr":"{\"title\":\"Ananya - a Named-Entity-Recognition (NER) system for Sinhala language\",\"authors\":\"S. A. P. M. Manamini, A. F. Ahamed, R. Rajapakshe, G. H. A. Reemal, Sanath Jayasena, G. Dias, Surangika Ranathunga\",\"doi\":\"10.1109/MERCON.2016.7480111\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Named-Entity-Recognition (NER) is one of the major tasks under Natural Language Processing, which is widely used in the fields of Computer Science and Computational Linguistics. However, the amount of prior research done on NER for Sinhala is very minimal. In this paper, we present data-driven techniques to detect Named Entities in Sinhala text, with the use of Conditional Random Fields (CRF) and Maximum Entropy (ME) statistical modeling methods. Results obtained from experiments indicate that CRF, which provided the highest accuracy for the same task for other languages outperforms ME in Sinhala NER as well. Furthermore, we identify different linguistic features such as orthographic word level and contextual information that are effective with both CRF and ME Algorithms.\",\"PeriodicalId\":184790,\"journal\":{\"name\":\"2016 Moratuwa Engineering Research Conference (MERCon)\",\"volume\":\"147 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2016-04-05\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"9\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2016 Moratuwa Engineering Research Conference (MERCon)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/MERCON.2016.7480111\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2016 Moratuwa Engineering Research Conference (MERCon)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/MERCON.2016.7480111","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 9

摘要

命名实体识别(NER)是自然语言处理的主要任务之一，在计算机科学和计算语言学等领域有着广泛的应用。然而，之前对僧伽罗人的NER进行的研究非常少。在本文中，我们提出了使用条件随机场(CRF)和最大熵(ME)统计建模方法来检测僧伽罗语文本中的命名实体的数据驱动技术。实验结果表明，CRF在其他语言的相同任务中提供了最高的准确率，也优于ME在僧伽罗语NER中的表现。此外，我们还确定了不同的语言特征，如正字法单词水平和上下文信息，这些特征在CRF和ME算法中都是有效的。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Ananya - a Named-Entity-Recognition (NER) system for Sinhala language

Named-Entity-Recognition (NER) is one of the major tasks under Natural Language Processing, which is widely used in the fields of Computer Science and Computational Linguistics. However, the amount of prior research done on NER for Sinhala is very minimal. In this paper, we present data-driven techniques to detect Named Entities in Sinhala text, with the use of Conditional Random Fields (CRF) and Maximum Entropy (ME) statistical modeling methods. Results obtained from experiments indicate that CRF, which provided the highest accuracy for the same task for other languages outperforms ME in Sinhala NER as well. Furthermore, we identify different linguistic features such as orthographic word level and contextual information that are effective with both CRF and ME Algorithms.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2016 Moratuwa Engineering Research Conference (MERCon)

自引率

0.00%

发文量