使用特征模板和束搜索策略的印度尼西亚移位-减少选区解析器

2018 5th International Conference on Advanced Informatics: Concept Theory and Applications (ICAICTA) Pub Date : 2018-08-01 DOI:10.1109/ICAICTA.2018.8541292

Robert Sebastian Herlim, A. Purwarianti

{"title":"使用特征模板和束搜索策略的印度尼西亚移位-减少选区解析器","authors":"Robert Sebastian Herlim, A. Purwarianti","doi":"10.1109/ICAICTA.2018.8541292","DOIUrl":null,"url":null,"abstract":"In natural language processing, the syntactic analysis process (such as constituency parsing) is required to understand word context in the sentence. We propose a modification on using binarization technique alternative and feature multiplication factors for shift-reduce constituency parser using beam search approach and structured learning algorithm. Our modification in binarization technique is inspired from assorted tagging schemes in NER, while the feature multiplication factors is used to scale up our scoring system for beam search algorithm. For evaluation, we mainly used the new INACL Treebank (consisting 11,356 and 4,457 instances for training and test set), resulted 50.3% in f1-score. Our parser also compared with previous work by using the same training and test set for IDN-Treebank, resulted 74.0% in f1-score.","PeriodicalId":184882,"journal":{"name":"2018 5th International Conference on Advanced Informatics: Concept Theory and Applications (ICAICTA)","volume":"117 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2018-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":"{\"title\":\"Indonesian Shift-Reduce Constituency Parser Using Feature Templates & Beam Search Strategy\",\"authors\":\"Robert Sebastian Herlim, A. Purwarianti\",\"doi\":\"10.1109/ICAICTA.2018.8541292\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In natural language processing, the syntactic analysis process (such as constituency parsing) is required to understand word context in the sentence. We propose a modification on using binarization technique alternative and feature multiplication factors for shift-reduce constituency parser using beam search approach and structured learning algorithm. Our modification in binarization technique is inspired from assorted tagging schemes in NER, while the feature multiplication factors is used to scale up our scoring system for beam search algorithm. For evaluation, we mainly used the new INACL Treebank (consisting 11,356 and 4,457 instances for training and test set), resulted 50.3% in f1-score. Our parser also compared with previous work by using the same training and test set for IDN-Treebank, resulted 74.0% in f1-score.\",\"PeriodicalId\":184882,\"journal\":{\"name\":\"2018 5th International Conference on Advanced Informatics: Concept Theory and Applications (ICAICTA)\",\"volume\":\"117 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2018-08-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"2\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2018 5th International Conference on Advanced Informatics: Concept Theory and Applications (ICAICTA)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICAICTA.2018.8541292\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2018 5th International Conference on Advanced Informatics: Concept Theory and Applications (ICAICTA)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICAICTA.2018.8541292","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 2

摘要

在自然语言处理中，需要句法分析过程(如成分分析)来理解句子中的单词上下文。我们提出了一种基于二值化技术和特征乘法因子的改进方法。我们对二值化技术的改进灵感来自于NER中的分类标记方案，而特征乘法因子用于扩展我们的波束搜索算法的评分系统。对于评估，我们主要使用新的INACL树库(包括11,356和4,457个实例作为训练集和测试集)，结果f1得分为50.3%。我们的解析器还使用IDN-Treebank相同的训练和测试集与之前的工作进行了比较，结果f1得分为74.0%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Indonesian Shift-Reduce Constituency Parser Using Feature Templates & Beam Search Strategy

In natural language processing, the syntactic analysis process (such as constituency parsing) is required to understand word context in the sentence. We propose a modification on using binarization technique alternative and feature multiplication factors for shift-reduce constituency parser using beam search approach and structured learning algorithm. Our modification in binarization technique is inspired from assorted tagging schemes in NER, while the feature multiplication factors is used to scale up our scoring system for beam search algorithm. For evaluation, we mainly used the new INACL Treebank (consisting 11,356 and 4,457 instances for training and test set), resulted 50.3% in f1-score. Our parser also compared with previous work by using the same training and test set for IDN-Treebank, resulted 74.0% in f1-score.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2018 5th International Conference on Advanced Informatics: Concept Theory and Applications (ICAICTA)

自引率

0.00%

发文量