Rule based chunker for Hindi

2016 2nd International Conference on Contemporary Computing and Informatics (IC3I) Pub Date : 2016-12-01 DOI:10.1109/IC3I.2016.7918005

S. Asopa, Pooja Asopa, Iti Mathur, Nisheeth Joshi

引用次数: 4

Abstract

In this research paper, a rule based chunker is developed and evaluated. For the development of the chunker, handcrafted linguistic rules for mainly noun, adverb, verb, adjective phrases and conjuncts were generated. Indian Languages Chunk Tagset is used for annotations. In order to evaluate, 500 sentences of Hindi language tagged by HMM tagger were considered and given as an input to our chunker. Precision, Recall and F-Measure for the system were calculated and found to be 79.68, 69.36 and 74.16 respectively.

查看原文本刊更多论文

基于规则的印地语分块器

本文开发了一种基于规则的分块器，并对其进行了评价。为了开发分块器，生成了主要用于名词、副词、动词、形容词短语和连词的手工语言规则。印度语言块标记集用于注释。为了评估，考虑了500个由HMM标记器标记的印地语句子，并将其作为我们的分块器的输入。系统的精密度、召回率和F-Measure分别为79.68、69.36和74.16。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2016 2nd International Conference on Contemporary Computing and Informatics (IC3I)

自引率

0.00%

发文量