Affix-augmented stem-based language model for persian

Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering(NLPKE-2010) Pub Date : 2010-09-30 DOI:10.1109/NLPKE.2010.5587823

Heshaam Faili, H. Ravanbakhsh

引用次数: 2

Abstract

Language modeling is used in many NLP applications like machine translation, POS tagging, speech recognition and information retrieval. It assigns a probability to a sequence of words. This task becomes a challenging problem for high inflectional languages. In this paper we investigate standard statistical language models on the Persian as an inflectional language. We propose two variations of morphological language models that rely on a morphological analyzer to manipulate the dataset before modeling. Then we discuss shortcoming of these models, and introduce a novel approach that exploits the structure of the language and produces more accurate. Experimental results are encouraging especially when we use n-gram models with small training dataset.

查看原文本刊更多论文

波斯语词缀增强词干语言模型

语言建模用于许多NLP应用，如机器翻译、词性标注、语音识别和信息检索。它为单词序列分配一个概率。对于高屈折变化的语言来说，这是一个具有挑战性的问题。本文研究了波斯语作为一种屈折变化语言的标准统计语言模型。我们提出了两种形态学语言模型的变体，它们依赖于形态学分析器在建模之前对数据集进行操作。然后讨论了这些模型的不足之处，并介绍了一种利用语言结构的新方法。实验结果令人鼓舞，特别是当我们使用n-gram模型和小训练数据集时。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering(NLPKE-2010)

自引率

0.00%

发文量