Stemming in the language modeling framework

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval Pub Date : 2003-07-28 DOI:10.1145/860435.860548

James Allan, G. Kumaran

引用次数: 17

Abstract

Stemming is the process of collapsing words into their morphological root. For example, the terms addicted, addicting, addictions, addictive, and addicts might be conflated to their stem, addict. Over the years, numerous studies [2, 3, 4] have considered stemming as an external process — either to be ignored or used as a pre-processing step. In this study, we try and provide a fresh perspective to stemming. We are motivated by the observation that stemming can be viewed as a form of smoothing, as a way of improving statistical estimates. This suggests that stemming could be directly incorporated into a language model, which is what we achieve in this paper. Detailed discussions are available in[1].

查看原文本刊更多论文

语言建模框架中的词干提取

词干是将单词分解成词根的过程。例如，上瘾的、上瘾的、上瘾的、上瘾的和成瘾的这些术语可能会被合并到它们的词干上，成瘾。多年来，许多研究[2,3,4]都认为词干提取是一个外部过程——要么被忽略，要么被用作预处理步骤。在这项研究中，我们试图为词干学提供一个新的视角。我们的动机是观察到词干可以被看作是平滑的一种形式，是改进统计估计的一种方式。这表明词干提取可以直接合并到语言模型中，这就是我们在本文中所实现的。详细讨论请参见[1]。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval

自引率

0.00%

发文量