Weakly supervised relevance feedback based on an improved language model

Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering(NLPKE-2010) Pub Date : 2010-09-30 DOI:10.1109/NLPKE.2010.5587859

Xinsheng Li, Si Li, Weiran Xu, Guang Chen, Jun Guo

引用次数: 0

Abstract

Relevance feedback, which traditionally uses the terms in the relevant documents to enrich the user's initial query, is an effective method for improving retrieval performance. This approach has another problem is that Relevance feedback assumes that most frequent terms in the feedback documents are useful for the retrieval. In fact, the reports of some experiments show that it does not hold in reality many expansion terms identified in traditional approaches are indeed unrelated to the query and harmful to the retrieval. In this paper, we propose to select better and more relevant documents with a clustering algorithm. And then we present an improved Language Model to help us identify the good terms from those relevant documents. Ours experiments on the 2008 TREC collection show that retrieval effectiveness can be much improved when the improved Language Model is used.

查看原文本刊更多论文

基于改进语言模型的弱监督相关反馈

相关反馈是一种提高检索性能的有效方法，传统上使用相关文档中的术语来丰富用户的初始查询。这种方法的另一个问题是相关性反馈假设反馈文档中最常见的术语对检索有用。事实上，一些实验报告表明，它在现实中并不成立，许多在传统方法中识别的扩展术语确实与查询无关，并且对检索有害。在本文中，我们提出了一种聚类算法来选择更好和更相关的文档。然后，我们提出了一个改进的语言模型，以帮助我们从这些相关文档中识别出好的术语。我们在2008年的TREC集合上的实验表明，使用改进的语言模型可以大大提高检索效率。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering(NLPKE-2010)

自引率

0.00%

发文量