Features of Disagreement Between Retrieval Effectiveness Measures

Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval Pub Date : 2015-08-09 DOI:10.1145/2766462.2767824

Timothy Jones, Paul Thomas, Falk Scholer, M. Sanderson

引用次数: 6

Abstract

Many IR effectiveness measures are motivated from intuition, theory, or user studies. In general, most effectiveness measures are well correlated with each other. But, what about where they don't correlate? Which rankings cause measures to disagree? Are these rankings predictable for particular pairs of measures? In this work, we examine how and where metrics disagree, and identify differences that should be considered when selecting metrics for use in evaluating retrieval systems.

查看原文本刊更多论文

检索有效性度量差异的特征

许多IR有效性度量是由直觉、理论或用户研究驱动的。一般来说，大多数有效性度量都是相互关联的。但是，如果它们不相关呢?哪些排名导致测量结果不一致?这些排名对于特定的衡量标准是可预测的吗?在这项工作中，我们检查度量不一致的方式和位置，并确定在选择用于评估检索系统的度量时应该考虑的差异。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval

自引率

0.00%

发文量