De-duplicating a large crowd-sourced catalogue of bibliographic records

Q Social Sciences

Program-Electronic Library and Information Systems Pub Date : 2016-03-21 DOI:10.1108/PROG-02-2015-0021

Ilija Subasic, N. Gvozdenovic, Kris Jack

引用次数: 0

Abstract

Purpose – The purpose of this paper is to describe a large-scale algorithm for generating a catalogue of scientific publication records (citations) from a crowd-sourced data, demonstrate how to learn an optimal combination of distance metrics for duplicate detection and introduce a parallel duplicate clustering algorithm. Design/methodology/approach – The authors developed the algorithm and compared it with state-of-the art systems tackling the same problem. The authors used benchmark data sets (3k data points) to test the effectiveness of our algorithm and a real-life data ( > 90 million) to test the efficiency and scalability of our algorithm. Findings – The authors show that duplicate detection can be improved by an additional step we call duplicate clustering. The authors also show how to improve the efficiency of map/reduce similarity calculation algorithm by introducing a sampling step. Finally, the authors find that the system is comparable to the state-of-the art systems for duplicate detection, a...

查看原文本刊更多论文

从大量的文献记录中删除重复的目录

目的-本文的目的是描述一种大规模算法，用于从众包数据中生成科学出版记录(引用)目录，演示如何学习用于重复检测的距离度量的最佳组合，并引入并行重复聚类算法。设计/方法论/方法-作者开发了算法，并将其与解决相同问题的最先进系统进行了比较。作者使用基准数据集(3k个数据点)来测试我们算法的有效性，并使用实际数据(> 9000万)来测试我们算法的效率和可扩展性。发现-作者表明，重复检测可以通过我们称之为重复聚类的额外步骤来改进。作者还介绍了如何通过引入采样步骤来提高map/reduce相似度计算算法的效率。最后，作者发现该系统可与最先进的重复检测系统相媲美。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Program-Electronic Library and Information Systems 工程技术-计算机：信息系统

CiteScore

1.30

自引率

0.00%

发文量

审稿时长

>12 weeks

期刊介绍： ■Automation of library and information services ■Storage and retrieval of all forms of electronic information ■Delivery of information to end users ■Database design and management ■Techniques for storing and distributing information ■Networking and communications technology ■The Internet ■User interface design ■Procurement of systems ■User training and support ■System evaluation