Algorithm of the longest commonly consecutive word for Plagiarism detection in text based document

2008 Third International Conference on Digital Information Management Pub Date : 2008-11-01 DOI:10.1109/ICDIM.2008.4746827

Agung Sediyono, K. Ku-Mahamud

引用次数: 12

Abstract

Plagiarism is a form of academic misconduct which has increased with the easy access to obtain information through electronic documents and the Internet. The problem of finding document plagiarism in full text document can be viewed as a problem of finding the longest common parts of strings. Moreover, the detection system has to be capable to determine and visualize not only the common parts but also the location of the common parts in both the source and the observed document. Unlike previous research, this paper proposes a numerical based comparison algorithm that is comparable in the computation time without loosing the word order of common parts. Based on the experiment, the proposed algorithm outperforms the suffix tree in the length of observed paragraph below one hundred words.

查看原文本刊更多论文

基于文本的文档中最长共同连续词检测算法

抄袭是一种学术不端行为，随着通过电子文档和互联网获取信息的便利，这种行为有所增加。全文文档中查找文档剽窃的问题可以看作是查找字符串最长公共部分的问题。此外，检测系统必须不仅能够确定和可视化所述共同部分，而且还能够确定所述共同部分在源文件和所观察文件中的位置。与以往的研究不同，本文提出了一种基于数值的比较算法，该算法在计算时间上具有可比性，并且不会丢失公共部分的词序。实验结果表明，该算法在100个单词以下的观察段落长度上优于后缀树算法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2008 Third International Conference on Digital Information Management

自引率

0.00%

发文量