Internet-scale Real-time Code Clone Search Via Multi-level Indexing

2011 18th Working Conference on Reverse Engineering Pub Date : 2011-10-17 DOI:10.1109/WCRE.2011.13

I. Keivanloo, J. Rilling, P. Charland

引用次数: 43

Abstract

Finding lines of code similar to a code fragment across large knowledge bases in fractions of a second is a new branch of code clone research also known as real-time code clone search. Among the requirements real-time code clone search has to meet are scalability, short response time, scalable incremental corpus updates, and support for type-1, type-2, and type-3 clones. We conducted a set of empirical studies on a large open source code corpus to gain insight about its characteristics. We used these results to design and optimize a multi-level indexing approach using hash table-based and binary search to improve Internet-scale real-time code clone search response time. Finally, we performed an evaluation on an Internet-scale corpus (1.5 million Java files and 266 MLOC). Our approach maintains a response time for 99.9% of clone searches in the microseconds range, while supporting the aforementioned requirements.

查看原文本刊更多论文

互联网规模的实时代码克隆搜索通过多层次索引

在不到一秒的时间内在大型知识库中找到与代码片段相似的代码行是代码克隆研究的一个新分支，也称为实时代码克隆搜索。实时代码克隆搜索必须满足的需求包括可伸缩性、短响应时间、可伸缩的增量语料库更新，以及对类型1、类型2和类型3克隆的支持。我们对一个大型开源代码语料库进行了一组实证研究，以深入了解其特征。我们利用这些结果设计并优化了基于哈希表和二进制搜索的多级索引方法，以改善互联网规模的实时代码克隆搜索响应时间。最后，我们对一个互联网规模的语料库(150万个Java文件和266个MLOC)进行了评估。我们的方法将99.9%的克隆搜索的响应时间保持在微秒范围内，同时支持上述需求。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2011 18th Working Conference on Reverse Engineering

自引率

0.00%

发文量