Optimal Hash List for Word Frequency Analysis

2010 International Conference on Web Information Systems and Mining Pub Date : 2010-10-23 DOI:10.1109/WISM.2010.59

Sheng-Lan Peng

引用次数: 0

Abstract

Word frequency analysis plays an essential role in many data mining tasks of large-scale data set based on text corpus, and hash list is a very simple but efficient structure for frequent pattern discovering. In this paper, a Poisson approximation approach is exploited to analyze the space efficiency of hash list under different parameters on probability. Based on our theoretical model, an optimal parameter setting for hash list is given. Experimental result of real data shows that hash list with the optimal parameter can reach minimum or nearly minimum memory cost.

查看原文本刊更多论文

词频分析的最优哈希表

词频分析在许多基于文本语料库的大规模数据集的数据挖掘任务中起着至关重要的作用，而哈希表是一种非常简单而有效的频繁模式发现结构。本文利用泊松近似方法分析了哈希表在不同参数下的空间效率。在理论模型的基础上，给出了哈希表的最优参数设置。实际数据的实验结果表明，采用最优参数的哈希表可以达到最小或接近最小的内存开销。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2010 International Conference on Web Information Systems and Mining

自引率

0.00%

发文量