MFWK-Means: Minkowski metric Fuzzy Weighted K-Means for high dimensional data clustering

2013 IEEE 14th International Conference on Information Reuse & Integration (IRI) Pub Date : 2013-10-24 DOI:10.1109/IRI.2013.6642535

L. Svetlova, B. Mirkin, H. Lei

引用次数: 6

Abstract

This paper presents a clustering algorithm, namely MFWK-Means, which is a novel extension of K-Means clustering to the case of fuzzy clusters and weighted features. First, the Weighted K-Means criterion utilizing Minkowski metric is adopted to solve the problem of feature selection for high dimensional data. Then, a further extension to the case of fuzzy clustering is presented to group datasets with natural fuzziness of cluster boundaries. Also, we adopt an intelligent version of K-Means, using Mirkin's method of Anomalous Pattern for initialization. Our new Minkowski metric Fuzzy Weighted K-Means (MFWK-Means) is experimentally validated on both benchmark datasets and synthetic datasets. MFWK-Means is shown to be competitive and more stable against noise in comparison with a variety of versions of K-Means based methods. Moreover, in most situations it reaches the highest clustering accuracy at wider intervals of Minkowski exponent.

查看原文本刊更多论文

MFWK-Means:用于高维数据聚类的Minkowski度量模糊加权K-Means

本文提出了一种聚类算法MFWK-Means，它是K-Means聚类在模糊聚类和加权特征情况下的新颖扩展。首先，采用基于Minkowski度量的加权K-Means准则解决高维数据的特征选择问题;然后，对模糊聚类的情况进行了进一步的扩展，利用聚类边界的自然模糊性对数据集进行分组。此外，我们还采用了智能版的K-Means，使用Mirkin的异常模式方法进行初始化。我们的新Minkowski度量模糊加权K-Means (MFWK-Means)在基准数据集和合成数据集上进行了实验验证。与各种版本的基于K-Means的方法相比，MFWK-Means具有竞争力，并且对噪声更稳定。在大多数情况下，该方法在较宽的闵可夫斯基指数区间内达到最高的聚类精度。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2013 IEEE 14th International Conference on Information Reuse & Integration (IRI)

自引率

0.00%

发文量