Sparse Fuzzy C-Means Clustering with Lasso Penalty

Symmetry Pub Date : 2024-09-13 DOI:10.3390/sym16091208

Shazia Parveen, Miin-Shen Yang

{"title":"Sparse Fuzzy C-Means Clustering with Lasso Penalty","authors":"Shazia Parveen, Miin-Shen Yang","doi":"10.3390/sym16091208","DOIUrl":null,"url":null,"abstract":"Clustering is a technique of grouping data into a homogeneous structure according to the similarity or dissimilarity measures between objects. In clustering, the fuzzy c-means (FCM) algorithm is the best-known and most commonly used method and is a fuzzy extension of k-means in which FCM has been widely used in various fields. Although FCM is a good clustering algorithm, it only treats data points with feature components under equal importance and has drawbacks for handling high-dimensional data. The rapid development of social media and data acquisition techniques has led to advanced methods of collecting and processing larger, complex, and high-dimensional data. However, with high-dimensional data, the number of dimensions is typically immaterial or irrelevant. For features to be sparse, the Lasso penalty is capable of being applied to feature weights. A solution for FCM with sparsity is sparse FCM (S-FCM) clustering. In this paper, we propose a new S-FCM, called S-FCM-Lasso, which is a new type of S-FCM based on the Lasso penalty. The irrelevant features can be diminished towards exactly zero and assigned zero weights for unnecessary characteristics by the proposed S-FCM-Lasso. Based on various clustering performance measures, we compare S-FCM-Lasso with the S-FCM and other existing sparse clustering algorithms on several numerical and real-life datasets. Comparisons and experimental results demonstrate that, in terms of these performance measures, the proposed S-FCM-Lasso performs better than S-FCM and existing sparse clustering algorithms. This validates the efficiency and usefulness of the proposed S-FCM-Lasso algorithm for high-dimensional datasets with sparsity.","PeriodicalId":501198,"journal":{"name":"Symmetry","volume":"3 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2024-09-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Symmetry","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.3390/sym16091208","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Clustering is a technique of grouping data into a homogeneous structure according to the similarity or dissimilarity measures between objects. In clustering, the fuzzy c-means (FCM) algorithm is the best-known and most commonly used method and is a fuzzy extension of k-means in which FCM has been widely used in various fields. Although FCM is a good clustering algorithm, it only treats data points with feature components under equal importance and has drawbacks for handling high-dimensional data. The rapid development of social media and data acquisition techniques has led to advanced methods of collecting and processing larger, complex, and high-dimensional data. However, with high-dimensional data, the number of dimensions is typically immaterial or irrelevant. For features to be sparse, the Lasso penalty is capable of being applied to feature weights. A solution for FCM with sparsity is sparse FCM (S-FCM) clustering. In this paper, we propose a new S-FCM, called S-FCM-Lasso, which is a new type of S-FCM based on the Lasso penalty. The irrelevant features can be diminished towards exactly zero and assigned zero weights for unnecessary characteristics by the proposed S-FCM-Lasso. Based on various clustering performance measures, we compare S-FCM-Lasso with the S-FCM and other existing sparse clustering algorithms on several numerical and real-life datasets. Comparisons and experimental results demonstrate that, in terms of these performance measures, the proposed S-FCM-Lasso performs better than S-FCM and existing sparse clustering algorithms. This validates the efficiency and usefulness of the proposed S-FCM-Lasso algorithm for high-dimensional datasets with sparsity.

查看原文本刊更多论文

带有 Lasso 惩罚的稀疏模糊 C-Means 聚类法

聚类（Clustering）是一种根据对象之间的相似性或不相似性度量将数据分组为同质结构的技术。在聚类中，模糊均值（FCM）算法是最著名、最常用的方法，它是 K 均值的模糊扩展，FCM 已被广泛应用于各个领域。虽然 FCM 是一种很好的聚类算法，但它只处理具有同等重要性特征成分的数据点，在处理高维数据时存在缺陷。社交媒体和数据采集技术的飞速发展带来了收集和处理更大、更复杂和更高维数据的先进方法。然而，对于高维数据，维数通常并不重要或无关紧要。对于稀疏特征，Lasso 惩罚可以应用于特征权重。稀疏 FCM（S-FCM）聚类是具有稀疏性的 FCM 的一种解决方案。本文提出了一种新的 S-FCM，称为 S-FCM-Lasso，它是一种基于 Lasso 惩罚的新型 S-FCM。本文提出的 S-FCM-Lasso 是一种基于 Lasso 惩罚的新型 S-FCM，它可以将不相关的特征减小为零，并为不必要的特征分配零权重。基于各种聚类性能指标，我们在多个数值和现实数据集上比较了 S-FCM-Lasso 与 S-FCM 和其他现有的稀疏聚类算法。比较和实验结果表明，就这些性能指标而言，所提出的 S-FCM-Lasso 比 S-FCM 和现有的稀疏聚类算法表现更好。这验证了所提出的 S-FCM-Lasso 算法在具有稀疏性的高维数据集上的效率和实用性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Symmetry

自引率

0.00%

发文量