Yosr Naïja, Salem Chakhar, Kaouther Blibech Sinaoui, R. Robbana
{"title":"Extension of Partitional Clustering Methods for Handling Mixed Data","authors":"Yosr Naïja, Salem Chakhar, Kaouther Blibech Sinaoui, R. Robbana","doi":"10.1109/ICDMW.2008.85","DOIUrl":null,"url":null,"abstract":"Clustering is an active research topic in data mining and different methods have been proposed in the literature. Most of these methods are based on the use of a distance measure defined either on numerical attributes or on categorical attributes. However, in fields such as road traffic and medicine, datasets are composed of numerical and categorical attributes. Recently, there have been several proposals to develop clustering methods that support mixed attributes. There are three basic categories of clustering methods: partitional methods, hierarchical methods and density-based methods. This paper proposes an extension of partitional clustering methods devoted to mixed attributes. The proposed extension looks to create several partitions by using numerical attributes-based clustering methods and then chooses the one that maximizes a measure---called ``homogeneity degree\"---of these partitions according to categorical attributes.","PeriodicalId":175955,"journal":{"name":"2008 IEEE International Conference on Data Mining Workshops","volume":"240 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2008-12-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"10","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2008 IEEE International Conference on Data Mining Workshops","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDMW.2008.85","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 10
Abstract
Clustering is an active research topic in data mining and different methods have been proposed in the literature. Most of these methods are based on the use of a distance measure defined either on numerical attributes or on categorical attributes. However, in fields such as road traffic and medicine, datasets are composed of numerical and categorical attributes. Recently, there have been several proposals to develop clustering methods that support mixed attributes. There are three basic categories of clustering methods: partitional methods, hierarchical methods and density-based methods. This paper proposes an extension of partitional clustering methods devoted to mixed attributes. The proposed extension looks to create several partitions by using numerical attributes-based clustering methods and then chooses the one that maximizes a measure---called ``homogeneity degree"---of these partitions according to categorical attributes.