{"title":"Probability-based Imputation Method for Fuzzy Cluster Analysis of Gene Expression Microarray Data","authors":"Thanh Le, T. Altman, K. Gardiner","doi":"10.1109/ITNG.2012.159","DOIUrl":null,"url":null,"abstract":"Fuzzy clustering has been widely used for analysis of gene expression micro array data. However, most fuzzy clustering algorithms require complete datasets and, because of technical limitations, most micro array datasets have missing values. To address this problem, we present a new algorithm where genes are clustered using the Fuzzy C-Means algorithm, followed by approximating the fuzzy partition by a probabilistic data distribution model which is then used to estimate the missing values in the dataset. Using distribution-based approach, our method is most appropriate for datasets where the data are nonuniform. We show that our method outperforms six popular imputation algorithms on uniform and nonuniform artificial datasets as well as real datasets with unknown data distribution model.","PeriodicalId":117236,"journal":{"name":"2012 Ninth International Conference on Information Technology - New Generations","volume":"75 3 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2012-04-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"7","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2012 Ninth International Conference on Information Technology - New Generations","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ITNG.2012.159","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 7
Abstract
Fuzzy clustering has been widely used for analysis of gene expression micro array data. However, most fuzzy clustering algorithms require complete datasets and, because of technical limitations, most micro array datasets have missing values. To address this problem, we present a new algorithm where genes are clustered using the Fuzzy C-Means algorithm, followed by approximating the fuzzy partition by a probabilistic data distribution model which is then used to estimate the missing values in the dataset. Using distribution-based approach, our method is most appropriate for datasets where the data are nonuniform. We show that our method outperforms six popular imputation algorithms on uniform and nonuniform artificial datasets as well as real datasets with unknown data distribution model.