Christelle Kemda Ngueda, Julia Palm, Flavia Remo, André Scherag, Lutz Leistritz
{"title":"One-sample missing DNA-methylation value imputation.","authors":"Christelle Kemda Ngueda, Julia Palm, Flavia Remo, André Scherag, Lutz Leistritz","doi":"10.1186/s12859-025-06154-9","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Currently, the most popular methods for missing DNA-methylation value imputation rely on exploiting methylation patterns across multiple samples from the same population. However, if there is significant variability between individuals or limited data available, these methods might produce biased results. This situation has prompted researchers to seek alternative approaches for handling single-sample data, particularly in the context of personalized medicine. Accordingly, we propose One-Sample Methyl Imputation (OSMI), an imputation method that can also be used in single-sample applications.</p><p><strong>Results: </strong>The proposed method in single-subject cases yielded an average imputation accuracy of RMSE = 0.2713 (95%-CI from 0.2696 to 0.2730) in β-value units (range: 0-1) based on real 450 K BeadChip data sets of 3,402 individuals. It is possible to take the affiliation of individual CpGs to CpG islands into account during the imputation of missing methylation values. This improves the imputation accuracy. In addition, the accuracy of imputation depends in general on the density of CpG sites on DNA-methylation microarrays and increases as the CpG site density increases. OSMI has low memory and computational requirements.</p><p><strong>Conclusions: </strong>OSMI uses a single methylome to impute missing values quickly at very low memory constraints. Its imputation accuracy is inferior to other methods if multiple samples are available and these samples are reasonably similar, but OSMI represents a useful addition to the imputation toolbox for the case of single-sample applications.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"26 1","pages":"143"},"PeriodicalIF":3.3000,"publicationDate":"2025-05-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12126866/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"BMC Bioinformatics","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1186/s12859-025-06154-9","RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"BIOCHEMICAL RESEARCH METHODS","Score":null,"Total":0}
引用次数: 0
Abstract
Background: Currently, the most popular methods for missing DNA-methylation value imputation rely on exploiting methylation patterns across multiple samples from the same population. However, if there is significant variability between individuals or limited data available, these methods might produce biased results. This situation has prompted researchers to seek alternative approaches for handling single-sample data, particularly in the context of personalized medicine. Accordingly, we propose One-Sample Methyl Imputation (OSMI), an imputation method that can also be used in single-sample applications.
Results: The proposed method in single-subject cases yielded an average imputation accuracy of RMSE = 0.2713 (95%-CI from 0.2696 to 0.2730) in β-value units (range: 0-1) based on real 450 K BeadChip data sets of 3,402 individuals. It is possible to take the affiliation of individual CpGs to CpG islands into account during the imputation of missing methylation values. This improves the imputation accuracy. In addition, the accuracy of imputation depends in general on the density of CpG sites on DNA-methylation microarrays and increases as the CpG site density increases. OSMI has low memory and computational requirements.
Conclusions: OSMI uses a single methylome to impute missing values quickly at very low memory constraints. Its imputation accuracy is inferior to other methods if multiple samples are available and these samples are reasonably similar, but OSMI represents a useful addition to the imputation toolbox for the case of single-sample applications.
期刊介绍:
BMC Bioinformatics is an open access, peer-reviewed journal that considers articles on all aspects of the development, testing and novel application of computational and statistical methods for the modeling and analysis of all kinds of biological data, as well as other areas of computational biology.
BMC Bioinformatics is part of the BMC series which publishes subject-specific journals focused on the needs of individual research communities across all areas of biology and medicine. We offer an efficient, fair and friendly peer review service, and are committed to publishing all sound science, provided that there is some advance in knowledge presented by the work.