Asmaa M. Hassan, Safaa M. Naeem, Mohamed A. A. Eldosoky, Mai S. Mabrouk
{"title":"Multi-omics-based Machine Learning for the Subtype Classification of Breast Cancer","authors":"Asmaa M. Hassan, Safaa M. Naeem, Mohamed A. A. Eldosoky, Mai S. Mabrouk","doi":"10.1007/s13369-024-09341-7","DOIUrl":null,"url":null,"abstract":"<p>Cancer is a complicated disease that produces deregulatory changes in cellular activities (such as proteins). Data from these levels must be integrated into multi-omics analyses to better understand cancer and its progression. Deep learning approaches have recently helped with multi-omics analysis of cancer data. Breast cancer is a prevalent form of cancer among women, resulting from a multitude of clinical, lifestyle, social, and economic factors. The goal of this study was to predict breast cancer using several machine learning methods. We applied the architecture for mono-omics data analysis of the Cancer Genome Atlas Breast Cancer datasets in our analytical investigation. The following classifiers were used: random forest, partial least squares, Naive Bayes, decision trees, neural networks, and Lasso regularization. They were used and evaluated using the area under the curve metric. The random forest classifier and the Lasso regularization classifier achieved the highest area under the curve values of 0.99 each. These areas under the curve values were obtained using the mono-omics data employed in this investigation. The random forest and Lasso regularization classifiers achieved the maximum prediction accuracy, showing that they are appropriate for this problem. For all mono-omics classification models used in this paper, random forest and Lasso regression offer the best results for all metrics (precision, recall, and F1 score). The integration of various risk factors in breast cancer prediction modeling can aid in early diagnosis and treatment, utilizing data collection, storage, and intelligent systems for disease management. The integration of diverse risk factors in breast cancer prediction modeling holds promise for early diagnosis and treatment. Leveraging data collection, storage, and intelligent systems can further enhance disease management strategies, ultimately contributing to improved patient outcomes.</p>","PeriodicalId":8109,"journal":{"name":"Arabian Journal for Science and Engineering","volume":"2 1","pages":""},"PeriodicalIF":2.9000,"publicationDate":"2024-09-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Arabian Journal for Science and Engineering","FirstCategoryId":"103","ListUrlMain":"https://doi.org/10.1007/s13369-024-09341-7","RegionNum":4,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"Multidisciplinary","Score":null,"Total":0}
引用次数: 0
Abstract
Cancer is a complicated disease that produces deregulatory changes in cellular activities (such as proteins). Data from these levels must be integrated into multi-omics analyses to better understand cancer and its progression. Deep learning approaches have recently helped with multi-omics analysis of cancer data. Breast cancer is a prevalent form of cancer among women, resulting from a multitude of clinical, lifestyle, social, and economic factors. The goal of this study was to predict breast cancer using several machine learning methods. We applied the architecture for mono-omics data analysis of the Cancer Genome Atlas Breast Cancer datasets in our analytical investigation. The following classifiers were used: random forest, partial least squares, Naive Bayes, decision trees, neural networks, and Lasso regularization. They were used and evaluated using the area under the curve metric. The random forest classifier and the Lasso regularization classifier achieved the highest area under the curve values of 0.99 each. These areas under the curve values were obtained using the mono-omics data employed in this investigation. The random forest and Lasso regularization classifiers achieved the maximum prediction accuracy, showing that they are appropriate for this problem. For all mono-omics classification models used in this paper, random forest and Lasso regression offer the best results for all metrics (precision, recall, and F1 score). The integration of various risk factors in breast cancer prediction modeling can aid in early diagnosis and treatment, utilizing data collection, storage, and intelligent systems for disease management. The integration of diverse risk factors in breast cancer prediction modeling holds promise for early diagnosis and treatment. Leveraging data collection, storage, and intelligent systems can further enhance disease management strategies, ultimately contributing to improved patient outcomes.
期刊介绍:
King Fahd University of Petroleum & Minerals (KFUPM) partnered with Springer to publish the Arabian Journal for Science and Engineering (AJSE).
AJSE, which has been published by KFUPM since 1975, is a recognized national, regional and international journal that provides a great opportunity for the dissemination of research advances from the Kingdom of Saudi Arabia, MENA and the world.