A Machine Learning-Based Framework for Accurate and Early Diagnosis of Liver Diseases: A Comprehensive Study on Feature Selection, Data Imbalance, and Algorithmic Performance
IF 5 2区 计算机科学Q1 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE
{"title":"A Machine Learning-Based Framework for Accurate and Early Diagnosis of Liver Diseases: A Comprehensive Study on Feature Selection, Data Imbalance, and Algorithmic Performance","authors":"Attique Ur Rehman, Wasi Haider Butt, Tahir Muhammad Ali, Sabeen Javaid, Maram Fahaad Almufareh, Mamoona Humayun, Hameedur Rahman, Azka Mir, Momina Shaheen","doi":"10.1155/2024/6111312","DOIUrl":null,"url":null,"abstract":"<div>\n <p>The liver is the largest organ of the human body with more than 500 vital functions. In recent decades, a large number of liver patients have been reported with diseases such as cirrhosis, fibrosis, or other liver disorders. There is a need for effective, early, and accurate identification of individuals suffering from such disease so that the person may recover before the disease spreads and becomes fatal. For this, applications of machine learning are playing a significant role. Despite the advancements, existing systems remain inconsistent in performance due to limited feature selection and data imbalance. In this article, we reviewed 58 articles extracted from 5 different electronic repositories published from January 2015 to 2023. After a systematic and protocol-based review, we answered 6 research questions about machine learning algorithms. The identification of effective feature selection techniques, data imbalance management techniques, accurate machine learning algorithms, a list of available data sets with their URLs and characteristics, and feature importance based on usage has been identified for diagnosing liver disease. The reason to select this research question is, in any machine learning framework, the role of dimensionality reduction, data imbalance management, machine learning algorithm with its accuracy, and data itself is very significant. Based on the conducted review, a framework, machine learning-based liver disease diagnosis (MaLLiDD), has been proposed and validated using three datasets. The proposed framework classified liver disorders with 99.56%, 76.56%, and 76.11% accuracy. In conclusion, this article addressed six research questions by identifying effective feature selection techniques, data imbalance management techniques, algorithms, datasets, and feature importance based on usage. It also demonstrated a high accuracy with the framework for early diagnosis, marking a significant advancement.</p>\n </div>","PeriodicalId":14089,"journal":{"name":"International Journal of Intelligent Systems","volume":null,"pages":null},"PeriodicalIF":5.0000,"publicationDate":"2024-06-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1155/2024/6111312","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"International Journal of Intelligent Systems","FirstCategoryId":"94","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.1155/2024/6111312","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0
Abstract
The liver is the largest organ of the human body with more than 500 vital functions. In recent decades, a large number of liver patients have been reported with diseases such as cirrhosis, fibrosis, or other liver disorders. There is a need for effective, early, and accurate identification of individuals suffering from such disease so that the person may recover before the disease spreads and becomes fatal. For this, applications of machine learning are playing a significant role. Despite the advancements, existing systems remain inconsistent in performance due to limited feature selection and data imbalance. In this article, we reviewed 58 articles extracted from 5 different electronic repositories published from January 2015 to 2023. After a systematic and protocol-based review, we answered 6 research questions about machine learning algorithms. The identification of effective feature selection techniques, data imbalance management techniques, accurate machine learning algorithms, a list of available data sets with their URLs and characteristics, and feature importance based on usage has been identified for diagnosing liver disease. The reason to select this research question is, in any machine learning framework, the role of dimensionality reduction, data imbalance management, machine learning algorithm with its accuracy, and data itself is very significant. Based on the conducted review, a framework, machine learning-based liver disease diagnosis (MaLLiDD), has been proposed and validated using three datasets. The proposed framework classified liver disorders with 99.56%, 76.56%, and 76.11% accuracy. In conclusion, this article addressed six research questions by identifying effective feature selection techniques, data imbalance management techniques, algorithms, datasets, and feature importance based on usage. It also demonstrated a high accuracy with the framework for early diagnosis, marking a significant advancement.
期刊介绍:
The International Journal of Intelligent Systems serves as a forum for individuals interested in tapping into the vast theories based on intelligent systems construction. With its peer-reviewed format, the journal explores several fascinating editorials written by today''s experts in the field. Because new developments are being introduced each day, there''s much to be learned — examination, analysis creation, information retrieval, man–computer interactions, and more. The International Journal of Intelligent Systems uses charts and illustrations to demonstrate these ground-breaking issues, and encourages readers to share their thoughts and experiences.