Wenqi Wu, Xiaohao Hu, Linyang Yan, Zhiyin Li, Bo Li, Xinpeng Chen, Zexun Lin, Huiqiong Zeng, Chun Li, Yingqian Mo, Yalin Wu, Qingwen Wang
{"title":"Development and Validation of a Cost-Effective Machine Learning Model for Screening Potential Rheumatoid Arthritis in Primary Healthcare Clinics.","authors":"Wenqi Wu, Xiaohao Hu, Linyang Yan, Zhiyin Li, Bo Li, Xinpeng Chen, Zexun Lin, Huiqiong Zeng, Chun Li, Yingqian Mo, Yalin Wu, Qingwen Wang","doi":"10.2147/JIR.S487595","DOIUrl":null,"url":null,"abstract":"<p><strong>Objective: </strong>In primary healthcare, diagnosing rheumatoid arthritis (RA) is challenging due to a general lack of in-depth knowledge of RA by general practitioners (GPs) and the lack of effective tools, leading to high rates of missed diagnosis. This study focuses on a screening model for primary healthcare, aiming to improve early RA screening accuracy and efficiency at a relatively lower cost, reducing delays in GPs' recognition of RA.</p><p><strong>Methods: </strong>We randomly selected 2106 participants from the RA group or combined control group (comprising healthy individuals and patients with non-RA rheumatic diseases) at Peking University Shenzhen Hospital as the developing cohort. Guided by experienced rheumatologists, we built a comprehensive database with 26 clinical features. Using 10 classical machine learning algorithms, we developed screening models. Evaluation metrics determined the best model. Employing multivariatelogistic regression results and the best-performing model to identify the least costly features, ensuring applicability in primary healthcare clinics. Subsequently, we retrained and validated our proposed model based on two primary healthcare validation cohorts.</p><p><strong>Results: </strong>In experiments, the algorithms achieved over 88% accuracy on training and test sets. Random Forest (RF) excelled with 96.20% (95% CI 95.39% to 97.02%) accuracy, 96.22% (95% CI 95.40% to 97.03%) specificity, 96.18% (95% CI 95.37% to 97.00%) sensitivity, and 96.20% (95% CI 95.39% to 97.02%) Areas Under Curves (AUC). A meticulous feature selection identified 11 key features for RA screening. In an external test on two primary healthcare datasets with these features, RF demonstrated an accuracy of 88.435% (95% CI 85.55% to 91.32%), sensitivity of 98.55% (95% CI 97.47% to 99.63%), specificity of 85.56% (95% CI 82.39% to 88.73%), and an AUC of 92.055% (95% CI 89.62% to 94.49%).</p><p><strong>Conclusion: </strong>The screening model excels in automating prompt identification of RA in primary healthcare, improving the early detection of RA, and reducing delays and associated costs. Our findings contribute positively and are poised to elevate prospective RA management, fostering improvements in healthcare sector responsiveness and resource efficiency.</p>","PeriodicalId":16107,"journal":{"name":"Journal of Inflammation Research","volume":"18 ","pages":"1511-1522"},"PeriodicalIF":4.2000,"publicationDate":"2025-02-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11804240/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Inflammation Research","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.2147/JIR.S487595","RegionNum":2,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/1/1 0:00:00","PubModel":"eCollection","JCR":"Q2","JCRName":"IMMUNOLOGY","Score":null,"Total":0}
引用次数: 0
Abstract
Objective: In primary healthcare, diagnosing rheumatoid arthritis (RA) is challenging due to a general lack of in-depth knowledge of RA by general practitioners (GPs) and the lack of effective tools, leading to high rates of missed diagnosis. This study focuses on a screening model for primary healthcare, aiming to improve early RA screening accuracy and efficiency at a relatively lower cost, reducing delays in GPs' recognition of RA.
Methods: We randomly selected 2106 participants from the RA group or combined control group (comprising healthy individuals and patients with non-RA rheumatic diseases) at Peking University Shenzhen Hospital as the developing cohort. Guided by experienced rheumatologists, we built a comprehensive database with 26 clinical features. Using 10 classical machine learning algorithms, we developed screening models. Evaluation metrics determined the best model. Employing multivariatelogistic regression results and the best-performing model to identify the least costly features, ensuring applicability in primary healthcare clinics. Subsequently, we retrained and validated our proposed model based on two primary healthcare validation cohorts.
Results: In experiments, the algorithms achieved over 88% accuracy on training and test sets. Random Forest (RF) excelled with 96.20% (95% CI 95.39% to 97.02%) accuracy, 96.22% (95% CI 95.40% to 97.03%) specificity, 96.18% (95% CI 95.37% to 97.00%) sensitivity, and 96.20% (95% CI 95.39% to 97.02%) Areas Under Curves (AUC). A meticulous feature selection identified 11 key features for RA screening. In an external test on two primary healthcare datasets with these features, RF demonstrated an accuracy of 88.435% (95% CI 85.55% to 91.32%), sensitivity of 98.55% (95% CI 97.47% to 99.63%), specificity of 85.56% (95% CI 82.39% to 88.73%), and an AUC of 92.055% (95% CI 89.62% to 94.49%).
Conclusion: The screening model excels in automating prompt identification of RA in primary healthcare, improving the early detection of RA, and reducing delays and associated costs. Our findings contribute positively and are poised to elevate prospective RA management, fostering improvements in healthcare sector responsiveness and resource efficiency.
期刊介绍:
An international, peer-reviewed, open access, online journal that welcomes laboratory and clinical findings on the molecular basis, cell biology and pharmacology of inflammation.