Raja Waqar Ahmed Khan, Hamayun Shaheen, Muhammad Ejaz Ul Islam Dar, Tariq Habib, Muhammad Manzoor, Syed Waseem Gillani, Abeer Al-Andal, John Oluwafemi Ayoola, Muhammad Waheed
{"title":"A data-driven approach to forest health assessment through multivariate analysis and machine learning techniques.","authors":"Raja Waqar Ahmed Khan, Hamayun Shaheen, Muhammad Ejaz Ul Islam Dar, Tariq Habib, Muhammad Manzoor, Syed Waseem Gillani, Abeer Al-Andal, John Oluwafemi Ayoola, Muhammad Waheed","doi":"10.1186/s12870-025-06937-5","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Himalayan forests are fragile, rich in biodiversity, and face increasing threats from anthropogenic pressures and climate change. Assessing their health is critical for sustainable forest management. This study integrated ecological indicators (tree density, size, regeneration, deforestation, slope, grazing, and erosion) with machine learning (ML) to classify forest health and identify key drivers across 37 Western Himalayan sites. Principal component analysis (PCA) reduced data dimensionality, highlighting major ecological gradients. K-means clustering was used to group forests into three distinct classes based on ecological characteristics, due to its efficiency in identifying natural patterns within multivariate data. ML models, including Decision Tree (DT), Random Forest (RF), and Support Vector Machine (SVM) were trained and validated using an 80:20 train-test split and 5-fold cross-validation.</p><p><strong>Results: </strong>PCA revealed that elevation, disturbance, and regeneration explained 74.3% variance. Forest health varied across sites, with 10 categorized as healthy, 19 as moderate, and 8 as unhealthy. Forest regeneration was highly skewed (2.67) and leptokurtic (9.8), with few sites showing high seedling abundance, while deforestation (mean = 294 stumps ha<sup>-1</sup>) indicated uneven human impact. Among ML models, RF showed the best performance with a mean accuracy of 0.83, Kappa 0.87, and balanced accuracy 0.88. SVM followed with 0.75 accuracy, Kappa 0.70, and balanced accuracy 0.81. DT performed lowest with 0.66 accuracy and Kappa 0.45. Cross-validation confirmed RF's highest mean accuracy (90.3%), followed by SVM (88.1%) and DT (65.1%). RF-based feature importance analysis showed tree DBH, height, regeneration rate, soil erosion, and tree density as key ecological drivers of forest health.</p><p><strong>Conclusions: </strong>This study highlights ML-driven classification as a precise, scalable, and objective tool for large-scale forest health assessments. Conservation efforts should prioritize degraded forests through afforestation, slope stabilization, controlled grazing, erosion management, and continuous ecosystem monitoring. Future studies should integrate high-resolution remote sensing (e.g., Landsat, Sentinel-2) and climate datasets (e.g., temperature, precipitation, and drought indices) to enhance predictive capabilities and support long-term forest management planning. The findings underscore the value of data-driven approaches, establishing machine learning as an effective tool to enhance forest monitoring and support evidence-based forest conservation and management in the Himalayas.</p>","PeriodicalId":9198,"journal":{"name":"BMC Plant Biology","volume":"25 1","pages":"915"},"PeriodicalIF":4.3000,"publicationDate":"2025-07-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12261679/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"BMC Plant Biology","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1186/s12870-025-06937-5","RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"PLANT SCIENCES","Score":null,"Total":0}
引用次数: 0
Abstract
Background: Himalayan forests are fragile, rich in biodiversity, and face increasing threats from anthropogenic pressures and climate change. Assessing their health is critical for sustainable forest management. This study integrated ecological indicators (tree density, size, regeneration, deforestation, slope, grazing, and erosion) with machine learning (ML) to classify forest health and identify key drivers across 37 Western Himalayan sites. Principal component analysis (PCA) reduced data dimensionality, highlighting major ecological gradients. K-means clustering was used to group forests into three distinct classes based on ecological characteristics, due to its efficiency in identifying natural patterns within multivariate data. ML models, including Decision Tree (DT), Random Forest (RF), and Support Vector Machine (SVM) were trained and validated using an 80:20 train-test split and 5-fold cross-validation.
Results: PCA revealed that elevation, disturbance, and regeneration explained 74.3% variance. Forest health varied across sites, with 10 categorized as healthy, 19 as moderate, and 8 as unhealthy. Forest regeneration was highly skewed (2.67) and leptokurtic (9.8), with few sites showing high seedling abundance, while deforestation (mean = 294 stumps ha-1) indicated uneven human impact. Among ML models, RF showed the best performance with a mean accuracy of 0.83, Kappa 0.87, and balanced accuracy 0.88. SVM followed with 0.75 accuracy, Kappa 0.70, and balanced accuracy 0.81. DT performed lowest with 0.66 accuracy and Kappa 0.45. Cross-validation confirmed RF's highest mean accuracy (90.3%), followed by SVM (88.1%) and DT (65.1%). RF-based feature importance analysis showed tree DBH, height, regeneration rate, soil erosion, and tree density as key ecological drivers of forest health.
Conclusions: This study highlights ML-driven classification as a precise, scalable, and objective tool for large-scale forest health assessments. Conservation efforts should prioritize degraded forests through afforestation, slope stabilization, controlled grazing, erosion management, and continuous ecosystem monitoring. Future studies should integrate high-resolution remote sensing (e.g., Landsat, Sentinel-2) and climate datasets (e.g., temperature, precipitation, and drought indices) to enhance predictive capabilities and support long-term forest management planning. The findings underscore the value of data-driven approaches, establishing machine learning as an effective tool to enhance forest monitoring and support evidence-based forest conservation and management in the Himalayas.
期刊介绍:
BMC Plant Biology is an open access, peer-reviewed journal that considers articles on all aspects of plant biology, including molecular, cellular, tissue, organ and whole organism research.