{"title":"Interpretable machine learning classification of cedar and cypress pollen on routine Durham slides for environmental exposure assessment.","authors":"Nobuyoshi Suzuki, Kenjiro Sugiyama, Katsuhiko Kobayashi, Hisakuni Fukuoka, Yutaka Takumi","doi":"10.3389/falgy.2026.1805985","DOIUrl":null,"url":null,"abstract":"<p><p>Accurate discrimination of <i>Cryptomeria japonica</i> (cedar) and <i>Chamaecyparis obtusa</i> (cypress) pollen on routine Durham slides is clinically and environmentally important, because local pollen counts influence patient visits and regional exposure assessment. However, manual counting is time-consuming and often complicated by debris, air bubbles, and burst pollen, which is morphologically distinct from intact grains. In this methodological proof-of-concept study, we evaluated whether machine learning could classify particles cropped from real-world Durham slides collected under routine field conditions. We collected five routine slides from two sites in Nagano Prefecture and obtained 1,480 particle images categorized into five classes: cedar, cypress, burst cedar, burst cypress, and dust/miscellaneous artifacts. We extracted interpretable morphological and textural descriptors and trained a support vector machine using nested 5-fold stratified cross-validation. Using an optimized, interpretable feature set, the model achieved a macro-F1 score of 0.833 ± 0.025 and an overall accuracy of 0.863. Classification of intact cedar and cypress pollen was good, whereas discrimination between the two burst classes was more difficult. Misclassifications were concentrated mainly between cedar and cypress, and between burst cedar and morphologically similar particles such as burst cypress or dust. In contrast, dust was rarely misclassified as intact cedar or intact cypress. One-vs.-rest ROC-AUC values were high across classes (0.93-0.98), although performance was lower for burst pollen than for intact pollen. Among the extracted descriptors, size-related features, particularly area and radius, contributed most strongly to classification. These findings show that interpretable machine learning can distinguish intact cedar and cypress pollen on routine Durham slides under real-world conditions, while burst pollen remains a major source of classification difficulty. Further refinement of feature design for burst pollen will be necessary for routine application.</p>","PeriodicalId":73062,"journal":{"name":"Frontiers in allergy","volume":"7 ","pages":"1805985"},"PeriodicalIF":3.4000,"publicationDate":"2026-05-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13189800/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Frontiers in allergy","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.3389/falgy.2026.1805985","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/1/1 0:00:00","PubModel":"eCollection","JCR":"Q2","JCRName":"ALLERGY","Score":null,"Total":0}
引用次数: 0
Abstract
Accurate discrimination of Cryptomeria japonica (cedar) and Chamaecyparis obtusa (cypress) pollen on routine Durham slides is clinically and environmentally important, because local pollen counts influence patient visits and regional exposure assessment. However, manual counting is time-consuming and often complicated by debris, air bubbles, and burst pollen, which is morphologically distinct from intact grains. In this methodological proof-of-concept study, we evaluated whether machine learning could classify particles cropped from real-world Durham slides collected under routine field conditions. We collected five routine slides from two sites in Nagano Prefecture and obtained 1,480 particle images categorized into five classes: cedar, cypress, burst cedar, burst cypress, and dust/miscellaneous artifacts. We extracted interpretable morphological and textural descriptors and trained a support vector machine using nested 5-fold stratified cross-validation. Using an optimized, interpretable feature set, the model achieved a macro-F1 score of 0.833 ± 0.025 and an overall accuracy of 0.863. Classification of intact cedar and cypress pollen was good, whereas discrimination between the two burst classes was more difficult. Misclassifications were concentrated mainly between cedar and cypress, and between burst cedar and morphologically similar particles such as burst cypress or dust. In contrast, dust was rarely misclassified as intact cedar or intact cypress. One-vs.-rest ROC-AUC values were high across classes (0.93-0.98), although performance was lower for burst pollen than for intact pollen. Among the extracted descriptors, size-related features, particularly area and radius, contributed most strongly to classification. These findings show that interpretable machine learning can distinguish intact cedar and cypress pollen on routine Durham slides under real-world conditions, while burst pollen remains a major source of classification difficulty. Further refinement of feature design for burst pollen will be necessary for routine application.