Interpretable machine learning classification of cedar and cypress pollen on routine Durham slides for environmental exposure assessment.

IF 3.4 Q2 ALLERGY
Frontiers in allergy Pub Date : 2026-05-07 eCollection Date: 2026-01-01 DOI:10.3389/falgy.2026.1805985
Nobuyoshi Suzuki, Kenjiro Sugiyama, Katsuhiko Kobayashi, Hisakuni Fukuoka, Yutaka Takumi
{"title":"Interpretable machine learning classification of cedar and cypress pollen on routine Durham slides for environmental exposure assessment.","authors":"Nobuyoshi Suzuki, Kenjiro Sugiyama, Katsuhiko Kobayashi, Hisakuni Fukuoka, Yutaka Takumi","doi":"10.3389/falgy.2026.1805985","DOIUrl":null,"url":null,"abstract":"<p><p>Accurate discrimination of <i>Cryptomeria japonica</i> (cedar) and <i>Chamaecyparis obtusa</i> (cypress) pollen on routine Durham slides is clinically and environmentally important, because local pollen counts influence patient visits and regional exposure assessment. However, manual counting is time-consuming and often complicated by debris, air bubbles, and burst pollen, which is morphologically distinct from intact grains. In this methodological proof-of-concept study, we evaluated whether machine learning could classify particles cropped from real-world Durham slides collected under routine field conditions. We collected five routine slides from two sites in Nagano Prefecture and obtained 1,480 particle images categorized into five classes: cedar, cypress, burst cedar, burst cypress, and dust/miscellaneous artifacts. We extracted interpretable morphological and textural descriptors and trained a support vector machine using nested 5-fold stratified cross-validation. Using an optimized, interpretable feature set, the model achieved a macro-F1 score of 0.833 ± 0.025 and an overall accuracy of 0.863. Classification of intact cedar and cypress pollen was good, whereas discrimination between the two burst classes was more difficult. Misclassifications were concentrated mainly between cedar and cypress, and between burst cedar and morphologically similar particles such as burst cypress or dust. In contrast, dust was rarely misclassified as intact cedar or intact cypress. One-vs.-rest ROC-AUC values were high across classes (0.93-0.98), although performance was lower for burst pollen than for intact pollen. Among the extracted descriptors, size-related features, particularly area and radius, contributed most strongly to classification. These findings show that interpretable machine learning can distinguish intact cedar and cypress pollen on routine Durham slides under real-world conditions, while burst pollen remains a major source of classification difficulty. Further refinement of feature design for burst pollen will be necessary for routine application.</p>","PeriodicalId":73062,"journal":{"name":"Frontiers in allergy","volume":"7 ","pages":"1805985"},"PeriodicalIF":3.4000,"publicationDate":"2026-05-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13189800/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Frontiers in allergy","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.3389/falgy.2026.1805985","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/1/1 0:00:00","PubModel":"eCollection","JCR":"Q2","JCRName":"ALLERGY","Score":null,"Total":0}
引用次数: 0

Abstract

Accurate discrimination of Cryptomeria japonica (cedar) and Chamaecyparis obtusa (cypress) pollen on routine Durham slides is clinically and environmentally important, because local pollen counts influence patient visits and regional exposure assessment. However, manual counting is time-consuming and often complicated by debris, air bubbles, and burst pollen, which is morphologically distinct from intact grains. In this methodological proof-of-concept study, we evaluated whether machine learning could classify particles cropped from real-world Durham slides collected under routine field conditions. We collected five routine slides from two sites in Nagano Prefecture and obtained 1,480 particle images categorized into five classes: cedar, cypress, burst cedar, burst cypress, and dust/miscellaneous artifacts. We extracted interpretable morphological and textural descriptors and trained a support vector machine using nested 5-fold stratified cross-validation. Using an optimized, interpretable feature set, the model achieved a macro-F1 score of 0.833 ± 0.025 and an overall accuracy of 0.863. Classification of intact cedar and cypress pollen was good, whereas discrimination between the two burst classes was more difficult. Misclassifications were concentrated mainly between cedar and cypress, and between burst cedar and morphologically similar particles such as burst cypress or dust. In contrast, dust was rarely misclassified as intact cedar or intact cypress. One-vs.-rest ROC-AUC values were high across classes (0.93-0.98), although performance was lower for burst pollen than for intact pollen. Among the extracted descriptors, size-related features, particularly area and radius, contributed most strongly to classification. These findings show that interpretable machine learning can distinguish intact cedar and cypress pollen on routine Durham slides under real-world conditions, while burst pollen remains a major source of classification difficulty. Further refinement of feature design for burst pollen will be necessary for routine application.

用于环境暴露评估的常规达勒姆载玻片上雪松和柏树花粉的可解释机器学习分类。
在常规达勒姆载玻片上准确鉴别杉木(Cryptomeria japonica)和柏树(Chamaecyparis obtusa)花粉具有重要的临床和环境意义,因为局部花粉计数影响患者就诊和区域暴露评估。然而,人工计数是费时的,并且经常被碎片、气泡和破裂的花粉所复杂,这些花粉在形态上与完整的谷物不同。在这项方法学概念验证研究中,我们评估了机器学习是否可以对常规现场条件下收集的真实达勒姆载玻片中的颗粒进行分类。我们从长野县的两个地点收集了5张常规幻灯片,获得了1480张颗粒图像,分为5类:雪松、柏树、爆发雪松、爆发柏树和灰尘/杂项文物。我们提取了可解释的形态和纹理描述符,并使用嵌套的5倍分层交叉验证训练了一个支持向量机。利用优化后的可解释特征集,该模型的宏观f1得分为0.833±0.025,整体准确率为0.863。完整的雪松和柏树花粉分类较好,而两个突发类的区分较困难。错误分类主要集中在雪松和柏树之间,以及雪松与柏树或粉尘等形态相似的颗粒之间。相比之下,灰尘很少被错误地分类为完整的雪松或完整的柏树。One-vs。其余各组的ROC-AUC值较高(0.93-0.98),但破碎花粉的性能低于完整花粉。在提取的描述符中,与尺寸相关的特征,特别是面积和半径,对分类的贡献最大。这些发现表明,在现实条件下,可解释的机器学习可以区分常规达勒姆载玻片上完整的雪松和柏树花粉,而爆裂花粉仍然是分类困难的主要来源。在常规应用中,有必要进一步改进爆裂花粉的特征设计。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
CiteScore
2.80
自引率
0.00%
发文量
0
审稿时长
12 weeks
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书