Enhancing spectral interpretability and band selection prior to prediction model development via CAKE.

IF 6.7 1区 化学 Q1 CHEMISTRY, ANALYTICAL
Talanta Pub Date : 2026-12-01 Epub Date: 2026-06-19 DOI:10.1016/j.talanta.2026.130168
Minh-Quan Nguyen, Mizuki Tsuta, Mito Kokawa
{"title":"Enhancing spectral interpretability and band selection prior to prediction model development via CAKE.","authors":"Minh-Quan Nguyen, Mizuki Tsuta, Mito Kokawa","doi":"10.1016/j.talanta.2026.130168","DOIUrl":null,"url":null,"abstract":"<p><p>A direct, singular causal relationship between the objective variable and spectral data underpins reliable and robust prediction models that avoid spurious correlations. However, machine learning models often lack causal interpretability due to their \"black-box\" nature. To address this, we developed the Causal Analysis via Kernel Estimation (CAKE), a framework that reveals single-component bands using only variables from the calibration model. CAKE employs an information-theoretic approach by calculating the differences in mutual information between variables and regression residuals to indicate causal direction. A Kernel Density Estimation (KDE) classifier further distinguishes causal structures prone to spurious correlations. The framework was optimized on simulated data and validated on real-world spectral measurements. Seven causal structures were constructed for the simulated data, while real-world data came from near-infrared and fluorescence spectroscopy of three-solvent mixtures (dimethyl sulfoxide, ethylene glycol, and glycerol). Three causality types - spectral bands with a single direct cause, hidden causes, and confounders - were distinctly characterized by density functions within the framework, enabling the accurate classification of all single-component bands into their respective causal structures. Compared with post-hoc approaches, such as external model validation and common variable selection indices, CAKE identifies reliable spectral bands that exclude spurious correlations across diverse spectroscopic applications while operating independently of the calibration process and requiring no prior knowledge of pure spectra. Although CAKE is designed as a pre-calibration method rather than a prediction-optimized variable selector, it achieves improved predictive accuracy, demonstrating that causal reliability and practical performance are not mutually exclusive.</p>","PeriodicalId":435,"journal":{"name":"Talanta","volume":"310 ","pages":"130168"},"PeriodicalIF":6.7000,"publicationDate":"2026-12-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Talanta","FirstCategoryId":"92","ListUrlMain":"https://doi.org/10.1016/j.talanta.2026.130168","RegionNum":1,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/6/19 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"CHEMISTRY, ANALYTICAL","Score":null,"Total":0}
引用次数: 0

Abstract

A direct, singular causal relationship between the objective variable and spectral data underpins reliable and robust prediction models that avoid spurious correlations. However, machine learning models often lack causal interpretability due to their "black-box" nature. To address this, we developed the Causal Analysis via Kernel Estimation (CAKE), a framework that reveals single-component bands using only variables from the calibration model. CAKE employs an information-theoretic approach by calculating the differences in mutual information between variables and regression residuals to indicate causal direction. A Kernel Density Estimation (KDE) classifier further distinguishes causal structures prone to spurious correlations. The framework was optimized on simulated data and validated on real-world spectral measurements. Seven causal structures were constructed for the simulated data, while real-world data came from near-infrared and fluorescence spectroscopy of three-solvent mixtures (dimethyl sulfoxide, ethylene glycol, and glycerol). Three causality types - spectral bands with a single direct cause, hidden causes, and confounders - were distinctly characterized by density functions within the framework, enabling the accurate classification of all single-component bands into their respective causal structures. Compared with post-hoc approaches, such as external model validation and common variable selection indices, CAKE identifies reliable spectral bands that exclude spurious correlations across diverse spectroscopic applications while operating independently of the calibration process and requiring no prior knowledge of pure spectra. Although CAKE is designed as a pre-calibration method rather than a prediction-optimized variable selector, it achieves improved predictive accuracy, demonstrating that causal reliability and practical performance are not mutually exclusive.

在利用CAKE开发预测模型之前,提高光谱可解释性和波段选择。
客观变量和光谱数据之间的直接的、单一的因果关系支撑了可靠和稳健的预测模型,避免了虚假的相关性。然而,由于机器学习模型的“黑箱”性质,它们往往缺乏因果可解释性。为了解决这个问题,我们开发了通过核估计的因果分析(CAKE),这是一个仅使用校准模型中的变量来揭示单分量波段的框架。CAKE采用信息论方法,通过计算变量之间互信息的差异和回归残差来指示因果方向。核密度估计(KDE)分类器进一步区分容易产生伪相关的因果结构。该框架在模拟数据上进行了优化,并在实际光谱测量上进行了验证。模拟数据构建了7个因果结构,而真实数据来自三溶剂混合物(二甲基亚砜、乙二醇和甘油)的近红外和荧光光谱。三种因果关系类型-具有单一直接原因、隐藏原因和混杂原因的光谱带-在框架内的密度函数中具有明显的特征,从而能够将所有单组分波段准确分类到各自的因果结构中。与外部模型验证和常用变量选择指数等事后方法相比,CAKE识别出可靠的光谱带,排除了不同光谱应用中的虚假相关性,同时独立于校准过程运行,不需要对纯光谱有任何先验知识。虽然CAKE被设计为一种预校准方法,而不是预测优化变量选择器,但它实现了更高的预测精度,表明因果可靠性和实际性能并不相互排斥。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Talanta
Talanta 化学-分析化学
CiteScore
12.30
自引率
4.90%
发文量
861
审稿时长
29 days
期刊介绍: Talanta provides a forum for the publication of original research papers, short communications, and critical reviews in all branches of pure and applied analytical chemistry. Papers are evaluated based on established guidelines, including the fundamental nature of the study, scientific novelty, substantial improvement or advantage over existing technology or methods, and demonstrated analytical applicability. Original research papers on fundamental studies, and on novel sensor and instrumentation developments, are encouraged. Novel or improved applications in areas such as clinical and biological chemistry, environmental analysis, geochemistry, materials science and engineering, and analytical platforms for omics development are welcome. Analytical performance of methods should be determined, including interference and matrix effects, and methods should be validated by comparison with a standard method, or analysis of a certified reference material. Simple spiking recoveries may not be sufficient. The developed method should especially comprise information on selectivity, sensitivity, detection limits, accuracy, and reliability. However, applying official validation or robustness studies to a routine method or technique does not necessarily constitute novelty. Proper statistical treatment of the data should be provided. Relevant literature should be cited, including related publications by the authors, and authors should discuss how their proposed methodology compares with previously reported methods.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书