{"title":"Enhancing spectral interpretability and band selection prior to prediction model development via CAKE.","authors":"Minh-Quan Nguyen, Mizuki Tsuta, Mito Kokawa","doi":"10.1016/j.talanta.2026.130168","DOIUrl":null,"url":null,"abstract":"<p><p>A direct, singular causal relationship between the objective variable and spectral data underpins reliable and robust prediction models that avoid spurious correlations. However, machine learning models often lack causal interpretability due to their \"black-box\" nature. To address this, we developed the Causal Analysis via Kernel Estimation (CAKE), a framework that reveals single-component bands using only variables from the calibration model. CAKE employs an information-theoretic approach by calculating the differences in mutual information between variables and regression residuals to indicate causal direction. A Kernel Density Estimation (KDE) classifier further distinguishes causal structures prone to spurious correlations. The framework was optimized on simulated data and validated on real-world spectral measurements. Seven causal structures were constructed for the simulated data, while real-world data came from near-infrared and fluorescence spectroscopy of three-solvent mixtures (dimethyl sulfoxide, ethylene glycol, and glycerol). Three causality types - spectral bands with a single direct cause, hidden causes, and confounders - were distinctly characterized by density functions within the framework, enabling the accurate classification of all single-component bands into their respective causal structures. Compared with post-hoc approaches, such as external model validation and common variable selection indices, CAKE identifies reliable spectral bands that exclude spurious correlations across diverse spectroscopic applications while operating independently of the calibration process and requiring no prior knowledge of pure spectra. Although CAKE is designed as a pre-calibration method rather than a prediction-optimized variable selector, it achieves improved predictive accuracy, demonstrating that causal reliability and practical performance are not mutually exclusive.</p>","PeriodicalId":435,"journal":{"name":"Talanta","volume":"310 ","pages":"130168"},"PeriodicalIF":6.7000,"publicationDate":"2026-12-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Talanta","FirstCategoryId":"92","ListUrlMain":"https://doi.org/10.1016/j.talanta.2026.130168","RegionNum":1,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/6/19 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"CHEMISTRY, ANALYTICAL","Score":null,"Total":0}
引用次数: 0
Abstract
A direct, singular causal relationship between the objective variable and spectral data underpins reliable and robust prediction models that avoid spurious correlations. However, machine learning models often lack causal interpretability due to their "black-box" nature. To address this, we developed the Causal Analysis via Kernel Estimation (CAKE), a framework that reveals single-component bands using only variables from the calibration model. CAKE employs an information-theoretic approach by calculating the differences in mutual information between variables and regression residuals to indicate causal direction. A Kernel Density Estimation (KDE) classifier further distinguishes causal structures prone to spurious correlations. The framework was optimized on simulated data and validated on real-world spectral measurements. Seven causal structures were constructed for the simulated data, while real-world data came from near-infrared and fluorescence spectroscopy of three-solvent mixtures (dimethyl sulfoxide, ethylene glycol, and glycerol). Three causality types - spectral bands with a single direct cause, hidden causes, and confounders - were distinctly characterized by density functions within the framework, enabling the accurate classification of all single-component bands into their respective causal structures. Compared with post-hoc approaches, such as external model validation and common variable selection indices, CAKE identifies reliable spectral bands that exclude spurious correlations across diverse spectroscopic applications while operating independently of the calibration process and requiring no prior knowledge of pure spectra. Although CAKE is designed as a pre-calibration method rather than a prediction-optimized variable selector, it achieves improved predictive accuracy, demonstrating that causal reliability and practical performance are not mutually exclusive.
期刊介绍:
Talanta provides a forum for the publication of original research papers, short communications, and critical reviews in all branches of pure and applied analytical chemistry. Papers are evaluated based on established guidelines, including the fundamental nature of the study, scientific novelty, substantial improvement or advantage over existing technology or methods, and demonstrated analytical applicability. Original research papers on fundamental studies, and on novel sensor and instrumentation developments, are encouraged. Novel or improved applications in areas such as clinical and biological chemistry, environmental analysis, geochemistry, materials science and engineering, and analytical platforms for omics development are welcome.
Analytical performance of methods should be determined, including interference and matrix effects, and methods should be validated by comparison with a standard method, or analysis of a certified reference material. Simple spiking recoveries may not be sufficient. The developed method should especially comprise information on selectivity, sensitivity, detection limits, accuracy, and reliability. However, applying official validation or robustness studies to a routine method or technique does not necessarily constitute novelty. Proper statistical treatment of the data should be provided. Relevant literature should be cited, including related publications by the authors, and authors should discuss how their proposed methodology compares with previously reported methods.