Aiman Mahmood, Muhammad Usama Bukhari, Hafsa Afzal, Misbah Sultana, Fatima Rasool, Muzaffar Abbas, Nadeem Irfan Bukhari
{"title":"Effect of Modeling and Transformation Approaches on Predictability and Error Metrics of Design of Experiment-Based Optimization","authors":"Aiman Mahmood, Muhammad Usama Bukhari, Hafsa Afzal, Misbah Sultana, Fatima Rasool, Muzaffar Abbas, Nadeem Irfan Bukhari","doi":"10.1002/cem.70176","DOIUrl":"https://doi.org/10.1002/cem.70176","url":null,"abstract":"<div>\u0000 \u0000 <p>Design of Experiment (DoE) is widely employed for formulation and process optimization. Current project studied the effect of different models and response data transformations, beyond recommended by DoE tool on prediction. DoE tool was used on a case study of central composite design-based multifactor response surface methodology having three factors (X1, X2, and X3) and two responses (Y1 and Y2). Artificial neural network (ANN) software, InForm Version 5, was employed to empirically compare findings. DoE tool recommended quadratic and linear models for prediction, without data transformation for Y1 and Y2. Nevertheless, for Y1 and Y2, natural and base-10 log transformations improved the model performance, as indicated by reduced SSE by 6.0% (Y1) and 3.7% (Y2) and RMSE by 16.1% (Y1) and 2.5% (Y2). With the above transformations, model re-evaluation indicated that DoE recommended quadratic and linear models, respectively for Y1 and Y2 were best suited. With recommended and data driven models, the pairs of quadratic-log and linear-base-10 log- transformations improved model predictability; increased <i>R</i><sup>2</sup> by 3.6%, reduced root mean squared error (RMSE) by 16.1% and decreased predicted residual error squares (PRESS) by 19.7% (Y1) and improved <i>R</i><sup>2</sup> by 0.4%, reduced RMSE by 2.5%, and decreased sum of squared error by 3.7% (Y2). Optimized factor levels generated with above model-transformation pairs for Y1 and Y2 were close to ANN-predicted levels. Indeed, DoE recommendation for data transformation (if required) and model selection must stand out from nonrecommended ones, but present results indicated otherwise warranting enhancement of features in DoE tools.</p>\u0000 </div>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 9","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-08-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148848860","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Evaluation of ATR-FTIR–Based Asphalt Binder Classification Models Under Sample-Level and Cross-Source Conditions","authors":"Juntao Jiao, Erhu Yan, Huisen Xia, Yan Gong, Xinyue Xu, Tingting Xie","doi":"10.1002/cem.70166","DOIUrl":"https://doi.org/10.1002/cem.70166","url":null,"abstract":"<div>\u0000 \u0000 <p>ATR-FTIR spectroscopy combined with chemometric modelling is increasingly used for asphalt binder classification, but model reliability can be overestimated when repeated spectra, related parent samples, and instrument/source differences are not controlled during validation. This study evaluates ATR-FTIR–based asphalt binder classification models under sample-level and cross-source conditions using a multi-source spectral database organized into measurement records, parent samples, and final split units. Binary classification of base/non-SBS and SBS-modified binders was used as the case task. Candidate preprocessing-model pipelines were evaluated at the parent sample–level, and two cross-source validation directions were used to test stability across source conditions. Random spectrum-level splitting produced optimistic performance estimates. SNV+LinearSVM achieved the highest internal main BA (0.9482), but its minimum cross-source BA decreased to 0.5621. In contrast, the selected SNV+PLS-DA–style pipeline achieved comparable internal performance (main BA = 0.9428) while maintaining a higher minimum-direction BA (0.9554) and a smaller direction gap BA (0.0237). Diagnostic experiments showed that cross-source stability depends on both sufficient training sample support and adequate source representation. Latent-variable analysis, coefficient/VIP spectra, and window masking indicated that the model used multiple spectral regions rather than a single SBS marker peak. Source-predictability analysis further showed that source effects were attenuated but not eliminated. The results support a sample-level and cross-source evaluation strategy for reproducible ATR-FTIR chemometric classification.</p>\u0000 </div>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 9","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-08-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148783710","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Comparative Assessment of Similarity Metrics for 1H NMR Spectral Matching: Evidence for the Robustness of the Normalized Square Dot Product","authors":"Fabien Torralba, Guillaume Hoffmann, Emmanuel Cassin, Gwladys Charpentier, Séverine Sechet, Audrey Gaucher, Guillaume Rousselot, Olivier Aroule, Christophe Morell, Asma Bourafai-Aziez","doi":"10.1002/cem.70173","DOIUrl":"https://doi.org/10.1002/cem.70173","url":null,"abstract":"<div>\u0000 \u0000 <p>Automated comparison of <sup>1</sup>H NMR fingerprints is increasingly used for the quality control and authentication of plant extracts; however, there is insufficient guidance on the most reliable similarity metrics for large and heterogeneous spectral datasets. This paper presents a systematic comparison of similarity metrics for <sup>1</sup>H NMR spectra using a large experimental dataset. For this purpose, a large database of 4574 <sup>1</sup>H NMR spectra was constructed (available on request). Spectra were examined both over the full chemical shift range and within three windows: aliphatic (0–3 ppm), aromatic (6–10 ppm) and combined (0–3 + 6–10 ppm). Classical vector-based metrics (Pearson correlation, cosine similarity, Euclidean distance, dot products, normalized area by sum), binary indices (Jaccard, Dice, Russell-Rao) and three recently introduced measures, including the normalized square dot product (NSDP), were applied to all spectrum pairs. Performance was assessed using retrieval statistics (Top 1/Top 10) and bootstrap resampling. Cosine similarity, Pearson correlation, NSDP and a subtraction-based metric yielded high mean similarity values (typically ≥ 0.94), low variance and robust Top 1/Top 10 rankings across all spectral windows, whereas Euclidean distance, raw or squared dot products and binary indices were strongly affected by intensity scaling, threshold choice and peak overlap. The aliphatic region provided the strongest discrimination, whereas combining aliphatic and aromatic regions improved overall robustness of spectral matching. These results support NSDP, cosine similarity, Pearson correlation and subtraction similarity as default metrics for automated <sup>1</sup>H NMR spectral matching in plant extract quality control, dereplication and library-search workflows.</p>\u0000 </div>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 9","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148710500","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Novelty Assessment Method for Streaming Samples Based on Slope Entropy","authors":"Zhonghai He, Haoxiang Zhang, Dongliang Bai, Xiaofang Zhang","doi":"10.1002/cem.70175","DOIUrl":"https://doi.org/10.1002/cem.70175","url":null,"abstract":"<div>\u0000 \u0000 <p>During spectral detection operations, when the streaming samples under test differ significantly from the modeling samples, the prediction error of the spectral model tends to increase. In such cases, it becomes necessary to incorporate novel samples into the modeling set for model updating. A common approach to assess sample novelty involves calculating the spectral residual, typically using the <i>Q</i> residual. However, existing methods generally employ a single distance scalar to evaluate the novelty of multidimensional spectra. This approach suffers from the issue that residuals at different positions are aggregated into a single value, which is overly integrative. Consequently, variations at different locations may lead to the same novelty value. Therefore, a window-based method that evaluates data in segments is more reasonable. The proposed method segments the residual dimensionality using a windowing approach, discretizes the residual values, and classifies and counts the patterns of residual variations within each window. The slope entropy derived from the residual is then used to represent sample novelty. The effectiveness of the proposed method is demonstrated through both simulated and real-world data.</p>\u0000 </div>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 8","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-08-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148753565","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Minsu Son, Hyoju Kim, Hyungjun Kim, Youngho Jin, Jaeoh Kim
{"title":"Reliable Identification and Out-of-Library Detection in Mass Spectra","authors":"Minsu Son, Hyoju Kim, Hyungjun Kim, Youngho Jin, Jaeoh Kim","doi":"10.1002/cem.70174","DOIUrl":"https://doi.org/10.1002/cem.70174","url":null,"abstract":"<p>Accurate molecular identification from mass spectra is essential in analytical workflows, yet conventional library search typically returns the closest match even when the true compound is absent, leading to overconfident false positives. We propose a probabilistic framework that supports both reliable identification of compounds represented in a reference library and model-based flagging of spectra that may not be adequately represented in the reference library. Spectra are modeled within a Bayesian nonparametric method that does not predefine the number of clusters; instead, the model can allocate new clusters when incoming spectra are insufficiently explained by existing references. This property provides a statistical mechanism for flagging potentially unseen compounds while maintaining coherent grouping of known ones. Experiments on large-scale electron ionization libraries demonstrate stable performance under diverse noise conditions, consistent grouping of replicate spectra, and the tendency of spectra absent from the reference database to form new clusters. The framework supports reliable identification while reducing forced matches for spectra that are poorly represented in the reference data.</p>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 8","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-08-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1002/cem.70174","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148752760","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Yingxia Li, Jiajing Zhao, Huan Xue, Nuo Han, Xihui Bian
{"title":"Snake Optimization Algorithm for Variable Selection in Spectral Quantitative Analysis of Complex Samples","authors":"Yingxia Li, Jiajing Zhao, Huan Xue, Nuo Han, Xihui Bian","doi":"10.1002/cem.70172","DOIUrl":"https://doi.org/10.1002/cem.70172","url":null,"abstract":"<div>\u0000 \u0000 <p>The discretized snake optimization algorithm was first proposed as a variable selection method to reduce irrelevant variables and enhance the prediction accuracy of complex samples. In discretized snake optimization (SO), the positions of the snakes were updated, and three transfer functions, V-shaped, arctangent and sigmoid functions, were introduced and compared for discretization. The partial least squares (PLS) model was built using the spectral variables selected by discretized SO. The performance of snake population, transfer functions in SO, and distribution of selected variables for different methods are investigated. To verify the feasibility of SO-PLS, the predictive accuracy of SO-PLS was compared with full-spectrum PLS, uninformative variable elimination-PLS (UVE-PLS), Monte Carlo UVE-PLS (MCUVE-PLS), randomization test-PLS (RT-PLS), grey wolf optimizer-PLS (GWO-PLS), and whale optimization algorithm-PLS (WOA-PLS) models on four complex sample datasets. The results indicate that the V-shaped function is the best transfer function. Compared with the other variable selection methods, SO-PLS uses the least number of variables and gets the best prediction accuracy. Furthermore, compared to PLS, the SO-PLS model reduced the root mean squared error of prediction (RMSEP) by 52%, 43%, 38%, and 23% for predicting protein, sugar, alcohol, and fat in wheat, orange juice, wine, and cocoa bean dataset, respectively. The conresponding correlation coefficients (<i>R</i>) increased from 0.8942, 0.7375, 0.9984, and 0.8121 to 0.9782, 0.8935, 0.9996, and 0.8920, respectively.</p>\u0000 </div>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 8","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-08-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148752618","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Fred T. G. White, Geert Roelof van der Ploeg, Anna Heintz-Buschart, Lemeng Dong, Harro J. Bouwmeester, Age K. Smilde, Johan A. Westerhuis
{"title":"Supervised Restricted Data Fusion With Common, Local, and Distinct Components","authors":"Fred T. G. White, Geert Roelof van der Ploeg, Anna Heintz-Buschart, Lemeng Dong, Harro J. Bouwmeester, Age K. Smilde, Johan A. Westerhuis","doi":"10.1002/cem.70171","DOIUrl":"https://doi.org/10.1002/cem.70171","url":null,"abstract":"<p>In multi-block data, the dominant sources of variation are not always most relevant to a response of interest, meaning that purely exploratory decompositions may fail to recover subtle but important response-associated structure. We introduce PESCAR, a supervised extension of <i>Penalised Exponential Simultaneous Component Analysis</i> (PESCA) that incorporates response information directly into the estimation of common, local and distinct (CLD) structure across multiple data blocks. This allows simultaneous multiblock decomposition and response variable-influenced recovery of latent structure. Through simulation studies, we show that PESCAR can detect weak response-related components across a range of settings, including different noise levels and model-rank mis-specification. Applied to a real multi-omics dataset, PESCAR recovers biologically meaningful response-associated patterns and retains interpretable block structure. We further demonstrate that sparsity in the fitted loading matrices admits a hypergraph-based interpretability layer, summarising overlapping support patterns across components and blocks. These results show that direct incorporation of response information into multiblock decomposition can improve detection of subtle relevant signal and facilitate interpretation in complex systems.</p>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 8","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1002/cem.70171","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148647689","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Gaussian Mixture Modelling and DBSCAN for Reclassifying Legacy pXRF Compositional Data: A Chemometric Workflow Applied to Archaeological Ceramics","authors":"Meligkotsidou Loukia, Liritzis Ioannis","doi":"10.1002/cem.70170","DOIUrl":"https://doi.org/10.1002/cem.70170","url":null,"abstract":"<div>\u0000 \u0000 <p>Legacy portable X-ray fluorescence (pXRF) compositional datasets in archaeometry are often interpreted using principal component analysis and hierarchical clustering, which may incompletely resolve internal structure or outlier behaviour. Here, we reanalyse a published pXRF dataset of 125 archaeological ceramics and experimental clay briquettes using two complementary chemometric classification approaches: Gaussian mixture modelling (GMM) with Bayesian information criterion (BIC) selection, and density-based spatial clustering of applications with noise (DBSCAN). Element/Si ratios were log-transformed prior to analysis. GMM on the first six principal components (retaining 77% of total variance) identified two statistically supported clusters (BIC = −120.7), with probabilistic allocation of samples. DBSCAN (eps = 20, minPts = 18) independently confirmed a two-cluster structure and further distinguished two chemically meaningful outlier subsets, primarily separated by log(S/Si) and log(P/Si) ratios. Both methods show that compositional variability is compatible with local clay sources, clay mixing and firing-temperature effects, rather than non-local imports. Crucially, experimental briquettes co-cluster with archaeological specimens, and firing-induced shifts (e.g., DS8 at 700°C vs. 900°C) are captured as outlier movement. The study demonstrates that probabilistic model-based clustering and density-based unsupervised learning can extract structured technological variability from legacy pXRF data beyond conventional exploratory methods. This chemometric workflow is transferable to other semi-quantitative compositional datasets where new measurements are infeasible.</p>\u0000 </div>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 8","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-07-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148617383","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Adrián Gómez-Sánchez, Berta Torres-Cobos, Rodrigo Rocha de Oliveira
{"title":"Local Asymmetric Least Squares (LAsLS) for Noisy and Complex Baseline Correction","authors":"Adrián Gómez-Sánchez, Berta Torres-Cobos, Rodrigo Rocha de Oliveira","doi":"10.1002/cem.70169","DOIUrl":"https://doi.org/10.1002/cem.70169","url":null,"abstract":"<p>The asymmetric least squares (AsLS) method is commonly used and works well when baseline characteristics are relatively uniform, but can be less accurate for non-uniform, complex, or noisy baselines. Here, we present local AsLS (LAsLS), a local extension of AsLS that preserves the original penalized least-squares framework while allowing asymmetric weights and smoothing penalties to vary across user-defined regions of the signal. This local parameterization provides more accurate baseline estimation for heterogeneous baselines. The method is versatile and can be applied to a variety of techniques, such as chromatography, Raman spectroscopy, and mid-infrared spectroscopy. To test the method under controlled conditions, we compared LAsLS with traditional AsLS using simulated chromatogram-like signals that mimic complex baseline scenarios. Across four cases with heterogeneous baselines, AsLS baseline rRMSE was 2.6–5.7 times higher than LAsLS, and peak quantification rRMSE was 1.4–3.5 times higher. In a controlled PLSR quantitation example, LAsLS also reduced external prediction rRMSE compared with AsLS (11.3% vs. 17.9%). We provide open-source MATLAB implementations with a graphical interface to facilitate method adoption.</p>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 8","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-07-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1002/cem.70169","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148616549","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"EPSpec: An Evidence-Guided, Prior-Retrieval Agent for Near-Infrared Spectral Band Selection","authors":"Shenghao Gu, Mingjian Hong","doi":"10.1002/cem.70167","DOIUrl":"https://doi.org/10.1002/cem.70167","url":null,"abstract":"<div>\u0000 \u0000 <p>To address band selection in high-dimensional continuous near-infrared spectra for quantitative analysis, the evidence-guided, prior-retrieval spectral interval selection (EPSpec) framework is proposed. The method partitions the full spectrum along the wavelength into a set of contiguous candidate intervals and computes multidimensional statistical indicators for each interval to construct a structured evidence table. At the same time, it introduces mechanistic prior knowledge based on the vibrational band evidence index (VIBEX) functional group–band prior library and task-relevant mapping. Subsequently, under the joint guidance of interval evidence and mechanistic prior knowledge, a large language model ranks the candidate intervals by importance and outputs interpretable rationales. Finally, band selection and modeling are completed through Top-K subspectrum search and nested cross-validation without information leakage. Experimental results on three near-infrared regression datasets show that EPSpec combined with partial least squares regression (EPSpec+PLSR) consistently improves upon full-spectrum PLSR while substantially reducing the number of input wavelength variables and maintaining relatively low model complexity. Specifically, on the shootout, soil, and corn datasets, EPSpec+PLSR increases the coefficient of determination (<i>R</i><sup>2</sup>) from 0.9298, 0.9428, and 0.8092 to 0.9372, 0.9657, and 0.9273, respectively, and reduces the root mean square error of prediction (RMSEP) from 4.4856, 2.5369, and 0.2833 to 4.2272, 1.9538, and 0.1671, respectively. Compared with traditional variable selection methods, EPSpec+PLSR achieves overall comparable or locally slightly better results. In particular, on the soil dataset, it achieves the best result while retaining only about 85/1050 wavelengths on average. Ablation experiments further show that interval-level statistical evidence and mechanistic priors are key to stable performance gains, and that sliding-window partitioning can further improve predictive performance.</p>\u0000 </div>","PeriodicalId":15274,"journal":{"name":"Journal of Chemometrics","volume":"40 8","pages":""},"PeriodicalIF":2.0,"publicationDate":"2026-07-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148616624","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}