{"title":"The accuracy and replicability of data analysis performed with large language models","authors":"Justus Eaglesmith, Tim Johnson, Robert W. Walker","doi":"10.1080/00031305.2026.2688943","DOIUrl":"https://doi.org/10.1080/00031305.2026.2688943","url":null,"abstract":"","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"97 1","pages":"1-4"},"PeriodicalIF":1.8,"publicationDate":"2026-06-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148288074","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Foundations of Multiple Regression and Analysis of Variance","authors":"Jun Wu","doi":"10.1080/00031305.2026.2684948","DOIUrl":"https://doi.org/10.1080/00031305.2026.2684948","url":null,"abstract":"","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"4 1","pages":"1-2"},"PeriodicalIF":1.8,"publicationDate":"2026-06-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148288076","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Are Statistical Methods Obsolete in the Era of Deep Learning? A Study of ODE Inverse Problems","authors":"Skyler Wu, Shihao Yang, S. C. Kou","doi":"10.1080/00031305.2026.2669086","DOIUrl":"https://doi.org/10.1080/00031305.2026.2669086","url":null,"abstract":"In the era of AI, neural networks have become increasingly popular for modeling, inference, and prediction, largely due to their potential for universal approximation. With the proliferation of such deep learning models, a question arises: are leaner statistical methods still relevant? To shed insight on this question, we employ the mechanistic nonlinear ordinary differential equation (ODE) inverse problem as a testbed, using the physics-informed neural network (PINN) as a representative of the deep learning paradigm and manifold-constrained Gaussian process inference (MAGI) as a representative of statistically principled methods. Through case studies involving the SEIR model from epidemiology and the Lorenz model from chaotic dynamics, we demonstrate that statistical methods are far from obsolete, especially when working with sparse and noisy observations. On tasks such as parameter inference and trajectory reconstruction, statistically principled methods consistently achieve lower bias and variance, while using far fewer parameters and requiring less hyperparameter tuning. Statistical methods can also decisively outperform deep learning models on out-of-sample future prediction, where the absence of relevant data often leads overparameterized models astray. Additionally, we find that statistically principled approaches are more robust to accumulation of numerical imprecision and can represent the underlying system more faithfully to the true governing ODEs.","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"35 1","pages":"1-16"},"PeriodicalIF":1.8,"publicationDate":"2026-05-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148287385","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Irantzu Barrio, Javier Roca-Pardiñas, Cristobal Esteban, Maria Durban
{"title":"Proposal of a general framework to categorize continuous predictor variables","authors":"Irantzu Barrio, Javier Roca-Pardiñas, Cristobal Esteban, Maria Durban","doi":"10.1080/00031305.2026.2658195","DOIUrl":"https://doi.org/10.1080/00031305.2026.2658195","url":null,"abstract":"The use of discretized variables in the development of prediction models is a common practice, in part because the decision-making process is more natural when it is based on rules created from segmented models. Although this practice is perhaps more common in medicine, it is extensible to any area of knowledge where a predictive model helps in decision-making. Therefore, providing researchers with a useful and valid categorization method could be a relevant issue when developing prediction models.In this paper, we propose a new general methodology that can be applied to categorize a predictor variable in any regression model where the response variable belongs to the exponential family distribution. Furthermore, it can be applied in any multivariate context, allowing to categorize more than one continuous covariate simultaneously. In addition, a computationally very efficient method is proposed to obtain the optimal number of categories, based on a pseudo-BIC proposal. Several simulation studies have been conducted in which the efficiency of the method with respect to both the location and the number of estimated cut-off points is shown.Finally, the categorization proposal has been illustrated in a real data set of 543 patients with chronic obstructive pulmonary disease (COPD) from Galdakao Hospital’s five outpatient respiratory clinics, who were followed up for 10 years. Exercise capacity is known to be an important predictor of adverse events in patients with COPD. In this paper, we applied the proposed methodology to jointly categorize the continuous variables six-minute walking test and forced expiratory volume in one second in a multiple Poisson generalized additive model for the response variable rate of the number of hospital admissions by years of follow-up. The location and number of cut-off points obtained were clinically validated as being in line with the categorizations used in the literature.","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"21 1","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-04-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147681792","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Wenlong Ji, Weizhe Yuan, Emily Getzen, Kyunghyun Cho, Michael I. Jordan, Song Mei, Jason Weston, Weijie J. Su, Jing Xu, Linjun Zhang
{"title":"An Overview of Large Language Models for Statisticians","authors":"Wenlong Ji, Weizhe Yuan, Emily Getzen, Kyunghyun Cho, Michael I. Jordan, Song Mei, Jason Weston, Weijie J. Su, Jing Xu, Linjun Zhang","doi":"10.1080/00031305.2026.2657480","DOIUrl":"https://doi.org/10.1080/00031305.2026.2657480","url":null,"abstract":"Large Language Models (LLMs) have emerged as transformative tools in artificial intelligence (AI), exhibiting remarkable capabilities across diverse tasks such as text generation, reasoning, and decision-making. While their success has primarily been driven by advances in computational power and deep learning architectures, emerging problems—in areas such as uncertainty quantification, decision-making, causal inference, and distribution shift—require a deeper engagement with the field of statistics. This paper explores potential areas where statisticians can make important contributions to the development of LLMs, particularly those that aim to engender trustworthiness and transparency for human users. Thus, we focus on issues such as uncertainty quantification, interpretability, algorithmic fairness, privacy, watermarking and model adaptation. We also consider possible roles for LLMs in statistical analysis. By bridging AI and statistics, we aim to foster a deeper collaboration that advances both the theoretical foundations and practical applications of LLMs, ultimately shaping their role in addressing complex societal challenges.","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"1 1","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-04-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147695434","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"A Markov Inequality-Inspired Stochastic Bound for Tail Events","authors":"Joan del Castillo, Pedro Puig","doi":"10.1080/00031305.2026.2657483","DOIUrl":"https://doi.org/10.1080/00031305.2026.2657483","url":null,"abstract":"This article presents methods for estimating extreme probabilities, beyond the range of the observations. These methods are model-free and applicable to almost any sample size. They are grounded in order statistics theory and have a wide range of applications, as they simply require the assumption of a finite expectation. Even in cases when a particular risk model exists, the new methods provide clarity, security and simplicity. The methodology is applicable to the behavior of financial markets, and the results are comparable to those provided by extreme value theory.","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"95 4 1","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-04-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147681793","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Taylor’s Theorem and Mean Value Theorem for Random Functions and Random Variables","authors":"Yifan Yang, Xiaoyu Zhou, Ming Wang","doi":"10.1080/00031305.2026.2657484","DOIUrl":"https://doi.org/10.1080/00031305.2026.2657484","url":null,"abstract":"This study addresses the often-overlooked issue of measurability at intermediate points when applying Taylor’s theorems to random functions and random vectors (e.g., likelihood functions with respect to estimators) in statistics. Classical Taylor-related theorems were originally developed for deterministic settings. Consequently, they do not directly extend to stochastic functions and variables and do not inherently guarantee the measurability of intermediate points. In statistical contexts, applying these theorems without properly accounting for randomness can lead to analyses that lack well-defined probabilistic interpretations. Elementary approaches, such as pointwise constructions, are insufficient for handling random quantities and establishing measurable intermediate points. Moreover, some statistical literature has implicitly disregarded this issue, often neglecting the stochastic nature of the problem and assuming that intermediate points are measurable. To address this gap, we develop multivariate Taylor’s and mean value theorems tailored for random functions and random variables under mild assumptions. We provide illustrative examples demonstrating the applicability of our results to commonly used statistical methods, including maximum likelihood estimation, <i>M</i>-estimation, and profile estimation. Our findings contribute a rigorous foundation for the applications of Taylor expansions in statistics.","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"15 1","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-04-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147684728","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Statistical inference for spatio-temporal autoregressive models of covariates with additive measurement errors","authors":"Zhensheng Huang, Yueyi Wu, Weihao Yu","doi":"10.1080/00031305.2026.2656374","DOIUrl":"https://doi.org/10.1080/00031305.2026.2656374","url":null,"abstract":"In this paper, we address the statistical inference problem proposed for the sparse spatio-temporal autoregressive models with additive measurement error when the number of spatial nodes exceeds the number of temporal observations. We use the improved Yule-Walker estimation method, adding the bagging algorithm to the estimation process to solve the over-identification problem. The simulation-extrapolation (SIMEX) method is used to reduce the influence of additive measurement error and we confirm the feasibility of empirical likelihood method to establish confidence intervals for model coefficients. Furthermore, some simulations and real examples are carried out to evaluate the finite sample performance.","PeriodicalId":50801,"journal":{"name":"American Statistician","volume":"20 1","pages":""},"PeriodicalIF":1.8,"publicationDate":"2026-04-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147648958","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}