BiometrikaPub Date : 2026-06-11eCollection Date: 2026-01-01DOI: 10.1093/biomet/asag036
Rebecca Farina, Eric Tchetgen Tchetgen, Arun Kumar Kuchibhotla
{"title":"Doubly robust and efficient calibration of prediction sets for right-censored time-to-event outcomes.","authors":"Rebecca Farina, Eric Tchetgen Tchetgen, Arun Kumar Kuchibhotla","doi":"10.1093/biomet/asag036","DOIUrl":"10.1093/biomet/asag036","url":null,"abstract":"<p><p>Our objective is to construct well-calibrated prediction sets for a time-to-event outcome subject to right censoring with guaranteed coverage. Inspired by modern conformal inference, our approach avoids the need for a well-specified parametric or semiparametric survival model. Unlike existing conformal methods for survival data, which assume Type-I censoring with fully observed censoring times, we consider the more common right-censoring setting in which only the censoring time or the event time is observed, whichever comes first. Under a standard conditional independence censoring condition, we propose and analyse several lower prediction bounds for the survival time of a future observation, including inverse-probability-of-censoring weighting and its augmented version based on the semiparametric efficient influence function for the relevant marginal quantile of the outcome, accounting for dependent censoring. We formally establish asymptotic coverage guarantees for the proposed methods and demonstrate, both theoretically and through empirical experiments, that the augmented approach substantially improves efficiency over all other proposed methods. Specifically, its coverage error bound is doubly robust, and therefore of second order, thus ensuring that it is asymptotically negligible relative to the coverage error of the other methods.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"113 3","pages":"asag036"},"PeriodicalIF":2.8,"publicationDate":"2026-06-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13433195/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148672997","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2026-04-16eCollection Date: 2026-01-01DOI: 10.1093/biomet/asag023
Caitrin Murphy, Eric Laber, Rhonda Merwin, Brian Reich, Jake Koerner
{"title":"Functional principal component analysis forsparse censored data.","authors":"Caitrin Murphy, Eric Laber, Rhonda Merwin, Brian Reich, Jake Koerner","doi":"10.1093/biomet/asag023","DOIUrl":"https://doi.org/10.1093/biomet/asag023","url":null,"abstract":"<p><p>Functional principal component analysis is a key tool in the study of functional data, driving both exploratory analyses and feature construction for use in formal modelling and testing procedures. However, existing methods do not apply when functional observations are censored; for example, when the measurement instrument only supports recordings within a prespecified interval, thereby truncating values outside this range to the nearest boundary. A naïve application of existing methods, without correction for instrument-induced censoring, introduces bias into the estimators of the mean, covariance and functional principal component scores. We extend the functional principal component analysis framework to accommodate noisy and potentially sparse censored functional data. Local loglikelihood maximization is used to recover smooth estimates of the mean and covariance surface that are representative of the latent process's mean and covariance functions. The covariance-smoothing procedure yields a positive semidefinite covariance surface, computed without the need to retroactively remove negative eigenvalues in the covariance-operator decomposition. Additionally, we construct a predictor of the scores, conditional on the censored functional data, and demonstrate its use in the generalized functional linear model. Convergence rates for the proposed estimators are established. In simulation experiments, the proposed method yields improved predictive performance and lower bias than existing alternatives. We illustrate its practical value in a study aimed at classifying eating disorder diagnoses in individuals with type 1 diabetes, using censored functional blood glucose data.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"113 2","pages":"asag023"},"PeriodicalIF":2.8,"publicationDate":"2026-04-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13245921/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148216277","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2026-04-04eCollection Date: 2026-01-01DOI: 10.1093/biomet/asag025
Yonghoon Lee, Edgar Dobriban, Eric J Tchetgen Tchetgen
{"title":"Finding distributions that differ, with false discovery rate control.","authors":"Yonghoon Lee, Edgar Dobriban, Eric J Tchetgen Tchetgen","doi":"10.1093/biomet/asag025","DOIUrl":"10.1093/biomet/asag025","url":null,"abstract":"<p><p>We consider the problem of comparing a reference distribution with several other distributions. Given a sample from both the reference and the comparison groups, we aim to identify the comparison groups whose distributions differ from that of the reference group. Viewing this as a multiple-testing problem, we introduce a methodology that provides exact, distribution-free control of the false discovery rate. To do so, we introduce the concept of <i>batch conformal p -values</i> and demonstrate that they satisfy positive regression dependence across the groups Benjamini & Yekutieli (2001), thereby enabling control of the false discovery rate through the Benjamini-Hochberg procedure. The proof of positive regression dependence introduces a novel technique for the inductive construction of rank vectors with almost-sure dominance under exchangeability. We evaluate the performance of the proposed procedure through simulations. Despite being distribution-free, in some cases it shows performance comparable to methods with knowledge of the data-generating normal distribution, and it further has more power than direct approaches based on conformal out-of-distribution detection. Furthermore, we illustrate our methods on a hepatitis C treatment dataset, where they identify patient groups with large treatment effects, and on the Current Population Survey dataset, where they identify subpopulations with long working hours.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"113 2","pages":"asag025"},"PeriodicalIF":2.8,"publicationDate":"2026-04-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13245922/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148210078","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2026-03-24eCollection Date: 2026-01-01DOI: 10.1093/biomet/asag020
Steven Winter, Omar Melikechi, David B Dunson
{"title":"Sequential Gibbs posteriors with applications to principal component analysis.","authors":"Steven Winter, Omar Melikechi, David B Dunson","doi":"10.1093/biomet/asag020","DOIUrl":"https://doi.org/10.1093/biomet/asag020","url":null,"abstract":"<p><p>Gibbs posteriors are proportional to a prior distribution multiplied by an exponentiated loss function, with a key tuning parameter that weights the information in the loss relative to the prior and provides control of posterior uncertainty. Gibbs posteriors provide a principled framework for likelihood-free Bayesian inference; however, in many situations, the inclusion of a single tuning parameter inevitably leads to poor uncertainty quantification. In particular, regardless of the value of the parameter, credible regions are far from attaining nominal frequentist coverage, even in large samples. We propose a sequential extension to Gibbs posteriors to address this problem. We prove that the proposed sequential posterior exhibits concentration and satisfies a Bernstein-von Mises theorem, which holds under easily verifiable conditions in Euclidean space and on manifolds. As a by-product, we obtain the first Bernstein-von Mises theorem for traditional likelihood-based Bayesian posteriors on manifolds. All methods are illustrated with an application to principal component analysis.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"113 2","pages":"asag020"},"PeriodicalIF":2.8,"publicationDate":"2026-03-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13229586/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148155328","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2026-03-13eCollection Date: 2026-01-01DOI: 10.1093/biomet/asag017
A McClean, Y Li, S Bae, M McAdams DeMarco, I Díaz, W Wu
{"title":"Comparing causal parameters with many treatments and positivity violations.","authors":"A McClean, Y Li, S Bae, M McAdams DeMarco, I Díaz, W Wu","doi":"10.1093/biomet/asag017","DOIUrl":"https://doi.org/10.1093/biomet/asag017","url":null,"abstract":"<p><p>Comparing outcomes across treatments is essential in medicine and public policy. To do so, researchers typically estimate a set of parameters, possibly counterfactual, each targeting adifferent treatment. Treatment-specific means are commonly used, but their identification requires a positivity assumption: every subject has a nonzero probability of receiving each treatment. This assumption is often implausible, especially when treatment can take many values. Causal parameters based on dynamic stochastic interventions offer robustness to positivity violations. However, comparing these parameters may fail to reflect the effects of the underlying target treatments because the parameters can depend on outcomes under nontarget treatments. To clarify when two parameters targeting different treatments yield a useful comparison of treatment efficacy, we propose a comparability criterion: if the conditional treatment-specific mean for one treatment is greater than that for another, then the corresponding causal parameter should also be greater. Many standard parameters fail to satisfy this criterion, but we show that only a mild positivity assumption is needed to identify parameters that yield useful comparisons. We then provide two simple examples that satisfy this criterion and are identifiable under the milder positivity assumption: trimmed and smooth-trimmed treatment-specific means with multivalued treatments. For smooth-trimmed treatment-specific means, we develop doubly robust-style estimators that attain parametric convergence rates under nonparametric conditions. We illustrate our methods with an analysis of dialysis providers in New York State.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"113 2","pages":"asag017"},"PeriodicalIF":2.8,"publicationDate":"2026-03-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13229585/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148155333","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2026-01-01Epub Date: 2025-07-10DOI: 10.1093/biomet/asaf047
B Ren, F Ferrari, S Fortini, S Ventz, L Trippa
{"title":"Leveraging External Data for Testing Experimental Therapies with Biomarker Interactions in Randomized Clinical Trials.","authors":"B Ren, F Ferrari, S Fortini, S Ventz, L Trippa","doi":"10.1093/biomet/asaf047","DOIUrl":"10.1093/biomet/asaf047","url":null,"abstract":"<p><p>In oncology the efficacy of novel therapeutics often differs across patient subgroups, and these variations are difficult to predict during the initial phases of the drug development process. The relation between the power of randomized clinical trials and heterogeneous treatment effects has been discussed by several authors. In particular, false negative results are likely to occur when the treatment effects concentrate in a subpopulation but the study design did not account for potential heterogeneous treatment effects. The use of external data from completed clinical studies and electronic health records has the potential to improve decision-making throughout the development of new therapeutics, from early-stage trials to registration. Here we discuss the use of external data to evaluate experimental treatments with potential heterogeneous treatment effects. We introduce a permutation procedure to test, at the completion of a randomized clinical trial, the null hypothesis that the experimental therapy does not improve the primary outcomes in any subpopulation. The permutation test leverages the available external data to increase power. Also, the procedure controls the false positive rate at the desired <math><mi>α</mi></math> -level without restrictive assumptions on the external data, for example, in scenarios with unmeasured confounders, different pre-treatment patient profiles in the trial population compared to the external data, and other discrepancies between the trial and the external data. We illustrate that the permutation test is optimal according to an interpretable criteria and discuss examples based on asymptotic results and simulations, followed by a retrospective analysis of individual patient-level data from a collection of glioblastoma clinical trials.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"113 1","pages":""},"PeriodicalIF":2.8,"publicationDate":"2026-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13147469/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147833145","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2025-10-29DOI: 10.1093/biomet/asaf073
Ting-Hsuan Chang, Zijian Guo, Daniel Malinsky
{"title":"Post-selection inference for causal effects after causal discovery.","authors":"Ting-Hsuan Chang, Zijian Guo, Daniel Malinsky","doi":"10.1093/biomet/asaf073","DOIUrl":"10.1093/biomet/asaf073","url":null,"abstract":"<p><p>Algorithms for constraint-based causal discovery select graphical causal models among a space of possible candidates (e.g., all directed acyclic graphs) by executing a sequence of conditional independence tests. These may be used to inform the estimation of causal effects (e.g., average treatment effects) when there is uncertainty about which covariates ought to be adjusted for, or which variables act as confounders versus mediators. However, naively using the data twice, for model selection and estimation, would lead to invalid confidence intervals. Moreover, if the selected graph is incorrect, the inferential claims may apply to a selected functional that is distinct from the actual causal effect. We propose an approach to post-selection inference that is based on a resampling and screening procedure, which essentially performs causal discovery multiple times with randomly varying intermediate test statistics. Then, an estimate of the target causal effect and corresponding confidence sets are constructed from a union of individual graph-based estimates and intervals. We show that this construction has asymptotically correct coverage for the true causal effect parameter. Importantly, the guarantee holds for a fixed population-level effect, not a data-dependent or selection-dependent quantity. Most of our exposition focuses on the PC-algorithm for learning directed acyclic graphs and the multivariate Gaussian case for simplicity, but the approach is general and modular, so it may be used with other conditional independence based discovery algorithms and distributional families.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":" ","pages":""},"PeriodicalIF":2.8,"publicationDate":"2025-10-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12849794/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"146084027","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2025-10-06eCollection Date: 2025-01-01DOI: 10.1093/biomet/asaf067
E Smucler, J M Robins, A Rotnitzky
{"title":"On the asymptotic validity of confidence sets for linear functionals of solutions to integral equations.","authors":"E Smucler, J M Robins, A Rotnitzky","doi":"10.1093/biomet/asaf067","DOIUrl":"10.1093/biomet/asaf067","url":null,"abstract":"<p><p>This paper examines the construction of confidence sets for parameters defined as linear functionals of a function of [Formula: see text] and [Formula: see text] whose conditional mean given [Formula: see text] and [Formula: see text] equals the conditional mean of another variable [Formula: see text] given [Formula: see text] and [Formula: see text]. Many estimands of interest in causal inference can be expressed in this form, including the average treatment effect in proximal causal inference and treatment effect contrasts in instrumental variable models. We derive a necessary condition for a confidence set to be uniformly valid over a model that allows for the dependence between [Formula: see text] and [Formula: see text] given [Formula: see text] to be arbitrarily weak. We show that, for any such confidence set, there must exist some laws in the model under which, with high probability, the confidence set has a diameter greater than or equal to the diameter of the parameter's range. In particular, consistent with the weak instrument literature, Wald confidence intervals are not uniformly valid over the aforementioned model when the parameter's range is infinite. Furthermore, we argue that inverting the score test, a successful approach in that literature, generally fails for the broader class of parameters considered here. We present a method for constructing uniformly valid confidence sets when all variables, but possibly [Formula: see text], are binary, discuss its limitations and emphasize that developing valid confidence sets for the class of parameters considered here remains an open problem.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"112 4","pages":"asaf067"},"PeriodicalIF":2.8,"publicationDate":"2025-10-06","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12614171/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145538917","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2025-07-30eCollection Date: 2025-01-01DOI: 10.1093/biomet/asaf059
Michael W Robbins, Lane Burgette
{"title":"Resampling methods with multiply imputed data.","authors":"Michael W Robbins, Lane Burgette","doi":"10.1093/biomet/asaf059","DOIUrl":"10.1093/biomet/asaf059","url":null,"abstract":"<p><p>Resampling techniques have become increasingly popular for estimation of uncertainty. However, data are often fraught with missing values that are commonly imputed to facilitate analysis. This article addresses the issue of using resampling methods such as a jackknife or bootstrap in conjunction with imputations that have been sampled stochastically, in the vein of multiple imputation. We derive the theory needed to illustrate two key points regarding the use of resampling methods in lieu of traditional combining rules. First, imputations should be independently generated multiple times within each replicate group of a jackknife or bootstrap. Second, the number of multiply imputed datasets per replicate group must dramatically exceed the number of replicate groups for a jackknife; however, this is not the case in a bootstrap approach. We also discuss bias-adjusted analogues of the jackknife and bootstrap that are argued to require fewer imputed datasets. A simulation study is provided to support these theoretical conclusions.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":"112 4","pages":"asaf059"},"PeriodicalIF":2.8,"publicationDate":"2025-07-30","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12614170/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145538956","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
BiometrikaPub Date : 2025-07-21DOI: 10.1093/biomet/asaf053
Yinxiang Wu, Hyunseung Kang, Ting Ye
{"title":"A More Robust Approach to Multivariable Mendelian Randomization.","authors":"Yinxiang Wu, Hyunseung Kang, Ting Ye","doi":"10.1093/biomet/asaf053","DOIUrl":"10.1093/biomet/asaf053","url":null,"abstract":"<p><p>Multivariable Mendelian randomization (MVMR) uses genetic variants as instrumental variables to infer the direct effects of multiple exposures on an outcome. However, unlike univariable Mendelian randomization, MVMR often faces greater challenges with many weak instruments, which can lead to bias not necessarily toward zero and inflation of type I errors. In this work, we introduce a new asymptotic regime that allows exposures to have varying degrees of instrument strength, providing a more accurate theoretical framework for studying MVMR estimators. Under this regime, our analysis of the widely used multivariable inverse-variance weighted method shows that it is often biased and tends to produce misleadingly narrow confidence intervals in the presence of many weak instruments. To address this, we propose a simple, closed-form modification to the multivariable inverse-variance weighted estimator to reduce bias from weak instruments, and additionally introduce a novel spectral regularization technique to improve finite-sample performance. We show that the resulting spectral-regularized estimator remains consistent and asymptotically normal under many weak instruments. Through simulations and real data applications, we demonstrate that our proposed estimator and asymptotic framework can enhance the robustness of MVMR analyses.</p>","PeriodicalId":9001,"journal":{"name":"Biometrika","volume":" ","pages":""},"PeriodicalIF":2.8,"publicationDate":"2025-07-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12335017/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"144815700","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}