{"title":"Reflections on the I-squared index for measuring inconsistency in meta-analysis.","authors":"Julian P T Higgins, José A López-López","doi":"10.1017/rsm.2025.10062","DOIUrl":"10.1017/rsm.2025.10062","url":null,"abstract":"<p><p>The I-squared index was proposed in 2002 as a measure to help understand the consistency of study results in a meta-analysis. It was developed to overcome some of the limitations of existing measures, principally the chi-squared test for heterogeneity and the between-study variance as estimated in a random-effects meta-analysis. I-squared measures approximately the proportion of total variability in results that is due to true heterogeneity rather than random error; it is also conveniently interpreted as a measure of inconsistency in the results of the studies. The index has become extremely widely used, although it is often misinterpreted as an absolute measure of the amount of heterogeneity, which it is not. Here, we discuss the I-squared index and the different ways it can be defined, computed, and interpreted. We introduce a new interpretation of I-squared as a weighted sum of squares, which we propose may be helpful when setting up simulation studies. We discuss some of the extensions and repurposes that have been proposed for I-squared and offer some recommendations on the appropriate use of the index in practice.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"389-402"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126227/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147696931","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Suzanne Jak, Mike W-L Cheung, Selcuk Acar, Reuben Kindred
{"title":"Evaluating differences in latent means across studies: Extending meta-analytic confirmatory factor analysis with the analysis of means.","authors":"Suzanne Jak, Mike W-L Cheung, Selcuk Acar, Reuben Kindred","doi":"10.1017/rsm.2025.10057","DOIUrl":"10.1017/rsm.2025.10057","url":null,"abstract":"<p><p>Meta-analytic confirmatory factor analysis (CFA) is a type of meta-analytic structural equation modeling (MASEM) that is useful for evaluating the factor structure of measurement scales based on data from multiple studies. Modeling the factor structure is just one example of the many potentially interesting research questions. Analyzing covariance matrices allows for the evaluation of measurement properties across studies, such as whether indicators are functioning the same across studies. For example, are some indicators more indicative of the common factor in certain types of studies than in others? The additional analysis of means of the observed variables opens up many other research questions to consider such as: \"Are there mean differences in mental health between clinical and non-clinical samples?\" To answer such questions, it is necessary to analyze both the covariance and the mean structure of the indicators. In this paper, we present, illustrate, and evaluate a method to incorporate the means of variables in the MASEM analyses of such datasets. We focus on meta-analytic CFA, with the aim of testing differences in latent means across studies. We provide illustrations of the comparison of latent means across groups of studies using two empirical datasets, for which data and analysis scripts are provided online. The performance of the new model was tested in a small-scale simulation study. The results showed adequate performance under the tested conditions. Finally, we discuss how the proposed method relates to other analysis options such as multigroup or multilevel structural equation modeling.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"498-516"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126220/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147697120","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Strategizing AI utilization for psychological literature screening: A comparative analysis of machine learning algorithms and key factors to consider.","authors":"Lars König, Steffen Zitzmann, Martin Hecht","doi":"10.1017/rsm.2025.10053","DOIUrl":"10.1017/rsm.2025.10053","url":null,"abstract":"<p><p>With the rapid growth of scholarly literature, efficient artificial intelligence (AI)-aided abstract screening tools are becoming increasingly important. This study evaluated 10 different machine learning (ML) algorithms used in AI-aided screening tools for ordering abstracts according to their estimated relevance. We focused on assessing their performance in terms of the number of abstracts required to screen to achieve a sufficient detection rate of relevant articles. Our evaluation included articles screened with diverse inclusion and exclusion criteria. Crucially, we examined how characteristics of the screening data-such as the proportion of relevant articles, the overall frequency of abstracts, and the amount of training data-impacted algorithm effectiveness. Our findings provide valuable insights for researchers across disciplines, highlighting key factors to consider when selecting an ML algorithm and determining a stopping point for AI-aided screening. Specifically, we observed that the algorithm combining the logistic regression (LR) classifier with the sentence-bidirectional encoder representations from transformers (SBERT) feature extractor outperformed other algorithms, demonstrating both the highest efficiency and the lowest variability in performance. Nonetheless, the algorithm's performance varied across experimental conditions. Building on these findings, we discuss the results and provide practical recommendations to assist users in the AI-aided screening process.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"451-482"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126213/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147696940","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Bayes factor hypothesis testing in meta-analyses: Practical advantages and methodological considerations.","authors":"Joris Mulder, Robbie C M van Aert","doi":"10.1017/rsm.2025.10060","DOIUrl":"10.1017/rsm.2025.10060","url":null,"abstract":"<p><p>Bayesian hypothesis testing via Bayes factors offers a principled alternative to classical <i>p</i>-value methods in meta-analysis, particularly suited to its cumulative and sequential nature. Unlike <i>p</i>-values, Bayes factors allow for quantifying support both for and against the existence of an effect, facilitate ongoing evidence monitoring, and maintain coherent long-run behavior as additional studies are incorporated. Recent theoretical developments further show how Bayes factors can flexibly control Type I error rates through connections to e-value theory. Despite these advantages, their use remains limited in the meta-analytic literature. This article provides a critical overview of their theoretical properties, methodological considerations-such as prior sensitivity-and practical advantages for evidence synthesis. Two illustrative applications are provided: one on statistical learning in individuals with language impairments, and another on seroma incidence following post-operative exercise in breast cancer patients. New tools supporting these methods are available in the open-source R package BFpack.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"589-623"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126231/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147697107","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Lasai Barreñada, Bavo De Cock Campo, Laure Wynants, Ben Van Calster
{"title":"Clustered flexible calibration plots for binary outcomes using random effects modeling.","authors":"Lasai Barreñada, Bavo De Cock Campo, Laure Wynants, Ben Van Calster","doi":"10.1017/rsm.2025.10046","DOIUrl":"10.1017/rsm.2025.10046","url":null,"abstract":"<p><p>Evaluation of clinical prediction models across multiple clusters, whether centers or datasets, is becoming increasingly common. A comprehensive evaluation includes an assessment of the agreement between the estimated risks and the observed outcomes, also known as calibration. Calibration is of utmost importance for clinical decision making with prediction models, and it often varies between clusters. We present three approaches to take clustering into account when evaluating calibration: (1) clustered group calibration (CG-C), (2) two-stage meta-analysis calibration (2MA-C), and (3) mixed model calibration (MIX-C), which can obtain flexible calibration plots with random effects modeling and provide confidence interval (CI) and prediction interval (PI). As a case example, we externally validate a model to estimate the risk that an ovarian tumor is malignant in multiple centers (<i>N</i> = 2489). We also conduct a simulation study and a synthetic data study generated from a true clustered dataset to evaluate the methods. In the simulation study, MIX-C and 2MA-C (splines) gave estimated curves closest to the true overall curve. In the synthetic data study, MIX-C produced cluster-specific curves closest to the truth. Coverage of the PI across the plot was best for 2MA-C with splines. We recommend using 2MA-C with splines to estimate the overall curve and 95% PI and MIX-C for cluster-specific curves, especially when the sample size per cluster is limited. We provide ready-to-use code to construct summary flexible calibration curves, with CI and PI to assess heterogeneity in calibration across datasets or centers.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"567-588"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126218/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147697123","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Péter Mátrai, Tamás Kói, Zoltán Sipos, Nelli Farkas
{"title":"Assessing the properties of the prediction interval in random-effects meta-analysis.","authors":"Péter Mátrai, Tamás Kói, Zoltán Sipos, Nelli Farkas","doi":"10.1017/rsm.2025.10055","DOIUrl":"10.1017/rsm.2025.10055","url":null,"abstract":"<p><p>Random-effects meta-analysis is a widely applied methodology to synthesize research findings of studies related to a specific scientific question. Besides estimating the mean effect, an important aim of the meta-analysis is to summarize the heterogeneity, that is, the variation in the underlying effects caused by the differences in study circumstances. The prediction interval is frequently used for this purpose: a 95% prediction interval contains the true effect of a similar new study in 95% of the cases when it is constructed, or in other words, it covers 95% of the true effects distribution on average in repeated sampling. In this article, after providing a clear mathematical background, we present an extensive simulation investigating the performance of all frequentist prediction interval methods published to date. The work focuses on the distribution of the coverage probabilities and how these distributions change depending on the amount of heterogeneity and the number of involved studies. Although the single requirement that a prediction interval has to fulfill is to keep a nominal coverage probability on average, we demonstrate why the distribution of coverages should not be disregarded. We show that for meta-analyses with small number of studies, this distribution has an unideal, asymmetric shape. We argue that assessing only the mean coverage can easily lead to misunderstanding and misinterpretation. The length of the intervals and the robustness of the methods concerning the non-normality of the true effects are also investigated.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"517-537"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126221/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147697104","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Ziren Jiang, Jialing Liu, Weili He, Joseph Cappelleri, Satrajit Roychoudhury, Yong Chen, Haitao Chu
{"title":"The hazards of using hazard ratios from proportional hazard models in indirect treatment comparisons.","authors":"Ziren Jiang, Jialing Liu, Weili He, Joseph Cappelleri, Satrajit Roychoudhury, Yong Chen, Haitao Chu","doi":"10.1017/rsm.2025.10059","DOIUrl":"10.1017/rsm.2025.10059","url":null,"abstract":"<p><p>Indirect treatment comparison (ITC) is widely used to estimate the comparative effectiveness of treatments when head-to-head trials are unavailable. For the typical scenario of anchored ITC where one trial compares drug A to drug C (AC trial) and another compares drug B to drug C (BC trial), the comparative effectiveness of drugs A versus B is calculated by subtracting (or dividing) the relative treatment effect of A versus C in the AC trial by that of B versus C in the BC trial, assuming the covariate distributions in both trials are balanced. This operation is valid only if the chosen effect measure is transitive, that is, in a three-arm randomized trial of drugs A, B, and C, the direct treatment effect of A versus B equals the indirect treatment effect of A versus B through their comparisons to C. For survival outcomes, many ITCs use the hazard ratio (HR) as the effect measure. In this article, we demonstrate that HR is generally not transitive and should be used with caution. As more reliable alternatives, we recommend effect measures with better transitivity properties: the restricted mean survival time (RMST) difference, the landmark survival probability difference (or ratio) at a prespecified time point, and the average hazard with survival weights (AH-SW) difference.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"483-497"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126216/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147696937","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Aggregating and analysing clinical trials data from multiple public registers using R package ctrdata.","authors":"Ralf Herold","doi":"10.1017/rsm.2025.10061","DOIUrl":"10.1017/rsm.2025.10061","url":null,"abstract":"<p><p>The ctrdata package has been created to boost the use of data available in public registers of clinical trials. It enables user-friendly, reproducible workflows to identify trials of interest, download protocol- and results-related data, and conduct sophisticated analyses, across multiple registers and trials. ctrdata works in the widely used R environment, and its databases can be used with other tools. The package is open source with a permissive licence, to facilitate collaboration.This report provides an overview of ctrdata, including its implementation, cases of interest to researchers in public health, medicines, and regulatory science, as well as potential limitations and further developments. At this time, ctrdata works with the European Union (EU) Clinical Trials Information System (CTIS), the EU Clinical Trials Register (EUCTR), the US Clinicaltrials.Gov (CTGOV), and the ISRCTN-the UK's Clinical Study Registry. The registers are complementary in scope and scientific value, yet differences in data models, variable definitions, search parametrisations, and retrieval options hamper efficient scientific workflows, calling for a scientific-technical, programmatic solution and driving the development of ctrdata.By employing ctrdata to comprehensively use and easily leverage trial register data, researchers can effectively address a variety of questions, gain useful insights into evolving policies and practices of drug development, and inform further clinical research. Patients and their organisations, developers, policymakers, and other interested parties can build on ctrdata to create solutions for their use cases.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"624-656"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126229/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147697135","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Transforming evidence synthesis: A systematic review of the evolution of automated meta-analysis in the age of AI.","authors":"Lingbo Li, Anuradha Mathrani, Teo Susnjak","doi":"10.1017/rsm.2025.10065","DOIUrl":"10.1017/rsm.2025.10065","url":null,"abstract":"<p><p>Exponential growth in scientific literature has heightened the demand for efficient evidence-based synthesis, driving the rise of the field of automated meta-analysis (AMA) powered by natural language processing and machine learning. This PRISMA systematic review introduces a structured framework for assessing the current state of AMA, based on screening 13,216 papers (2006-2024) and analyzing 61 studies across diverse domains. Findings reveal a predominant focus on automating data processing (52.5%), such as extraction and statistical modeling, while only 16.4% address advanced synthesis stages. Just one study (approximately 2%) explored preliminary full-process automation, highlighting a critical gap that limits AMA's capacity for comprehensive synthesis. Despite recent breakthroughs in large language models and advanced AI, their integration into statistical modeling and higher-order synthesis, such as heterogeneity assessment and bias evaluation, remains underdeveloped. This has constrained AMA's potential for fully autonomous meta-analysis (MA). From our dataset spanning medical (67.2%) and non-medical (32.8%) applications, we found that AMA has exhibited distinct implementation patterns and varying degrees of effectiveness in actually improving efficiency, scalability, and reproducibility. While automation has enhanced specific meta-analytic tasks, achieving seamless, end-to-end automation remains an open challenge. As AI systems advance in reasoning and contextual understanding, addressing these gaps is now imperative. Future efforts must focus on bridging automation across all MA stages, refining interpretability, and ensuring methodological robustness to fully realize AMA's potential for scalable, domain-agnostic synthesis.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":"17 3","pages":"403-450"},"PeriodicalIF":8.0,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13126215/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147696943","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Meta-analytic-predictive priors based on a single study.","authors":"Christian Röver, Tim Friede","doi":"10.1017/rsm.2026.10081","DOIUrl":"10.1017/rsm.2026.10081","url":null,"abstract":"<p><p>Meta-analytic-predictive (MAP) priors have been proposed as a generic approach to deriving informative prior distributions, where external empirical data are processed to learn about certain parameter distributions. The use of MAP priors is also closely related to shrinkage estimation (also sometimes referred to as <i>dynamic borrowing</i>). A potentially odd situation arises when the external data consist only of <i>a single study</i>. Conceptually, this is not a problem, it only implies that certain prior assumptions gain in importance and need to be specified with particular care. We outline this important, not uncommon special case and demonstrate its implementation and interpretation based on the normal-normal hierarchical model. The approach is illustrated using example applications in clinical medicine.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"1-19"},"PeriodicalIF":8.0,"publicationDate":"2026-03-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13311358/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147502632","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}