Gerit Wagner, Julian Prester, Roman Lukyanenko, Guy Paré
{"title":"Data management in literature reviews: The C5-DM Framework.","authors":"Gerit Wagner, Julian Prester, Roman Lukyanenko, Guy Paré","doi":"10.1017/rsm.2026.10091","DOIUrl":"10.1017/rsm.2026.10091","url":null,"abstract":"<p><p>Effective data management is essential for tasks involving decisions based on data, including knowledge synthesis and literature reviews. Despite this, how to carry out data management in literature reviews effectively remains unclear. With the increasing volume of research papers and the expansion of computational techniques for processing data (e.g., machine learning or large language models), it becomes imperative to consider data management as a crucial element for the advancement of literature review practices and tools. Presently, there are shortcomings related to (1) handling the growth of research to be synthesized, (2) addressing data quality issues when applying computational techniques or facilitating the verification of content produced by generative artificial intelligence, (3) enabling efficient reuse of datasets and innovative recombination of tools, and (4) facilitating transparent collaboration across heterogeneous review teams. To address these shortcomings, we develop the C5-DM Framework with conceptual principles to address data management challenges across five areas relevant to literature reviews: data conceptualization, collection, curation, control, and consumption. Methodological guidance for researchers with respect to these five areas is necessary to reduce errors, save time on repetitive tasks, and allow review teams to develop insightful syntheses.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"976-998"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147697157","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Personalized treatment hierarchies in Bayesian network meta-analysis.","authors":"Augustine Wigle, Erica E M Moodie","doi":"10.1017/rsm.2026.10089","DOIUrl":"10.1017/rsm.2026.10089","url":null,"abstract":"<p><p>Network meta-analysis (NMA) is an increasingly popular evidence synthesis tool that can provide a ranking of competing treatments, also known as a treatment hierarchy. Treatment-covariate interactions (TCIs) can be included in NMA models to allow relative treatment effects to vary with covariate values. We show that in an NMA model that includes TCIs, treatment hierarchies should be created with a particular covariate profile in mind. We outline the typical approach for creating a treatment hierarchy in standard Bayesian NMA and show how a treatment hierarchy for a particular covariate profile can be created from an NMA model that estimates TCIs. We demonstrate our methods using a real network of studies for the treatment of major depressive disorder.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"1018-1024"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147830765","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Literature-based meta-analysis of adverse events accounting for heterogeneous follow-up duration in oncology clinical trials.","authors":"Sumika Kawaguchi, Satoshi Hattori","doi":"10.1017/rsm.2026.10083","DOIUrl":"10.1017/rsm.2026.10083","url":null,"abstract":"<p><p>It is difficult to understand the safety profile of drugs based on a single clinical trial since clinical trials are often designed to prove efficacies, and sample size is not powered for safety assessment. Thus, meta-analysis would be a valuable tool to infer the safety profiles utilizing multiple studies. Individual clinical trials usually report the incidence proportions of adverse events (AEs) observed in the study. The follow-up duration may be study-specific, and furthermore different between the treatment groups within a single study. It often occurs in oncology clinical trials and if this is the case, it is hard to interpret the aggregated relative risk of AEs and compare the risk of AEs between the treatment groups with the standard meta-analysis techniques. The progression-free survival or the overall survival is often used as the primary endpoint in oncology clinical trials and the Kaplan-Meier estimates of the survival functions for the primary endpoint are often demonstrated graphically, which give us information of the follow-up duration of the AEs. We propose novel meta-analysis methods for AEs that address differences in follow-up durations by efficiently utilizing the Kaplan-Meier estimates of the primary endpoint. We adapt our approach using both simulated data and real data from a meta-analysis of bevacizumab. Simulation studies demonstrate that the proposed methods perform well when follow-up time differs between trials and groups.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"884-911"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147632065","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Angelika Eisele-Metzger, Judith-Lisa Lieberum, Markus Toews, Waldemar Siemens, Felix Heilmeyer, Christian Haverkamp, Daniel Boehringer, Joerg J Meerpohl
{"title":"Response to: Five methodological considerations for validating LLMs in risk of bias assessment.","authors":"Angelika Eisele-Metzger, Judith-Lisa Lieberum, Markus Toews, Waldemar Siemens, Felix Heilmeyer, Christian Haverkamp, Daniel Boehringer, Joerg J Meerpohl","doi":"10.1017/rsm.2026.10102","DOIUrl":"10.1017/rsm.2026.10102","url":null,"abstract":"","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"1057-1059"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148196680","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"A novel visualization approach for network meta-analysis: The plate plot and the nmaplateplot R package.","authors":"Yanan Ren, Zhenxun Wang, Lifeng Lin, Shanshan Zhao, Haitao Chu","doi":"10.1017/rsm.2026.10088","DOIUrl":"10.1017/rsm.2026.10088","url":null,"abstract":"<p><p>Network meta-analysis (NMA) provides a powerful framework for synthesizing evidence across multiple interventions, accommodating both direct and indirect comparisons. However, effectively visualizing the complex, multidimensional results, such as effect magnitudes, uncertainty, <i>p</i>-values, and treatment rankings, remains a significant challenge. Outputs such as relative treatment effects, uncertainty, statistical significance, and treatment rankings are often reported separately, making it difficult for researchers and stakeholders to synthesize findings efficiently. We introduce <i>plate plot</i>, an innovative approach for visualizing key outcomes from NMA in a single, compact format. It enables simultaneous display of point estimates, confidence or credible intervals, significance levels, and surface under the cumulative ranking curve values, thereby facilitating clearer interpretation and communication of NMA findings. Using an example dataset, we demonstrate how the <i>plate plot</i> displays multiple relevant metrics to compare the efficacy and acceptability of various antidepressant interventions in a single, intuitive plot. The <i>plate plot</i>, generated effortlessly via the open-source <i>nmaplateplot</i> R package, enables users to generate customizable, publication-ready graphics with minimal programming. This tool enhances the ability to holistically evaluate and interpret complex comparative effectiveness data, supporting better-informed decision-making in research and clinical practice.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"1045-1053"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13289567/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147626481","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Meta-analytic pooling of intraclass correlation coefficient estimates.","authors":"Bethany H Bhat, S Natasha Beretvas","doi":"10.1017/rsm.2026.10077","DOIUrl":"10.1017/rsm.2026.10077","url":null,"abstract":"<p><p>Intraclass correlation coefficient (ICC) estimates are necessary for several statistical techniques. Researchers need accurate ICC estimates when conducting prospective power analyses for clustered data scenarios. In addition, meta-analysts require reasonable ICC values when adjusting effect size estimates to account for clustered primary study data or to correct for psychometric artifacts when using the ICC as a reliability measure. The validity of these analyses hinges on the accuracy of the ICC estimate. Beyond these secondary analyses, ICC estimates have been used as the focal outcome of meta-analysis itself to obtain a pooled measure of agreement, reliability, or the influence of a cluster's effect. This study evaluates how well meta-analytically pooled ICC estimates recover the population ICC parameter value when using different ICC variance formulas as the inverse variance weights used in the pooling. We found that the variance formula that uses a normalizing transformation performs best across most conditions.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"850-883"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147571485","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Till J Adam, Salma A S Abosabie, Max Dittmer, Elise Wolf, Sara A Abosabie, Clara Behnke, Felix Baier, Annabelle Weickmann, Ludwig Köser, Christoph U Correll, Niklas Rutsch
{"title":"Prompt engineering of large language models for paper screening in medical meta-analyses and systematic reviews: A prospective comparative study.","authors":"Till J Adam, Salma A S Abosabie, Max Dittmer, Elise Wolf, Sara A Abosabie, Clara Behnke, Felix Baier, Annabelle Weickmann, Ludwig Köser, Christoph U Correll, Niklas Rutsch","doi":"10.1017/rsm.2026.10093","DOIUrl":"10.1017/rsm.2026.10093","url":null,"abstract":"<p><p>Interest in large language models (LLMs) as a tool for meta-analyses and systematic reviews (MA/SRs) is growing. We prospectively developed 515 unique prompts by predefined screening-related categories and tested with open-access LLMs (Llama, Mistral) against four gold-standard MA/SRs from different medical fields published after the LLMs' training cut-offs, using a Python-based pipeline. Heterogeneity between prompts was quantified, and hypothetical workload/cost reduction with top-performing prompts calculated. Across 12,360 pipeline runs, LLMs versus MA/SRs reached average recall/sensitivity = 83.6 ± 17.0%, precision = 18.5 ± 15.6%, specificity = 36.6 ± 23.7% F1-score = 27.6 ± 17.2%, and accuracy = 61.1 ± 11.0%. F1-scores were significantly higher when prompts focused on methods (0.78 ± 0.40%), explicitly mentioned MA/SR screening (0.81 ± 0.37%), included the comparison MA/SR's title (5.64 ± 0.37%) or selection criteria (8.05 ± 0.68%), and with more LLM parameters (70b = 4.48 ± 0.31%, 123b = 7.77 ± 0.31%), but lower when screening abstracts instead of titles (-3.67 ± 0.28%). In LLM-base preselection, top-performing F1-score prompts (recall/sensitivity = 72.2%, specificity = 66.1%, precision = 28.6%) would reduce screening demands by 34.5%-37.5%, saving 8.4-8.8 weeks of work and 17,592-18,552. Recall/sensitivity increased with less MA/SR information contrasting F1-score results, which highlights a recall/sensitivity-precision/specificity trade-off. F1-score increased with detailed MA/SR information, while recall/sensitivity increased with shorter, zeroshot prompts. We provide the first prospectively assessed prompt engineering framework for early-stage LLM-based paper screening across medical fields. The publicly available Python pipeline and full prompt list used here support further development of LLM-based evidence synthesis.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"939-956"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147669309","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Trevor Riley, Sarah Young, Avery Paxton, Lukas Wallrich, Kaitlyn Hair, Matthew Grainger
{"title":"CiteSource: An R package for data-driven search strategy development and enhanced evidence synthesis reporting.","authors":"Trevor Riley, Sarah Young, Avery Paxton, Lukas Wallrich, Kaitlyn Hair, Matthew Grainger","doi":"10.1017/rsm.2026.10084","DOIUrl":"10.1017/rsm.2026.10084","url":null,"abstract":"<p><p>Evidence synthesis findings hinge upon well-designed, effective search strategies. When developing these strategies, evidence synthesis teams make multiple decisions (e.g., selecting information sources, developing search string architecture, and picking supplementary search methods) that directly affect the breadth of discovered evidence and thus evidence synthesis outcomes. Despite the number of decisions required when developing search strategies, limited guidance exists to inform these decisions using a data-driven approach. To help address this gap, we developed CiteSource, an R package and accompanying Shiny application, that supports data-driven search strategy development and reporting. CiteSource allows users to assign and retain metadata across three custom fields: <i>source, label</i>, and <i>string</i> to indicate where the records were found, what method or string was used to find them, and whether they were included after screening. CiteSource allows users to visually map the overlap between sets of records, create data summaries of citation records, and export citation records with the newly assigned metadata. CiteSource's analysis and visualization outputs can be harnessed for a variety of use cases, such as optimizing literature source selection, honing and understanding the effectiveness of search strings, and evaluating the impacts of literature sources and supplementary search methods. Overall, CiteSource provides a tool for evidence synthesizers to make informed data-driven decisions that boost the efficiency, rigor, and transparency of search strategies and associated reporting.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"1026-1044"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147621248","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Zihan Song, Shan Huang, Ngeemasara Thapa, Xin Zhang, Byung-Kwon Park, Jie Lu, Wenyang Li, Wenbin Liu, Bei Zhan, Jianfei Li
{"title":"Large language model-based paper classification framework with key-insight extraction and confidence-weighted voting.","authors":"Zihan Song, Shan Huang, Ngeemasara Thapa, Xin Zhang, Byung-Kwon Park, Jie Lu, Wenyang Li, Wenbin Liu, Bei Zhan, Jianfei Li","doi":"10.1017/rsm.2026.10094","DOIUrl":"10.1017/rsm.2026.10094","url":null,"abstract":"<p><p>Systematic reviews (SRs) are critical for evidence-based research but are time-consuming and labor-intensive. The rapid expansion of academic publications further challenges the performance and applicability of existing screening and classification methods. While large language models (LLMs) present new opportunities for automation, limited research has examined whether they can achieve classification performance comparable to human reviewers in large-scale, multi-class settings. With the goal of improving classification performance, we proposed an LLM-based framework that leverages full-text key-insight extraction to enhance literature classification. We constructed a manually curated dataset of 900 articles from 17 published SRs to quantitatively evaluate the classification capabilities of LLMs. The results provided empirical evidence of LLMs' potential in supporting large-scale SRs and introduced a practical pathway for improving efficiency and reliability in evidence synthesis. Empirical results showed that key-insight-based classification (KBC) significantly outperforms abstract-based classification (ABC). We implemented a confidence-weighted voting (CWV) mechanism using multiple LLMs to improve robustness. The CWV method achieved the highest macro <i>F</i>1-score of 0.796, substantially exceeding KBC (0.732), ABC (0.676), and unsupervised K-means clustering (0.446). By employing zero-shot LLMs, our approach demonstrated the potential for enhanced adaptability across diverse domains and classification tasks without requiring fine-tuning, demonstrating that a carefully designed pipeline can enable LLMs to achieve classification performance comparable to human reviewers.</p>","PeriodicalId":226,"journal":{"name":"Research Synthesis Methods","volume":" ","pages":"999-1017"},"PeriodicalIF":8.0,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147758474","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}