GeneGenie: enhancing biomedical question-answering with agentic graphs.

IF 7.3 2区 生物学 Q1 BIOCHEMICAL RESEARCH METHODS
Mahmoud Gamal Abdelsalam, Abdulaziz H El-Safty, Amine Zaidi, Tanvir Alam
{"title":"GeneGenie: enhancing biomedical question-answering with agentic graphs.","authors":"Mahmoud Gamal Abdelsalam, Abdulaziz H El-Safty, Amine Zaidi, Tanvir Alam","doi":"10.1093/bib/bbag430","DOIUrl":null,"url":null,"abstract":"<p><p>Large language models (LLMs) have revolutionized biomedical research, yet they remain prone to hallucinations and struggle with the precise, multi-hop reasoning required for biomedical analysis. To bridge this gap between generative capability of AI model and factual rigor, this article introduces GeneGenie, a model-agnostic, multi-agent framework built upon a directed acyclic graph architecture. Unlike static prompting strategies, GeneGenie implements a deterministic five-node pipeline that orchestrates query planning, intelligent retrieval-augmented generation across curated databases (GenCC, HGNC, and UniProt), and the dynamic execution of bioinformatics tools, including NCBI E-Utilities and local BLAST+. We evaluated the system using the updated 16-module GeneTuring benchmark, comprising 1600 question-answer pairs. The experimental design compared six state-of-the-art models-including GPT-4o, Claude Sonnet 4.5, and Gemini 2.5 Pro-operating in a standalone \"Direct Mode\" versus the agentic \"Graph Mode.\" The results demonstrate that the graph-based architecture consistently outperforms single-model baselines across all metrics. Notably, among the six selected LLM models we explored, Gemini 2.5 Pro achieved the highest performance, correctly answering 1158 questions (72.375% accuracy), compared with the best baseline score of only 15.8%. Furthermore, our evaluation utilized an \"LLM-as-Judge\" semantic assessment, revealing that the agentic approach significantly enhances not only lexical accuracy but also the completeness and factual grounding of responses. While limitations remain in named entity recognition for protein-coding genes, GeneGenie establishes a robust, reproducible paradigm for future biomedical AI systems, proving that tool-augmented orchestration is superior to reliance on raw model scale alone.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3000,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13440129/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Briefings in bioinformatics","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1093/bib/bbag430","RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"BIOCHEMICAL RESEARCH METHODS","Score":null,"Total":0}
引用次数: 0

Abstract

Large language models (LLMs) have revolutionized biomedical research, yet they remain prone to hallucinations and struggle with the precise, multi-hop reasoning required for biomedical analysis. To bridge this gap between generative capability of AI model and factual rigor, this article introduces GeneGenie, a model-agnostic, multi-agent framework built upon a directed acyclic graph architecture. Unlike static prompting strategies, GeneGenie implements a deterministic five-node pipeline that orchestrates query planning, intelligent retrieval-augmented generation across curated databases (GenCC, HGNC, and UniProt), and the dynamic execution of bioinformatics tools, including NCBI E-Utilities and local BLAST+. We evaluated the system using the updated 16-module GeneTuring benchmark, comprising 1600 question-answer pairs. The experimental design compared six state-of-the-art models-including GPT-4o, Claude Sonnet 4.5, and Gemini 2.5 Pro-operating in a standalone "Direct Mode" versus the agentic "Graph Mode." The results demonstrate that the graph-based architecture consistently outperforms single-model baselines across all metrics. Notably, among the six selected LLM models we explored, Gemini 2.5 Pro achieved the highest performance, correctly answering 1158 questions (72.375% accuracy), compared with the best baseline score of only 15.8%. Furthermore, our evaluation utilized an "LLM-as-Judge" semantic assessment, revealing that the agentic approach significantly enhances not only lexical accuracy but also the completeness and factual grounding of responses. While limitations remain in named entity recognition for protein-coding genes, GeneGenie establishes a robust, reproducible paradigm for future biomedical AI systems, proving that tool-augmented orchestration is superior to reliance on raw model scale alone.

GeneGenie:通过代理图形增强生物医学问答。
大型语言模型(llm)已经彻底改变了生物医学研究,但它们仍然容易产生幻觉,并且难以实现生物医学分析所需的精确、多跳推理。为了弥合人工智能模型的生成能力与事实严谨性之间的差距,本文介绍了GeneGenie,这是一个基于有向无循环图架构的模型不可知的多智能体框架。与静态提示策略不同,GeneGenie实现了一个确定性的五节点管道,该管道协调查询计划,跨策划数据库(GenCC, HGNC和UniProt)的智能检索增强生成,以及生物信息学工具的动态执行,包括NCBI E-Utilities和本地BLAST+。我们使用更新的16模块GeneTuring基准来评估系统,包括1600个问答对。实验设计比较了六种最先进的模型——包括gpt - 40、Claude Sonnet 4.5和Gemini 2.5 pro——在独立的“直接模式”和代理的“图形模式”下运行。结果表明,基于图的体系结构在所有度量中始终优于单一模型基线。值得注意的是,在我们探索的六个选择的LLM模型中,Gemini 2.5 Pro取得了最高的性能,正确回答了1158个问题(准确率为72.375%),而最佳基线得分仅为15.8%。此外,我们的评估采用了“法学硕士作为法官”的语义评估,表明代理方法不仅显著提高了词汇准确性,而且显著提高了回答的完整性和事实基础。虽然在蛋白质编码基因的命名实体识别方面仍然存在局限性,但GeneGenie为未来的生物医学人工智能系统建立了一个强大的、可重复的范例,证明了工具增强编排优于仅依赖原始模型规模。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Briefings in bioinformatics
Briefings in bioinformatics 生物-生化研究方法
CiteScore
13.20
自引率
13.70%
发文量
549
审稿时长
6 months
期刊介绍: Briefings in Bioinformatics is an international journal serving as a platform for researchers and educators in the life sciences. It also appeals to mathematicians, statisticians, and computer scientists applying their expertise to biological challenges. The journal focuses on reviews tailored for users of databases and analytical tools in contemporary genetics, molecular and systems biology. It stands out by offering practical assistance and guidance to non-specialists in computerized methodologies. Covering a wide range from introductory concepts to specific protocols and analyses, the papers address bacterial, plant, fungal, animal, and human data. The journal's detailed subject areas include genetic studies of phenotypes and genotypes, mapping, DNA sequencing, expression profiling, gene expression studies, microarrays, alignment methods, protein profiles and HMMs, lipids, metabolic and signaling pathways, structure determination and function prediction, phylogenetic studies, and education and training.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书