Evaluating Adherence to Canadian Radiology Guidelines for Incidental Hepatobiliary Findings Using RAG-Enabled LLMs.

IF 2.9 3区 医学 Q2 RADIOLOGY, NUCLEAR MEDICINE & MEDICAL IMAGING
Nicholas Dietrich, Brett Stubbert
{"title":"Evaluating Adherence to Canadian Radiology Guidelines for Incidental Hepatobiliary Findings Using RAG-Enabled LLMs.","authors":"Nicholas Dietrich, Brett Stubbert","doi":"10.1177/08465371251323124","DOIUrl":null,"url":null,"abstract":"<p><p><b>Purpose:</b> Large language models (LLMs) have the potential to support clinical decision-making but often lack training on the latest clinical guidelines. Retrieval-augmented generation (RAG) may enhance guideline adherence by dynamically integrating external information. This study evaluates the performance of two LLMs, GPT-4o and o1-mini, with and without RAG, in adhering to Canadian radiology guidelines for incidental hepatobiliary findings. <b>Methods:</b> A customized RAG architecture was developed to integrate guideline-based recommendations into LLM prompts. Clinical cases were curated and used to prompt models with and without RAG. Primary analyses assessed the rate of guideline adherence with comparisons made between LLMs with and without RAG. Secondary analyses evaluated reading ease, grade level, and response times for generated outputs. <b>Results:</b> A total of 319 clinical cases were evaluated. Adherence rates were 81.7% for GPT-4o without RAG, 97.2% for GPT-4o with RAG, 79.3% for o1-mini without RAG, and 95.1% for o1-mini with RAG. Model performance differed significantly across groups, with RAG-enabled configurations outperforming their non-RAG counterparts (<i>P</i> < .05). RAG-enabled models demonstrated improved reading ease and lower grade level scores; however, all model outputs remained at advanced comprehension levels. Response times for RAG-enabled models increased slightly due to additional retrieval processing but remained clinically acceptable. <b>Conclusions:</b> RAG-enabled LLMs significantly improved adherence to Canadian radiology guidelines for incidental hepatobiliary findings without compromising readability or response times. This approach holds promise for advancing evidence-based care and warrants further validation across broader clinical settings.</p>","PeriodicalId":55290,"journal":{"name":"Canadian Association of Radiologists Journal-Journal De L Association Canadienne Des Radiologistes","volume":" ","pages":"8465371251323124"},"PeriodicalIF":2.9000,"publicationDate":"2025-02-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Canadian Association of Radiologists Journal-Journal De L Association Canadienne Des Radiologistes","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1177/08465371251323124","RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"RADIOLOGY, NUCLEAR MEDICINE & MEDICAL IMAGING","Score":null,"Total":0}
引用次数: 0

Abstract

Purpose: Large language models (LLMs) have the potential to support clinical decision-making but often lack training on the latest clinical guidelines. Retrieval-augmented generation (RAG) may enhance guideline adherence by dynamically integrating external information. This study evaluates the performance of two LLMs, GPT-4o and o1-mini, with and without RAG, in adhering to Canadian radiology guidelines for incidental hepatobiliary findings. Methods: A customized RAG architecture was developed to integrate guideline-based recommendations into LLM prompts. Clinical cases were curated and used to prompt models with and without RAG. Primary analyses assessed the rate of guideline adherence with comparisons made between LLMs with and without RAG. Secondary analyses evaluated reading ease, grade level, and response times for generated outputs. Results: A total of 319 clinical cases were evaluated. Adherence rates were 81.7% for GPT-4o without RAG, 97.2% for GPT-4o with RAG, 79.3% for o1-mini without RAG, and 95.1% for o1-mini with RAG. Model performance differed significantly across groups, with RAG-enabled configurations outperforming their non-RAG counterparts (P < .05). RAG-enabled models demonstrated improved reading ease and lower grade level scores; however, all model outputs remained at advanced comprehension levels. Response times for RAG-enabled models increased slightly due to additional retrieval processing but remained clinically acceptable. Conclusions: RAG-enabled LLMs significantly improved adherence to Canadian radiology guidelines for incidental hepatobiliary findings without compromising readability or response times. This approach holds promise for advancing evidence-based care and warrants further validation across broader clinical settings.

使用 RAG Enabled LLMs 评估加拿大放射学指南对偶然肝胆发现的遵循情况。
目的:大型语言模型(llm)具有支持临床决策的潜力,但往往缺乏最新临床指南的培训。检索增强生成(RAG)可以通过动态整合外部信息来增强指南的依从性。本研究评估了两种LLMs, gpt - 40和01 -mini,有无RAG,在遵守加拿大放射学指南中偶发肝胆发现的表现。方法:开发了一个定制的RAG架构,将基于指南的建议集成到LLM提示中。整理临床病例,用于有无RAG提示模型。初步分析通过比较有RAG和没有RAG的llm来评估指南依从率。二次分析评估了阅读难度、年级水平和生成输出的响应时间。结果:共评估319例临床病例。无RAG的gpt - 40依从率为81.7%,有RAG的gpt - 40依从率为97.2%,无RAG的o1-mini依从率为79.3%,有RAG的o1-mini依从率为95.1%。模型性能在各组之间差异显著,启用rag的配置优于未启用rag的配置(P < 0.05)。启用rag的模型显示了阅读易用性的提高和较低的年级水平分数;然而,所有模型输出仍然处于高级理解水平。由于额外的检索处理,启用rag的模型的响应时间略有增加,但在临床上仍然可以接受。结论:RAG-enabled LLMs显著提高了加拿大放射学指南对偶发肝胆发现的依从性,而不影响可读性或反应时间。这种方法有望推进循证护理,并在更广泛的临床环境中得到进一步验证。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
CiteScore
6.20
自引率
12.90%
发文量
98
审稿时长
6-12 weeks
期刊介绍: The Canadian Association of Radiologists Journal is a peer-reviewed, Medline-indexed publication that presents a broad scientific review of radiology in Canada. The Journal covers such topics as abdominal imaging, cardiovascular radiology, computed tomography, continuing professional development, education and training, gastrointestinal radiology, health policy and practice, magnetic resonance imaging, musculoskeletal radiology, neuroradiology, nuclear medicine, pediatric radiology, radiology history, radiology practice guidelines and advisories, thoracic and cardiac imaging, trauma and emergency room imaging, ultrasonography, and vascular and interventional radiology. Article types considered for publication include original research articles, critically appraised topics, review articles, guest editorials, pictorial essays, technical notes, and letter to the Editor.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信