Evaluating Adherence to Canadian Radiology Guidelines for Incidental Hepatobiliary Findings Using RAG-Enabled LLMs.

IF 2.9 3区医学 Q2 RADIOLOGY, NUCLEAR MEDICINE & MEDICAL IMAGING

Canadian Association of Radiologists Journal-Journal De L Association Canadienne Des Radiologistes Pub Date : 2025-02-27 DOI:10.1177/08465371251323124

Nicholas Dietrich, Brett Stubbert

{"title":"Evaluating Adherence to Canadian Radiology Guidelines for Incidental Hepatobiliary Findings Using RAG-Enabled LLMs.","authors":"Nicholas Dietrich, Brett Stubbert","doi":"10.1177/08465371251323124","DOIUrl":null,"url":null,"abstract":"Purpose: Large language models (LLMs) have the potential to support clinical decision-making but often lack training on the latest clinical guidelines. Retrieval-augmented generation (RAG) may enhance guideline adherence by dynamically integrating external information. This study evaluates the performance of two LLMs, GPT-4o and o1-mini, with and without RAG, in adhering to Canadian radiology guidelines for incidental hepatobiliary findings. Methods: A customized RAG architecture was developed to integrate guideline-based recommendations into LLM prompts. Clinical cases were curated and used to prompt models with and without RAG. Primary analyses assessed the rate of guideline adherence with comparisons made between LLMs with and without RAG. Secondary analyses evaluated reading ease, grade level, and response times for generated outputs. Results: A total of 319 clinical cases were evaluated. Adherence rates were 81.7% for GPT-4o without RAG, 97.2% for GPT-4o with RAG, 79.3% for o1-mini without RAG, and 95.1% for o1-mini with RAG. Model performance differed significantly across groups, with RAG-enabled configurations outperforming their non-RAG counterparts (P < .05). RAG-enabled models demonstrated improved reading ease and lower grade level scores; however, all model outputs remained at advanced comprehension levels. Response times for RAG-enabled models increased slightly due to additional retrieval processing but remained clinically acceptable. Conclusions: RAG-enabled LLMs significantly improved adherence to Canadian radiology guidelines for incidental hepatobiliary findings without compromising readability or response times. This approach holds promise for advancing evidence-based care and warrants further validation across broader clinical settings.","PeriodicalId":55290,"journal":{"name":"Canadian Association of Radiologists Journal-Journal De L Association Canadienne Des Radiologistes","volume":" ","pages":"8465371251323124"},"PeriodicalIF":2.9000,"publicationDate":"2025-02-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Canadian Association of Radiologists Journal-Journal De L Association Canadienne Des Radiologistes","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1177/08465371251323124","RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"RADIOLOGY, NUCLEAR MEDICINE & MEDICAL IMAGING","Score":null,"Total":0}

引用次数: 0

Abstract

Purpose: Large language models (LLMs) have the potential to support clinical decision-making but often lack training on the latest clinical guidelines. Retrieval-augmented generation (RAG) may enhance guideline adherence by dynamically integrating external information. This study evaluates the performance of two LLMs, GPT-4o and o1-mini, with and without RAG, in adhering to Canadian radiology guidelines for incidental hepatobiliary findings. Methods: A customized RAG architecture was developed to integrate guideline-based recommendations into LLM prompts. Clinical cases were curated and used to prompt models with and without RAG. Primary analyses assessed the rate of guideline adherence with comparisons made between LLMs with and without RAG. Secondary analyses evaluated reading ease, grade level, and response times for generated outputs. Results: A total of 319 clinical cases were evaluated. Adherence rates were 81.7% for GPT-4o without RAG, 97.2% for GPT-4o with RAG, 79.3% for o1-mini without RAG, and 95.1% for o1-mini with RAG. Model performance differed significantly across groups, with RAG-enabled configurations outperforming their non-RAG counterparts (P < .05). RAG-enabled models demonstrated improved reading ease and lower grade level scores; however, all model outputs remained at advanced comprehension levels. Response times for RAG-enabled models increased slightly due to additional retrieval processing but remained clinically acceptable. Conclusions: RAG-enabled LLMs significantly improved adherence to Canadian radiology guidelines for incidental hepatobiliary findings without compromising readability or response times. This approach holds promise for advancing evidence-based care and warrants further validation across broader clinical settings.

查看原文本刊更多论文

使用 RAG Enabled LLMs 评估加拿大放射学指南对偶然肝胆发现的遵循情况。

目的：大型语言模型（llm）具有支持临床决策的潜力，但往往缺乏最新临床指南的培训。检索增强生成（RAG）可以通过动态整合外部信息来增强指南的依从性。本研究评估了两种LLMs， gpt - 40和01 -mini，有无RAG，在遵守加拿大放射学指南中偶发肝胆发现的表现。方法：开发了一个定制的RAG架构，将基于指南的建议集成到LLM提示中。整理临床病例，用于有无RAG提示模型。初步分析通过比较有RAG和没有RAG的llm来评估指南依从率。二次分析评估了阅读难度、年级水平和生成输出的响应时间。结果：共评估319例临床病例。无RAG的gpt - 40依从率为81.7%，有RAG的gpt - 40依从率为97.2%，无RAG的o1-mini依从率为79.3%，有RAG的o1-mini依从率为95.1%。模型性能在各组之间差异显著，启用rag的配置优于未启用rag的配置（P < 0.05）。启用rag的模型显示了阅读易用性的提高和较低的年级水平分数；然而，所有模型输出仍然处于高级理解水平。由于额外的检索处理，启用rag的模型的响应时间略有增加，但在临床上仍然可以接受。结论：RAG-enabled LLMs显著提高了加拿大放射学指南对偶发肝胆发现的依从性，而不影响可读性或反应时间。这种方法有望推进循证护理，并在更广泛的临床环境中得到进一步验证。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Canadian Association of Radiologists Journal-Journal De L Association Canadienne Des Radiologistes 医学-核医学

CiteScore

6.20

自引率

12.90%

发文量

审稿时长

6-12 weeks

期刊介绍： The Canadian Association of Radiologists Journal is a peer-reviewed, Medline-indexed publication that presents a broad scientific review of radiology in Canada. The Journal covers such topics as abdominal imaging, cardiovascular radiology, computed tomography, continuing professional development, education and training, gastrointestinal radiology, health policy and practice, magnetic resonance imaging, musculoskeletal radiology, neuroradiology, nuclear medicine, pediatric radiology, radiology history, radiology practice guidelines and advisories, thoracic and cardiac imaging, trauma and emergency room imaging, ultrasonography, and vascular and interventional radiology. Article types considered for publication include original research articles, critically appraised topics, review articles, guest editorials, pictorial essays, technical notes, and letter to the Editor.