Education Research: Quality of Narrative Feedback Generated by a Large Language Model Compared With Expert Faculty for Case-Based Learning in Neurology Education.

Neurology. Education Pub Date : 2026-05-27 eCollection Date: 2026-06-01 DOI:10.1212/NE9.0000000000200320
Hannah Fruitman, Sasha Severin, Atikul Miah, Christina Gao, Haelynn Gim, Carolyn Qian, Sang-O Park, Kelly Hou, Edward L Kong, Benjamin Cook, Jasmin Le, Brandon Stretton, John Maddison, Liam G McCoy, Luke Collins, Andrew Vanlint, Rudy Goh, Matthew Arnold, Aye Thant, Rani Priyanka Vasireddy, Doris Kung, Ashley M Paul, Haatem Reda, Tamara B Kaplan, Adam Karp, Galina Gheihman
{"title":"Education Research: Quality of Narrative Feedback Generated by a Large Language Model Compared With Expert Faculty for Case-Based Learning in Neurology Education.","authors":"Hannah Fruitman, Sasha Severin, Atikul Miah, Christina Gao, Haelynn Gim, Carolyn Qian, Sang-O Park, Kelly Hou, Edward L Kong, Benjamin Cook, Jasmin Le, Brandon Stretton, John Maddison, Liam G McCoy, Luke Collins, Andrew Vanlint, Rudy Goh, Matthew Arnold, Aye Thant, Rani Priyanka Vasireddy, Doris Kung, Ashley M Paul, Haatem Reda, Tamara B Kaplan, Adam Karp, Galina Gheihman","doi":"10.1212/NE9.0000000000200320","DOIUrl":null,"url":null,"abstract":"<p><strong>Background and objectives: </strong>Neurology learners often receive limited feedback in clinical settings because of workflow constraints, variability in supervision, and competing clinical demands. Artificial intelligence, including large language models (LLMs) may help address these gaps and provide clinical learners with effective formative feedback by generating real-time, case-specific feedback during neurology case-based learning (CBL). The aim of this study was to examine how the quality of LLM-generated feedback compares with human expert-generated feedback in neurology CBL.</p><p><strong>Methods: </strong>In this exploratory quantitative study, student participants undertook LLM-enabled interactive cases on the TEACHABLE platform, which included history gathering, physical examination elements, and ordering diagnostic testing. Participants were clinical-level students recruited from 2 medical institutions. Case transcripts were recorded and analyzed for feedback generation, which was provided by an LLM and human experts in 2 components: history taking/physical examination elements (H&P) and assessment and plan (A&P). Feedback characteristics including sentence count, word count, and reference to case key learning points were summarized and compared. Feedback quality was scored by blinded experts using the QuAL and EFeCT instruments. Results were compared for the H&P and A&P components of the case interactions.</p><p><strong>Results: </strong>Four student participants completed 5 interactive cases each, generating 20 total transcripts for feedback. Word and sentence number were similar among LLM-generated and expert-generated feedback, except for a greater word length in expert-generated A&P feedback. Regarding H&P, the LLM commented on the key learning points in 20/20 (100%) of the cases as compared with 39/60 (65%) for the human experts. For A&P, the LLM feedback discussed key points in 20/20 (100%) cases as compared with 39/40 (97.5%) for the human experts. The LLM feedback had no medical inaccuracies. QuAL and EFeCT scores were significantly greater for the LLM as compared with human experts for the H&P component, but not significantly different for the A&P component.</p><p><strong>Discussion: </strong>LLMs provided with key learning points can generate timely, quality feedback on case-based interactions in a manner comparable with human experts. A hybrid framework combining LLM-generated feedback with faculty input may offer high-quality and equitably accessible formative feedback at scale. These pilot findings are limited by a small sample size and experimental setting.</p>","PeriodicalId":520085,"journal":{"name":"Neurology. Education","volume":"5 2","pages":"e200320"},"PeriodicalIF":0.0000,"publicationDate":"2026-05-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13220965/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Neurology. Education","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1212/NE9.0000000000200320","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/6/1 0:00:00","PubModel":"eCollection","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0

Abstract

Background and objectives: Neurology learners often receive limited feedback in clinical settings because of workflow constraints, variability in supervision, and competing clinical demands. Artificial intelligence, including large language models (LLMs) may help address these gaps and provide clinical learners with effective formative feedback by generating real-time, case-specific feedback during neurology case-based learning (CBL). The aim of this study was to examine how the quality of LLM-generated feedback compares with human expert-generated feedback in neurology CBL.

Methods: In this exploratory quantitative study, student participants undertook LLM-enabled interactive cases on the TEACHABLE platform, which included history gathering, physical examination elements, and ordering diagnostic testing. Participants were clinical-level students recruited from 2 medical institutions. Case transcripts were recorded and analyzed for feedback generation, which was provided by an LLM and human experts in 2 components: history taking/physical examination elements (H&P) and assessment and plan (A&P). Feedback characteristics including sentence count, word count, and reference to case key learning points were summarized and compared. Feedback quality was scored by blinded experts using the QuAL and EFeCT instruments. Results were compared for the H&P and A&P components of the case interactions.

Results: Four student participants completed 5 interactive cases each, generating 20 total transcripts for feedback. Word and sentence number were similar among LLM-generated and expert-generated feedback, except for a greater word length in expert-generated A&P feedback. Regarding H&P, the LLM commented on the key learning points in 20/20 (100%) of the cases as compared with 39/60 (65%) for the human experts. For A&P, the LLM feedback discussed key points in 20/20 (100%) cases as compared with 39/40 (97.5%) for the human experts. The LLM feedback had no medical inaccuracies. QuAL and EFeCT scores were significantly greater for the LLM as compared with human experts for the H&P component, but not significantly different for the A&P component.

Discussion: LLMs provided with key learning points can generate timely, quality feedback on case-based interactions in a manner comparable with human experts. A hybrid framework combining LLM-generated feedback with faculty input may offer high-quality and equitably accessible formative feedback at scale. These pilot findings are limited by a small sample size and experimental setting.

教育研究:神经学教育案例学习中大型语言模型与专家教师叙事反馈的质量比较。
背景和目的:由于工作流程的限制、监督的可变性和临床需求的竞争,神经学学习者在临床环境中经常得到有限的反馈。包括大型语言模型(llm)在内的人工智能可以帮助解决这些差距,并通过在神经病学基于案例的学习(CBL)过程中生成实时的、针对特定病例的反馈,为临床学习者提供有效的形成性反馈。本研究的目的是检验在神经学CBL中llm生成的反馈与人类专家生成的反馈的质量如何。方法:在这一探索性定量研究中,学生参与者在teeable平台上进行了法学硕士支持的互动案例,包括病史收集、体检要素和订购诊断测试。研究对象为来自2所医疗机构的临床级学生。记录并分析病例记录以生成反馈,反馈由法学硕士和人类专家提供,分为两部分:历史记录/体检要素(H&P)和评估与计划(A&P)。总结和比较反馈特征,包括句子数、单词数和参考案例重点学习点。反馈质量由盲法专家使用QuAL和effect工具评分。结果比较了病例相互作用的H&P和A&P组成部分。结果:4名学生参与者每人完成5个互动案例,共生成20份成绩单用于反馈。除了专家生成的A&P反馈的单词长度更大外,llm生成的反馈和专家生成的反馈的单词和句子数量相似。在H&P方面,法学硕士在20/20(100%)的案例中评论了关键的学习点,而人类专家的这一比例为39/60(65%)。对于A&P, LLM的反馈在20/20(100%)的案例中讨论了关键点,而人类专家的反馈为39/40(97.5%)。法学硕士的反馈没有医学上的不准确。与人类专家相比,法学硕士在H&P部分的QuAL和effect得分显著更高,但在A&P部分没有显著差异。讨论:提供关键学习点的法学硕士可以以与人类专家相当的方式,对基于案例的交互产生及时、高质量的反馈。将llm生成的反馈与教师输入相结合的混合框架可以提供高质量且公平可访问的大规模形成性反馈。这些初步研究结果受到样本量小和实验环境的限制。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
自引率
0.00%
发文量
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书