Dylan Owens, Danh Q Nguyen, Michael Dohopolski, Justin F Rousseau, Eric D Peterson, Ann Marie Navar
{"title":"Accuracy of Large Language Models to Identify Stroke Subtypes Within Unstructured Electronic Health Record Data.","authors":"Dylan Owens, Danh Q Nguyen, Michael Dohopolski, Justin F Rousseau, Eric D Peterson, Ann Marie Navar","doi":"10.1161/STROKEAHA.125.051993","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>While <i>International Classification of Diseases, Tenth Revision</i> codes suffice for identifying stroke events in surveillance, accurately classifying stroke types and subtypes using electronic health records remains challenging due to limitations in structured data. This often necessitates manual review of clinical documentation. This study evaluated whether a large language model, Generative Pre-Trained Transformer 4 Omni (GPT-4o), can accurately identify stroke types and subtypes from unstructured clinical notes.</p><p><strong>Methods: </strong>We implemented a retrieval-augmented generation framework with GPT-4o to classify stroke types (ischemic versus hemorrhagic) and ischemic stroke subtypes using electronic health records data. The American Heart Association Get With The Guidelines-Stroke registry served as the gold standard. Model development used a 20% subset of Get With The Guidelines-Stroke-linked data from UT Southwestern Medical Center (UTSW), with the remaining 80% reserved for testing. External validation used data from the Parkland Health and Hospital System (PHHS). A total of 4123 stroke hospitalizations from January 2019 to August 2023 were included (UTSW: n=2047; PHHS: n=2076). Three prompting strategies-zero-shot chain-of-thought, expert-guided, and instruction-based-were evaluated. Predictions of GPT-4os were compared with classifications made by trained abstractors contributing to the Get With The Guidelines-Stroke registry.</p><p><strong>Results: </strong>In the external validation set, 79.6% of patients had ischemic stroke and 20.4% hemorrhagic. GPT-4o achieved 98% accuracy (95% CI, 0.97-0.99) in classifying stroke type, where accuracy reflects the overall proportion of correctly classified patients. Sensitivity was 0.98 (95% CI, 0.97-0.99), and specificity was 0.97 (95% CI, 0.96-0.98). For ischemic stroke subtypes, sensitivity ranged from 0.40 (95% CI, 0.31-0.49) for cryptogenic to 0.95 (95% CI, 0.93-0.97) for small-vessel occlusion. Specificity ranged from 0.94 (95% CI, 0.92-0.96) for large-artery atherosclerosis to 0.98 (95% CI, 0.97-0.99) for cardioembolism. Zero-shot chain-of-thought prompting-requiring minimal human input-performed comparably to more labor-intensive strategies. Consistency analysis revealed <i>></i>99% agreement across repeated queries.</p><p><strong>Conclusions: </strong>GPT-4o demonstrated strong accuracy in classifying stroke types but faced challenges with ischemic subtypes.</p>","PeriodicalId":21989,"journal":{"name":"Stroke","volume":" ","pages":"2966-2975"},"PeriodicalIF":8.9000,"publicationDate":"2025-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12313299/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Stroke","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1161/STROKEAHA.125.051993","RegionNum":1,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/7/25 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"CLINICAL NEUROLOGY","Score":null,"Total":0}
引用次数: 0
Abstract
Background: While International Classification of Diseases, Tenth Revision codes suffice for identifying stroke events in surveillance, accurately classifying stroke types and subtypes using electronic health records remains challenging due to limitations in structured data. This often necessitates manual review of clinical documentation. This study evaluated whether a large language model, Generative Pre-Trained Transformer 4 Omni (GPT-4o), can accurately identify stroke types and subtypes from unstructured clinical notes.
Methods: We implemented a retrieval-augmented generation framework with GPT-4o to classify stroke types (ischemic versus hemorrhagic) and ischemic stroke subtypes using electronic health records data. The American Heart Association Get With The Guidelines-Stroke registry served as the gold standard. Model development used a 20% subset of Get With The Guidelines-Stroke-linked data from UT Southwestern Medical Center (UTSW), with the remaining 80% reserved for testing. External validation used data from the Parkland Health and Hospital System (PHHS). A total of 4123 stroke hospitalizations from January 2019 to August 2023 were included (UTSW: n=2047; PHHS: n=2076). Three prompting strategies-zero-shot chain-of-thought, expert-guided, and instruction-based-were evaluated. Predictions of GPT-4os were compared with classifications made by trained abstractors contributing to the Get With The Guidelines-Stroke registry.
Results: In the external validation set, 79.6% of patients had ischemic stroke and 20.4% hemorrhagic. GPT-4o achieved 98% accuracy (95% CI, 0.97-0.99) in classifying stroke type, where accuracy reflects the overall proportion of correctly classified patients. Sensitivity was 0.98 (95% CI, 0.97-0.99), and specificity was 0.97 (95% CI, 0.96-0.98). For ischemic stroke subtypes, sensitivity ranged from 0.40 (95% CI, 0.31-0.49) for cryptogenic to 0.95 (95% CI, 0.93-0.97) for small-vessel occlusion. Specificity ranged from 0.94 (95% CI, 0.92-0.96) for large-artery atherosclerosis to 0.98 (95% CI, 0.97-0.99) for cardioembolism. Zero-shot chain-of-thought prompting-requiring minimal human input-performed comparably to more labor-intensive strategies. Consistency analysis revealed >99% agreement across repeated queries.
Conclusions: GPT-4o demonstrated strong accuracy in classifying stroke types but faced challenges with ischemic subtypes.
期刊介绍:
Stroke is a monthly publication that collates reports of clinical and basic investigation of any aspect of the cerebral circulation and its diseases. The publication covers a wide range of disciplines including anesthesiology, critical care medicine, epidemiology, internal medicine, neurology, neuro-ophthalmology, neuropathology, neuropsychology, neurosurgery, nuclear medicine, nursing, radiology, rehabilitation, speech pathology, vascular physiology, and vascular surgery.
The audience of Stroke includes neurologists, basic scientists, cardiologists, vascular surgeons, internists, interventionalists, neurosurgeons, nurses, and physiatrists.
Stroke is indexed in Biological Abstracts, BIOSIS, CAB Abstracts, Chemical Abstracts, CINAHL, Current Contents, Embase, MEDLINE, and Science Citation Index Expanded.