{"title":"\"Reports in Medical Illustration (REMIL) in Musculoskeletal Radiology: An Evaluation of Evolving AI Models\".","authors":"V Umamaheswara Reddy, Nsl Susmitha, Rajesh Botchu","doi":"10.1016/j.acra.2026.08.082","DOIUrl":null,"url":null,"abstract":"<p><strong>Objective: </strong>Radiology reports remain predominantly text-based, requiring clinicians and patients to mentally reconstruct imaging findings. Reports in Medical Illustration (REMIL) represent an emerging approach in which artificial intelligence (AI) generates simplified visual summaries directly from report text. This study aimed to evaluate the feasibility, anatomical accuracy, and clinical utility of AI-generated REMIL in musculoskeletal (MSK) radiology.</p><p><strong>Methods: </strong>Twenty-five MSK imaging cases were selected. Identical report text and standardized prompts were provided to three premium multimodal AI systems-ChatGPT (GPT-4 with DALL-E 3), Perplexity AI, and Google Gemini 3.0 Pro. Each model generated a representative illustration based solely on the report description. Two fellowship-trained musculoskeletal radiologists independently assessed each illustration for anatomical accuracy and clinical usefulness. Errors were categorized as minor or major, and image-generation time was recorded.</p><p><strong>Results: </strong>Google Gemini 3.0 Pro demonstrated the most consistent performance, producing anatomically accurate illustrations in approximately 40-42% of cases and clinically useful images in 60-65% of cases, whereas ChatGPT and Perplexity AI frequently generated visually plausible images with substantial anatomical inaccuracies. Major errors were observed across all models, particularly in complex cases involving multiple anatomical structures or imaging planes. Simpler cases with a single dominant abnormality were illustrated more accurately by all models.</p><p><strong>Conclusion: </strong>AI-generated REMIL holds promise as an adjunctive tool for enhancing communication in MSK radiology by providing rapid visual summaries of imaging findings. However, current AI models exhibit inconsistent anatomical accuracy and are not yet reliable for unsupervised clinical use. REMIL should therefore be implemented only with radiologist validation.</p>","PeriodicalId":50928,"journal":{"name":"Academic Radiology","volume":" ","pages":""},"PeriodicalIF":4.7000,"publicationDate":"2026-09-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Academic Radiology","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1016/j.acra.2026.08.082","RegionNum":2,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"RADIOLOGY, NUCLEAR MEDICINE & MEDICAL IMAGING","Score":null,"Total":0}
引用次数: 0
Abstract
Objective: Radiology reports remain predominantly text-based, requiring clinicians and patients to mentally reconstruct imaging findings. Reports in Medical Illustration (REMIL) represent an emerging approach in which artificial intelligence (AI) generates simplified visual summaries directly from report text. This study aimed to evaluate the feasibility, anatomical accuracy, and clinical utility of AI-generated REMIL in musculoskeletal (MSK) radiology.
Methods: Twenty-five MSK imaging cases were selected. Identical report text and standardized prompts were provided to three premium multimodal AI systems-ChatGPT (GPT-4 with DALL-E 3), Perplexity AI, and Google Gemini 3.0 Pro. Each model generated a representative illustration based solely on the report description. Two fellowship-trained musculoskeletal radiologists independently assessed each illustration for anatomical accuracy and clinical usefulness. Errors were categorized as minor or major, and image-generation time was recorded.
Results: Google Gemini 3.0 Pro demonstrated the most consistent performance, producing anatomically accurate illustrations in approximately 40-42% of cases and clinically useful images in 60-65% of cases, whereas ChatGPT and Perplexity AI frequently generated visually plausible images with substantial anatomical inaccuracies. Major errors were observed across all models, particularly in complex cases involving multiple anatomical structures or imaging planes. Simpler cases with a single dominant abnormality were illustrated more accurately by all models.
Conclusion: AI-generated REMIL holds promise as an adjunctive tool for enhancing communication in MSK radiology by providing rapid visual summaries of imaging findings. However, current AI models exhibit inconsistent anatomical accuracy and are not yet reliable for unsupervised clinical use. REMIL should therefore be implemented only with radiologist validation.
期刊介绍:
Academic Radiology publishes original reports of clinical and laboratory investigations in diagnostic imaging, the diagnostic use of radioactive isotopes, computed tomography, positron emission tomography, magnetic resonance imaging, ultrasound, digital subtraction angiography, image-guided interventions and related techniques. It also includes brief technical reports describing original observations, techniques, and instrumental developments; state-of-the-art reports on clinical issues, new technology and other topics of current medical importance; meta-analyses; scientific studies and opinions on radiologic education; and letters to the Editor.