Performance of ChatGPT on Solving Orthopedic Board-Style Questions: A Comparative Analysis of ChatGPT 3.5 and ChatGPT 4.

IF 1.9 2区医学 Q2 ORTHOPEDICS

Clinics in Orthopedic Surgery Pub Date : 2024-08-01 Epub Date: 2024-03-07 DOI:10.4055/cios23179

Sung Eun Kim, Ji Han Lee, Byung Sun Choi, Hyuk-Soo Han, Myung Chul Lee, Du Hyun Ro

{"title":"Performance of ChatGPT on Solving Orthopedic Board-Style Questions: A Comparative Analysis of ChatGPT 3.5 and ChatGPT 4.","authors":"Sung Eun Kim, Ji Han Lee, Byung Sun Choi, Hyuk-Soo Han, Myung Chul Lee, Du Hyun Ro","doi":"10.4055/cios23179","DOIUrl":null,"url":null,"abstract":"Background: The application of artificial intelligence and large language models in the medical field requires an evaluation of their accuracy in providing medical information. This study aimed to assess the performance of Chat Generative Pre-trained Transformer (ChatGPT) models 3.5 and 4 in solving orthopedic board-style questions.Methods: A total of 160 text-only questions from the Orthopedic Surgery Department at Seoul National University Hospital, conforming to the format of the Korean Orthopedic Association board certification examinations, were input into the ChatGPT 3.5 and ChatGPT 4 programs. The questions were divided into 11 subcategories. The accuracy rates of the initial answers provided by Chat GPT 3.5 and ChatGPT 4 were analyzed. In addition, inconsistency rates of answers were evaluated by regenerating the responses.Results: ChatGPT 3.5 answered 37.5% of the questions correctly, while ChatGPT 4 showed an accuracy rate of 60.0% (p < 0.001). ChatGPT 4 demonstrated superior performance across most subcategories, except for the tumor-related questions. The rates of inconsistency in answers were 47.5% for ChatGPT 3.5 and 9.4% for ChatGPT 4.Conclusions: ChatGPT 4 showed the ability to pass orthopedic board-style examinations, outperforming ChatGPT 3.5 in accuracy rate. However, inconsistencies in response generation and instances of incorrect answers with misleading explanations require caution when applying ChatGPT in clinical settings or for educational purposes.","PeriodicalId":47648,"journal":{"name":"Clinics in Orthopedic Surgery","volume":"16 4","pages":"669-673"},"PeriodicalIF":1.9000,"publicationDate":"2024-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11262944/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Clinics in Orthopedic Surgery","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.4055/cios23179","RegionNum":2,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2024/3/7 0:00:00","PubModel":"Epub","JCR":"Q2","JCRName":"ORTHOPEDICS","Score":null,"Total":0}

引用次数: 0

Abstract

Background: The application of artificial intelligence and large language models in the medical field requires an evaluation of their accuracy in providing medical information. This study aimed to assess the performance of Chat Generative Pre-trained Transformer (ChatGPT) models 3.5 and 4 in solving orthopedic board-style questions.

Methods: A total of 160 text-only questions from the Orthopedic Surgery Department at Seoul National University Hospital, conforming to the format of the Korean Orthopedic Association board certification examinations, were input into the ChatGPT 3.5 and ChatGPT 4 programs. The questions were divided into 11 subcategories. The accuracy rates of the initial answers provided by Chat GPT 3.5 and ChatGPT 4 were analyzed. In addition, inconsistency rates of answers were evaluated by regenerating the responses.

Results: ChatGPT 3.5 answered 37.5% of the questions correctly, while ChatGPT 4 showed an accuracy rate of 60.0% (p < 0.001). ChatGPT 4 demonstrated superior performance across most subcategories, except for the tumor-related questions. The rates of inconsistency in answers were 47.5% for ChatGPT 3.5 and 9.4% for ChatGPT 4.

Conclusions: ChatGPT 4 showed the ability to pass orthopedic board-style examinations, outperforming ChatGPT 3.5 in accuracy rate. However, inconsistencies in response generation and instances of incorrect answers with misleading explanations require caution when applying ChatGPT in clinical settings or for educational purposes.

查看原文本刊更多论文

ChatGPT 在解决骨科 Board 类型问题上的性能：ChatGPT 3.5 和 ChatGPT 4 的比较分析。

背景：在医疗领域应用人工智能和大型语言模型需要对其提供医疗信息的准确性进行评估。本研究旨在评估 Chat Generative Pre-trained Transformer（ChatGPT）3.5 和 4 模型在解决骨科板式问题时的性能：在 ChatGPT 3.5 和 ChatGPT 4 程序中输入了来自首尔国立大学医院骨外科的 160 道纯文本试题，这些试题符合韩国骨科协会委员会认证考试的格式。这些问题被分为 11 个子类别。对 Chat GPT 3.5 和 ChatGPT 4 所提供的初始答案的准确率进行了分析。此外，还通过重新生成答案评估了答案的不一致率：结果：ChatGPT 3.5 回答正确率为 37.5%，而 ChatGPT 4 的正确率为 60.0%（p < 0.001）。除肿瘤相关问题外，ChatGPT 4 在大多数子类别中都表现优异。ChatGPT 3.5 的答案不一致率为 47.5%，ChatGPT 4 为 9.4%：结论：ChatGPT 4 显示了通过骨科委员会式考试的能力，在准确率方面优于 ChatGPT 3.5。然而，在将 ChatGPT 应用于临床环境或教育目的时，需要注意在生成答案时的不一致性以及错误答案和误导性解释的情况。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊