{"title":"Can ChatGPT be a knowledgeable and accurate ethical accountant? Assessing the overall usefulness of ChatGPT’s answers to ethical dilemmas","authors":"Fábio Albuquerque , Paula Gomes dos Santos","doi":"10.1016/j.jaccedu.2025.101000","DOIUrl":null,"url":null,"abstract":"<div><div>This paper aims to assess the overall usefulness of ChatGPT regarding ethical dilemmas in accounting by exploring the tool’s ability to produce accurate answers, as well as reliably and properly identify the relevant legal frameworks in its justification. This study employs a quasi-experimental method and an exploratory perspective to evaluate ChatGPT 4.0′s responses to questions from the Portuguese Order of Certified Accountants exams. Consistency and robustness tests were conducted by varying and repeating the prompts, and the outputs (ChatGPT responses) were assessed through a content analysis method by comparing them with official answer keys (accuracy) and their justification (legal frameworks). The findings indicate that, throughout the process, ChatGPT’s performance is influenced by prompt design and repetition, and it consistently lacks accuracy, namely when professional judgment is required. Besides, although it is generally capable of recognising the relevant legal frameworks, its justifications were usually vague. Moreover, whenever more precise answers were provided, evidence of hallucinations was commonly found. Therefore, while the responses are often persuasive and structured, their reliability is limited, raising concerns about the model’s epistemic soundness and reducing its overall usefulness, which highlights the need for users to exercise caution when relying on their outputs. To the best of the authors’ knowledge, this is the first research to analyse in depth the characteristics within ChatGPT responses, specifically regarding professional ethics in an accounting proficiency exam. This study offers empirical evidence on ChatGPT’s limitations and potential as a supplementary tool in accounting ethics. It contributes to ongoing debates about the role of generative AI in professional education and practice, providing insights for educators and practitioners regarding its responsible and cautious integration. While promising as a discussion aid, ChatGPT still cannot replace the critical thinking and contextual reasoning required for ethical decision-making.</div></div>","PeriodicalId":35578,"journal":{"name":"Journal of Accounting Education","volume":"73 ","pages":"Article 101000"},"PeriodicalIF":0.0000,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Accounting Education","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S074857512500051X","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/11/9 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"Social Sciences","Score":null,"Total":0}
引用次数: 0
Abstract
This paper aims to assess the overall usefulness of ChatGPT regarding ethical dilemmas in accounting by exploring the tool’s ability to produce accurate answers, as well as reliably and properly identify the relevant legal frameworks in its justification. This study employs a quasi-experimental method and an exploratory perspective to evaluate ChatGPT 4.0′s responses to questions from the Portuguese Order of Certified Accountants exams. Consistency and robustness tests were conducted by varying and repeating the prompts, and the outputs (ChatGPT responses) were assessed through a content analysis method by comparing them with official answer keys (accuracy) and their justification (legal frameworks). The findings indicate that, throughout the process, ChatGPT’s performance is influenced by prompt design and repetition, and it consistently lacks accuracy, namely when professional judgment is required. Besides, although it is generally capable of recognising the relevant legal frameworks, its justifications were usually vague. Moreover, whenever more precise answers were provided, evidence of hallucinations was commonly found. Therefore, while the responses are often persuasive and structured, their reliability is limited, raising concerns about the model’s epistemic soundness and reducing its overall usefulness, which highlights the need for users to exercise caution when relying on their outputs. To the best of the authors’ knowledge, this is the first research to analyse in depth the characteristics within ChatGPT responses, specifically regarding professional ethics in an accounting proficiency exam. This study offers empirical evidence on ChatGPT’s limitations and potential as a supplementary tool in accounting ethics. It contributes to ongoing debates about the role of generative AI in professional education and practice, providing insights for educators and practitioners regarding its responsible and cautious integration. While promising as a discussion aid, ChatGPT still cannot replace the critical thinking and contextual reasoning required for ethical decision-making.
期刊介绍:
The Journal of Accounting Education (JAEd) is a refereed journal dedicated to promoting and publishing research on accounting education issues and to improving the quality of accounting education worldwide. The Journal provides a vehicle for making results of empirical studies available to educators and for exchanging ideas, instructional resources, and best practices that help improve accounting education. The Journal includes four sections: a Main Articles Section, a Teaching and Educational Notes Section, an Educational Case Section, and a Best Practices Section. Manuscripts published in the Main Articles Section generally present results of empirical studies, although non-empirical papers (such as policy-related or essay papers) are sometimes published in this section. Papers published in the Teaching and Educational Notes Section include short empirical pieces (e.g., replications) as well as instructional resources that are not properly categorized as cases, which are published in a separate Case Section. Note: as part of the Teaching Note accompany educational cases, authors must include implementation guidance (based on actual case usage) and evidence regarding the efficacy of the case vis-a-vis a listing of educational objectives associated with the case. To meet the efficacy requirement, authors must include direct assessment (e.g grades by case requirement/objective or pre-post tests). Although interesting and encouraged, student perceptions (surveys) are considered indirect assessment and do not meet the efficacy requirement. The case must have been used more than once in a course to avoid potential anomalies and to vet the case before submission. Authors may be asked to collect additional data, depending on course size/circumstances.