LLMs in Automated Essay Evaluation: A Case Study

Proceedings of the AAAI Symposium Series Pub Date : 2024-05-20 DOI:10.1609/aaaiss.v3i1.31193

Milan Kostic, Hans Friedrich Witschel, Knut Hinkelmann, Maja Spahic-Bogdanovic

引用次数: 0

Abstract

This study delves into the application of large language models (LLMs), such as ChatGPT-4, for the automated evaluation of student essays, with a focus on a case study conducted at the Swiss Institute of Business Administration. It explores the effectiveness of LLMs in assessing German-language student transfer assignments, and contrasts their performance with traditional evaluations by human lecturers. The primary findings highlight the challenges faced by LLMs in terms of accurately grading complex texts according to predefined categories and providing detailed feedback. This research illuminates the gap between the capabilities of LLMs and the nuanced requirements of student essay evaluation. The conclusion emphasizes the necessity for ongoing research and development in the area of LLM technology to improve the accuracy, reliability, and consistency of automated essay assessments in educational contexts.

查看原文本刊更多论文

自动论文评估法学硕士：案例研究

本研究深入探讨了大型语言模型（LLM）（如 ChatGPT-4）在学生论文自动评估中的应用，重点是在瑞士工商管理学院进行的一项案例研究。研究探讨了 LLM 在评估德语学生转学作业方面的有效性，并将其表现与传统的人工讲师评估进行了对比。主要研究结果凸显了法学硕士在根据预定义类别对复杂文本进行准确分级和提供详细反馈方面所面临的挑战。这项研究揭示了法律硕士的能力与学生论文评价的细微要求之间的差距。结论强调，有必要在法律硕士技术领域不断进行研究和开发，以提高教育背景下自动论文评估的准确性、可靠性和一致性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the AAAI Symposium Series

自引率

0.00%

发文量