{"title":"Towards the implementation of automated scoring in international large-scale assessments: Scalability and quality control","authors":"Ji Yoon Jung, Lillian Tyack, Matthias von Davier","doi":"10.1016/j.caeai.2025.100375","DOIUrl":null,"url":null,"abstract":"<div><div>Even before the age of artificial intelligence, automated scoring received considerable attention in educational measurement. However, its application to constructed response (CR) items in international large-scale assessments (ILSAs) has remained a challenge, primarily due to the difficulty of handling multilingual responses spanning many languages. This study addresses this challenge by investigating two machine learning approaches — supervised and unsupervised learning — for scoring multilingual responses. We explored various scoring methods to assess three science CR items from TIMSS 2023 across all participating countries and 42 languages. The results showed that the supervised learning approach, particularly combining multiple machine translations with artificial neural networks (MMT_ANNs), showed comparable performance to human scoring. The MMT_ANN model demonstrated impressive accuracy, correctly classifying up to 94.88% of responses across all languages and countries. This remarkable performance can be attributed to MMT_ANNs providing more suitable translations at both individual response and language levels. Furthermore, MMT_ANNs consistently generated accurate scores for identical or borderline responses within and across countries. These findings indicate the potential of automated scoring as an accurate and cost-effective measure for quality control in ILSAs, reducing the need to hire additional human raters to ensure scoring reliability.</div></div>","PeriodicalId":34469,"journal":{"name":"Computers and Education Artificial Intelligence","volume":"8 ","pages":"Article 100375"},"PeriodicalIF":0.0000,"publicationDate":"2025-01-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Computers and Education Artificial Intelligence","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2666920X25000153","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"Social Sciences","Score":null,"Total":0}
引用次数: 0
Abstract
Even before the age of artificial intelligence, automated scoring received considerable attention in educational measurement. However, its application to constructed response (CR) items in international large-scale assessments (ILSAs) has remained a challenge, primarily due to the difficulty of handling multilingual responses spanning many languages. This study addresses this challenge by investigating two machine learning approaches — supervised and unsupervised learning — for scoring multilingual responses. We explored various scoring methods to assess three science CR items from TIMSS 2023 across all participating countries and 42 languages. The results showed that the supervised learning approach, particularly combining multiple machine translations with artificial neural networks (MMT_ANNs), showed comparable performance to human scoring. The MMT_ANN model demonstrated impressive accuracy, correctly classifying up to 94.88% of responses across all languages and countries. This remarkable performance can be attributed to MMT_ANNs providing more suitable translations at both individual response and language levels. Furthermore, MMT_ANNs consistently generated accurate scores for identical or borderline responses within and across countries. These findings indicate the potential of automated scoring as an accurate and cost-effective measure for quality control in ILSAs, reducing the need to hire additional human raters to ensure scoring reliability.