Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-23DOI: 10.1016/j.asw.2026.101090
Qiao Gan, Benjamin Adams
{"title":"From keystrokes to scores: Toward a multidimensional predictive model of writing evaluation by humans and large language models across linguistic, cognitive, and social dimensions","authors":"Qiao Gan, Benjamin Adams","doi":"10.1016/j.asw.2026.101090","DOIUrl":"10.1016/j.asw.2026.101090","url":null,"abstract":"<div><div>Automated writing evaluation (AWE) has traditionally emphasized textual features such as vocabulary and syntax, while often overlooking writers’ social identities and cognitive behaviors – factors central to understanding writing as a multidimensional construct. With the increasing integration of large language models (LLMs) into AWE, questions remain about how their assessments align with human judgments and the sources of potential divergences. This study investigates how linguistic (e.g., lexical diversity), cognitive (e.g., pausing behavior), and social (e.g., gender) factors covary with essay scores assigned by human raters and LLMs. We analyzed 4245 argumentative essays paired with demographic metadata and keystroke-logging data, using correlation analyses, random forest models, and regression-based approaches to examine relationships among writer characteristics, writing-process features, textual features, and essay scores. Results showed moderate agreement between human and LLM scores, but the two scoring systems exhibited different patterns of association with linguistic, cognitive, and social variables. These findings suggest that human and LLM evaluations rely on partially different cues and demonstrate how socio-cognitive metadata can be used to examine the factors associated with writing assessment decisions. By moving beyond text-only comparisons, this approach provides a complementary lens for understanding why and how human and machine judgments converge or diverge.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101090"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553717","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-30DOI: 10.1016/j.asw.2026.101091
Wenlei Zhu, Yao Zheng
{"title":"Navigating the emotional journey of writing feedback: A scoping review of empirical studies (2005–2025)","authors":"Wenlei Zhu, Yao Zheng","doi":"10.1016/j.asw.2026.101091","DOIUrl":"10.1016/j.asw.2026.101091","url":null,"abstract":"<div><div>Feedback is a common pedagogical approach to enhance students’ writing performance. Research on feedback emotion has recently gained growing interest and sparked a series of empirical studies. To timely update research progress and guide future research in this promising area, this study provides a scoping review of 61 empirical investigations published in academic journals and ProQuest Dissertations & Theses from 2005 to 2025. Iterative content and thematic analyses of these publications were conducted to scrutinize the conceptualizations of emotions, research scopes, contexts, and methodological characteristics regarding approaches, instruments, reporting of methodological rigor, and study durations. The findings reveal that emotions in writing feedback are fundamentally context-sensitive and dynamic. Furthermore, the literature shows a concentrated interest in describing participants’ emotions within single-source, written feedback situations, while relying heavily on self-report instruments across qualitative, quantitative, and mixed-methods studies. This review provides several empirically grounded suggestions for future research based on the results.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101091"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553725","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-24DOI: 10.1016/j.asw.2026.101089
Li-Jen Wang, Ying-Tien Wu, Teng-Yao Cheng
{"title":"Measuring engagement in technology-supported collaborative argumentative writing: Development and validation of the T-CLEA scale","authors":"Li-Jen Wang, Ying-Tien Wu, Teng-Yao Cheng","doi":"10.1016/j.asw.2026.101089","DOIUrl":"10.1016/j.asw.2026.101089","url":null,"abstract":"<div><div>Current moves toward AI-supported and collaborative writing environments highlight the need for validated instruments that assess learners’ engagement in technology-mediated argumentative writing. This study developed and validated the Technology-supported Collaborative Learning Engagement Assessment (T-CLEA) scale, a bilingual English-Mandarin self-report instrument for university EFL courses involving digitally supported collaborative argumentative tasks on Sustainable Development Goal topics. T-CLEA comprises three task-specific dimensions: English Writing Challenges (EWC), Student Issues on Global Topics (SIG), and Perceptions of Collaborative Argumentation (PCA). Items were generated from relevant theories and empirical studies, refined through expert review, translated and back-translated, and pilot tested. Data from 529 undergraduates were split into two independent samples: 227 for exploratory factor analysis and 302 for confirmatory factor analysis. Results supported an 18-item, three-factor structure with satisfactory model fit (χ²/df = 2.78, CFI =.943, TLI =.935, RMSEA =.059, SRMR =.059). Internal consistency was good across dimensions (Cronbach’s α =.87–.89; CR =.87–.89), and convergent and discriminant validity evidence supported the construct structure. T-CLEA offers a context-sensitive diagnostic tool for examining engagement, writing quality, and process-based indicators in technology-supported EAP/EFL writing contexts.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101089"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553720","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Automated integrity or automated injustice? Turnitin’s AI detection in multilingual writing contexts","authors":"Kgabo Bridget Maphoto, Louie Giray, Kershnee Sevnarayan","doi":"10.1016/j.asw.2026.101086","DOIUrl":"10.1016/j.asw.2026.101086","url":null,"abstract":"<div><div>This article critically examines the advantages and limitations of Turnitin’s AI writing detector in second language (L2) writing assessment. Although Turnitin and similar detectors (e.g., GPTZero, Copyleaks, Writer, and the OpenAI Classifier) claim high accuracy in identifying AI-generated text, recent evidence reveals systemic biases against multilingual writers, whose syntactic regularity and formulaic phrasing are frequently misclassified as AI-generated. This article draws on recent empirical studies and argues that overreliance on such tools in some institutional contexts may contribute to epistemic injustice, false accusations, and the reinforcement of deficit ideologies in L2 contexts. This article draws on recent empirical studies and argues that overreliance on such tools in some institutional contexts may contribute to epistemic injustice, false accusations, and the reinforcement of deficit ideologies in L2 contexts. Rather than treating detection scores as verdicts of misconduct, we propose a dialogic, process-oriented framework that repositions Turnitin’s output as a formative cue, one that, when triangulated with drafts, reflective commentaries, and oral explanations, can encourage authorship transparency and pedagogical trust. The article thus advances a linguistic approach to L2 assessment that acknowledges AI’s presence while safeguarding student agency, integrity, and equity in writing evaluation.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101086"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553723","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-26DOI: 10.1016/j.asw.2026.101082
Scott Crossley, Langdon Holmes, Wesley Morris
{"title":"Assessing the reliability and validity of large language models in automatic essay scoring","authors":"Scott Crossley, Langdon Holmes, Wesley Morris","doi":"10.1016/j.asw.2026.101082","DOIUrl":"10.1016/j.asw.2026.101082","url":null,"abstract":"<div><div>With the advent of artificial intelligence, large language model (LLM) based Automated Essay Scoring (AES) systems have been developed that can consistently make human-like decisions that do not depend fully on surface level linguistic features. However, research into the use of LLM-based AES systems is limited and little is known about the reliability, agreement, or validity of the systems. The goal of this study was to provide evidence for the reliability, agreement, and validity of LLM-based AES systems in a standardized writing assessment used for secondary school students. Both representation and generative LLM-based AES systems were developed to score persuasive essays and assessed for reliability. Then the agreement of the developed AES systems with human raters was assessed through correlational analyses. We used extrinsic convergent validation approaches to examine if the human and LLM scores correlated with linguistic components. Results indicate strong reliability and agreement for the LLM scores. In terms of convergent validity, initial correlational analyses indicated that the representation LLM AES system showed differential correlations with the human scores in terms of a text length and type-token ratio component. This result contrasts with the correlational results from the generative LLM AES model, which indicated no differences in associations between the model and human scores with regards to the linguistic components.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101082"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553719","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-15DOI: 10.1016/j.asw.2026.101083
Mikyung Kim Wolf, Michael Suhan, Paul Deane, Lorraine Sova, Jean Alderman
{"title":"Young L2 students’ use of an AI-assisted writing assessment and feedback tool: An exploratory study in multiple settings","authors":"Mikyung Kim Wolf, Michael Suhan, Paul Deane, Lorraine Sova, Jean Alderman","doi":"10.1016/j.asw.2026.101083","DOIUrl":"10.1016/j.asw.2026.101083","url":null,"abstract":"<div><div>This study investigated how young L2 learners engaged with an AI-assisted writing assessment feedback tool that employed the GPT-4o model. In particular, we focused on students’ interactions with the AI chatbot to seek feedback and their subsequent use of the chatbot’s responses. Conducted as part of a prototype development and usability testing efforts, the study involved eight teachers and their EFL/ESL students (<em>N</em> = 206) from upper elementary and middle school grades in Hong Kong, South Korea, Turkiye, and the United States. Students completed three Opinion writing tasks using the tool. Students’ chat messages were coded, and textual analyses were conducted to examine students’ revisions. Results indicated considerable variation in students’ usage with the chatbot. During the outlining stage, students primarily sought support with content development and translation; during revision, they focused on correction and content-related feedback. Students’ incorporation of AI feedback was evidenced by increased text length and more token additions than deletions or replacements. Both teachers and students rated the tool’s translation (L1 support) and personalized, interactive feedback features as highly useful. Implications for L2 writing assessment and future research are discussed.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101083"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553724","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-25DOI: 10.1016/j.asw.2026.101088
Michele Flammia, Joe Oyler, Ariel Sykes, Noriko Takahashi, Astha Singh, Abraham Onuorah, Evgeny Chukharev, Alina Reznitskaya
{"title":"Argument chains as a tool to improve automated scoring of argumentative writing","authors":"Michele Flammia, Joe Oyler, Ariel Sykes, Noriko Takahashi, Astha Singh, Abraham Onuorah, Evgeny Chukharev, Alina Reznitskaya","doi":"10.1016/j.asw.2026.101088","DOIUrl":"10.1016/j.asw.2026.101088","url":null,"abstract":"<div><div>The study introduces argument chains, a theory-driven analytic tool we developed to improve the automated assessment of argumentative writing. By attending to both the structure and content of written arguments, the argument-chains approach overcomes key limitations of structural analyses commonly used in writing assessment. Drawing on a dataset of 471 argumentative essays written by Grade 5 students, we used argument chains to reconstruct students’ multi-step reasoning, identify implicit assumptions, and assign acceptability and relevance scores to individual propositions. This analysis enabled us to pinpoint specific weaknesses in argumentation at the level of individual propositions, generating precise diagnostic information to support actionable instructional feedback. It also helped us evaluate multiple essay-level dimensions of student performance, including overall argument quality, acceptability, relevance, and consideration of opposing perspectives. Reliability studies indicated high agreement between human raters (Quadratic Weighted Kappa =.84) and high agreement between human raters and large language models (Quadratic Weighted Kappa =.81). Most discrepancies involved assumed propositions, highlighting an area that requires further investigation. While this work is in its early stages, we suggest that argument chains have the potential, with further automation, to become a broadly applicable tool for assessing argumentative writing.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101088"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553718","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01DOI: 10.1016/j.asw.2026.101081
Andrew Potter, Yu Tian, Renu Balyan, Maria Goldshtein, Laura K. Allen, Danielle S. McNamara
{"title":"Source integration in source-based writing: From theoretical modeling to predictive analytics","authors":"Andrew Potter, Yu Tian, Renu Balyan, Maria Goldshtein, Laura K. Allen, Danielle S. McNamara","doi":"10.1016/j.asw.2026.101081","DOIUrl":"10.1016/j.asw.2026.101081","url":null,"abstract":"<div><div>Effective source integration is a complex but essential skill in academic writing; yet it remains difficult to teach and evaluate. Assessing source integration has important implications for formative feedback and instructional practice, but existing approaches face limitations, particularly in automated systems. The purpose of this study was to develop and validate an automated measure of source integration by extracting linguistic features from student essays and modeling them as a latent construct with confirmatory factor analysis. Predictive validity with human scores and generalizability across datasets using measurement invariance was tested and examined. The source integration construct included linguistic features related to citation, quotation, plagiarism, and semantic overlap with the source text. Results indicated strong alignment with human ratings (<em>β</em> = .81, <em>R²</em> = .65) and evidence of structural consistency across a new dataset with novel prompts and sources. Predictive utility analyses showed that the latent construct improved machine learning models and enhanced agreement with human ratings when paired with BERT embeddings. GPT-5.2 produced interpretable justifications but lower scoring reliability. These findings suggest that a source integration construct grounded in linguistic features can complement modern AI methods, providing a foundation for formative feedback on source-based writing that addresses issues of fairness and interpretability.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101081"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553726","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-19DOI: 10.1016/j.asw.2026.101087
Li Ruyang, Cheng Zhang, Hedi Ye
{"title":"When does GenAI feedback support learning? Trust calibration and verification in L2 writing assessment","authors":"Li Ruyang, Cheng Zhang, Hedi Ye","doi":"10.1016/j.asw.2026.101087","DOIUrl":"10.1016/j.asw.2026.101087","url":null,"abstract":"<div><div>GenAI feedback can support L2 writing, but its benefits vary across classrooms. This study examines when GenAI-mediated formative feedback may support learning within classroom assessment routines. Using a QUAL-dominant embedded mixed-methods multiple-case design in two university EFL writing classes, we examined how assessment enactment was associated with students’ trust calibration and verification practices. Across cases, clearer norms, learning-oriented accountability, and teacher positioning of AI as contingent input were associated with more frequent verification, deeper revisions, and stronger transfer. Under higher assessment pressure and weaker verification routines, students more often accepted AI feedback uncritically, showed less revision reasoning, and demonstrated weaker transfer, patterns consistent with what we term learning displacement. Using RAFE as a preliminary analytic framework, the study traces how classroom conditions, assessment enactment, trust calibration, and appropriation trajectories were related across the two cases. The findings suggest that verification-oriented norm design is an important condition for more sustainable GenAI-supported L2 writing development.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101087"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553721","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Assessing WritingPub Date : 2026-07-01Epub Date: 2026-06-16DOI: 10.1016/j.asw.2026.101085
Kelly D. Hartwell, Laura L. Aull
{"title":"When AI reframes what ‘counts’ as ‘good’ writing","authors":"Kelly D. Hartwell, Laura L. Aull","doi":"10.1016/j.asw.2026.101085","DOIUrl":"10.1016/j.asw.2026.101085","url":null,"abstract":"<div><div>This editorial introduces the 2026 Tools & Tech Forum, which examines how generative AI writing technologies are transforming writing assessment and pedagogy. Across the contributions, a central theme emerges: AI systems do not merely assist writers but actively shape what becomes recognizable as assessable and valued writing. This raises important questions about validity, authorship, language bias, and assessment justice. Collectively, the reviews invite readers to consider AI tools as rhetorical and assessment infrastructures that reframe the constructs and consequences of writing assessment in contemporary educational contexts.</div></div>","PeriodicalId":46865,"journal":{"name":"Assessing Writing","volume":"69 ","pages":"Article 101085"},"PeriodicalIF":7.6,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148553722","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":1,"RegionCategory":"文学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}