A study on the reliability of ChatGPT-based English writing assessment: Internal and external perspectives
Downloads
Published
Copyright (c) 2026 Boyu Wang, Yunjia Zhang

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.
Accepted: 20 July, 2026
Abstract
With the advancement of big data and artificial intelligence, natural language processing (NLP) has been increasingly integrated into educational assessment, facilitating a shift from human to automated scoring in English writing assessment. This study investigates how prompt design, guided by the TELeR taxonomy (Santu & Feng, 2023), and sample training affects ChatGPT-4's scoring reliability. A sample of 120 IELTS-style essays was scored across four analytic dimensions--Task Response (TR), Coherence and Cohesion (CC), Lexical Resource (LR), and Grammatical Range and Accuracy (GRA)--as well as holistically. Internal (intra-rater) reliability was measured via ICC(3,1) across three scoring rounds, while external (inter-rater) reliability was assessed by integrating ChatGPT-4 as an additional rater into a panel of nine human scorers using ICC(2,1). The results revealed that ChatGPT-4's internal consistency was moderate and was markedly improved by detailed prompts on content-related dimensions (TR/CC), but showed minimal gains on linguistic dimensions (LR/GRA). Training with human-scored exemplars under the better prompt reduced discrepancies with human ratings for external reliability, yet slightly decreased internal consistency for most dimensions. These findings position TELeR as a framework linking prompt engineering to psychometric theory and underscore the value of dimension-specific prompt design for reliable AI-based writing assessment.
Keywords: ChatGPT, Human Rating, Writing Assessment, Reliability


