A study on the reliability of ChatGPT-based English writing assessment: Internal and external perspectives

Downloads

Published

2026-08-19

Section: Regular Articles


Received: 28 September, 2025
Accepted: 20 July, 2026

Authors

  • Boyu Wang Email ORCiD Huazhong University of Science and Technology, Hubei, China
  • Yunjia Zhang Email ORCiD Sichuan International Studies University, China
DOI: https://doi.org/10.29140/lea.2026.103386

Abstract

With the advancement of big data and artificial intelligence, natural language processing (NLP) has been increasingly integrated into educational assessment, facilitating a shift from human to automated scoring in English writing assessment. This study investigates how prompt design, guided by the TELeR taxonomy (Santu & Feng, 2023), and sample training affects ChatGPT-4's scoring reliability. A sample of 120 IELTS-style essays was scored across four analytic dimensions--Task Response (TR), Coherence and Cohesion (CC), Lexical Resource (LR), and Grammatical Range and Accuracy (GRA)--as well as holistically. Internal (intra-rater) reliability was measured via ICC(3,1) across three scoring rounds, while external (inter-rater) reliability was assessed by integrating ChatGPT-4 as an additional rater into a panel of nine human scorers using ICC(2,1). The results revealed that ChatGPT-4's internal consistency was moderate and was markedly improved by detailed prompts on content-related dimensions (TR/CC), but showed minimal gains on linguistic dimensions (LR/GRA). Training with human-scored exemplars under the better prompt reduced discrepancies with human ratings for external reliability, yet slightly decreased internal consistency for most dimensions. These findings position TELeR as a framework linking prompt engineering to psychometric theory and underscore the value of dimension-specific prompt design for reliable AI-based writing assessment.


Keywords: ChatGPT, Human Rating, Writing Assessment, Reliability

Suggested Citation:

Wang, B., & Zhang, Y. (2026). A study on the reliability of ChatGPT-based English writing assessment: Internal and external perspectives. Language Education & Assessment, 9, 103386. https://doi.org/10.29140/lea.2026.103386