An exploratory comparison of ChatGPT-generated and human-made Japanese reading comprehension test items

Downloads

Published

2026-04-24

Section: Regular Articles


Received: 19 February, 2026
Accepted: 4 April, 2026

Authors

  • Gilbert Dizon Email ORCiD Kansai University, Japan
  • Ryo Kurose Email ORCiD Kyoto University of Foreign Studies, Japan
DOI: https://doi.org/10.29140/lea.2026.104152

Abstract

While existing research involving generative artificial intelligence (GenAI) has focused primarily on English, little is known about the technology’s impact on other languages, particularly in the context of second language (L2) assessment. Thus, this study examined the potential of GenAI to support non-English language assessment by comparing expert-developed reading comprehension test items and ChatGPT-generated test materials based on the Japanese Language Proficiency Test (JLPT). Japanese university language instructors (N = 16) blindly evaluated both sets of materials in terms of naturalness, attractiveness, and overall quality. Descriptive statistics and the Mann-Whitney U test revealed no significant differences between the two types of reading assessment materials. Qualitative responses indicated some concerns regarding the naturalness of the expressions, cohesion, and answer choice quality, but no major perceived differences overall. These findings suggest that GenAI may serve as a useful supplementary tool for developing JLPT preparation materials, potentially reducing time and cost burdens. However, human expertise remains necessary to refine and ensure the quality of GenAI-generated assessment materials.


Keywords: Generative Artificial Intelligence, Japanese as a second language, Japanese language proficiency test, languages other than English

Suggested Citation:

Dizon, G., & Kurose, R. (2026). An exploratory comparison of ChatGPT-generated and human-made Japanese reading comprehension test items. Language Education & Assessment, 9, 104152. https://doi.org/10.29140/lea.2026.104152