An exploratory comparison of ChatGPT-generated and human-made Japanese reading comprehension test items
Downloads
Published
Copyright (c) 2026 Gilbert Dizon, Ryo Kurose

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.
Accepted: 4 April, 2026
Abstract
While existing research involving generative artificial intelligence (GenAI) has focused primarily on English, little is known about the technology’s impact on other languages, particularly in the context of second language (L2) assessment. Thus, this study examined the potential of GenAI to support non-English language assessment by comparing expert-developed reading comprehension test items and ChatGPT-generated test materials based on the Japanese Language Proficiency Test (JLPT). Japanese university language instructors (N = 16) blindly evaluated both sets of materials in terms of naturalness, attractiveness, and overall quality. Descriptive statistics and the Mann-Whitney U test revealed no significant differences between the two types of reading assessment materials. Qualitative responses indicated some concerns regarding the naturalness of the expressions, cohesion, and answer choice quality, but no major perceived differences overall. These findings suggest that GenAI may serve as a useful supplementary tool for developing JLPT preparation materials, potentially reducing time and cost burdens. However, human expertise remains necessary to refine and ensure the quality of GenAI-generated assessment materials.
Keywords: Generative Artificial Intelligence, Japanese as a second language, Japanese language proficiency test, languages other than English


