Streamlining learner corpus development with LLMs and NLP

Downloads

Published

2026-06-03

Section: Articles

Authors

  • Gavin Brooks Email Kyoto University of Foreign Studies, Japan
DOI: https://doi.org/10.29140/jct.v2n1.103562

Abstract

Learner corpus projects increasingly require scalable, reliable cleaning and annotation in order to develop corpora that are ecologically valid and span more learner contexts. This paper evaluates a hybrid pipeline that combines deterministic NLP preprocessing with a multi-LLM consensus classifier for context-dependent decisions. Using a pilot of 25 IELTS essays sampled from a 1,329-essay corpus, algorithms were used to handle document matching, anonymisation, segmentation, and mechanical typo correction, while three commercial LLMs independently classified off-list tokens (spelling errors, proper nouns, technical terms, and related categories). Automatic corrections were applied only when the models agreed, while disagreements triggered human-in-the-loop review. Deterministic plus targeted LLM use improved document matching and segmentation, while the LLMs were able to successfully clean 97.4% of the off-list words in the corpus. Results indicate that a pipeline that employs a multi-LLM consensus, coupled with algorithmic preprocessing and selective human oversight, can provide a practical, cost-conscious path to scaling learner corpus development while preserving methodological transparency.


Keywords: Corpus linguistics, learner corpora, LLMs, NLP

Suggested Citation:

Brooks, G. (2026). Streamlining learner corpus development with LLMs and NLP. JALTCALL Trends, 2(1), 103562. https://doi.org/10.29140/jct.v2n1.103562