Comparing ASR system accuracy on non-native English read speech across L1 backgrounds
Downloads
Published
Copyright (c) 2026 Michael McGuire

This work is licensed under a Creative Commons Attribution 4.0 International License.
Abstract
This study assesses five cutting-edge ASR systems' accuracy in the recognition of non-native accented English read speech from six different L1 backgrounds. Speech samples were taken from the L2-ARCTIC non-native English speech corpus, containing recordings of 24 speakers from Arabic, Chinese, Hindi, Korean, Spanish, and Vietnamese L1s. Four US English speakers were included as a control. Five ASR systems were tested: AssemblyAI's Universal-2, Deepgram's Nova-2, RevAI's V2, Speechmatics' Ursa-2, and Whisper large-v3 by OpenAI. ASR accuracy was measured using the Match Error Rate (MER) metric. Overall, the study found much lower error rates for all five selected ASR systems than systems used in previous studies. Results showed that Whisper and AssemblyAI performed significantly better than other models, with no significant difference between them. Hindi speakers showed lower error rates while Chinese and Vietnamese showed notably higher error rates across all systems.
Keywords: automatic speech recognition, learner English, non-native speech, CALL

