Speech-to-text model comparison using XLS-R, XLSR-53, and Wav2Vec 2.0
Abstract
Automatic speech recognition (ASR) systems have achieved significant progress in recent years; however, their performance remains limited for low-resource languages such as Indonesian. Multilingual ASR models are often expected to generalize across languages, yet they frequently underperform when applied to underrepresented languages without sufficient adaptation. This study presents a comparative evaluation of three ASR models—Wav2Vec 2.0, XLS-R, and XLSR-53—on Indonesian speech to analyze the impact of monolingual fine-tuning versus multilingual pretraining. The evaluation was conducted using approximately 28 hours of validated Indonesian speech from the Common Voice Corpus version 13. Model performance was assessed using word error rate (WER) without employing any external language model to ensure a fair comparison. Experimental results demonstrate that Wav2Vec 2.0, which is fine-tuned specifically for Indonesian, achieves substantially lower WER compared to the multilingual models. Qualitative analysis further confirms that multilingual models exhibit higher omission and substitution errors. These findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian. The results provide practical guidance for deploying ASR systems in low-resource language scenarios and highlight the importance of targeted model adaptation.
Keywords
Speech recognition; Wav2Vec 2.0; Word error rate; XLS-R; XLSR-53
Full Text:
PDFDOI: https://doi.org/10.11591/eei.v15i4.11431
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Bulletin of Electrical Engineering and Informatics (BEEI)
ISSN: 2089-3191
,
e-ISSN: 2302-9285
This journal is published by the
Institute of Advanced Engineering and Science (IAES)
.