Friday 14 March 2025
Researchers have made significant strides in developing automatic speech recognition (ASR) technology for low-resource languages, such as Armenian and Georgian. These languages are often spoken by small communities and lack the vast amounts of data required to train sophisticated AI models. To address this issue, a team has expanded their datasets using various methods, including crowd-sourcing, audiobooks, and pseudo-labeling.
The researchers began by increasing the amount of labeled data available for Armenian, which is notoriously limited. They organized two data collection events at universities in Armenia, where participants were invited to read aloud and record texts from a common voice corpus. This effort added 48 hours of validated speech data to their dataset, bringing the total to around 50 hours.
In addition to this, they used YouTube videos as a source of unlabeled data. They curated high-quality content, focusing on interviews, shows, and podcasts in Armenian, and removed non-Armenian language segments. This process yielded a cleaned dataset of approximately 145 hours, which was then used for iterative pseudo-labeling.
The team also turned to audiobooks as another source of unlabeled data. They partnered with Grqaser, a digital library of Armenian audiobooks, to access over 20 hours of audio-text pairs. By combining these datasets, they were able to train ASR models that achieved impressive results.
To evaluate the effectiveness of their methods, the researchers conducted an ablation study. This involved measuring the statistical significance of each data extension approach by comparing its impact on model performance. The results showed that paid crowd-sourcing offered the best balance between cost and quality, outperforming volunteer crowd-sourcing, open-source audiobooks, and unlabeled data.
The team’s top-performing ASR models, trained on their enriched datasets, consistently surpassed existing baselines. For Armenian, they achieved a word error rate (WER) of 5.73%, while for Georgian, the WER was 10.59%. These results demonstrate significant improvements in ASR performance for low-resource languages.
The study’s findings have important implications for language technology and global inclusivity. By developing effective data collection strategies for low-resource languages, researchers can improve access to information and communication tools for marginalized communities. The team’s work also highlights the potential of pseudo-labeling as a cost-effective method for leveraging unlabeled data in ASR development.
Cite this article: “Expanding Language Data to Improve Automatic Speech Recognition for Low-Resource Languages”, The Science Archive, 2025.
Automatic Speech Recognition, Low-Resource Languages, Armenian, Georgian, Data Collection, Crowdsourcing, Audiobooks, Pseudo-Labeling, Language Technology, Global Inclusivity







