Bangla Text-to-Speech System Achieves High-Quality Speech Synthesis with Minimal Training Data

Saturday 22 March 2025


Researchers have made significant progress in developing a Text-to-Speech (TTS) system that can generate high-quality speech outputs for low-resource languages like Bangla. The study, published recently, introduces BnTTS, the first framework for Bangla speaker adaptation-based TTS, designed to bridge the gap in Bangla speech synthesis using minimal training data.


The team developed a comprehensive dataset acquisition framework to collect and refine high-quality speech data with aligned transcripts. This framework leverages advanced speech processing models and carefully designed algorithms to process raw audio inputs and generate refined audio outputs with word-aligned transcripts.


To evaluate the performance of BnTTS, the researchers employed a range of objective metrics, including Character Error Rate (CER), SpeechBERTScore, Speaker Encoder Cosine Similarity (SECS), and Duration Equality Score. They also conducted subjective evaluations using a five-point rating scale to assess naturalness, clarity, fluency, consistency, and emotional expressiveness.


The results show that BnTTS outperforms existing TTS systems in Bangla speech synthesis, achieving superior performance in terms of intelligibility, naturalness, and speaker fidelity. The system’s ability to adapt to different speakers and languages demonstrates its potential for real-world applications.


One of the key challenges in developing a TTS system is collecting high-quality training data. In this study, the researchers used a combination of manual review and automated processing techniques to filter out low-quality audio segments and ensure that the final dataset consisted only of high-fidelity speech samples.


The team also developed an evaluation platform specifically designed for subjective assessment of TTS systems. This platform features anonymity, comprehensive evaluation criteria, and a streamlined user interface, allowing evaluators to rate audio samples efficiently and accurately.


The study highlights the importance of considering speaker adaptation in TTS systems, particularly for low-resource languages where large-scale datasets are scarce. The results demonstrate that BnTTS can effectively adapt to different speakers and languages, making it a promising solution for real-world applications.


In addition to its technical advancements, the study also underscores the significance of human evaluation in assessing the quality of speech synthesis outputs. By combining objective metrics with subjective evaluations, researchers can gain a more comprehensive understanding of the system’s performance and identify areas for improvement.


The development of BnTTS has significant implications for industries such as education, healthcare, and entertainment, where high-quality speech synthesis is crucial for effective communication.


Cite this article: “Bangla Text-to-Speech System Achieves High-Quality Speech Synthesis with Minimal Training Data”, The Science Archive, 2025.


Text-To-Speech, Bangla Language, Speech Synthesis, Tts System, Speaker Adaptation, Low-Resource Languages, Naturalness, Intelligibility, Speaker Fidelity, Objective Metrics


Reference: Mohammad Jahid Ibna Basher, Md Kowsher, Md Saiful Islam, Rabindra Nath Nandi, Nusrat Jahan Prottasha, Mehadi Hasan Menon, Tareq Al Muntasir, Shammur Absar Chowdhury, Firoj Alam, Niloofar Yousefi, et al., “BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting” (2025).


Leave a Reply