Sunday 06 April 2025
The art of evaluating speech synthesis has long been a topic of debate in the tech community. For years, researchers have relied on subjective measures like mean opinion scores (MOS) to gauge the quality and naturalness of synthesized speech. However, these methods have their limitations, and a new report from Erica Cooper and colleagues aims to provide some much-needed clarity.
The authors argue that many papers about speech synthesis rely too heavily on basic MOS tests, which are often conducted in isolation without considering the specific use case or target audience. This approach can lead to misleading results, as listeners may not be qualified to evaluate the quality of synthesized speech for a particular application.
Cooper and her team suggest that researchers should instead focus on designing evaluations that are more closely tied to the intended use of the synthesis technology. For example, if a system is designed to help language learners practice pronunciation, then the evaluation should include listeners who are familiar with the target language and have some experience with language learning.
The report also emphasizes the importance of using multiple evaluation metrics to get a more complete picture of a system’s performance. This might include not only subjective measures like MOS, but also objective metrics such as automatic quality predictors (AQPs) and statistical analysis.
AQPs, in particular, are gaining popularity as a way to assess speech synthesis quality. These algorithms use machine learning models to predict the quality of synthesized speech based on various acoustic features. While they can be useful, the authors caution that AQPs should not be relied upon exclusively, as they may not generalize well across different domains or languages.
The report also highlights some common pitfalls in speech synthesis evaluation, such as inadequate listener qualification and insufficient statistical power. To avoid these issues, researchers should strive to recruit a diverse pool of listeners and conduct thorough statistical analysis to ensure that their results are reliable and generalizable.
In addition to providing guidance on best practices for evaluating speech synthesis, the report also calls attention to some of the limitations of current evaluation methods. For example, MOS scores may not be directly comparable across different papers or experiments, as they can be influenced by a variety of factors such as listener fatigue and testing conditions.
Overall, the authors’ recommendations aim to promote more rigorous and meaningful evaluations of speech synthesis technology. By taking a more nuanced approach that considers the specific use case and target audience, researchers can develop more effective systems that better meet the needs of users.
Cite this article: “Evaluating Speech Synthesis: A Guide to Best Practices”, The Science Archive, 2025.
Speech Synthesis, Evaluation Methods, Mean Opinion Scores, Mos, Quality Predictors, Aqps, Machine Learning, Listener Qualification, Statistical Analysis, Speech Technology.







