Advances in Objective Speech Quality Assessment Models

Thursday 27 March 2025


The quest for a speech quality assessment model that can accurately predict human perception across languages has been an ongoing challenge in the field of audio processing. Researchers have long struggled to develop models that can effectively capture the complexities of speech degradation, noise, and linguistic characteristics. Recently, two teams of scientists made significant strides in this area by developing new architectures for objective speech quality assessment.


The first model, developed by a team at Quality and Usability Lab, Technische Universität Berlin, is based on a convolutional neural network (CNN) known as NISQA. This model uses Mel-spectrograms to extract features from audio signals and then applies a self-attention mechanism to capture long-range dependencies in the data. The results show that NISQA achieves strong correlations with human subjective ratings across multiple languages, including English, Mandarin, French, German, Swedish, and Dutch.


The second model, developed by a team at Deutsches Forschungszentrum für Künstliche Intelligenz (German Research Center for Artificial Intelligence), is based on an audio spectrogram transformer (AST) architecture. This model uses a self-attention mechanism to process the input audio signal in parallel, allowing it to capture complex patterns and relationships between different frequency bands.


The results of both models are strikingly similar, with each achieving strong correlations with human subjective ratings across multiple languages. However, there are some notable differences between the two models. For example, NISQA seems to perform better in certain languages, such as Mandarin, where its Mel-spectrogram features appear to capture the unique tonal characteristics of the language.


On the other hand, AST appears to be more robust in handling noise and distortion in certain languages, such as Dutch. This suggests that the self-attention mechanism in AST is particularly effective at capturing complex patterns in noisy data.


One of the most interesting aspects of these results is the insight they provide into the complexities of speech quality assessment. For example, the study reveals that different languages exhibit distinct patterns of degradation and noise, which can affect the performance of speech quality models. This highlights the importance of developing models that are not only effective but also linguistically diverse.


The implications of these findings are significant for a wide range of applications, from voice assistants to speech recognition systems. By developing more accurate and robust speech quality assessment models, researchers can improve the overall performance of these systems and enhance the user experience.


Cite this article: “Advances in Objective Speech Quality Assessment Models”, The Science Archive, 2025.


Speech Quality, Audio Processing, Neural Networks, Convolutional Neural Network, Self-Attention Mechanism, Mel-Spectrograms, Audio Spectrogram Transformer, Language, Noise, Distortion.


Reference: Wafaa Wardah, Tuğçe Melike Koçak Büyüktaş, Kirill Shchegelskiy, Sebastian Möller, Robert P. Spang, “Language Barriers: Evaluating Cross-Lingual Performance of CNN and Transformer Architectures for Speech Quality Estimation” (2025).


Leave a Reply