Neural Codec Source Tracing (NCST) Benchmark Advances Audio Deepfake Detection

Thursday 06 March 2025


Researchers have made significant progress in developing a comprehensive benchmark for evaluating audio deepfake detection models, which are designed to identify manipulated audio recordings. The new benchmark, known as Neural Codec Source Tracing (NCST), aims to improve the robustness and accuracy of these models in detecting not only fake audio but also unseen real audio.


The NCST benchmark is built upon a dataset called ST-Codecfake, which contains audio samples generated by 11 state-of-the-art neural codec methods and out-of-distribution test samples. The dataset was created to mimic real-world scenarios where deepfake audio may be used to manipulate audio recordings.


To evaluate the performance of the NCST models, researchers conducted two types of experiments: in-distribution (ID) close-set evaluation and out-of-distribution (OOD) open-set evaluation. In the ID experiment, the models were tested on their ability to classify real and fake audio samples generated by known neural codec methods. The results showed that the models performed well in identifying fake audio, but struggled with distinguishing between real and unseen real audio.


The OOD experiment, on the other hand, aimed to test the models’ ability to detect novel fake audio algorithms not present in the training dataset. The results revealed that the models were able to identify these novel fake audio methods with high accuracy, but still had difficulty classifying unseen real audio.


One of the key findings from this research is that the NCST models tend to produce excessively high-confidence logits for known ID samples, making it challenging for them to distinguish between unknown OOD samples. This highlights the need for further improvement in the robustness and generalizability of deepfake detection models.


The development of the NCST benchmark and ST-Codecfake dataset has significant implications for the field of audio deepfake detection. It provides a standardized framework for evaluating the performance of different models, allowing researchers to compare and improve their methods. Additionally, it highlights the importance of considering unseen real audio in the evaluation process, as these samples can be easily misclassified by current models.


As the use of deepfake technology becomes increasingly prevalent, the development of robust and accurate detection methods is crucial for ensuring the integrity of audio recordings. The NCST benchmark and ST-Codecfake dataset are essential steps towards achieving this goal, providing a foundation for future research in this area.


Cite this article: “Neural Codec Source Tracing (NCST) Benchmark Advances Audio Deepfake Detection”, The Science Archive, 2025.


Audio Deepfake Detection, Neural Codec Methods, St-Codecfake Dataset, Ncst Benchmark, Audio Recordings, Machine Learning Models, Fake Audio, Unseen Real Audio, Robustness, Generalizability


Reference: Yuankun Xie, Xiaopeng Wang, Zhiyong Wang, Ruibo Fu, Zhengqi Wen, Songjun Cao, Long Ma, Chenxing Li, Haonnan Cheng, Long Ye, “Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition” (2025).


Leave a Reply