Wednesday 12 March 2025
The world of machine learning is fraught with danger, where malicious actors can secretly manipulate datasets to manipulate models and wreak havoc on systems. This type of attack is known as dataset poisoning, and it’s a growing concern in the field.
One approach to detecting these attacks is to use statistical methods to identify anomalies in the data. However, this can be challenging, especially when dealing with large and complex datasets. A new paper published in the Journal of Machine Learning Research proposes an alternative solution: using conformal prediction to detect dataset poisoning attacks.
Conformal prediction is a statistical technique that provides a way to quantify the uncertainty associated with a model’s predictions. In the context of detecting dataset poisoning, it can be used to identify instances where the data appears to be anomalous or inconsistent with the rest of the dataset.
The authors propose a new method called Conformal Separability Test (CST), which uses conformal prediction to detect dataset poisoning attacks. The CST works by constructing a statistical model of the clean data, and then using this model to identify instances where the data appears to be anomalous or inconsistent with the rest of the dataset.
To evaluate the effectiveness of the CST, the authors conducted a series of experiments using real-world datasets and simulated attacks. They found that the CST was able to detect most of the poisoned samples with high accuracy, even when the attack rate was as low as 1%.
The authors also compared the performance of the CST with other state-of-the-art methods for detecting dataset poisoning attacks. They found that the CST outperformed these methods in terms of detection accuracy and robustness.
One of the key advantages of the CST is its ability to detect subtle attacks, where the attacker has only slightly modified a small number of samples. This can be particularly challenging for other methods, which may rely on more dramatic changes to the data to trigger an alarm.
The authors also highlight some potential limitations of their approach. For example, they note that the CST may not perform well in situations where the attacker is able to create multiple poisoned samples that are highly similar to each other. They also suggest that future work could focus on developing more robust methods for handling noise and outliers in the data.
Overall, the paper provides a promising new approach to detecting dataset poisoning attacks, which has significant implications for the security of machine learning systems. By using conformal prediction to identify anomalies in the data, the CST offers a powerful tool for defending against these types of attacks.
Cite this article: “Detecting Dataset Poisoning Attacks with Conformal Prediction”, The Science Archive, 2025.
Machine Learning, Dataset Poisoning, Conformal Prediction, Statistical Methods, Anomaly Detection, Data Security, Model Manipulation, Attacks, Robustness, Accuracy







