Thursday 06 March 2025
The quest for reliable data has long been a challenge in the field of vertical federated learning, where multiple parties contribute their own datasets to train models without sharing sensitive information. The problem lies in dealing with non-overlapping samples, which can lead to poor model performance and reduced accuracy.
Researchers have attempted to address this issue by imputing missing attributes and predicting pseudo-labels for these non-overlapping samples. However, this approach is often plagued by noise and uncertainty, which can negatively impact the overall performance of the model.
A new framework, proposed by a team of scientists, aims to tackle this challenge head-on. The Reliable Imputed-Sample Assisted (RISA) framework combines mean imputation with self-training to predict pseudo-labels for non-overlapping samples. By leveraging evidence theory, RISA also estimates the uncertainty associated with each sample, allowing for more reliable predictions.
The approach is based on the idea that not all missing values are created equal. By identifying and selecting only the most reliable imputed samples, RISA can improve the overall quality of the data and reduce the impact of noise. This is achieved through the use of evidence theory, which provides a mathematical framework for combining multiple sources of information to estimate uncertainty.
In addition to its improved performance, the RISA framework also offers several practical benefits. For example, it requires minimal additional computational resources and can be easily integrated into existing vertical federated learning architectures.
The researchers tested their approach on two widely used datasets – CIFAR-10 and Criteo – with impressive results. The RISA framework consistently outperformed other methods, achieving significant gains in accuracy, especially when the number of overlapping samples was limited.
While there is still much work to be done in this area, the RISA framework represents a significant step forward in the development of reliable data imputation and pseudo-label prediction techniques for vertical federated learning. By addressing the challenges posed by non-overlapping samples, researchers can move closer to achieving more accurate and robust models that better serve the needs of their users.
The potential applications of this technology are vast, ranging from improved healthcare outcomes to enhanced customer experiences in e-commerce. As the demand for data-driven insights continues to grow, so too does the need for innovative solutions like RISA that can help unlock the full potential of vertical federated learning.
Cite this article: “Reliable Imputation and Pseudo-Label Prediction in Vertical Federated Learning”, The Science Archive, 2025.
Vertical Federated Learning, Data Imputation, Pseudo-Label Prediction, Evidence Theory, Uncertainty Estimation, Non-Overlapping Samples, Model Performance, Accuracy Improvement, Computational Efficiency, Machine Learning







