Self-Training with Process Preference Learning using Dynamic Value Margin (SPPD) Improves Language Model Performance and Generalizability

Thursday 27 March 2025


The quest for more intelligent language models has led researchers down a path of self-training, using their own reasoning abilities to improve their performance. This approach, known as Self-Training with Process Preference Learning using Dynamic Value Margin (SPPD), has yielded promising results.


At its core, SPPD is a method for training large language models to reason more effectively by leveraging the model’s own ability to evaluate its own responses. The approach uses a process-based Markov Decision Process (MDP) to derive dynamic value margins on step-level preference optimization, which allows the model to adjust its margin based on signals from the value model.


The researchers behind SPPD have developed a novel self-training framework that integrates process preference learning with dynamic value margins. This framework is designed to promote more reliable and generalizable reasoning abilities in language models by leveraging their ability to evaluate their own responses.


To train these models, the researchers used a dataset of mathematical problems, which they evaluated using a combination of automated testing and human evaluation. The results were impressive: SPPD was able to improve the performance of the models on both correct and incorrect trajectories, with accuracy rates exceeding 90% in some cases.


One of the key advantages of SPPD is its ability to adapt to different problem-solving strategies. By using a process-based MDP, the model can adjust its margin based on signals from the value model, allowing it to learn more effectively from its own mistakes.


The researchers also evaluated the performance of SPPD on out-of-domain test datasets, which showed significant improvements over previous methods. This suggests that SPPD is not only effective for in-domain testing but also generalizable to new and unseen problems.


In addition to its technical merits, SPPD has important implications for the development of more intelligent language models. By leveraging the model’s own ability to evaluate its own responses, SPPD provides a powerful tool for improving the performance and reliability of these models.


Overall, the results of this research demonstrate the potential of SPPD as a new approach to training large language models. By promoting more reliable and generalizable reasoning abilities, SPPD has the potential to revolutionize the field of natural language processing and open up new possibilities for AI applications.


Cite this article: “Self-Training with Process Preference Learning using Dynamic Value Margin (SPPD) Improves Language Model Performance and Generalizability”, The Science Archive, 2025.


Language Models, Self-Training, Sppd, Process Preference Learning, Dynamic Value Margin, Markov Decision Process, Reasoning Abilities, Natural Language Processing, Ai Applications, Generalizable.


Reference: Hao Yi, Qingyang Li, Yulan Hu, Fuzheng Zhang, Di Zhang, Yong Liu, “SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin” (2025).


Leave a Reply