Thursday 20 March 2025
The quest for a more efficient and safe way to learn from human feedback has been an ongoing challenge in the field of artificial intelligence. Researchers have attempted to address this issue by developing methods that allow humans to actively intervene during training, providing corrective feedback to ensure the AI agent learns from its mistakes. However, these approaches often rely on manual labeling or require significant computational resources.
A new method proposed by a team of researchers seeks to alleviate these limitations by introducing a proxy value function that can express human intent. This approach, called Proxy Value Propagation (PVP), allows humans to actively intervene during training and provides a more efficient way to learn from feedback.
The PVP method works by labeling state-action pairs in the human demonstration with high values, indicating correct actions taken by the human. Conversely, agent actions that are intervened upon receive low values. Through this process, the proxy value function induces a policy that faithfully emulates human behavior. The TD-learning framework is used to propagate labeled values to other unlabeled data generated from agents’ exploration.
The researchers tested PVP in four virtual simulated environments: MetaDrive Safety Benchmark, CARLA Town01, Grand Theft Auto V (GTA V), and MiniGrid Two Room. In each environment, they compared the performance of PVP with a baseline method, HACO, which also utilizes human feedback but relies on manual labeling.
The results showed that PVP outperformed HACO in all four environments, achieving better success rates and reduced safety costs. In the MetaDrive Safety Benchmark task, for example, PVP demonstrated a 20% improvement in success rate compared to HACO. Similarly, in the GTA V environment, PVP achieved a 15% reduction in safety cost.
The researchers also explored the impact of different human input devices on PVP’s performance. They found that using a steering wheel or keyboard as an input device resulted in better performance compared to using a gamepad. This suggests that the type of input device used can influence the effectiveness of human feedback in AI training.
PVP’s ability to learn from human feedback without manual labeling makes it an attractive solution for applications where safety and efficiency are critical, such as autonomous driving or robotic control. The method’s flexibility also allows it to be adapted to various environments and tasks, making it a promising approach for addressing the challenges of human-AI collaboration.
In addition to its potential practical applications, PVP’s design provides valuable insights into the nature of human-AI interaction.
Cite this article: “Efficient Human-AI Collaboration Through Proxy Value Propagation”, The Science Archive, 2025.
Artificial Intelligence, Machine Learning, Human-Computer Interaction, Autonomous Driving, Robotic Control, Feedback, Reinforcement Learning, Proxy Value Function, Td-Learning Framework, Human-Ai Collaboration







