AI System Matches Human Accuracy in Math Problem-Solving

Wednesday 12 March 2025


For decades, researchers have been trying to crack the code of human intelligence. One of the most significant challenges is understanding how humans solve complex mathematical problems. Now, a team of scientists has made a breakthrough in developing an AI system that can judge math problems as accurately as a human.


The new system, called PairJudge RM, uses a unique approach to evaluate the correctness of math solutions. Unlike traditional systems that rely on scoring and ranking, PairJudge RM judges two candidate solutions simultaneously, eliminating the need for scoring and enabling cross-validation of solutions through parallel judgment.


To develop this system, researchers created a large-scale dataset of 432K pairwise judgments derived from NuminaMath and annotated using gemini-1.5-flash. This dataset is then used to train the PairJudge RM model.


The model’s performance was tested on two benchmark datasets: MATH-500, which covers a wide range of mathematical concepts, and Olympiad Bench, which includes challenging problems from international math olympiads. The results showed that PairJudge RM significantly outperformed baseline reward models in selecting the best candidate solution.


One of the key advantages of PairJudge RM is its ability to eliminate arbitrary and inconsistent scores assigned by traditional reward models. By judging two candidate solutions simultaneously, the system can detect errors and inconsistencies more effectively than a single-scoring approach.


The researchers also explored alternative tournament strategies, such as round-robin and Swiss-system tournaments, to improve the performance of PairJudge RM. These experiments showed that different tournament designs can lead to varying levels of accuracy and efficiency.


To further improve the model’s performance, the researchers suggest exploring larger model capacities, more data scaling dimensions, and long-COT base models. They also plan to release their codebase, model checkpoints, and dataset used in this work to facilitate replication of experiments.


The development of PairJudge RM has significant implications for artificial intelligence research and applications. The system’s ability to accurately judge math problems can be extended to other domains, such as natural language processing and computer vision. Moreover, the approach can be adapted to other problem-solving tasks that require complex reasoning and judgment.


Overall, the creation of PairJudge RM represents a major step forward in developing AI systems that can rival human intelligence in solving complex mathematical problems. By leveraging this technology, researchers can unlock new possibilities for artificial intelligence and improve our understanding of human cognition.


Cite this article: “AI System Matches Human Accuracy in Math Problem-Solving”, The Science Archive, 2025.


Artificial Intelligence, Math Problems, Pairjudge Rm, Human Intelligence, Machine Learning, Natural Language Processing, Computer Vision, Complex Reasoning, Judgment, Cognitive Science


Reference: Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, Juanzi Li, “PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament” (2025).


Leave a Reply