Monday 31 March 2025
Researchers have long sought to improve the performance of large language models (LLMs) in legal domains, where nuanced understanding and precise application of complex regulations are crucial. A new benchmark, LexRAG, aims to bridge this gap by providing a comprehensive platform for evaluating LLMs in multi-turn legal consultations.
LexRAG is built upon a dataset of 1,013 conversations between lawyers and clients, covering various aspects of Chinese law. These dialogues have been annotated with relevant legal articles, allowing researchers to assess the ability of LLMs to retrieve and generate accurate responses.
The benchmark includes two primary tasks: conversational knowledge retrieval and response generation. In the first task, LLMs are tasked with identifying relevant legal articles given a multi-turn conversation context. This requires not only understanding the client’s query but also recalling specific provisions from a vast body of law.
The second task involves generating responses to client queries based on the retrieved legal information. This demands not only linguistic proficiency but also a deep comprehension of legal concepts and their applications.
To evaluate LLMs in these tasks, LexRAG provides a set of metrics, including recall, precision, and nDCG (normalized discounted cumulative gain). These metrics allow researchers to assess the performance of different models under various settings, such as zero-shot learning or retrieval-augmented generation.
LexRAG’s creators have also developed an open-source toolkit, LexiT, which provides a comprehensive implementation of LLM components tailored for legal domains. This includes tools for processing and analyzing dialogues, as well as metrics for evaluating model performance.
The development of LexRAG represents a significant step forward in the evaluation of LLMs for legal applications. By providing a standardized platform for testing and comparing models, researchers can better understand the strengths and limitations of existing approaches and develop more effective solutions for supporting lawyers and clients.
In addition to improving the performance of LLMs, LexRAG has the potential to advance the development of intelligent judicial technologies. As AI-powered systems become increasingly prevalent in legal practice, it is essential that they be designed with both accuracy and nuance in mind. By evaluating LLMs in a more comprehensive and realistic manner, researchers can create systems that better serve the needs of lawyers, clients, and the legal system as a whole.
The availability of LexRAG and LexiT marks an important milestone in the pursuit of more effective LLM-based solutions for legal domains.
Cite this article: “LexRAG: A Comprehensive Benchmark for Evaluating Large Language Models in Legal Domains”, The Science Archive, 2025.
Language Models, Large Language Models, Legal Domains, Benchmark, Lexrag, Conversational Knowledge Retrieval, Response Generation, Open-Source Toolkit, Lexit, Intelligent Judicial Technologies.







