Wednesday 09 April 2025
A team of researchers has created a dataset that could revolutionize the way we search for legal information in Portuguese. The dataset, called JurisTCU, is made up of 16,045 judicial documents from Brazil’s Federal Court of Accounts, along with 150 queries annotated with relevance judgments.
The problem with searching for legal information is that it can be a daunting task, even for experts. Legal language is often complex and technical, making it difficult to find the right information quickly. This is where JurisTCU comes in. By providing a dataset of judicial documents and relevant queries, researchers hope to improve the efficiency and accuracy of legal information retrieval.
One of the key features of JurisTCU is its use of relevance judgments. These are annotations that indicate how well each query matches with the corresponding document. This allows researchers to evaluate the performance of different search algorithms and identify areas for improvement.
The dataset also includes three types of queries: real user keyword-based queries, synthetic keyword-based queries, and synthetic question-based queries. These queries cover a range of topics, from procurement law to healthcare policy. By analyzing how well these queries match with the relevant documents, researchers can gain insights into what makes a good search algorithm.
JurisTCU has already been used in 14 experiments using lexical search (document expansion methods) and semantic search (BERT-based and OpenAI embeddings). The results show that document expansion methods significantly improve the performance of standard BM25 search on this dataset. This is particularly true for short keyword-based queries, where improvements exceeded 45% in P@10, R@10, and nDCG@10 metrics.
The use of semantic search also produced promising results. OpenAI models performed the best, with improvements of approximately 70% in P@10, R@10, and nDCG@10 metrics for short keyword-based queries. This suggests that these dense embeddings capture semantic relationships in this domain, surpassing the reliance on lexical terms.
The implications of JurisTCU are significant. By improving the efficiency and accuracy of legal information retrieval, researchers hope to make it easier for citizens to access important documents and information. This could have a major impact on public policy, as well as the daily lives of individuals and businesses.
In the future, researchers plan to continue developing and refining JurisTCU. They hope to add more data and queries to the dataset, as well as explore new methods for evaluating search performance.
Cite this article: “Unlocking Legal Insights: A Brazilian Portuguese Information Retrieval Dataset and its Applications in Jurisprudential Search”, The Science Archive, 2025.
Portuguese, Legal Information, Search, Dataset, Juristcu, Judicial Documents, Relevance Judgments, Semantic Search, Lexical Search, Natural Language Processing







