Thursday 06 March 2025
Scientists have been working tirelessly to unravel the mysteries of the global research infrastructure, a complex network of databases and repositories that store information about scientific publications, datasets, and funding sources. The latest breakthrough comes from the INFORMATE Project, which has made significant strides in understanding how this infrastructure impacts our ability to identify and access publicly funded research.
At its core, the problem is one of metadata – the data that describes other data. In the case of scientific publications, metadata includes information such as author names, funding sources, and keywords. But when it comes to tracking down specific research papers or datasets, this metadata can be incomplete, inconsistent, or simply missing.
The INFORMATE Project set out to tackle this issue by comparing three major sources of metadata: the National Science Foundation’s (NSF) Public Access Repository (PAR), CHORUS, a database that aggregates data from various repositories, and Crossref, a global network of publishers. By analyzing over 400,000 awards and corresponding publications, researchers were able to identify significant disparities in the completeness and accuracy of metadata across these sources.
One of the most striking findings was the sheer volume of missing metadata. In some cases, as many as 60% of award-related publications were not included in PAR, while CHORUS and Crossref had their own sets of incomplete data. This highlights the need for more standardized and automated processes for capturing and sharing metadata.
The study also revealed that the timing of metadata availability plays a crucial role in its completeness. For example, awards with effective dates between 2014 and 2016 were less likely to have corresponding publications included in PAR or CHORUS. In contrast, awards from 2017 onwards showed significant increases in metadata coverage across all three sources.
These findings have important implications for researchers, policymakers, and funders alike. By better understanding the strengths and weaknesses of these metadata sources, scientists can develop more effective strategies for tracking down specific research papers and datasets, ultimately advancing scientific discovery and knowledge sharing.
The INFORMATE Project’s work also underscores the need for increased transparency and collaboration across the global research infrastructure. By leveraging the insights gained from this study, researchers and policymakers can work together to design more efficient and effective systems for capturing and sharing metadata, ultimately benefiting the entire scientific community.
Cite this article: “Unlocking the Secrets of Scientific Metadata: INFORMATE Project Breakthrough”, The Science Archive, 2025.
Metadata, Research Infrastructure, Funding Sources, Scientific Publications, Datasets, National Science Foundation, Public Access Repository, Chorus, Crossref, Data Sharing, Transparency







