Monday 10 March 2025
A major challenge facing researchers is the lack of discoverability, accessibility and reusability of open research software. This software, which is used in countless scientific studies, often remains hidden within the manuscript of research papers, making it difficult for others to build upon or reproduce the findings.
To address this issue, a new project called SoFAIR has been launched, aimed at making software more Findable, Accessible, Interoperable and Reusable (FAIR). The project involves developing a machine-learned workflow that can identify research software assets from within research manuscripts, validate them with authors and register them with persistent identifiers.
The solution is being built on top of existing open-source tools and scholarly infrastructures. CORE, a comprehensive bibliographic database, will deploy the extended GROBID software to extract software mentions from research papers. Once identified, these mentions will be enriched with additional descriptive metadata by processing documentation from code repositories.
The next step is to disambiguate and validate the discovered mentions through author validation. This involves routing newly identified but not yet validated software assets via a free-to-use repository dashboard service to institutional repository managers, who can then approve their automated routing for validation by authors located at their institution.
Once validated, the repository will issue an asset registration request to Software Heritage, which permanently archives the new software asset and issues a permanent identifier. This will enable repositories to expose information linking software assets with research outputs that mention them within their OAI-PMH feed to aggregators and do so in an interoperable fashion.
To test the efficacy of the solution, SoFAIR is conducting two disciplinary use cases: one in life sciences, conducted in cooperation with Europe PMC, and another in digital humanities, conducted in cooperation with DARIAH. An additional multidisciplinary use case will be conducted in cooperation with HAL repository.
The project’s key innovations lie in its ability to address the whole software assets management lifecycle, from identification to registration and archival. It also applies machine learning models for software mentions extraction and disambiguation, which have been shown to outperform traditional approaches.
The scalability of SoFAIR lies in its integration within open scholarly infrastructures, providing a repository-centric solution that can be applied consistently to both pre-existing and newly deposited content from anywhere within the global open repositories network. This approach not only ensures fast and wide adoption but also provides tangible pathways to impact.
Cite this article: “FAIRification of Open Research Software”, The Science Archive, 2025.
Open Research Software, Fair, Sofair, Machine Learning, Software Asset Management, Discovery, Accessibility, Reusability, Bibliographic Database, Persistent Identifiers, Scholarly Infrastructures







