Wednesday 26 March 2025
The quest for transparency in clinical trials has been an ongoing challenge, with researchers and clinicians struggling to navigate the complexities of data sharing statements (DSS). These statements are supposed to provide a clear picture of how individual participant data (IPD) will be shared after a trial is completed. However, the reality is often far from transparent.
A recent study published in a medical journal has shed some light on this issue by evaluating the effectiveness of pre-trained language models in categorizing DSS. The researchers used three different BERT-based language models – SciBERT, BioBERT, and BlueBERT – to analyze a dataset of over 30,000 clinical trial records from ClinicalTrials.gov.
The results were striking: while the original categorical labels provided by ClinicalTrials.gov were often inconsistent with the actual text of the DSS, the pre-trained language models were able to accurately categorize IPD sharing intentions with an average accuracy rate of around 83%. This suggests that there is valuable information hidden in the textual descriptions of DSS that can be leveraged to improve data discovery and reuse.
One of the key limitations of the study was its reliance on a dataset that could be exported from ClinicalTrials.gov. In reality, many DSS include links to additional web-based platforms or direct users to contact the respective research teams to ask for IPD access. This makes it difficult to assess whether IPD is truly shared, as real data sharing may fall short of declarations.
Despite this limitation, the study’s findings have significant implications for the medical community. The ability to automatically categorize DSS could support researchers in efficiently identifying datasets that align with their study goals. This, in turn, could accelerate the pace of scientific discovery and improve patient care.
The study also highlights the potential benefits of using pre-trained language models in other areas of biomedical research. For example, these models could be used to analyze large amounts of medical literature, identify trends and patterns, and even aid in the development of personalized treatment plans.
As researchers continue to grapple with the complexities of DSS, it is clear that there is no one-size-fits-all solution. However, by leveraging the power of machine learning and natural language processing, they may be able to unlock new insights and improve data sharing practices.
Cite this article: “Machine Learning Models Shine Light on Clinical Trial Data Sharing Statements”, The Science Archive, 2025.
Clinical Trials, Transparency, Data Sharing Statements, Individual Participant Data, Language Models, Bert, Scibert, Biobert, Bluebert, Natural Language Processing, Machine Learning







