Fact-Checking the Facts: Evaluating AIs Ability to Verify Plain Language Summaries of Scientific Research

Wednesday 09 April 2025


A new framework for evaluating the accuracy of plain language summaries has been developed, aiming to improve communication between healthcare professionals and patients. The issue of hallucination – where information is added to a summary that isn’t present in the original scientific abstract – has long plagued the field.


Researchers have created a dataset called PLAINFACT, containing 200 pairs of plain language summaries and their corresponding scientific abstracts. Each pair was annotated by human experts to identify sentences that require external information for verification, and those that can be validated solely from the abstract. This dataset will serve as the foundation for training machine learning models to evaluate the factuality of plain language summaries.


The new framework, called PLAINQAFACT, uses a combination of classification and answer extraction techniques to assess the accuracy of plain language summaries. In the first stage, a classifier is trained to determine whether a sentence or summary requires external information for verification. If it does, an answer extractor is then used to retrieve relevant information from external sources.


The team tested PLAINQAFACT on both sentence- and summary-level evaluations, comparing its performance with existing factuality evaluation metrics. The results showed that PLAINQAFACT outperformed these metrics in detecting hallucinations and assessing the accuracy of plain language summaries.


To further evaluate the effectiveness of PLAINQAFACT, researchers conducted a pilot study using the FactPICO dataset, which contains human-labeled plain language summaries of randomized controlled trials (RCTs). The team found that existing factuality evaluation metrics struggled to accurately assess the added information in these summaries, with some even decreasing in performance as more external information was removed.


However, PLAINQAFACT demonstrated improved performance when removing factual and non-factual added information from the summaries. This suggests that the framework is better equipped to handle the complexities of plain language summarization and identify instances of hallucination.


The development of PLAINQAFACT has significant implications for healthcare communication. By improving the accuracy of plain language summaries, clinicians can provide patients with more reliable and trustworthy information about their conditions and treatments. Additionally, this technology could be applied to other fields where clear communication is critical, such as finance or education.


As researchers continue to refine PLAINQAFACT, its potential applications will only continue to grow. With the ability to accurately evaluate the accuracy of plain language summaries, we can take a crucial step towards improving communication and promoting better health outcomes for all.


Cite this article: “Fact-Checking the Facts: Evaluating AIs Ability to Verify Plain Language Summaries of Scientific Research”, The Science Archive, 2025.


Healthcare, Plain Language, Summaries, Accuracy, Factuality, Communication, Machine Learning, Hallucination, Evaluation Metrics, Healthcare Professionals


Reference: Zhiwen You, Yue Guo, “PlainQAFact: Automatic Factuality Evaluation Metric for Biomedical Plain Language Summaries Generation” (2025).


Leave a Reply