Fact-Checking in the Age of AI: A Systematic Examination of Large Language Models Capacity to Detect Veracity

Wednesday 09 April 2025


Fact-checking has become a crucial part of our digital lives, as we increasingly rely on online sources for news and information. But how reliable are these fact-checkers? A recent study set out to investigate the capabilities of large language models (LLMs) in identifying true or false claims.


The researchers used five different LLMs – ChatGPT-4, Llama 3 (70B), Llama 3.1 (405B), Claude 3.5 Sonnet, and Google Gemini – to evaluate their ability to classify statements as true, false, or mixed. They drew on a dataset of over 16,000 fact-checked claims from the ClaimsKG knowledge graph, covering topics such as politics, health, and social issues.


The results showed that while LLMs can be effective in identifying false claims, they are less accurate when it comes to verifying true statements. In fact, some models performed worse than others, with ChatGPT-4 struggling to accurately identify true information. The researchers suggest that this may be due to the model’s training data, which is biased towards labeling more false statements.


Another interesting finding was that LLMs were better at identifying false claims on certain topics, such as COVID-19 and American political controversies. This could be because these topics have a higher proportion of false information, making it easier for the models to detect. However, this also raises concerns about the potential biases in these models, which could perpetuate existing social and political divides.


The study highlights the importance of understanding the limitations of AI-powered fact-checking tools. While they can be useful in identifying false claims, they are not a panacea for the misinformation problem. In fact, the researchers argue that LLMs may even exacerbate the issue by spreading false information or reinforcing existing biases.


To improve the accuracy and reliability of LLM-based fact-checking, the study suggests several strategies. One approach is to fine-tune the models on specific datasets or topics, allowing them to better adapt to different contexts. Another strategy is to develop more diverse training data, which would help reduce biases and improve performance across a range of topics.


The findings also underscore the need for human oversight and intervention in the fact-checking process. While AI can be useful in identifying potential false claims, it is ultimately up to humans to verify and validate information.


Cite this article: “Fact-Checking in the Age of AI: A Systematic Examination of Large Language Models Capacity to Detect Veracity”, The Science Archive, 2025.


Here Are The Keywords: Large Language Models, Fact-Checking, Artificial Intelligence, Misinformation, Bias, Accuracy, Reliability, Training Data, Fine-Tuning, Human Oversight


Reference: Elizaveta Kuznetsova, Ilaria Vitulano, Mykola Makhortykh, Martha Stolze, Tomas Nagy, Victoria Vziatysheva, “Fact-checking with Generative AI: A Systematic Cross-Topic Examination of LLMs Capacity to Detect Veracity of Political Information” (2025).


Leave a Reply