Wednesday 05 March 2025
The threat of deception detection has long been a concern for various fields, including law enforcement, psychology, and computer science. With the increasing reliance on automated systems for credibility assessment, researchers have been exploring ways to subvert these methods using adversarial attacks. Recently, a study published in the journal Nature Human Behaviour shed light on the effectiveness of these attacks in deceiving both human judges and machine learning models.
The researchers created a dataset of 505 truthful and fabricated autobiographical stories, which were then used to train a large language model to rewrite deceptive statements to make them appear truthful. The team tested the rewritten texts against human judgments and two machine learning models: a fine-tuned language model and a simple n-gram model.
When the adversarial attacks were targeted at humans, they found that the rewritten texts significantly reduced the accuracy of human credibility assessments, with judges’ ratings dropping to chance level. This was not surprising, as humans are notoriously bad at detecting deception. What was more concerning was the impact on machine learning models. The researchers discovered that when the attacks were aligned with their target – whether it was a human or machine-based assessment – both human and machine judgments dropped to the chance level.
This finding has significant implications for various fields. In law enforcement, adversarial attacks could compromise the effectiveness of automated deception detection systems, potentially leading to false convictions or wrongful acquittals. In psychology, this raises concerns about the reliability of human credibility assessments, which are often used in research and forensic settings.
The study also highlights the need for more robustness against targeted modifications in deception detection approaches. The researchers suggest that incorporating adversarial attacks into deception research could provide new insights into how humans think a model makes its decisions, as well as informing theory on human deception and deception detection.
Furthermore, the findings underscore the importance of considering the target audience when designing adversarial attacks. When attacks were not aligned with their target, both human and machine judgments improved significantly. This suggests that attackers may need to adapt their strategies depending on whether they are targeting humans or machines.
The study’s results also raise questions about the reliability of automated deception detection systems currently in use. Can these systems be compromised by targeted adversarial attacks? Should law enforcement agencies and researchers be more cautious when relying on these systems?
In the end, this research highlights the need for a more nuanced understanding of deception detection and its limitations.
Cite this article: “Deception Detection Under Threat: Adversarial Attacks Compromise Human and Machine Judgments”, The Science Archive, 2025.
Deception, Detection, Machine Learning, Language Models, Adversarial Attacks, Credibility Assessment, Human Judgments, Accuracy, Reliability, Security







