Saturday 22 March 2025
Detecting inconsistencies in eyewitness testimony is a crucial task for investigators and researchers alike. Inconsistencies can arise from various factors, including differences in perception, memory lapses, or even intentional deception. However, traditional approaches to detecting these inconsistencies often fail to capture the nuanced nature of human language.
A new approach, introduced by researchers, tackles this challenge head-on by developing a framework that identifies contextually related incongruences between eyewitness testimonies. The framework, dubbed INTEND, uses a combination of techniques to detect contradictions and extract conflicting spans from the testimonies.
The researchers created a comprehensive dataset called MIND (MultI-eyewitNess Deception), consisting of 2,979 pairs of contextually related answers designed to capture both explicit and implicit contradictions. The dataset was used to train two models, one based on large language models like LLaMA-3 (8B) and another using GPT-4o mini.
The framework’s performance was evaluated using three metrics: precision, recall, and F1-score. The results showed that INTEND significantly outperformed the baseline models, achieving a margin of +5.63% in detection accuracy. Furthermore, the framework demonstrated improved performance when compared to fine-tuning and prompt-tuning techniques on MLMs (Masked Language Models) and LLMs (Large Language Models).
The researchers also analyzed the performance of INTEND using different numbers of reasoning hops, finding that a three-hop configuration achieved the best results. This setup allowed the model to focus on each subtask independently, including identifying fine-grained key details, inferring incongruity reasons, and extracting conflicting spans.
In addition to its improved accuracy, INTEND’s framework also showed promising results when evaluated by human evaluators. The evaluators assessed the contradictions based on three criteria: contradiction clarity, logical exclusivity, and context relevance. The results indicated that INTEND accurately captured the contradictory spans in most cases, even when the testimonies contained subtle differences.
While the results are impressive, there is still room for improvement. The framework struggled with detecting inconsistencies in logically consistent narratives, a common issue in many real-world scenarios. Additionally, the models’ performance varied depending on the specific context and the number of reasoning hops used.
Despite these limitations, INTEND’s framework represents a significant step forward in developing more accurate methods for detecting inconsistencies in eyewitness testimony.
Cite this article: “INTEND: A Framework for Identifying Contextually Related Inconsistencies in Eyewitness Testimony”, The Science Archive, 2025.
Eyewitness Testimony, Inconsistency Detection, Intend Framework, Mind Dataset, Language Models, Llama, Gpt, Precision, Recall, F1-Score, Contradiction Detection







