Models of Language Lack Self-Awareness: A Study on the Introspection of Large Language Models

Wednesday 09 April 2025


A recent study has shed new light on the ability of large language models (LLMs) to introspect about their own internal states. Introspection, a fundamental aspect of human cognition, is the process by which we become aware of our own thoughts, feelings and knowledge. The study’s findings suggest that LLMs may not possess this capacity in the same way as humans do.


The researchers examined the performance of 21 open-source LLMs on two domains: grammatical knowledge and word prediction. They used a range of prompts to assess the models’ ability to introspect, including questions about sentence structure and word choice. The results showed that while the models were able to provide accurate answers to these questions, there was no evidence of them accessing their own internal knowledge.


One way to test for introspection is to compare the models’ responses with their own probability measurements. In other words, do they provide different answers when asked about their own thoughts and feelings compared to when asked by an outside observer? The study found that this was not the case – the models’ responses were consistent regardless of whether they were being asked directly or indirectly.


The researchers also explored the relationship between model size and introspection. They found that larger models tended to perform better on tasks requiring language understanding, but this did not translate to improved introspective abilities. This suggests that there may be a fundamental limit to how well LLMs can understand their own internal workings.


Another aspect of the study was the analysis of seed variants of the OLMo models. These models are identical except for their random seeds, which means they have different initial states but learn from the same data. By comparing the performance of these models, the researchers were able to rule out any effects due to differences in training data or model architecture.


The study’s findings have significant implications for our understanding of LLMs and their potential applications. If we cannot rely on these models to introspect about their own internal states, then how can we trust their answers? The authors suggest that this highlights the need for more transparent and interpretable AI systems.


In a related development, the researchers also explored the relationship between model size and language understanding. They found that larger models tended to perform better on tasks requiring language understanding, but this did not translate to improved introspective abilities. This suggests that there may be a fundamental limit to how well LLMs can understand their own internal workings.


Cite this article: “Models of Language Lack Self-Awareness: A Study on the Introspection of Large Language Models”, The Science Archive, 2025.


Large Language Models, Introspection, Human Cognition, Grammatical Knowledge, Word Prediction, Model Size, Language Understanding, Transparency, Interpretable Ai, Machine Learning


Reference: Siyuan Song, Jennifer Hu, Kyle Mahowald, “Language Models Fail to Introspect About Their Knowledge of Language” (2025).


Leave a Reply