Friday 21 March 2025
Detecting hallucinations in language models has become a pressing issue as these powerful tools are increasingly integrated into our daily lives. While they can generate human-like responses, they often produce false information or fictional events, which can have serious consequences if taken at face value.
Researchers have long known that hallucinations are linked to model uncertainty, but most methods focus on capturing aleatoric uncertainty – the randomness inherent in the data. However, a new approach has emerged, injecting noise into hidden unit activations during sampling to capture epistemic uncertainty – the uncertainty stemming from the model’s own limitations and biases.
This innovative method has been shown to significantly improve hallucination detection, particularly when combined with existing measures of aleatoric uncertainty. By perturbing intermediate representations, the model is forced to consider alternative explanations for its outputs, making it more likely to reject incorrect responses.
One of the key advantages of this approach is its simplicity. Unlike Bayesian methods that require complex computations and retraining, noise injection can be applied in a training-free paradigm, allowing researchers to quickly test and evaluate their models.
The technique has been tested on a range of large language models, including Llama-2-7B-chat and Gemma-2-it. Results show that noise injection improves hallucination detection across multiple datasets, with some models achieving AUROC scores as high as 79%.
But how does it work? In essence, the approach is based on the idea that a model’s uncertainty about its own predictions can be leveraged to detect hallucinations. By introducing random noise into the model’s internal workings, researchers create an environment where the model must consider multiple plausible explanations for its outputs.
This has several benefits. Firstly, it allows the model to generate more diverse and creative responses, which can be useful in certain applications such as language translation or text summarization. Secondly, it forces the model to think more critically about its own predictions, making it less likely to produce hallucinations.
The implications of this work are far-reaching. As language models become increasingly prevalent in our daily lives, the need for accurate and reliable information is crucial. By detecting hallucinations and rejecting incorrect responses, these models can provide users with more trustworthy and informed outputs.
In addition, this approach could have significant benefits for tasks such as question answering or text classification, where accuracy is paramount. By injecting noise into the model’s internal workings, researchers may be able to improve overall performance and reduce errors.
Cite this article: “Detecting Hallucinations in Language Models: A Novel Approach Using Noise Injection”, The Science Archive, 2025.
Language Models, Hallucinations, Uncertainty, Noise Injection, Epistemic Uncertainty, Aleatoric Uncertainty, Model Limitations, Biases, Training-Free Paradigm, Auroc Scores.







