Malicious Memory Injection Attacks on Large Language Models: A New Threat to AI Security?

Sunday 06 April 2025


Researchers have developed a new method for hacking language models, which could potentially be used to manipulate self-driving cars, medical diagnosis systems, and other AI-powered technologies.


The team of scientists created an attack called MINJA, short for Memory Injection Neural Attack. It works by injecting malicious data into the memory of language models, causing them to produce incorrect or misleading outputs. This can happen when a user interacts with a language model, such as asking it a question or providing input.


To conduct their research, the scientists used several different types of language models, including ones designed for tasks like generating text and answering questions. They created attack queries that were specifically designed to target each type of model and test how well they could inject malicious data into its memory.


The results showed that MINJA was able to successfully inject malicious data into all of the language models tested, with a success rate of over 90% in most cases. This means that if an attacker were to use MINJA against a real-world language model, they would likely be able to manipulate it and produce incorrect or misleading outputs.


The researchers also found that some language models were more vulnerable to attack than others. For example, the ones designed for text generation were more susceptible to MINJA than those designed for question-answering. This suggests that attackers may need to tailor their attacks specifically to the type of language model they are targeting in order to be successful.


The implications of this research are significant. If an attacker were able to use MINJA against a self-driving car’s language model, it could potentially cause the car to make incorrect decisions or take unsafe actions. Similarly, if an attacker were able to use MINJA against a medical diagnosis system, it could potentially lead to misdiagnoses or delayed treatment.


The researchers are calling for more attention to be paid to the security of language models and for developers to implement additional safeguards to prevent attacks like MINJA. They also hope that their research will inspire others to explore ways to improve the security of AI-powered technologies.


One potential solution is to use techniques like encryption and secure communication protocols to protect language models from attack. Another approach could be to design language models with built-in security features, such as the ability to detect and reject malicious input.


Overall, the discovery of MINJA highlights the importance of ensuring the security of language models and other AI-powered technologies.


Cite this article: “Malicious Memory Injection Attacks on Large Language Models: A New Threat to AI Security?”, The Science Archive, 2025.


Ai-Powered Technologies, Hacking, Language Models, Minja Attack, Memory Injection, Neural Networks, Security Threats, Machine Learning, Cybersecurity, Artificial Intelligence


Reference: Shen Dong, Shaochen Xu, Pengfei He, Yige Li, Jiliang Tang, Tianming Liu, Hui Liu, Zhen Xiang, “A Practical Memory Injection Attack against LLM Agents” (2025).


Leave a Reply