Simulating Multi-Turn Jailbreak Attacks on Large Language Models

Friday 14 March 2025


Recent advancements in large language models have led to significant concerns about their potential misuse for malicious purposes, such as generating harmful content or manipulating people’s opinions. To address these concerns, a team of researchers has developed a novel framework that simulates multi-turn jailbreak attacks on these models.


The concept of jailbreaking refers to the ability of an attacker to manipulate a language model into producing a response that is not aligned with its intended purpose or safety guidelines. This can be done by crafting a series of carefully designed queries that exploit vulnerabilities in the model’s training data or algorithmic biases.


The researchers’ framework, called Siren, consists of three stages: training set construction, post-training fine-tuning, and interactions between the attacking and target language models. The first stage involves constructing a dataset of turn-level feedback from the attacking model to help it learn how to generate effective queries. In the second stage, the attacking model is fine-tuned using supervised learning with direct preference optimization.


The third stage is where Siren really shines. Here, the attacking model interacts with the target language model across multiple turns, generating a series of queries that are designed to evade the target’s safety mechanisms and manipulate its responses. The researchers demonstrate that Siren can achieve an attack success rate of up to 90% against certain language models.


One of the key innovations behind Siren is its ability to decompose complex queries into more manageable sub-questions, allowing the attacking model to craft a series of carefully designed prompts that are tailored to the specific vulnerabilities of the target model. This approach allows Siren to simulate real-world human jailbreak behaviors, where attackers may use subtle and nuanced language to manipulate the model’s responses.


The researchers’ results demonstrate that Siren is capable of outperforming existing single-turn attack methods, which rely on static patterns or predefined logical chains to generate adversarial prompts. Instead, Siren uses a learning-based approach that adapts to the target model’s responses in real-time, allowing it to evolve and improve its attacks over multiple turns.


The implications of this research are significant. While language models have many potential benefits, such as improving customer service or generating creative content, they also pose risks if not properly secured against malicious attacks. The development of Siren highlights the need for researchers and developers to prioritize safety and security when designing these models, and to develop robust defenses against advanced multi-turn jailbreak attacks.


Cite this article: “Simulating Multi-Turn Jailbreak Attacks on Large Language Models”, The Science Archive, 2025.


Language Models, Jailbreaking, Large Language Models, Malicious Purposes, Harmful Content, Opinion Manipulation, Siren Framework, Multi-Turn Attacks, Attack Success Rate, Fine-Tuning.


Reference: Yi Zhao, Youzhi Zhang, “Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors” (2025).


Leave a Reply