Egida: A Dataset for Assessing Language Model Safety

Thursday 27 March 2025


The quest for safe and reliable language models has taken a significant step forward with the development of Egida, a dataset designed to test the safety alignment of large language models. By analyzing the performance of various models on this dataset, researchers have gained valuable insights into how these powerful tools can be improved.


The main goal of Egida is to assess the ability of language models to generate safe responses when faced with harmful or offensive prompts. This is a crucial challenge in today’s digital age, where AI-powered chatbots and virtual assistants are increasingly being used to interact with humans. If these models are not properly trained to recognize and respond to unsafe input, they can potentially spread misinformation, promote harmful behavior, or even facilitate illegal activities.


To create Egida, researchers compiled a dataset of over 27,000 prompts, each designed to test the model’s ability to generate safe responses in various scenarios. These prompts were then categorized into six different safety topics, including violence, discrimination, and harm reduction. The dataset also included a range of attack styles, such as sarcasm, irony, and ambiguity, to further challenge the models.


The Egida dataset was used to train four state-of-the-art language models: Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.1-70B-Instruct, Qwen2.5-7B-Instruct, and Qwen2.5-72B-Instruct. Each model was evaluated on its ability to generate safe responses when presented with a range of prompts from the Egida dataset.


The results were impressive, with all four models showing significant improvements in safety alignment after being trained on the Egida dataset. However, the researchers also discovered that larger models did not necessarily perform better than smaller ones. In fact, the Meta-Llama-3.1-70B-Instruct model, which has a modest size of 70 billion parameters, outperformed its larger counterparts in some cases.


The study also explored the effect of adding safe data to the training process. Surprisingly, the results showed that incorporating safe data actually degraded the performance of the models, making them less effective at generating safe responses. This finding highlights the importance of carefully designing and evaluating language model training datasets to ensure that they promote safe and responsible behavior.


The Egida dataset is an important step forward in the development of safe and reliable language models.


Cite this article: “Egida: A Dataset for Assessing Language Model Safety”, The Science Archive, 2025.


Language Models, Safety Alignment, Egida Dataset, Large Language Models, Ai-Powered Chatbots, Virtual Assistants, Unsafe Input, Misinformation, Harmful Behavior, Illegal Activities, Meta-Llama-3.1-8B-Instruct, Meta-


Reference: Dario Garcia-Gasulla, Adrian Tormos, Anna Arias-Duart, Daniel Hinjos, Oscar Molina-Sedano, Ashwin Kumar Gururajan, Maria Eugenia Cardello, “Efficient Safety Retrofitting Against Jailbreaking for LLMs” (2025).


Leave a Reply