Improving Safety in Vision Language Models with Activation Shift Disentanglement and Calibration

Thursday 27 March 2025


Artificial Intelligence has made tremendous progress in recent years, and one of its most significant applications is in Vision Language Models (VLMs). These models have enabled computers to understand and process visual data like images and videos, allowing them to perform tasks that were previously the exclusive domain of humans.


However, as VLMs have become more sophisticated, they’ve also become more vulnerable to malicious attacks. Cybercriminals are finding ways to exploit these models by tricking them into producing harmful or offensive responses. This is a major concern, especially in applications where safety and ethics are critical, such as healthcare and finance.


Researchers have been working to address this issue by developing techniques to improve the safety of VLMs. One approach is to use defensive prompting techniques that guide the model to focus on specific aspects of the input data. Another method is to train the model on datasets specifically designed to promote safe and ethical behavior.


But a new study has taken a different approach, targeting the root cause of the problem: the way VLMs perceive safety itself. The researchers have discovered that these models tend to overestimate the safety of harmful inputs when they’re presented in a visual format. This is because the visual modality can introduce biases and distortions that affect the model’s understanding of safety.


To address this issue, the researchers have developed a technique called Activation Shift Disentanglement and Calibration (ShiftDC). This method involves decomposing the activation patterns in the VLM to identify the specific components that are responsible for the distortion. The distorted components are then removed or calibrated to restore the model’s original alignment with safety.


The results of this study are promising, showing significant improvements in the performance of VLMs on safety benchmarks. The models were able to better distinguish between safe and harmful inputs, reducing the attack success rate by a substantial margin.


But what does this mean for us? In practical terms, ShiftDC could be used to improve the safety of AI-powered applications in various industries. For example, in healthcare, VLMs could be trained on patient data to generate accurate diagnoses while avoiding harmful or offensive responses. In finance, the models could be used to analyze financial transactions and detect potential fraud without inadvertently promoting illegal activities.


The development of ShiftDC is a significant step forward in addressing the safety concerns surrounding VLMs.


Cite this article: “Improving Safety in Vision Language Models with Activation Shift Disentanglement and Calibration”, The Science Archive, 2025.


Vision Language Models, Artificial Intelligence, Safety, Ethics, Cybersecurity, Defensive Prompting, Shiftdc, Activation Patterns, Distortion, Calibration


Reference: Xiaohan Zou, Jian Kang, George Kesidis, Lu Lin, “Understanding and Rectifying Safety Perception Distortion in VLMs” (2025).


Leave a Reply