Limitations of Vision-Language Models Revealed: A Study on Image Transformations

Thursday 10 April 2025


For a long time, artificial intelligence (AI) has been getting better at understanding and generating human language. But despite this progress, AI systems have struggled to understand simple image transformations – like rotating an image or adjusting its brightness.


A recent study set out to investigate why this is the case. The researchers used two popular AI models, called CLIP and SigLIP, and tested their ability to recognize different types of image modifications.


The results were striking. Both models performed poorly when it came to recognizing simple transformations like rotating an image by 90 degrees or adjusting its brightness. In fact, the study found that these models are not able to identify a single correct example for most of the augmentations they were tested on.


This is surprising because AI models are designed to be robust and invariant to standard image transforms. But it seems that this invariance comes at the cost of explicit understanding of spatial relationships between objects in an image.


The researchers also found that these models struggle when it comes to identifying specific transformations, such as rotating an image or adjusting its brightness. They were able to identify some of the transformations correctly, but only by chance.


This study has important implications for the use of AI in applications like image editing and video processing. Currently, AI systems are not able to understand simple image transformations, which limits their ability to perform tasks that require these types of modifications.


The researchers suggest that future AI models should be designed with explicit understanding of spatial relationships between objects in an image. This could involve training AI models on datasets that include a wide range of images with different transformations applied.


Another approach would be to develop new techniques for training AI models that can learn to recognize and generate complex visual patterns, such as those found in natural scenes.


The study also highlights the importance of understanding how humans perceive and process visual information. By studying human perception and cognition, researchers may be able to develop more effective AI systems that are better able to understand and manipulate visual data.


Overall, this study sheds light on a important limitation of current AI models and suggests new directions for future research in this area.


Cite this article: “Limitations of Vision-Language Models Revealed: A Study on Image Transformations”, The Science Archive, 2025.


Artificial Intelligence, Image Recognition, Image Transformations, Spatial Relationships, Object Recognition, Image Editing, Video Processing, Computer Vision, Machine Learning, Visual Data.


Reference: Ahmad Mustafa Anis, Hasnain Ali, Saquib Sarfraz, “On the Limitations of Vision-Language Models in Understanding Image Transforms” (2025).


Leave a Reply