Thursday 10 April 2025
Artificial intelligence has long struggled to accurately depict action-based relationships in images, such as a mouse chasing a cat or a horse riding an astronaut. These scenarios are inherently complex and require a deep understanding of spatial orientation, pose, facial expression, and interaction between entities.
Researchers have made significant progress in addressing this challenge by developing text-to-image diffusion models that can generate photorealistic images from textual prompts. However, these models still fall short when it comes to generating counter-stereotypical action-based relationships, such as a mouse chasing a cat instead of the more common cat chasing a mouse.
A team of researchers has identified the root cause of this limitation and developed a novel approach to overcome it. They discovered that current text-to-image diffusion models have a strong bias towards frequent compositions, even when prompted to generate rare ones. This bias stems from the way these models are trained on large datasets that disproportionately feature common scenarios.
To address this issue, the researchers designed a new framework called Role-Bridging Decomposition, which leverages intermediate compositions to gradually teach rare relationships without requiring architectural modifications. They also created ActionBench, a comprehensive benchmark specifically designed to evaluate action-based relationship generation across both stereotypical and counter-stereotypical configurations.
The results are striking. When tested on the ActionBench dataset, the researchers’ approach significantly outperformed state-of-the-art models in generating accurate and diverse images that accurately depict rare action-based relationships. The model’s ability to generate intermediate compositions also enabled it to produce more realistic and nuanced interactions between entities.
One of the key insights from this study is that the quality of generated images is directly tied to the quality of the textual prompts used to train the models. By crafting more precise and relation-focused descriptions, researchers can significantly improve the accuracy and diversity of generated images.
The implications of this work are far-reaching, with potential applications in areas such as visual storytelling, creative content generation, and even robotics and computer vision. As AI continues to evolve, it is essential that we develop models that can accurately depict complex action-based relationships, enabling us to better understand and interact with the world around us.
The researchers’ approach provides a significant step forward in this direction, demonstrating the potential for text-to-image diffusion models to generate accurate and diverse images of action-based relationships. As the field continues to evolve, it will be exciting to see how these advances are applied in various domains, enabling new possibilities for creative expression and innovation.
Cite this article: “Uncovering Biases in Text-to-Image Generation: A Study on Counter-Stereotypical Action Relations”, The Science Archive, 2025.
Artificial Intelligence, Text-To-Image Diffusion Models, Action-Based Relationships, Image Generation, Counter-Stereotypical Scenarios, Bias, Role-Bridging Decomposition, Actionbench, Visual Storytelling, Creative Content Generation.







