Optimal Transport Regularization for Vision-Language Prompt Learning

Wednesday 09 April 2025


A new approach to adapting vision-language models, designed to improve their ability to learn and generalize across different tasks and datasets, has been proposed by researchers. The technique, which leverages the concept of optimal transport, aims to balance the model’s performance on both novel and base classes, thereby mitigating forgetting and enhancing adaptability.


The problem of catastrophic forgetting is a common issue in deep learning, where models tend to forget previously learned information when adapting to new tasks or datasets. This can be particularly detrimental for vision-language models, which are designed to learn from large-scale datasets and generalize well across different domains.


To address this challenge, the researchers developed an optimal transport (OT) regularization framework that enforces a structural consistency between pre-trained and fine-tuned models. The approach uses OT to model cross-instance relationships and preserve the embedding structure between pre-trained and fine-tuned models. This allows the model to adapt to new tasks while retaining its knowledge of previously learned concepts.


The proposed method, dubbed PromptOT, was evaluated on several benchmark datasets, including ImageNet, SUN397, DTD, EuroSAT, and UCF101. The results showed that PromptOT outperformed state-of-the-art methods in terms of base-to-novel generalization, cross-dataset evaluation, and domain generalization.


One of the key advantages of PromptOT is its ability to balance performance across both novel and base classes. This is achieved by using a regularization term that encourages the model to maintain a consistent embedding structure between pre-trained and fine-tuned models.


The researchers also found that PromptOT was able to adapt more effectively to new tasks, even when faced with limited training data. This is because the approach allows the model to learn from both novel and base classes, rather than relying solely on the latter.


The potential applications of PromptOT are vast, ranging from visual question answering to image captioning and beyond. By enabling vision-language models to adapt more effectively to new tasks and datasets, this technique has the potential to revolutionize the field of computer vision and natural language processing.


The researchers’ approach is not without its limitations, however. For example, the optimal transport regularization term can be computationally expensive to calculate, particularly for large-scale datasets. Additionally, the method may require careful tuning of hyperparameters to achieve optimal performance.


Despite these challenges, PromptOT represents a significant step forward in the development of adaptive vision-language models.


Cite this article: “Optimal Transport Regularization for Vision-Language Prompt Learning”, The Science Archive, 2025.


Vision-Language Models, Optimal Transport, Catastrophic Forgetting, Deep Learning, Image Captioning, Visual Question Answering, Computer Vision, Natural Language Processing, Regularization, Adaptation


Reference: Xiwen Chen, Wenhui Zhu, Peijie Qiu, Hao Wang, Huayu Li, Haiyu Wu, Aristeidis Sotiras, Yalin Wang, Abolfazl Razi, “Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation” (2025).


Leave a Reply