Wednesday 12 March 2025
Researchers have developed a new approach to medical image classification, one that relies on large vision-language models (LVLMs) to diagnose diseases more accurately and efficiently. This method, called CBVLM, has shown promising results in multiple medical datasets, outperforming traditional techniques without the need for extensive training or annotation.
The challenge facing medical imaging analysis is the availability of labeled data. Deep learning models require large amounts of annotated images to learn patterns and make predictions. However, collecting and labeling these images can be a time-consuming and costly process. CBVLM addresses this issue by leveraging LVLMs, which have been trained on vast amounts of text data and can generate explanations for their predictions.
In CBVLM, the LVLM is first prompted to identify specific concepts within an image, such as the presence or absence of certain features. This concept detection step allows the model to focus on relevant information and ignore irrelevant details. The LVLM then uses this information to make a diagnosis based on the detected concepts.
The benefits of CBVLM are twofold. First, it reduces the need for extensive annotation, as the LVLM can generate explanations for its predictions. Second, it enables more accurate diagnoses by focusing on specific concepts within an image, rather than trying to learn patterns from the entire image.
To test the effectiveness of CBVLM, researchers evaluated the method on four medical datasets: Derm7pt, SkinCon, CORDA, and DDR. Each dataset contained images with corresponding labels for various skin conditions or diseases. The results showed that CBVLM outperformed traditional methods without the need for extensive training or annotation.
The Derm7pt dataset, for example, contains images of skin lesions with corresponding diagnoses. By prompting the LVLM to identify specific concepts within an image, such as the presence of a blue- whitish veil or streaks, the model was able to accurately diagnose melanoma and other skin conditions.
Similarly, in the SkinCon dataset, which contains images of various skin conditions, CBVLM was able to correctly identify papules, plaques, and other lesions. In CORDA, a dataset containing chest X-rays with corresponding diagnoses, CBVLM accurately identified pneumonia and other respiratory diseases. Finally, in DDR, a dataset containing retinal images with corresponding diagnoses, CBVLM detected diabetic retinopathy and other eye conditions.
The implications of this research are significant.
Cite this article: “Advances in Medical Image Classification: Leveraging Large Vision-Language Models for Accurate Diagnoses”, The Science Archive, 2025.
Medical Image Classification, Deep Learning Models, Large Vision-Language Models, Lvlms, Cbvlm, Medical Imaging Analysis, Annotation, Skin Conditions, Chest X-Rays, Retinal Images, Diabetic Retinopathy







