Vision Transformers for Accurate Retinal Disease Diagnosis Using Optical Coherence Tomography Images

Wednesday 26 March 2025


As medical imaging technology continues to advance, researchers are working to improve the accuracy and efficiency of disease diagnosis using artificial intelligence (AI) and deep learning algorithms. A recent study published in a leading scientific journal highlights the potential of Vision Transformers (ViTs) for classifying optical coherence tomography (OCT) images, commonly used to diagnose retinal diseases such as diabetic macular edema (DME).


Traditional convolutional neural networks (CNNs) have been widely used for medical image analysis, but they often struggle with long-range dependencies and contextual relationships within images. ViTs, on the other hand, are designed to address these limitations by processing input sequences in parallel and using self-attention mechanisms.


The researchers compared the performance of pre-trained ViTs with scratch-trained models on a dataset of OCT images from four retinal pathologies: DME, Drusen, Choroidal Neovascularization (CNV), and Normal. The study found that both approaches achieved high accuracy, but scratch-trained models slightly outperformed pre-trained ones in larger datasets.


One of the key findings was that ViTs were able to learn robust features from OCT data even when the dataset was limited. This suggests that the self-attention mechanism can help the model focus on relevant regions within the image and ignore irrelevant background noise.


The study also showed that class-wise performance varied across different retinal pathologies. For example, DME images were well-predicted by both pre-trained and scratch-trained models, while Normal vs. Drusen classification was more challenging for all models.


These results have significant implications for the development of AI-powered diagnosis tools in ophthalmology. By leveraging ViTs and other transformer-based architectures, researchers may be able to improve the accuracy and speed of disease diagnosis using OCT imaging.


Furthermore, this study highlights the importance of domain-specific pre-training and data augmentation techniques when applying deep learning algorithms to medical images. As the field continues to evolve, it is essential to develop more effective methods for adapting these models to specific clinical applications.


Overall, this research demonstrates the potential of Vision Transformers for improving OCT image classification and highlights the need for further investigation into their application in ophthalmology and other medical imaging domains.


Cite this article: “Vision Transformers for Accurate Retinal Disease Diagnosis Using Optical Coherence Tomography Images”, The Science Archive, 2025.


Medical Imaging, Artificial Intelligence, Deep Learning, Vision Transformers, Oct Images, Retinal Diseases, Diabetic Macular Edema, Convolutional Neural Networks, Self-Attention Mechanism, Domain-Specific Pre-Training


Reference: Zihao Han, Philippe De Wilde, “OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging” (2025).


Leave a Reply