Limitations of Multimodal Large Language Models in Fundus Reading Skills Revealed by Comprehensive Evaluation Benchmark

Saturday 05 April 2025


The field of medical imaging has long relied on computers and artificial intelligence to analyze and diagnose diseases, but a new benchmark is pushing the boundaries of what’s possible. FunBench, a comprehensive evaluation platform for multimodal large language models (MLLMs), aims to assess their ability to interpret fundus images, a critical skill for ophthalmology.


Fundus imaging involves capturing high-resolution photos of the retina and optic nerve to detect diseases such as diabetic retinopathy, age-related macular degeneration, and glaucoma. While MLLMs have shown promise in medical image analysis, they often lack fine-grained details required to discriminate between dozens of fundus diseases.


The FunBench benchmark addresses this limitation by introducing a hierarchical task organization across four levels: modality perception, anatomy perception, lesion analysis, and disease diagnosis. The platform also offers three targeted evaluation modes: linear-probe based vision encoder (VE) evaluation, knowledge-prompted language model evaluation, and holistic evaluation.


Researchers evaluated nine open-source MLLMs on the FunBench benchmark, including popular models such as HuatuoGPT- Vision and Qwen2-7B. The results were striking – none of the models performed well in basic tasks like laterality recognition or disease diagnosis. In fact, many MLLMs struggled to recognize diseases at the long tail of the distribution.


The findings suggest that current MLLMs lack fundamental skills necessary for fundus reading and highlight the need for domain-specific training and improved language and vision encoders. The results also underscore the importance of developing strong ophthalmic language models.


One promising approach is to incorporate clinical guidelines into multimodal large language models, allowing them to adapt to specific medical contexts. For example, researchers have developed a model that uses adapting multi-modal large language model for prostate cancer PI-RADS scoring, which shows great potential in diagnostic accuracy.


The development of FunBench and similar benchmarks will likely accelerate the creation of more sophisticated MLLMs capable of interpreting complex medical images. As AI continues to transform healthcare, it’s essential to push the boundaries of what’s possible and ensure that these models are equipped to handle the nuances of medical diagnosis. By doing so, we can unlock new possibilities for early disease detection and improved patient outcomes.


Cite this article: “Limitations of Multimodal Large Language Models in Fundus Reading Skills Revealed by Comprehensive Evaluation Benchmark”, The Science Archive, 2025.


Medical Imaging, Artificial Intelligence, Fundus Images, Ophthalmology, Multimodal Large Language Models, Mllms, Benchmark, Disease Diagnosis, Clinical Guidelines, Prostate Cancer.


Reference: Qijie Wei, Kaiheng Qian, Xirong Li, “FunBench: Benchmarking Fundus Reading Skills of MLLMs” (2025).


Leave a Reply