Monday 24 March 2025
The deployment of large language models (LLMs) on mobile devices has opened up new possibilities for medical applications, allowing for enhanced privacy, security, and cost-efficiency by keeping sensitive health data local. However, the performance and accuracy of these on-device LLMs in real-world medical contexts remain underexplored.
A recent study aimed to address this gap by benchmarking publicly available on-device LLMs using the AMEGA dataset, evaluating their accuracy, computational efficiency, and thermal limitations across various mobile devices. The results indicate that compact general-purpose models like Phi-3 Mini achieve a strong balance between speed and accuracy, while medically fine-tuned models such as Med42 and Aloe attain the highest accuracy.
One of the key findings is that deploying LLMs on older devices remains feasible, with memory constraints posing a greater challenge than raw processing power. This highlights the need for more efficient inference and models tailored to real-world clinical reasoning.
The study’s authors used the AMEGA dataset, which comprises 20 clinical cases across 13 medical specialties. The benchmarking process involved evaluating LLMs on their ability to follow medical guidelines in realistic clinical scenarios, emphasizing diagnostic reasoning, treatment planning, and adherence to established protocols.
The results show that Med42 and Aloe, which are specifically fine-tuned for medical applications, outperform general-purpose models like Phi-3 Mini. However, the latter still demonstrate impressive accuracy, with Phi-3 Mini achieving a strong balance between speed and accuracy.
Another notable finding is that thermal limitations play a significant role in LLM performance on mobile devices. The study’s authors observed that models with higher computational requirements tend to generate more heat, which can lead to reduced performance over time.
The implications of this research are far-reaching, as it highlights the potential for on-device LLMs to transform healthcare by providing accurate and efficient diagnostic support. Furthermore, the findings underscore the need for further development of fine-tuned models specifically designed for medical applications.
In addition to improving accuracy, the study’s authors also identified areas for optimization, including more efficient inference methods and the use of specialized hardware accelerators. These advancements could enable LLMs to be deployed on a wider range of devices, expanding their potential impact in healthcare.
The future of on-device LLMs in medical applications is bright, with this study providing valuable insights into their performance and limitations.
Cite this article: “Evaluating Large Language Models for Medical Applications on Mobile Devices”, The Science Archive, 2025.
Large Language Models, Medical Applications, Mobile Devices, Accuracy, Computational Efficiency, Thermal Limitations, Fine-Tuning, Medical Guidelines, Diagnostic Reasoning, Treatment Planning







