Revolutionizing Outpatient Referral with Large Language Models

Wednesday 09 April 2025


The quest for a more effective way to manage outpatient referrals has long been a challenge in healthcare systems worldwide. In recent years, researchers have been exploring the potential of large language models (LLMs) to aid in this process. A new study published today sheds light on the performance of various LLMs in handling dynamic scenarios, where patients’ conditions and symptoms are constantly evolving.


The research team evaluated several popular LLMs, including Gemma, Mixtral-8x22B-Instruct, Llama-3.2-90B-Instruct, GPT-4o, and Chinese-centric models such as Qwen2.5-0.5B-Instruct and Qwen2.5-7B-Instruct. The models were tasked with recommending the most suitable department for patients based on their symptoms, medical history, and other relevant information.


The study found that while LLMs showed promise in static evaluation scenarios, where patient data was provided in advance, they struggled to adapt to dynamic situations. In these cases, human doctors outperformed the models, often requiring fewer questions and arriving at accurate diagnoses more quickly.


One of the key challenges faced by LLMs was their inability to effectively engage with patients and gather relevant information. This was particularly evident in scenarios where patients’ symptoms were vague or incomplete, requiring the model to make educated guesses about their condition.


The researchers also evaluated the performance of human participants in both static and dynamic evaluation tasks. In the former, doctors achieved an accuracy rate of 85%, while laypersons scored significantly lower at around 65%. In the dynamic task, human doctors outperformed LLMs, with some achieving an accuracy rate of over 90%.


The study’s findings have significant implications for the development of AI-powered outpatient referral systems. While LLMs may not yet be ready to replace human doctors in complex decision-making scenarios, they could potentially serve as useful tools for triaging patients and identifying the most appropriate departments.


To improve the performance of LLMs in dynamic situations, the researchers suggest incorporating more advanced natural language processing techniques and training data that better reflects real-world clinical scenarios. They also emphasize the importance of human oversight and feedback in fine-tuning these models.


Ultimately, this study highlights the need for a balanced approach to AI adoption in healthcare, one that combines the strengths of both human doctors and machine learning algorithms. By working together, we can develop more effective systems that improve patient outcomes and streamline the referral process.


Cite this article: “Revolutionizing Outpatient Referral with Large Language Models”, The Science Archive, 2025.


Large Language Models, Outpatient Referrals, Healthcare, Artificial Intelligence, Machine Learning, Natural Language Processing, Clinical Scenarios, Patient Outcomes, Referral Process, Medical Diagnosis.


Reference: Xiaoxiao Liu, Qingying Xiao, Junying Chen, Xiangyi Feng, Xiangbo Wu, Bairui Zhang, Xiang Wan, Jian Chang, Guangjun Yu, Yan Hu, et al., “Large Language Models for Outpatient Referral: Problem Definition, Benchmarking and Challenges” (2025).


Leave a Reply