Assessing the Performance of Large Language Models in Healthcare Chatbots: A Study on Menopause Advice

Thursday 20 March 2025


The quest for accurate and reliable health advice has taken a significant step forward with the emergence of Large Language Models (LLMs) in healthcare chatbots. These AI-powered assistants have been touted as a game-changer in providing patients with timely and effective guidance on various medical topics, including menopause.


A recent study published in a prominent scientific journal examined the performance of five publicly available LLM-based chatbots in responding to common questions about menopause. The researchers employed a novel evaluation framework, dubbed S.C.O.R.E., which assessed the chatbots’ safety, consensus, objectivity, reproducibility, and explainability.


The results were both impressive and concerning. On one hand, two of the chatbots, GPT-4o and Menopause Coach, consistently delivered precise and clinically aligned information, addressing topics such as treatment options, symptom management, and lifestyle modifications. These chatbots demonstrated a clear understanding of menopause-related issues and provided detailed explanations to support their responses.


On the other hand, three other chatbots – Gemini, Meta AI, and Copilot – struggled to provide comprehensive and empathetic responses. They often lacked depth in their explanations, failed to address nuanced psychosocial aspects of menopause, and occasionally offered contradictory or incomplete information.


The study’s findings also highlighted the importance of explainability in healthcare chatbots. While all chatbots provided accurate answers, some failed to justify their responses with reliable sources or clear citations. This lack of transparency raises concerns about the credibility of the advice being offered.


Another critical aspect of the study was its examination of objectivity and bias in the chatbots’ responses. The researchers found that while most chatbots maintained a neutral tone, some exhibited subtle biases related to insurance status and race. For instance, insured patients received more detailed information on treatment options, whereas those without insurance were provided with generic advice.


The S.C.O.R.E. framework proved valuable in identifying these shortcomings and highlighting the need for further development in healthcare chatbot evaluation. The study’s authors emphasized that a standardized approach to assessing LLM performance is essential for ensuring the reliability and trustworthiness of these AI-powered assistants.


As healthcare providers continue to explore the potential of LLM-based chatbots, this research serves as a crucial reminder of the importance of accuracy, empathy, and transparency in these interactions. By refining their evaluation methods and addressing the limitations identified in this study, developers can create more effective and trustworthy tools for patients seeking health advice online.


Cite this article: “Assessing the Performance of Large Language Models in Healthcare Chatbots: A Study on Menopause Advice”, The Science Archive, 2025.


Large Language Models, Healthcare Chatbots, Menopause, Ai-Powered Assistants, Scientific Journal, Chatbot Evaluation, Explainability, Objectivity, Bias, Standardized Approach


Reference: Roshini Deva, Manvi S, Jasmine Zhou, Elizabeth Britton Chahine, Agena Davenport-Nicholson, Nadi Nina Kaonga, Selen Bozkurt, Azra Ismail, “A Mixed-Methods Evaluation of LLM-Based Chatbots for Menopause” (2025).


Leave a Reply