{"id":55995,"date":"2026-05-30T07:24:33","date_gmt":"2026-05-30T07:24:33","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/55995\/"},"modified":"2026-05-30T07:24:33","modified_gmt":"2026-05-30T07:24:33","slug":"ai-chatbots-do-not-consistently-deliver-accurate-health-responses","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/55995\/","title":{"rendered":"AI Chatbots Do Not Consistently Deliver Accurate Health Responses"},"content":{"rendered":"<p>            <img loading=\"lazy\" decoding=\"async\" width=\"696\" height=\"392\" class=\"entry-thumb\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/Jun1_2025_GettyImages_1494104649_AIchatbot-696x392.jpg\"   alt=\"AI chatbot - Artificial Intelligence digital concept\" title=\"1494104649\"\/>Credit: ertigo3d \/ Getty Images<\/p>\n<p>A study into the responses of AI chatbots to everyday medical questions found that nearly 76% of the responses were accurate, according to investigators at Penn State University. Their research, <a href=\"https:\/\/arxiv.org\/pdf\/2506.13805\" target=\"_blank\" rel=\"noopener nofollow\">published as a preprint in arXiv<\/a>, found that while the AI chatbots could deliver useful information, the error rates remain high enough that they shouldn\u2019t replace physicians for either diagnosing or suggesting treatments.<\/p>\n<p>\u201cOur work focuses explicitly on healthcare scenarios that the average internet user might ask AI, which is a perspective that prior research into <a href=\"https:\/\/www.insideprecisionmedicine.com\/?s=large%20language%20models&amp;filter=&amp;page=null\" target=\"_blank\" rel=\"noopener nofollow\">large language models<\/a> (LLMs) and healthcare hasn\u2019t covered,\u201d said the study\u2019s senior author Amulya Yadav, PhD, associate professor of informatics and intelligent systems at Penn State. \u201cWe wanted to understand that if people are using LLMs like ChatGPT as a symptom health checker, like historically we\u2019ve used Google, how accurate is the LLM in answering those queries, and how harmful could those responses be?\u201d<\/p>\n<p>The study examined a shift in how people seek healthcare information online. According to the researchers, \u201cover half of U.S. adults consult online resources for medical advice,\u201d a trend that is particularly prominent in younger adults, with near one-quarter of people under 30 using AI for health-related guidance.<\/p>\n<p>The research was designed to evaluate AI performance in real-world everyday health communication, rather than in controlled testing environments that rely on medical licensing exams or clinical case studies. The researchers wrote that earlier studies \u201cfail to account for the unstructured and often ambiguous nature of general-purpose everyday health inquiries.\u201d<\/p>\n<p>To conduct the study, the Penn State team organized a weeklong \u201cDiagnose-a-thon\u201d competition involving 34 participants, including faculty, staff, undergraduate students and graduate students. Participants submitted 212 prompts describing real or imagined health concerns from both patient and physician perspectives. They were allowed to use one of four publicly accessible AI models: ChatGPT-4o, ChatGPT-3.5, Gemini-1.5 Pro or Llama3-8b.<\/p>\n<p>\u201cOne of the strengths of our study is we\u2019re essentially trying to replicate real-world usage of LLMs by telling participants to choose the LLM of their choice and use it as they would on a normal day,\u201d said lead author Bonam Mingole, a doctoral candidate in information sciences and technology at Penn State. \u201cThis type of participatory research is so important for understanding how the public uses AI in their daily life.\u201d<\/p>\n<p>Nine board-certified physicians evaluated the AI-generated responses to the prompts using a six-point scale to measure validity, quality of information, understanding and reasoning, and potential harm. About 76% of responses were considered accurate overall and showed that ChatGPT-4o was the most accurate at 84.62%, while Llama3-8b had the lowest at 50%.<\/p>\n<p>The results also showed significant differences across medical specialties. Obstetrics and gynecology and otolaryngology generated the strongest AI performance, with high validity scores and low harm scores. Internal medicine, neurology, and dermatology produced the weakest results, including lower validity and higher risk of harmful responses.<\/p>\n<p>The responses also shined a light on a continuing problem in healthcare research that leads to disparities in care. \u201cThe lower quality LLM-generated responses for underrepresented patient populations and rare medical conditions raises concerns about the potential of LLMs to inadvertently exacerbate existing healthcare disparities,\u201d the researchers wrote. They added that addressing these issues \u201crequires more than technical mechanisms, it calls for a broader commitment to equity in the data collection, model development, and evaluation processes.\u201d<\/p>\n<p>How the prompts were written also influenced AI performance. Queries between 60 and 250 characters produced the most accurate responses, while very specific prompts also improved output quality. Interestingly, the physician reviewers perceived greater risk of harm in responses generated from prompts that attempted to be written from a medical professional\u2019s perspective.<\/p>\n<p>To test whether specialized medical training could improve the answer generated by AI, the investigators enhanced the LLMs using Retrieval-Augmented Generation, by training them with medical textbooks, clinical guidelines and peer-reviewed research that is typically included in medical school curriculums.<\/p>\n<p>Performance of the LLMs trained in this way were mixed. The physician reviewers preferred the baseline versions of Gemini and Llama over their medically trained counterparts, while no meaningful preference emerged for the ChatGPT models.<\/p>\n<p>The researchers will now look to expand their study by collecting larger and more balanced datasets of crowdsourced health prompts and looking for ways to discourage overreliance on AI-generated medical advice.<\/p>\n<p>\u201cLike it or not, people will continue to use AI for diagnosing their health problems,\u201d said study co-author S. Shyam Sundar, PhD, a professor in the Center for Socially Responsible AI at Penn State. \u201cBy understanding their use patterns and testing the validity of AI performance, our project helps advance literacy on the best and worst uses of AI for medical advice.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"Credit: ertigo3d \/ Getty Images A study into the responses of AI chatbots to everyday medical questions found&hellip;\n","protected":false},"author":2,"featured_media":55996,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,32680,32681,2749,17767,360,4792,4793],"class_list":["post-55995","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-data-collection-public-health","tag-health-care-services","tag-news-features","tag-patient-care","tag-precision-medicine","tag-topics","tag-translational-research"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/55995","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=55995"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/55995\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/55996"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=55995"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=55995"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=55995"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}