{"id":55397,"date":"2026-05-29T16:34:20","date_gmt":"2026-05-29T16:34:20","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/55397\/"},"modified":"2026-05-29T16:34:20","modified_gmt":"2026-05-29T16:34:20","slug":"even-the-best-ai-chatbot-gets-health-questions-wrong-1-in-5-times-doctors-find","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/55397\/","title":{"rendered":"Even the Best AI Chatbot Gets Health Questions Wrong 1 in 5 Times, Doctors Find"},"content":{"rendered":"<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/AI-apps-1200x800.jpeg\" alt=\"AI Apps\"\/><\/p>\n<p class=\"post-featured-image-caption\">ChatGPT, Claude, and Gemini are among the most widely used AI apps used.  (\u00a9 prima91 &#8211; stock.adobe.com)<\/p>\n<p>Board-Certified Physicians Put Popular LLMs Through Their Paces, and Found Real Problems<\/p>\n<p>In a Nutshell<\/p>\n<p>Across more than 200 AI-generated medical responses, doctors found that about 76% were considered valid, meaning roughly 1 in 4 fell short.<\/p>\n<p>ChatGPT-4o performed best among the four AI models tested, while Llama3-8b performed worst, with doctors rating only half of its responses as valid.<\/p>\n<p>Adding a specialized medical knowledge library to the AI did not consistently improve results. For some models, doctors actually preferred the standard version.<\/p>\n<p>Mental health queries drew special concern from physicians, with some warning that AI responses in crisis situations could be actively dangerous.<\/p>\n<p class=\"wp-block-paragraph\">When people feel a strange pain or notice a worrying symptom, more and more of them are skipping the doctor\u2019s office and heading straight to an AI chatbot. It\u2019s fast, free, and available at 3 a.m. But a study suggests that convenience might come with a serious catch: even the best-performing AI gets medical questions wrong roughly one out of every five times.<\/p>\n<p class=\"wp-block-paragraph\">In a preprint study (not yet peer-reviewed) posted online by researchers from Penn State, four popular <a href=\"https:\/\/studyfinds.com\/tag\/chatbots\/\" type=\"post_tag\" id=\"93660\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">AI chatbots<\/a> were put to the test using real and imagined health concerns submitted by university students, staff, and faculty. A panel of nine board-certified physicians then graded the AI responses. Overall results were mixed: impressive enough to turn heads, but flawed enough to raise real concerns about what happens when someone acts on bad medical advice. <\/p>\n<p class=\"wp-block-paragraph\">Nearly one in four adults under 30 already use AI monthly for <a href=\"https:\/\/studyfinds.com\/dr-google-typically-wrong-when-diagnosing-medical-problems-online-symptom-checkers\/\" type=\"post\" id=\"25691\" rel=\"nofollow noopener\" target=\"_blank\">health-related guidance<\/a>, according to data cited <a href=\"https:\/\/arxiv.org\/pdf\/2506.13805\" type=\"link\" id=\"https:\/\/arxiv.org\/pdf\/2506.13805\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">in the paper<\/a>. Understanding what these tools get right (and wrong) is essential.<\/p>\n<p><img fetchpriority=\"high\" decoding=\"async\" width=\"1200\" height=\"673\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/Robot-hand-stethoscope-1200x673.jpeg\" alt=\"Robotic hand signifying artificial intelligence (AI) touching a stethoscope\"  \/><\/p>\n<p>The robot doctor may not be ready to see you just yet. (\u00a9  Slowlifetrader \u2013 stock.adobe.com)<\/p>\n<p>How Researchers Tested AI Chatbots on Health Questions<\/p>\n<p class=\"wp-block-paragraph\">Researchers organized a university-wide <a href=\"https:\/\/www.psu.edu\/news\/campus-life\/story\/competition-highlights-generative-ais-power-pitfalls-medical-diagnoses\" type=\"link\" id=\"https:\/\/www.psu.edu\/news\/campus-life\/story\/competition-highlights-generative-ais-power-pitfalls-medical-diagnoses\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">competition<\/a> in fall 2024. A total of 34 participants were invited to query one of four AI chatbots \u2014 ChatGPT-4o, ChatGPT-3.5, Gemini-1.5 Pro, and Llama3-8b \u2014 with health-related questions they might genuinely want answered. Participants could approach the task from one of three angles: as a patient describing personal symptoms, as a medical professional seeking diagnostic help, or through an out-of-the-box track that allowed for alternative medical query scenarios, such as analyzing images of handwritten prescriptions.<\/p>\n<p class=\"wp-block-paragraph\">Competition entries generated 212 AI responses in total. Those responses were then divided among a panel of nine board-certified physicians, each of whom graded them on four measures: how valid the information was, the quality of the information, how well the AI reasoned through the problem, and whether the response <a href=\"https:\/\/studyfinds.com\/ai-chatbots-may-be-worsening-mental-illnesses\/\" type=\"post\" id=\"150064\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">could cause harm<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">Gemini-1.5 Pro produced the largest share of responses, 140 out of 212, while Llama3-8b generated only 6. That imbalance matters when comparing models directly, and the researchers acknowledged it as a limitation.<\/p>\n<p>What Doctors Found When They Graded the AI Responses<\/p>\n<p class=\"wp-block-paragraph\">Across all four AI models, about 76% of responses were rated as valid by physicians. That sounds reasonable until the math flips: nearly one in four responses didn\u2019t make the cut. For <a href=\"https:\/\/studyfinds.com\/tag\/chatgpt\/\" type=\"post_tag\" id=\"93212\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">ChatGPT-4o<\/a>, the highest-performing model, validity hit 84.6%, still leaving more than 15% of answers falling short. Llama3-8b landed at the bottom, with only half its responses rated as valid.<\/p>\n<p class=\"wp-block-paragraph\">Which type of medical question was asked also mattered. Questions about obstetrics and gynecology scored the highest for accuracy, while neurology, internal medicine, and <a href=\"https:\/\/studyfinds.com\/dermatologists-finally-agree-major-study-reveals-skincare-ingredients-that-work\/\" type=\"post\" id=\"143505\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">dermatology<\/a> consistently ranked lower. Neurology cases in the study often involved rare conditions that are hard to diagnose under any circumstances, while dermatology relies heavily on visual examination \u2014 something a text-based chatbot simply cannot replicate.<\/p>\n<p class=\"wp-block-paragraph\">Prompt length turned out to be a factor, too. Very short questions and very long, detailed ones both produced weaker results. Best performance came from medium-length queries, somewhere between 60 and 250 characters. Medical professionals said in follow-up interviews that the more specific and focused the question, the better the AI tended to perform.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"757\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/ChatGPT-message-screen-1200x757.jpg\" alt=\"ChatGPT prompt on computer\"  \/><\/p>\n<p>ChatGPT-4o proved to be the best model among those tested in the study. (Bangla press\/Shutterstock)<\/p>\n<p>Adding a Medical Encyclopedia Didn\u2019t Always Help AI Chatbots<\/p>\n<p class=\"wp-block-paragraph\">One of the study\u2019s more surprising results involved a technique called Retrieval-Augmented Generation, or RAG, essentially giving the AI access to a curated library of medical textbooks, clinical guidelines, and research articles from a university medical school before it generates a response. Grounding the AI in vetted medical sources should, in theory, make its answers more reliable.<\/p>\n<p class=\"wp-block-paragraph\">Seven medical professionals were recruited to compare standard AI responses against RAG-enhanced ones, side by side. For Gemini-1.5 Pro and Llama3-8b, the medical professionals actually preferred the standard, unenhanced versions by a wide and statistically significant margin. For the <a href=\"https:\/\/studyfinds.com\/chatgpt-financial-advice\/\" type=\"post\" id=\"138412\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">ChatGPT models<\/a>, there was no significant difference either way.<\/p>\n<p class=\"wp-block-paragraph\">Researchers stopped short of declaring RAG unhelpful overall, noting that the results varied by model and that future research should explore the approach further.<\/p>\n<p>What Doctors Really Think About AI and Patient Safety<\/p>\n<p class=\"wp-block-paragraph\">Seven medical professionals who took part in the evaluation were also interviewed about their broader views on AI in medicine. On the positive side, they saw real potential for AI to improve health literacy: helping patients understand their conditions, explore possible explanations for symptoms, and feel more engaged in their own care. Several noted that AI could serve as a useful first step for people deciding whether a symptom warrants <a href=\"https:\/\/studyfinds.com\/healthcares-new-partisan-battleground-doctors-office\/\" type=\"post\" id=\"141230\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">a doctor\u2019s visit,<\/a> potentially easing the burden on overcrowded emergency rooms.<\/p>\n<p class=\"wp-block-paragraph\">Concerns ran equally deep. Every doctor interviewed raised worries about overreliance. One described the scenario of a parent being falsely reassured by an AI while their child was seriously ill. Another flagged the risk that patients from groups historically underrepresented in medical research might receive less accurate responses, potentially widening existing health disparities. <\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/studyfinds.com\/tag\/privacy\/\" type=\"post_tag\" id=\"1774\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Privacy<\/a> was another concern. Several doctors warned that people entering detailed personal health information into AI chatbots may be exposing themselves to serious data risks.<\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/studyfinds.com\/tag\/mental-health\/\" type=\"post_tag\" id=\"124\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Mental health<\/a> queries drew particular caution. Some interviewees said AI responses on mental health topics could be actively dangerous, with one suggesting that if an AI can\u2019t handle a mental health crisis responsibly, it simply shouldn\u2019t respond at all.<\/p>\n<p class=\"wp-block-paragraph\">A 20% error rate, the approximate failure rate for even <a href=\"https:\/\/studyfinds.com\/tag\/artificial-intelligence\/\" type=\"post_tag\" id=\"1101\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">the best AI model<\/a> in this study, would be considered unacceptable in almost any medical setting. Researchers are direct about this: these results are not a green light for using AI chatbots as a substitute for professional medical advice. People who rely on these tools for diagnosis or health decisions should treat them as a starting point for conversation, not a final answer. In medicine, \u201cpretty good\u201d has never been good enough.<\/p>\n<p class=\"wp-block-paragraph\">Disclaimer: This article is for general informational purposes only and does not constitute medical advice. The findings described come from a university competition in which crowdsourced health questions were posed to AI chatbots and evaluated by physicians; they reflect performance on a specific set of queries and may not predict how any AI tool will perform in all real-world situations. Always consult a qualified healthcare professional before making decisions about symptoms, diagnoses, or medical care.<\/p>\n<p>Paper Notes<\/p>\n<p>Limitations<\/p>\n<p class=\"wp-block-paragraph\">The study\u2019s dataset of 212 responses was unevenly distributed across the four AI models, with Gemini-1.5 Pro accounting for 140 responses and Llama3-8b contributing only 6. This imbalance limits the statistical power of direct model-to-model comparisons and reduces the ability to generalize findings for underrepresented models. Specialty-level analysis was restricted to medical categories with at least 10 entries, meaning several fields were excluded. The RAG pipeline was built from the curriculum of a single university medical school, which may not represent the full breadth of available medical knowledge. Researchers also note that the relatively small number of participants and the voluntary, competition-based format of data collection may affect how broadly the findings apply to the general public\u2019s everyday use of AI health tools.<\/p>\n<p>Funding and Disclosures<\/p>\n<p class=\"wp-block-paragraph\">No external funding sources or financial disclosures are identified in the paper\u2019s content. The study was conducted under institutional review board (IRB) approval. Participants in the follow-up interviews received a $60 Amazon e-gift card as compensation. Prize money was awarded to competition participants, with a first-place prize of $1,000, a second-place prize of $500, a third-place prize of $250, five consolation prizes of $50 each, and a separate $1,000 prize for the submission rated highest on harm assessment.<\/p>\n<p>Publication Details<\/p>\n<p class=\"wp-block-paragraph\">Authors: Bonam Mingole, Aditya Majumdar, Firdaus Ahmed Choudhury, Jennifer L. Kraschnewski, Shyam S. Sundar, Amulya Yadav \u2014 Pennsylvania State University and Penn State College of Medicine<\/p>\n<p class=\"wp-block-paragraph\">Paper Title: Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases<\/p>\n<p class=\"wp-block-paragraph\">Source: <a href=\"https:\/\/arxiv.org\/pdf\/2506.13805\" type=\"link\" id=\"https:\/\/arxiv.org\/pdf\/2506.13805\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">arXiv preprint<\/a> (submitted June 2025). This paper has not yet undergone formal peer review.<\/p>\n","protected":false},"excerpt":{"rendered":"ChatGPT, Claude, and Gemini are among the most widely used AI apps used. (\u00a9 prima91 &#8211; stock.adobe.com) Board-Certified&hellip;\n","protected":false},"author":2,"featured_media":55398,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[24,25,955,580,2536,1657,157],"class_list":["post-55397","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-ai","tag-artificial-intelligence","tag-chatbots","tag-chatgpt","tag-doctors","tag-health","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/55397","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=55397"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/55397\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/55398"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=55397"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=55397"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=55397"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}