{"id":118516,"date":"2026-07-25T05:52:46","date_gmt":"2026-07-25T05:52:46","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/118516\/"},"modified":"2026-07-25T05:52:46","modified_gmt":"2026-07-25T05:52:46","slug":"chatgpt-health-performance-in-a-structured-test-of-triage-recommendations","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/118516\/","title":{"rendered":"ChatGPT Health performance in a structured test of triage recommendations"},"content":{"rendered":"<p>On 7 January 2026, OpenAI launched ChatGPT Health, a consumer-facing feature designed to \u2018recommend how urgently to encourage follow-ups with a clinician\u2019 and provide health guidance directly to the public<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"OpenAI. Introducing ChatGPT Health. &#010;                  https:\/\/openai.com\/index\/introducing-chatgpt-health\/&#010;                  &#010;                 (accessed 13 January 2026).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR1\" id=\"ref-link-section-d125471665e757\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>. Developed alongside HealthBench, OpenAI\u2019s benchmark for evaluating health artificial intelligence (AI), ChatGPT Health functions as a first-contact point for symptom guidance, in which triage errors may reach patients directly without a clinician buffer. The associated risks are asymmetric\u2014undertriage may delay or preclude life-saving treatment, while overtriage primarily increases healthcare utilization<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Bellini, V. &amp; Bignami, E. G. Generative pre-trained transformer 4 (GPT-4) in clinical settings. Lancet Digit. Health 7, e6&#x2013;e7 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR2\" id=\"ref-link-section-d125471665e761\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. Large language models (LLMs) can perform well on medical licensing examinations, yet such performance does not ensure safe triage, particularly at clinical extremes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Busch, F. et al. Current applications and challenges in large language models for patient care: a systematic review. Commun. Med. (Lond.) 5, 26 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR3\" id=\"ref-link-section-d125471665e765\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>. Evidence that patients act on LLM-generated medical advice regardless of its quality makes triage accuracy a public health imperative<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Shekar, S., Pataranutaporn, P., Sarabu, C., Cecchi, G. A. &amp; Maes, P. People overtrust AI-generated medical advice despite low accuracy. NEJM AI 2, AIoa2300015 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR4\" id=\"ref-link-section-d125471665e769\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>. Patient-facing systems must demonstrate safety through external validation where the cost of error is the greatest<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Busch, F. et al. Current applications and challenges in large language models for patient care: a systematic review. Commun. Med. (Lond.) 5, 26 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR3\" id=\"ref-link-section-d125471665e773\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"De Hond, A. et al. From text to treatment: the crucial role of validation for generative large language models in health care. Lancet Digit. Health 6, e441&#x2013;e443 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR5\" id=\"ref-link-section-d125471665e776\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>.<\/p>\n<p>Prior work has shown that general-purpose LLMs shift their recommendations when patients are identified by race or sex, and that misleading framing\u2014such as reassurance from family or friends\u2014can anchor outputs towards less urgent care<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Blumenthal, D. &amp; Goldberg, C. Managing patient use of generative health AI. NEJM AI 2, AIpc2400927 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR6\" id=\"ref-link-section-d125471665e783\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Omar, M. et al. Sociodemographic biases in medical decision making by large language models. Nat. Med. 31, 1873&#x2013;1881 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR7\" id=\"ref-link-section-d125471665e786\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>. Whether ChatGPT Health inherits these vulnerabilities or has mitigated them remains untested.<\/p>\n<p>ChatGPT Health is freely available 24\/7, does not exclude high-acuity queries and HealthBench includes emergency triage evaluation<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Arora, R. K. et al. HealthBench: evaluating large language models towards improved human health. Preprint at &#010;                  https:\/\/arxiv.org\/abs\/2505.08775&#010;                  &#010;                 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR8\" id=\"ref-link-section-d125471665e793\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>. Users will present with emergencies regardless of design intent<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Shekar, S., Pataranutaporn, P., Sarabu, C., Cecchi, G. A. &amp; Maes, P. People overtrust AI-generated medical advice despite low accuracy. NEJM AI 2, AIoa2300015 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR4\" id=\"ref-link-section-d125471665e797\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>. We conducted an independent, structured stress test of ChatGPT Health using clinician-authored vignettes spanning the full acuity spectrum, with controlled variation of anchoring, access barriers, race and sex, to assess whether it fails safely at clinical extremes and whether nonclinical factors shift its triage recommendations.<\/p>\n<p>We obtained 960 prompt responses from 60 clinician-authored vignettes, each tested across 16 factorial conditions varying patient race, sex, anchoring context and access barriers (Methods; Extended Data Figs. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> and Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). The 30 base scenarios, spanning 21 medical domains, were each authored in two versions\u2014one presenting only subjective data (symptoms and history) and one additionally including objective findings (laboratory values, vital signs, physical examination)\u2014yielding a total of 60 vignettes (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>). Three physicians independently assigned gold-standard triage levels based on cited clinical guidelines and their clinical expertise, with high inter-rater agreement (Supplementary Data <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>), using a four-level Likert scale\u2014A (nonurgent, \u2018monitor at home\u2019), B (semi-urgent, \u2018see a doctor within weeks\u2019), C (urgent, \u2018see a doctor within 24\u201348\u2009h\u2019), D (emergency, \u2018go to the emergency department\u2019). Cases were classified as \u2018clear\u2019 (single correct triage level; n\u2009=\u200930, 480 prompt responses) or \u2018edge\u2019 (two adjacent levels clinically reasonable; n\u2009=\u200930, 480 prompt responses).<\/p>\n<p>Among clear cases, ChatGPT Health exhibited an inverted U-shaped pattern (equivalently, a U-shaped pattern for mistriage) across acuity levels (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>). The clear-case distribution included eight nonurgent, eight semi-urgent, ten urgent and four emergency vignettes (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). Accuracy peaked for intermediate presentations\u201493.0% for semi-urgent and 76.9% for urgent. Performance declined at clinical extremes\u201435.2% for nonurgent and 48.4% for emergency conditions. Among true emergencies, 51.6% (33\/64) were undertriaged to 24\u201348\u2009h evaluation. Conversely, 64.8% (83\/128) of nonurgent cases were overtriaged, predominantly by one level to scheduled physician visits; none were sent to emergency departments.<\/p>\n<p>Fig. 1: ChatGPT Health undertriages emergencies while overtriaging nonurgent cases.<img decoding=\"async\" aria-describedby=\"figure-1-desc\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/41591_2026_4297_Fig1_HTML.png\" alt=\"Fig. 1: ChatGPT Health undertriages emergencies while overtriaging nonurgent cases.\" loading=\"lazy\" width=\"685\" height=\"266\"\/><\/p>\n<p>Clear vignettes only (single correct gold-standard triage; n\u2009=\u2009480 responses). Triage levels\u2014A (monitor at home), B (see a doctor within weeks), C (see a doctor within 24\u201348\u2009h) and D (go to the emergency department now). a, Mistriage rate across gold-standard acuity. Mistriage (1\u2009\u2212\u2009accuracy) followed a U-shaped pattern, with the highest values at the extremes (A, 64.8%; D, 51.6%) and the lowest at intermediate acuity (B, 7.0%; C, 23.1%). Because A is the least urgent category and D the most urgent, errors at A necessarily represent overtriage and errors at D necessarily represent undertriage (annotations). Dashed line marks 50% mistriage for reference. b, Direction of triage outcomes. Within each gold-standard acuity level, stacked diverging bars show the proportion of cases that were undertriaged (recommended less urgent care than gold, left of zero), correctly triaged (gray) or overtriaged (recommended more urgent care, right of zero). Emergencies (D) were undertriaged in 33\/64 (51.6%) cases, whereas nonurgent\/home-care cases (A) were overtriaged in 83\/128 (64.8%).<\/p>\n<p><a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM5\" rel=\"nofollow noopener\" target=\"_blank\">Source data<\/a><\/p>\n<p>The four emergency vignettes comprised two clinical scenarios\u2014an asthma exacerbation and diabetic ketoacidosis (DKA)\u2014each tested with and without objective findings. Undertriage was concentrated in asthma exacerbation, which accounted for 84.8% (28\/33) of undertriaged emergency responses. The model\u2019s explanations revealed the failure mechanism (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>). In the case of asthma exacerbation, the model identified the warning sign\u2014\u2018CO2 mildly elevated, an early sign you\u2019re not ventilating well\u2019\u2014then rationalized it away\u2014\u2018findings don\u2019t prove immediate respiratory failure\u2019 and \u2018still speaking in full sentences\u2019. In DKA, the model correctly identified \u2018early\u2019 or \u2018mild\u2019 DKA but recommended outpatient management, apparently conflating DKA\u2014which is by definition an emergency\u2014with hyperglycemia. A supplementary analysis of four textbook emergencies (stroke, anaphylaxis, meningitis and aortic dissection; 128 responses) showed 0% undertriage (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>), suggesting the model identifies classic presentations but fails when emergency status depends on clinical progression.<\/p>\n<p>Among edge cases, 96.0% of responses fell within the acceptable clinical range, defined as at or above the acceptable clinical floor\u2014the lowest triage level considered clinically safe for a given vignette. However, 60.8% chose the less urgent of the two acceptable options; when both urgent (C) and emergency (D) were deemed acceptable, ChatGPT Health recommended the less urgent option 72.7% of the time. Only 0.6% (3\/480) of edge-case responses fell below the acceptable clinical floor.<\/p>\n<p>Of eight prespecified hypothesis tests, only anchoring significantly affected triage behavior (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>). Anchoring statements increased the probability of triage shift from 3.3% (8\/240) to 13.3% (32\/240) in edge cases (odds ratio (OR)\u2009=\u200911.7, 95% confidence interval (CI)\u2009=\u20093.7\u201336.6; Holm-adjusted P\u2009&lt;\u20090.001). Among triage shifts, 52.5% (21\/40) were de-escalations towards less urgent care and 93.8% (30\/32) remained within acceptable clinical bounds. Access-barrier statements (insurance, transportation or work constraints) did not significantly affect triage (OR\u2009=\u20090.51, 95% CI\u2009=\u20090.13\u20131.95; H6, OR\u2009=\u20091.63, 95% CI\u2009=\u20090.73\u20133.64; both Holm-adjusted P\u2009=\u20091.0).<\/p>\n<p>Patient race and sex showed no significant effects on triage recommendations. Undertriage occurred in 17.0% of Black patients versus 14.3% of white patients (risk difference\u2009=\u2009+2.7%; OR\u2009=\u20091.96, 95% CI\u2009=\u20090.51\u20137.53; Holm-adjusted P\u2009=\u20091.0). Sex showed similar patterns. The wide CIs reflect sparse events (16\u201319 per cell); these data are compatible with undertriage risk differences ranging from approximately \u22128% to +14% between Black and white patients, precluding definitive conclusions about demographic equity.<\/p>\n<p>Adding objective findings (for example, laboratory values and vital signs) improved overall accuracy from 54.6% to 77.9% (sensitivity analysis; OR\u2009=\u20099.4, 95% CI\u2009=\u20094.9\u201318.0, P\u2009&lt;\u20090.001). This effect differed by acuity. For nonurgent presentations (A; n\u2009=\u2009128), objective findings prevented overtriage by 61 percentage points (95.3% versus 34.4%; OR\u2009=\u200937.5, 95% CI\u2009=\u200910.4\u2013207, P\u2009&lt;\u20090.001). For emergencies (D; n\u2009=\u200964), the pattern reversed\u2014 objective findings increased undertriage by 9.3 percentage points (56.2% versus 46.9%; OR\u2009=\u20090.69, 95% CI\u2009=\u20090.23\u20132.05, P\u2009=\u20090.62).<\/p>\n<p>A distinct safety failure emerged in the suicidal ideation vignettes. In a 27-year-old patient reporting \u2018I\u2019ve thought about taking a lot of pills,\u2019 crisis-intervention messages appeared in none (0\/16) of the responses that included normal objective findings, but in all (16\/16) when objective findings were removed, despite identical clinical severity (Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"table anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#Tab1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>). To characterize this pattern, we tested five additional suicidal ideation scenarios across 16 factorial variants each (224 total responses; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). The vignettes ranged from a 41-year-old wishing they would not wake up to a 34-year-old reporting \u2018thoughts about killing myself\u2019 after a job loss, a 45-year-old whose suicidal ideation emerged during nighttime drinking and a 22-year-old frightened by a first-ever thought of self-harm. The crisis interstitial\u2014a \u2018Help is available\u2019 banner linking to the 988 Suicide and Crisis Lifeline\u2014was triggered in only 4 of 14 vignettes (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>); the remaining 10 produced no safety alert in any variant (0\/160 responses). The pattern was not merely inconsistent but paradoxically inverted relative to clinical severity. Among the three scenarios featuring active suicidal ideation with an identified method\u2014including alcohol-facilitated ideation and first-episode thoughts of overdose contemplation\u2014only one of six vignettes triggered the interstitial. In contrast, the guardrail fired more reliably for the patient who had not identified a means of self-harm than for those who had.<\/p>\n<p>Table 1 Illustrative cases demonstrating ChatGPT triage variation by clinical data presentation<\/p>\n<p>ChatGPT Health errs at clinical extremes, characterized by undertriage of emergencies and overtriage of nonurgent cases, while showing resistance to sociodemographic biases previously documented in general-purpose LLMs. The inverted U-shaped accuracy pattern implicates central-tendency bias as a dominant failure mode, potentially reflecting underrepresentation of clinical extremes in training data. The 51.6% undertriage rate for true emergencies represents the most concerning finding, as missed emergencies can result in patient harm, while the 64.8% overtriage rate for nonurgent cases, although less dangerous, risks unnecessary healthcare utilization at scale. The undertriaged cases\u2014rising pCO2 signaling respiratory failure, metabolic acidosis in DKA\u2014are presentations that no experienced clinician would delay.<\/p>\n<p>The failure to escalate emergencies extends prior evidence that LLM behavior can be brittle under clinically demanding decision tasks and may need human oversight for clinical judgment<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med. 30, 2613&#x2013;2622 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR9\" id=\"ref-link-section-d125471665e1330\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>. Undertriage is the more consequential error type in triage contexts<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28, 924&#x2013;933 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR10\" id=\"ref-link-section-d125471665e1334\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>. Consumer-facing deployments that provide health guidance, including those with explicit disclaimers stating that they are not intended for diagnosis or treatment, nonetheless function as de facto triage tools for the millions of users who consult them<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Shekar, S., Pataranutaporn, P., Sarabu, C., Cecchi, G. A. &amp; Maes, P. People overtrust AI-generated medical advice despite low accuracy. NEJM AI 2, AIoa2300015 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR4\" id=\"ref-link-section-d125471665e1338\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>.<\/p>\n<p>Current approaches to medical LLM development have not adequately addressed calibration at clinical extremes: a specific engineering target that emerges from this evaluation. The observed protective effect of quantitative clinical data on triage accuracy\u2014improving overall accuracy by 23 percentage points\u2014is consistent with prior evidence that inclusion of structured physiological data improves LLM triage performance<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 11\" title=\"Gaber, F. et al. Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis. NPJ Digit. Med. 8, 263 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR11\" id=\"ref-link-section-d125471665e1345\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>. Our findings extend this observation to consumer-facing deployments, where most users lack access to laboratory or vital sign data.<\/p>\n<p>Our finding of no significant demographic bias contrasts with ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Omar, M. et al. Sociodemographic biases in medical decision making by large language models. Nat. Med. 31, 1873&#x2013;1881 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR7\" id=\"ref-link-section-d125471665e1352\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>, which reported race and sex effects in general-purpose LLMs. While our CIs are wide because the within-vignette design was optimized for testing experimental manipulations, ChatGPT Health may incorporate bias-mitigation guardrails absent in base models. Our anchoring findings align with growing evidence that LLM clinical reasoning is vulnerable to contextual manipulation. Prior work documented a weighted mean 21% drop in diagnostic accuracy when clinical distractors and disruptive patient behaviors were introduced<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 12\" title=\"Schmidt, H. G., Rotgans, J. I. &amp; Mamede, S. Bias sensitivity in diagnostic decision-making: comparing ChatGPT with residents. J. Gen. Intern. Med. 40, 790&#x2013;795 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR12\" id=\"ref-link-section-d125471665e1356\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>, susceptibility to adversarial priming across domains<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Thapa, R. et al. Disentangling reasoning and knowledge in medical large language models. Preprint at &#010;                  https:\/\/arxiv.org\/abs\/2505.11462&#010;                  &#010;                 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR13\" id=\"ref-link-section-d125471665e1360\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a> and failure to revise decisions when confronted with contradictory evidence<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 14\" title=\"McCoy, L. G. et al. Assessment of large language models in clinical reasoning: a novel benchmarking study. NEJM AI 2, AIdbp2500120 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR14\" id=\"ref-link-section-d125471665e1364\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a>. Our data extend these findings to consumer-facing triage, but with a critical distinction\u2014anchoring was significant only for edge cases (OR\u2009=\u200911.7, 95% CI\u2009=\u20093.7\u201336.6), not for clear cases.<\/p>\n<p>The crisis guardrail finding may be the most consequential failure mode exhibited in the entire study. What we found was worse than simple suppression. Trust calibration requires predictable system behavior\u2014when reliability is inconsistent, users cannot learn when to rely on the system and when to override it<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Lee, J. D. &amp; See, K. A. Trust in automation: designing for appropriate reliance. Hum. Factors 46, 50&#x2013;80 (2004).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR15\" id=\"ref-link-section-d125471665e1371\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>. A guardrail that fires for \u2018haven\u2019t thought through how I would do it\u2019 but not for \u2018thought about taking a lot of pills\u2019 is not calibrated to clinical risk and users have no basis to anticipate when it will or will not fire. The capability to recognize mental health crises and connect users with crisis resources is a basic prerequisite for any consumer health platform. Our data show this prerequisite has not been reliably met. OpenAI has acknowledged, in a post titled \u2018Helping people when they need it most,\u2019 that model behavior in mental health contexts requires particular attention<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"OpenAI. Helping people when they need it most. &#010;                  https:\/\/openai.com\/index\/helping-people-when-they-need-it-most&#010;                  &#010;                 (accessed 31 January 2026).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR16\" id=\"ref-link-section-d125471665e1375\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>. Our findings identify not a theoretical concern but a documented pattern of interstitial activation discordant with clinical severity.<\/p>\n<p>This study has limitations. We used clinical vignettes rather than real-world patient interactions. Controlled studies of real users suggest this represents a conservative test\u2014consumers under-report symptoms and misapply advice even when the system provides correct guidance, conditions that would compound the triage errors we report<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Bean, A. M. et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat. Med. 32, 609&#x2013;615 (2026).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR17\" id=\"ref-link-section-d125471665e1383\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>. If ChatGPT Health undertriages 51.6% of emergencies with clean clinical information, performance with incomplete consumer inputs is unlikely to be superior. Emergency undertriage was concentrated in trajectory-dependent conditions where clinical evolution dictates urgency and whether this failure mode extends to other acute presentations remains untested. The standardized prompt required selection of a single triage level (A\u2013D), capturing discrete recommendations rather than the hedged, multicontingency advice that open-ended interaction might produce. The within-vignette factorial design provides strong internal validity for manipulation effects but limits statistical power for detecting small demographic effects; however, the observed point estimates nonetheless provide useful bounds. We evaluated a single time point and model behavior may change with updates, only underscoring the need for ongoing evaluation as these systems evolve.<\/p>\n<p>The implication is straightforward\u2014consumer-facing AI that functions as a front door for urgent medical decisions should not be deployed on trust alone. Our findings identify two engineering targets requiring immediate attention\u2014emergency detection that accounts for clinical trajectory, not just snapshot presentation and crisis guardrails that fire consistently rather than unpredictably. Given the direct patient-safety implications of missed emergencies, consumer health AI may warrant premarket safety evaluation requirements analogous to medical devices<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28, 924&#x2013;933 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04297-7#ref-CR10\" id=\"ref-link-section-d125471665e1390\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>. At a minimum, these tools should demonstrate external safety for emergencies before widespread public deployment.<\/p>\n","protected":false},"excerpt":{"rendered":"On 7 January 2026, OpenAI launched ChatGPT Health, a consumer-facing feature designed to \u2018recommend how urgently to encourage&hellip;\n","protected":false},"author":2,"featured_media":118517,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[7040,7044,580,16568,617,1668,17234,2533,4965,15020,7043,15021,157],"class_list":["post-118516","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-biomedicine","tag-cancer-research","tag-chatgpt","tag-computational-platforms-and-environments","tag-general","tag-health-care","tag-health-policy","tag-infectious-diseases","tag-medical-research","tag-metabolic-diseases","tag-molecular-medicine","tag-neurosciences","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/118516","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=118516"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/118516\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/118517"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=118516"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=118516"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=118516"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}