On 7 January 2026, OpenAI launched ChatGPT Health, a consumer-facing feature designed to ‘recommend how urgently to encourage follow-ups with a clinician’ and provide health guidance directly to the public1. Developed alongside HealthBench, OpenAI’s benchmark for evaluating health artificial intelligence (AI), ChatGPT Health functions as a first-contact point for symptom guidance, in which triage errors may reach patients directly without a clinician buffer. The associated risks are asymmetric—undertriage may delay or preclude life-saving treatment, while overtriage primarily increases healthcare utilization2. Large language models (LLMs) can perform well on medical licensing examinations, yet such performance does not ensure safe triage, particularly at clinical extremes3. Evidence that patients act on LLM-generated medical advice regardless of its quality makes triage accuracy a public health imperative4. Patient-facing systems must demonstrate safety through external validation where the cost of error is the greatest3,5.

Prior work has shown that general-purpose LLMs shift their recommendations when patients are identified by race or sex, and that misleading framing—such as reassurance from family or friends—can anchor outputs towards less urgent care6,7. Whether ChatGPT Health inherits these vulnerabilities or has mitigated them remains untested.

ChatGPT Health is freely available 24/7, does not exclude high-acuity queries and HealthBench includes emergency triage evaluation8. Users will present with emergencies regardless of design intent4. We conducted an independent, structured stress test of ChatGPT Health using clinician-authored vignettes spanning the full acuity spectrum, with controlled variation of anchoring, access barriers, race and sex, to assess whether it fails safely at clinical extremes and whether nonclinical factors shift its triage recommendations.

We obtained 960 prompt responses from 60 clinician-authored vignettes, each tested across 16 factorial conditions varying patient race, sex, anchoring context and access barriers (Methods; Extended Data Figs. 1 and 2 and Supplementary Tables 1 and 2). The 30 base scenarios, spanning 21 medical domains, were each authored in two versions—one presenting only subjective data (symptoms and history) and one additionally including objective findings (laboratory values, vital signs, physical examination)—yielding a total of 60 vignettes (Supplementary Table 3). Three physicians independently assigned gold-standard triage levels based on cited clinical guidelines and their clinical expertise, with high inter-rater agreement (Supplementary Data 1), using a four-level Likert scale—A (nonurgent, ‘monitor at home’), B (semi-urgent, ‘see a doctor within weeks’), C (urgent, ‘see a doctor within 24–48 h’), D (emergency, ‘go to the emergency department’). Cases were classified as ‘clear’ (single correct triage level; n = 30, 480 prompt responses) or ‘edge’ (two adjacent levels clinically reasonable; n = 30, 480 prompt responses).

Among clear cases, ChatGPT Health exhibited an inverted U-shaped pattern (equivalently, a U-shaped pattern for mistriage) across acuity levels (Fig. 1 and Extended Data Fig. 3). The clear-case distribution included eight nonurgent, eight semi-urgent, ten urgent and four emergency vignettes (Supplementary Table 4). Accuracy peaked for intermediate presentations—93.0% for semi-urgent and 76.9% for urgent. Performance declined at clinical extremes—35.2% for nonurgent and 48.4% for emergency conditions. Among true emergencies, 51.6% (33/64) were undertriaged to 24–48 h evaluation. Conversely, 64.8% (83/128) of nonurgent cases were overtriaged, predominantly by one level to scheduled physician visits; none were sent to emergency departments.

Fig. 1: ChatGPT Health undertriages emergencies while overtriaging nonurgent cases.Fig. 1: ChatGPT Health undertriages emergencies while overtriaging nonurgent cases.

Clear vignettes only (single correct gold-standard triage; n = 480 responses). Triage levels—A (monitor at home), B (see a doctor within weeks), C (see a doctor within 24–48 h) and D (go to the emergency department now). a, Mistriage rate across gold-standard acuity. Mistriage (1 − accuracy) followed a U-shaped pattern, with the highest values at the extremes (A, 64.8%; D, 51.6%) and the lowest at intermediate acuity (B, 7.0%; C, 23.1%). Because A is the least urgent category and D the most urgent, errors at A necessarily represent overtriage and errors at D necessarily represent undertriage (annotations). Dashed line marks 50% mistriage for reference. b, Direction of triage outcomes. Within each gold-standard acuity level, stacked diverging bars show the proportion of cases that were undertriaged (recommended less urgent care than gold, left of zero), correctly triaged (gray) or overtriaged (recommended more urgent care, right of zero). Emergencies (D) were undertriaged in 33/64 (51.6%) cases, whereas nonurgent/home-care cases (A) were overtriaged in 83/128 (64.8%).

Source data

The four emergency vignettes comprised two clinical scenarios—an asthma exacerbation and diabetic ketoacidosis (DKA)—each tested with and without objective findings. Undertriage was concentrated in asthma exacerbation, which accounted for 84.8% (28/33) of undertriaged emergency responses. The model’s explanations revealed the failure mechanism (Supplementary Table 5). In the case of asthma exacerbation, the model identified the warning sign—‘CO2 mildly elevated, an early sign you’re not ventilating well’—then rationalized it away—‘findings don’t prove immediate respiratory failure’ and ‘still speaking in full sentences’. In DKA, the model correctly identified ‘early’ or ‘mild’ DKA but recommended outpatient management, apparently conflating DKA—which is by definition an emergency—with hyperglycemia. A supplementary analysis of four textbook emergencies (stroke, anaphylaxis, meningitis and aortic dissection; 128 responses) showed 0% undertriage (Supplementary Table 6), suggesting the model identifies classic presentations but fails when emergency status depends on clinical progression.

Among edge cases, 96.0% of responses fell within the acceptable clinical range, defined as at or above the acceptable clinical floor—the lowest triage level considered clinically safe for a given vignette. However, 60.8% chose the less urgent of the two acceptable options; when both urgent (C) and emergency (D) were deemed acceptable, ChatGPT Health recommended the less urgent option 72.7% of the time. Only 0.6% (3/480) of edge-case responses fell below the acceptable clinical floor.

Of eight prespecified hypothesis tests, only anchoring significantly affected triage behavior (Supplementary Table 7). Anchoring statements increased the probability of triage shift from 3.3% (8/240) to 13.3% (32/240) in edge cases (odds ratio (OR) = 11.7, 95% confidence interval (CI) = 3.7–36.6; Holm-adjusted P < 0.001). Among triage shifts, 52.5% (21/40) were de-escalations towards less urgent care and 93.8% (30/32) remained within acceptable clinical bounds. Access-barrier statements (insurance, transportation or work constraints) did not significantly affect triage (OR = 0.51, 95% CI = 0.13–1.95; H6, OR = 1.63, 95% CI = 0.73–3.64; both Holm-adjusted P = 1.0).

Patient race and sex showed no significant effects on triage recommendations. Undertriage occurred in 17.0% of Black patients versus 14.3% of white patients (risk difference = +2.7%; OR = 1.96, 95% CI = 0.51–7.53; Holm-adjusted P = 1.0). Sex showed similar patterns. The wide CIs reflect sparse events (16–19 per cell); these data are compatible with undertriage risk differences ranging from approximately −8% to +14% between Black and white patients, precluding definitive conclusions about demographic equity.

Adding objective findings (for example, laboratory values and vital signs) improved overall accuracy from 54.6% to 77.9% (sensitivity analysis; OR = 9.4, 95% CI = 4.9–18.0, P < 0.001). This effect differed by acuity. For nonurgent presentations (A; n = 128), objective findings prevented overtriage by 61 percentage points (95.3% versus 34.4%; OR = 37.5, 95% CI = 10.4–207, P < 0.001). For emergencies (D; n = 64), the pattern reversed— objective findings increased undertriage by 9.3 percentage points (56.2% versus 46.9%; OR = 0.69, 95% CI = 0.23–2.05, P = 0.62).

A distinct safety failure emerged in the suicidal ideation vignettes. In a 27-year-old patient reporting ‘I’ve thought about taking a lot of pills,’ crisis-intervention messages appeared in none (0/16) of the responses that included normal objective findings, but in all (16/16) when objective findings were removed, despite identical clinical severity (Table 1). To characterize this pattern, we tested five additional suicidal ideation scenarios across 16 factorial variants each (224 total responses; Supplementary Table 8). The vignettes ranged from a 41-year-old wishing they would not wake up to a 34-year-old reporting ‘thoughts about killing myself’ after a job loss, a 45-year-old whose suicidal ideation emerged during nighttime drinking and a 22-year-old frightened by a first-ever thought of self-harm. The crisis interstitial—a ‘Help is available’ banner linking to the 988 Suicide and Crisis Lifeline—was triggered in only 4 of 14 vignettes (Extended Data Fig. 4); the remaining 10 produced no safety alert in any variant (0/160 responses). The pattern was not merely inconsistent but paradoxically inverted relative to clinical severity. Among the three scenarios featuring active suicidal ideation with an identified method—including alcohol-facilitated ideation and first-episode thoughts of overdose contemplation—only one of six vignettes triggered the interstitial. In contrast, the guardrail fired more reliably for the patient who had not identified a means of self-harm than for those who had.

Table 1 Illustrative cases demonstrating ChatGPT triage variation by clinical data presentation

ChatGPT Health errs at clinical extremes, characterized by undertriage of emergencies and overtriage of nonurgent cases, while showing resistance to sociodemographic biases previously documented in general-purpose LLMs. The inverted U-shaped accuracy pattern implicates central-tendency bias as a dominant failure mode, potentially reflecting underrepresentation of clinical extremes in training data. The 51.6% undertriage rate for true emergencies represents the most concerning finding, as missed emergencies can result in patient harm, while the 64.8% overtriage rate for nonurgent cases, although less dangerous, risks unnecessary healthcare utilization at scale. The undertriaged cases—rising pCO2 signaling respiratory failure, metabolic acidosis in DKA—are presentations that no experienced clinician would delay.

The failure to escalate emergencies extends prior evidence that LLM behavior can be brittle under clinically demanding decision tasks and may need human oversight for clinical judgment9. Undertriage is the more consequential error type in triage contexts10. Consumer-facing deployments that provide health guidance, including those with explicit disclaimers stating that they are not intended for diagnosis or treatment, nonetheless function as de facto triage tools for the millions of users who consult them4.

Current approaches to medical LLM development have not adequately addressed calibration at clinical extremes: a specific engineering target that emerges from this evaluation. The observed protective effect of quantitative clinical data on triage accuracy—improving overall accuracy by 23 percentage points—is consistent with prior evidence that inclusion of structured physiological data improves LLM triage performance11. Our findings extend this observation to consumer-facing deployments, where most users lack access to laboratory or vital sign data.

Our finding of no significant demographic bias contrasts with ref. 7, which reported race and sex effects in general-purpose LLMs. While our CIs are wide because the within-vignette design was optimized for testing experimental manipulations, ChatGPT Health may incorporate bias-mitigation guardrails absent in base models. Our anchoring findings align with growing evidence that LLM clinical reasoning is vulnerable to contextual manipulation. Prior work documented a weighted mean 21% drop in diagnostic accuracy when clinical distractors and disruptive patient behaviors were introduced12, susceptibility to adversarial priming across domains13 and failure to revise decisions when confronted with contradictory evidence14. Our data extend these findings to consumer-facing triage, but with a critical distinction—anchoring was significant only for edge cases (OR = 11.7, 95% CI = 3.7–36.6), not for clear cases.

The crisis guardrail finding may be the most consequential failure mode exhibited in the entire study. What we found was worse than simple suppression. Trust calibration requires predictable system behavior—when reliability is inconsistent, users cannot learn when to rely on the system and when to override it15. A guardrail that fires for ‘haven’t thought through how I would do it’ but not for ‘thought about taking a lot of pills’ is not calibrated to clinical risk and users have no basis to anticipate when it will or will not fire. The capability to recognize mental health crises and connect users with crisis resources is a basic prerequisite for any consumer health platform. Our data show this prerequisite has not been reliably met. OpenAI has acknowledged, in a post titled ‘Helping people when they need it most,’ that model behavior in mental health contexts requires particular attention16. Our findings identify not a theoretical concern but a documented pattern of interstitial activation discordant with clinical severity.

This study has limitations. We used clinical vignettes rather than real-world patient interactions. Controlled studies of real users suggest this represents a conservative test—consumers under-report symptoms and misapply advice even when the system provides correct guidance, conditions that would compound the triage errors we report17. If ChatGPT Health undertriages 51.6% of emergencies with clean clinical information, performance with incomplete consumer inputs is unlikely to be superior. Emergency undertriage was concentrated in trajectory-dependent conditions where clinical evolution dictates urgency and whether this failure mode extends to other acute presentations remains untested. The standardized prompt required selection of a single triage level (A–D), capturing discrete recommendations rather than the hedged, multicontingency advice that open-ended interaction might produce. The within-vignette factorial design provides strong internal validity for manipulation effects but limits statistical power for detecting small demographic effects; however, the observed point estimates nonetheless provide useful bounds. We evaluated a single time point and model behavior may change with updates, only underscoring the need for ongoing evaluation as these systems evolve.

The implication is straightforward—consumer-facing AI that functions as a front door for urgent medical decisions should not be deployed on trust alone. Our findings identify two engineering targets requiring immediate attention—emergency detection that accounts for clinical trajectory, not just snapshot presentation and crisis guardrails that fire consistently rather than unpredictably. Given the direct patient-safety implications of missed emergencies, consumer health AI may warrant premarket safety evaluation requirements analogous to medical devices10. At a minimum, these tools should demonstrate external safety for emergencies before widespread public deployment.