A consumer AI health tool undertriaged half the emergencies in a stress test
Researchers ran 960 responses through ChatGPT Health using clinician-written vignettes. Failures clustered at both extremes, and crisis safeguards activated unpredictably.
Triage is the decision that sits in front of every other decision in medicine: how urgently does this person need to be seen, and by whom. It is also the decision consumer AI health tools are most often asked to make, because it is the question people actually type. A study published in Nature Medicine this week put one such tool through a structured stress test, and the pattern of its failures is more interesting than the headline error rate.
The test
ChatGPT Health was launched in January 2026 as OpenAI's consumer health tool and has reached millions of users [s1]. Researchers built 60 clinician-authored vignettes spanning 21 clinical domains and ran each of them under 16 factorial conditions — systematically varying elements of how the case was presented — producing 960 total responses [s1].
Factorial design is the key methodological choice here. Rather than asking whether the system gets cases right on average, it asks which features of a presentation change the answer, which is the question that matters for a tool used by people who describe their symptoms in wildly different ways.
Failures at both ends
Performance followed an inverted U-shaped pattern, with the most dangerous failures concentrated at the clinical extremes: 35% for non-urgent presentations and 48% for emergency conditions [s1].
That shape is worth pausing on. A system that is worst in the middle would be reassuring — ambiguous cases are genuinely hard. A system that is worst at both ends is failing precisely where the correct answer should be least ambiguous.
Among gold-standard emergencies, the system undertriaged 52% of cases [s1]. The paper gives specifics: patients with diabetic ketoacidosis or impending respiratory failure were directed to 24–48 hour evaluation rather than to the emergency department, while classical emergencies such as stroke and anaphylaxis were triaged correctly [s1]. The implied pattern — recognisable emergency scripts handled well, physiologically urgent but less scripted ones handled badly — is exactly what one would expect from a system trained on text rather than on physiology.
Anchoring, and inconsistent safeguards
Two further findings concern robustness rather than accuracy.
When family or friends in the vignette minimised the patient's symptoms — a manipulation the authors describe as testing anchoring bias — triage recommendations shifted significantly in edge cases, with an odds ratio of 11.7 (95% CI 3.7–36.6), and the majority of shifts moved toward less urgent care [s1]. The confidence interval is wide, but its lower bound still implies a substantial effect.
Crisis-intervention messages activated unpredictably across presentations involving suicidal ideation, occurring more frequently when the patient described no specific method than when they did [s1]. That is the opposite of the ordering a clinician would use.
On demographics, patient race, sex and barriers to care did not show significant effects — though the authors note the confidence intervals did not exclude clinically meaningful differences [s1]. That is a null result with a stated limit on how much reassurance it can carry, and it should be read that way.
The authors' conclusion is that these findings raise safety concerns warranting prospective validation before consumer-scale deployment of AI triage systems [s1] — a deployment that, as the paper itself notes, has already happened [s1].
The measurement problem behind it
A second paper, published four days later in npj Digital Medicine, addresses why results like this are hard to interpret across systems: there has been no common yardstick [s2]. The authors introduce MedTriage, a benchmark for evaluating large models across diverse clinical scenarios, and describe a competition built on it using real-world clinician–patient dialogues from general hospitals and four specialised domains [s2]. They also report an enhanced model developed from the competition's insights, using what they term a "10 Relevant + 10 Random + Ensemble" strategy [s2].
The paper's own framing is that evaluation-driven training improved performance, and that priorities from here include data security, model generalisation, and legal and regulatory frameworks [s2]. It is a benchmark paper, not an outcomes paper, and it does not establish that better benchmark scores translate into safer triage — which is precisely the gap the Nature Medicine result illustrates.
What clinical deployment looks like when it works
For contrast, a prospective study published the same day in the Western Journal of Emergency Medicine randomised 18,000 adult patients in a high-volume emergency department equally between an AI-supported triage system and two traditional methods, the Emergency Severity Index and the Manchester Triage System [s3]. Compared with the Manchester Triage System, the AI-supported system was associated with significantly lower in-department mortality (odds ratio 0.39, 95% CI 0.32–0.47) [s3].
The authors are explicit that the single-centre design and short observation period limit generalisability, and that causal inferences could not be firmly established [s3].
The comparison to draw is not "AI triage works in hospitals but not for consumers." It is that the two settings differ in the thing that matters most: in an emergency department, a trained human sees the patient regardless of what the algorithm says. In a consumer app, the algorithm's output is frequently the last step before the person decides whether to go anywhere at all.
What to watch
The Nature Medicine authors call for prospective validation [s1]. That would mean following real users after they receive a recommendation, rather than scoring responses against vignettes — a study design that is harder, slower and, so far, absent from the published record for consumer health chatbots.
Sources
- [s1] ChatGPT Health performance in a structured test of triage recommendations. Nature Medicine, 23 February 2026. https://doi.org/10.1038/s41591-026-04297-7
- [s2] Advancing medical AI through benchmarking and competition for specialty triage. npj Digital Medicine, 27 February 2026. https://doi.org/10.1038/s41746-026-02433-8
- [s3] Impact of artificial intelligence-supported triage systems on emergency department management: a comparison of Infermedica, Emergency Severity Index, and Manchester Triage System. Western Journal of Emergency Medicine, 27 February 2026. https://doi.org/10.5811/westjem.48989
Sources
- ChatGPT Health performance in a structured test of triage recommendations — Nature Medicine , February 23, 2026
- Advancing medical AI through benchmarking and competition for specialty triage — npj Digital Medicine , February 27, 2026
- Impact of artificial intelligence-supported triage systems on emergency department management: a comparison of Infermedica, Emergency Severity Index, and Manchester Triage System — Western Journal of Emergency Medicine , February 27, 2026
AI chatbots ace medical scenarios alone. Give them to real people, and accuracy collapses.
A 1,298-person randomized study found people using chatbots to work through medical scenarios did no better than people without them — even though the same chatbots, tested alone, got the right answer most of the time.
When AI symptom-checkers disagree with patients, patients tend to walk away
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.
Most AI models that predict who will skip their medicines aren't ready for the clinic
A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.
Two FDA-cleared AI tools read prostate biopsies. A review flags what could go wrong.
A review of Paige Prostate Detect and Ibex Prostate Detect finds the tools mainly help less-specialized pathologists, and warns performance shifts when a tool meets a new population.