ANALYSIS

A consumer AI health tool undertriaged half the emergencies in a stress test

Researchers ran 960 responses through ChatGPT Health using clinician-written vignettes. Failures clustered at both extremes, and crisis safeguards activated unpredictably.

Triage is the decision that sits in front of every other decision in medicine: how urgently does this person need to be seen, and by whom. It is also the decision consumer AI health tools are most often asked to make, because it is the question people actually type. A study published in Nature Medicine this week put one such tool through a structured stress test, and the pattern of its failures is more interesting than the headline error rate.

The test

ChatGPT Health was launched in January 2026 as OpenAI's consumer health tool and has reached millions of users [s1]. Researchers built 60 clinician-authored vignettes spanning 21 clinical domains and ran each of them under 16 factorial conditions — systematically varying elements of how the case was presented — producing 960 total responses [s1].

Factorial design is the key methodological choice here. Rather than asking whether the system gets cases right on average, it asks which features of a presentation change the answer, which is the question that matters for a tool used by people who describe their symptoms in wildly different ways.

Failures at both ends

Performance followed an inverted U-shaped pattern, with the most dangerous failures concentrated at the clinical extremes: 35% for non-urgent presentations and 48% for emergency conditions [s1].

That shape is worth pausing on. A system that is worst in the middle would be reassuring — ambiguous cases are genuinely hard. A system that is worst at both ends is failing precisely where the correct answer should be least ambiguous.

Among gold-standard emergencies, the system undertriaged 52% of cases [s1]. The paper gives specifics: patients with diabetic ketoacidosis or impending respiratory failure were directed to 24–48 hour evaluation rather than to the emergency department, while classical emergencies such as stroke and anaphylaxis were triaged correctly [s1]. The implied pattern — recognisable emergency scripts handled well, physiologically urgent but less scripted ones handled badly — is exactly what one would expect from a system trained on text rather than on physiology.

Anchoring, and inconsistent safeguards

Two further findings concern robustness rather than accuracy.

When family or friends in the vignette minimised the patient's symptoms — a manipulation the authors describe as testing anchoring bias — triage recommendations shifted significantly in edge cases, with an odds ratio of 11.7 (95% CI 3.7–36.6), and the majority of shifts moved toward less urgent care [s1]. The confidence interval is wide, but its lower bound still implies a substantial effect.

Crisis-intervention messages activated unpredictably across presentations involving suicidal ideation, occurring more frequently when the patient described no specific method than when they did [s1]. That is the opposite of the ordering a clinician would use.

On demographics, patient race, sex and barriers to care did not show significant effects — though the authors note the confidence intervals did not exclude clinically meaningful differences [s1]. That is a null result with a stated limit on how much reassurance it can carry, and it should be read that way.

The authors' conclusion is that these findings raise safety concerns warranting prospective validation before consumer-scale deployment of AI triage systems [s1] — a deployment that, as the paper itself notes, has already happened [s1].

The measurement problem behind it

A second paper, published four days later in npj Digital Medicine, addresses why results like this are hard to interpret across systems: there has been no common yardstick [s2]. The authors introduce MedTriage, a benchmark for evaluating large models across diverse clinical scenarios, and describe a competition built on it using real-world clinician–patient dialogues from general hospitals and four specialised domains [s2]. They also report an enhanced model developed from the competition's insights, using what they term a "10 Relevant + 10 Random + Ensemble" strategy [s2].

The paper's own framing is that evaluation-driven training improved performance, and that priorities from here include data security, model generalisation, and legal and regulatory frameworks [s2]. It is a benchmark paper, not an outcomes paper, and it does not establish that better benchmark scores translate into safer triage — which is precisely the gap the Nature Medicine result illustrates.

What clinical deployment looks like when it works

For contrast, a prospective study published the same day in the Western Journal of Emergency Medicine randomised 18,000 adult patients in a high-volume emergency department equally between an AI-supported triage system and two traditional methods, the Emergency Severity Index and the Manchester Triage System [s3]. Compared with the Manchester Triage System, the AI-supported system was associated with significantly lower in-department mortality (odds ratio 0.39, 95% CI 0.32–0.47) [s3].

The authors are explicit that the single-centre design and short observation period limit generalisability, and that causal inferences could not be firmly established [s3].

The comparison to draw is not "AI triage works in hospitals but not for consumers." It is that the two settings differ in the thing that matters most: in an emergency department, a trained human sees the patient regardless of what the algorithm says. In a consumer app, the algorithm's output is frequently the last step before the person decides whether to go anywhere at all.

What to watch

The Nature Medicine authors call for prospective validation [s1]. That would mean following real users after they receive a recommendation, rather than scoring responses against vignettes — a study design that is harder, slower and, so far, absent from the published record for consumer health chatbots.

Sources

  • [s1] ChatGPT Health performance in a structured test of triage recommendations. Nature Medicine, 23 February 2026. https://doi.org/10.1038/s41591-026-04297-7
  • [s2] Advancing medical AI through benchmarking and competition for specialty triage. npj Digital Medicine, 27 February 2026. https://doi.org/10.1038/s41746-026-02433-8
  • [s3] Impact of artificial intelligence-supported triage systems on emergency department management: a comparison of Infermedica, Emergency Severity Index, and Manchester Triage System. Western Journal of Emergency Medicine, 27 February 2026. https://doi.org/10.5811/westjem.48989

Sources

  1. ChatGPT Health performance in a structured test of triage recommendationsNature Medicine , February 23, 2026
  2. Advancing medical AI through benchmarking and competition for specialty triagenpj Digital Medicine , February 27, 2026
  3. Impact of artificial intelligence-supported triage systems on emergency department management: a comparison of Infermedica, Emergency Severity Index, and Manchester Triage SystemWestern Journal of Emergency Medicine , February 27, 2026
Related coverage